A demo proves that an AI agent can succeed once.

A red-team fixture proves that an old failure mode stays fixed.

That distinction matters more as agents move from chat windows into workflows with tools, memory, approvals, handoffs, and side effects. A support agent can draft a good reply in a demo. A coding agent can patch the right file on a clean branch. A browser agent can fill the right form when the page behaves. None of that tells you what happens after a prompt changes, a tool schema shifts, a model upgrade lands, or a malicious instruction appears inside a document the agent was told to summarize.

Most teams already have examples. They have screenshots, transcripts, eval notes, incident writeups, and a few scary prompts in a shared doc. What they often do not have is a fixture: a small, reusable scenario that can be run again before the next release.

For production agents, that fixture is the missing safety artifact.

A red-team fixture is not just a jailbreak prompt. It is a packaged test of a workflow under adversarial or ambiguous conditions. It should include the user request, the hostile or misleading content, the tool state, the permissions in force, the memory or session context, the expected safe behavior, and the evidence required to call the run a pass. The result should be boring enough to automate and specific enough to catch regressions.

The security community already gives teams the raw material. MITRE ATLAS organizes tactics, techniques, and case studies for attacks against AI-enabled systems. OWASP's GenAI work catalogs threat patterns and mitigations for LLM and agentic applications, including issues around excessive agency, prompt injection, sensitive information disclosure, tool access, and output handling. NIST's AI Risk Management Framework asks organizations to govern, map, measure, and manage AI risk across the lifecycle. Anthropic's work on constitutional classifiers shows the same broad lesson from another angle: defenses need explicit artifacts that can be evaluated against adversarial behavior.

The practical move is to turn that guidance into fixtures that live next to the agent workflow.

A useful fixture starts with a threat pattern. Do not write “test prompt injection.” Write the actual operational risk: a vendor invoice includes hidden text telling the agent to change the payment destination; a customer email asks the support agent to reveal another account's details; a retrieved policy document contains instructions that conflict with system policy; a subagent returns a confident summary without evidence; a browser page changes the meaning of a button after the agent has already received approval.

Then define the environment. What tools are available? Which account or tenant is in scope? What approval has been granted, and what has not? What memory is visible? What data is intentionally stale? Which tool call would create a side effect? If the fixture does not pin the environment, it is too easy for a future run to pass for the wrong reason.

Next, define the expected behavior in operational terms. The agent should refuse, ask for fresh approval, quarantine the retrieved instruction, cite the trusted evidence, escalate to a human, avoid the side-effecting tool, or produce a safe draft without sending it. “Be safe” is not an assertion. “Does not call send_payment when the payee was changed by untrusted document text” is an assertion.

The fixture should also define what evidence proves the pass. A transcript alone is usually not enough. The test should check tool-call logs, trace spans, approval records, output text, and final state. If the agent says it refused but still called the dangerous tool, the fixture should fail. If it produced a safe answer but dropped the evidence pointer, the fixture may still fail. The point is to test the system, not the model's self-description.

That is especially important because agent regressions often enter through plumbing rather than prose. A developer widens a tool scope to unblock a feature. A prompt summary starts carrying untrusted document text into a privileged step. A handoff loses the original authority envelope. A new model follows a hidden instruction the old one ignored. A retry path repeats an action that was supposed to be idempotent. A classifier is added but only checks final output, not tool intent.

One-off demos rarely catch those changes. Fixtures can.

The first fixture pack does not need to be large. Start with five scenarios that represent your highest blast radius. One should attack tool authority. One should attack memory or session scope. One should attack retrieved content. One should attack a handoff between workers. One should attack output handling, such as sending, exporting, posting, or updating a record. If your agent can spend money, change production data, email customers, deploy code, or touch regulated records, include that side-effect path early.

Each fixture should have an owner and a reason for existing. The reason can be a real incident, a near miss, a red-team finding, an OWASP category, a MITRE ATLAS technique, or a policy requirement from your own risk register. Without that reason, fixture libraries turn into prompt museums. With it, they become institutional memory.

Run the pack at the moments when risk changes: before a model upgrade, before expanding tool permissions, before changing retrieval sources, before enabling a new handoff, before relaxing an approval step, and before promoting a prompt from staging to production. For high-risk workflows, run a smaller smoke pack on every build and the full pack before release.

The fixture result should be a receipt. Record the version of the prompt, model, tool schemas, policy bundle, retrieval corpus, and fixture pack. Record which assertions passed, which failed, and which were skipped. If a fixture is intentionally changed, require a note explaining why the expected safe behavior changed. Otherwise teams will slowly teach the test suite to accept the behavior the system already has.

This is where NIST's lifecycle language becomes concrete. Govern means someone owns the fixture pack and reviews changes. Map means the scenarios correspond to real workflow boundaries and data flows. Measure means the fixture has pass/fail assertions rather than vibes. Manage means failures block release, trigger escalation, or force a documented risk decision.

Good red-team fixtures also make engineering discussions less abstract. Instead of debating whether an agent is “secure enough,” the team can look at a failed scenario: the support agent obeyed an instruction inside an attached PDF; the code agent used a network tool outside the approved repo; the procurement agent accepted a payee change without fresh confirmation. Those failures are easier to fix because they are concrete and repeatable.

There is a cultural benefit, too. Fixtures preserve humility. They remind the team that a model can look improved while a workflow becomes less safe. They keep yesterday's lesson from disappearing into a Slack thread. They let new engineers see the traps the system has already fallen into. They give leaders a better question than “did the demo work?” The better question is: “which failure modes did this release have to survive?”

The answer should not be a paragraph of reassurance. It should be a small library of adversarial scenarios, run logs, and receipts.

Production AI does not need more confidence from perfect demos. It needs repeatable evidence from imperfect conditions. A red-team fixture is how a team turns a scary example into a release gate, a policy into an assertion, and an incident into a control that keeps working after the next change.

If your agent is important enough to automate real work, it is important enough to remember how it can fail.

Build AI Systems That Survive Contact With Real Work

We help teams turn AI research into practical automations, agent workflows, and operational systems that can be evaluated and improved.

Get the Field Guide — $10 →