Most agent failures are described as if nothing happened. The model got confused. The tool timed out. The workflow stalled. The run should be retried.
That framing is safe only when the agent was still inside the chat box. Once an agent can modify a CRM record, submit an order, send a customer email, update a ticket, change a calendar, trigger a deployment, or write into an operational database, “retry” becomes a dangerous word. The first attempt may have failed from the agent’s point of view while still succeeding somewhere else.
That is the gap a compensating-action plan is meant to close.
A retry answers, “Should we try this again?” A compensating-action plan answers a harder question: “If this action partially succeeded, how do we detect that, limit the damage, reverse what can be reversed, document what cannot, and decide whether a human must take over?”
Distributed-systems teams have lived with this problem for years. The Azure Architecture Center’s compensating transaction pattern is built around operations that span multiple steps, services, or data stores where strong transactional consistency is not realistic. If one step fails after earlier steps have already changed state, the system needs a reliable way to undo or mitigate the earlier work. The pattern also emphasizes two details that agent teams should steal immediately: record progress so recovery can resume, and make compensating steps idempotent because they may be retried too.
AI agents make the same problem feel new because the planner is probabilistic and the side effects are often wrapped in friendly language. A human sees “I could not finish that workflow” and assumes nothing happened. The backend may tell a different story: the invoice was drafted, the customer note was appended, the order was reserved, and the follow-up email failed. A second run might duplicate the note, reserve the order again, or send a message that no longer matches the state of the account.
Guardrails help, but they are not compensation. Input guardrails, output guardrails, tool guardrails, and tripwires are prevention layers. They reduce the chance that an agent starts the wrong action or emits unsafe output. They should absolutely exist. But a production design also needs to assume that allowed actions can still fail halfway through, that APIs can return ambiguous statuses, and that an apparently successful tool call can create downstream obligations.
The practical unit is a compensating-action contract for every meaningful side effect. Before giving an agent a tool, define the normal action, the evidence that proves whether it happened, the acceptable inverse or mitigation, and the authority required to run that inverse. For a ticket update, compensation might be an appended correction rather than deletion. For a payment action, it might be a void, refund, or escalation to finance. For an email, it may be impossible to “undo” the send, so the compensation is a follow-up correction plus a customer-service review. For a deployment, it may be rollback to a known version plus an incident note.
The important part is deciding this before the agent is live.
A good compensating-action plan has seven fields.
First, it has an action ledger entry. Every side effect gets a stable action ID, tool name, target resource, requested change, actor, timestamp, and correlation ID. Without that, the system is guessing about what it is trying to repair.
Second, it has a completion detector. The detector should not rely only on the agent’s final message. It should query the destination system, inspect receipts, or verify state through an independent read. “The tool returned an error” and “the business action did not happen” are not the same claim.
Third, it has an idempotency rule. The original action and the compensating action both need keys or state checks that make repeated execution safe. If compensation can duplicate harm, it is not a recovery mechanism; it is another incident path.
Fourth, it has an authorization boundary. The Model Context Protocol authorization specification is useful framing here because it treats protected resources, clients, resource owners, and access tokens as explicit roles. Compensation should not run with vague ambient power simply because the original tool had access. Reversing a payroll change, canceling an order, or rewriting a medical scheduling note may require a different scope, a fresh approval, or a different human role.
Fifth, it has a human-review threshold. Some compensation can run automatically when the state is clear and the impact is low. Other cases should pause. Ambiguous customer impact, financial movement, regulated records, irreversible communication, and conflicting source-of-truth reads are all reasons to stop the agent and put the case into an exception queue.
Sixth, it has trace evidence. OpenAI’s Agents SDK tracing documentation is a reminder that agent runs are not just model calls; they are sequences of traces and spans across tools, handoffs, and processors. The compensation should be traceable as part of the same operational story: what the agent intended, what it attempted, what external state changed, which detector fired, who approved the fix, and what the final state became.
Seventh, it has a final receipt. The user, operator, or downstream system should not receive a vague “handled” message. They should receive a completion receipt that says whether the original action completed, whether compensation ran, which resources were touched, and what remains unresolved.
This is also where governance stops being abstract. The NIST AI Risk Management Framework pushes teams to govern, map, measure, and manage AI risks. For agentic workflows, compensating actions are one concrete way to make that real. They turn risk management into a design question: which side effects are reversible, which are only mitigable, which require human review, and which should not be delegated to an agent yet?
The implementation does not need to start as a grand platform. Start with one high-value workflow and classify every tool call into three buckets: read-only, reversible side effect, and irreversible or externally visible side effect. Read-only calls need observability and access control. Reversible side effects need idempotent compensation. Irreversible side effects need stricter pre-approval, clearer receipts, and escalation paths.
Then test the ugly cases. Force the tool to time out after the remote system commits. Return a malformed response after a state change. Drop the network between step three and step four. Run the compensation twice. Revoke the agent’s original credential and see whether the recovery path still respects authorization. Ask an operator to reconstruct what happened from the trace alone. If they cannot, the agent is not production-ready; it is just lucky in demos.
The deeper lesson is simple: agents that act in the world inherit the old obligations of workflow systems, audit systems, and incident response. Better prompts will not remove those obligations. Larger context windows will not make partial side effects safe. A retry loop without compensation is optimism disguised as reliability.
Before an AI agent gets authority to change something important, ask for its compensating-action plan. If the team cannot explain how the system detects partial success, how it repairs or mitigates the result, who can approve the repair, and what evidence proves the final state, the agent should not have autonomous authority for that action yet.
The future of agent reliability will not be measured only by how often agents succeed. It will be measured by how cleanly they recover when success is partial, messy, and already written into someone else’s system.
Build AI Systems That Survive Contact With Real Work
We help teams turn AI research into practical automations, agent workflows, and operational systems that can be evaluated and improved.
Get the Field Guide — $10 →