A spinner is fine when a web form takes three seconds.
It is not fine when an AI agent is reconciling invoices, drafting a customer response, querying a vendor API, preparing a deployment, waiting on human approval, or handing work to a second agent. At that point, “working…” is not a status. It is a gap in the operating system.
The more useful pattern is a progress signal: a structured event that says what phase the agent is in, what it is waiting on, what evidence supports that state, whether the run is still authorized, whether cancellation is safe, and when the next transition should happen. A progress signal is not there to make the user feel better. It is there so the system can decide what to do while the agent is still alive.
That distinction matters because agent work is increasingly long-running and multi-step. A model may be fast, but the workflow around it is not. Tools have latency. APIs rate-limit. Browsers hang. Files change underneath the run. Subagents disagree. Human approvals expire. A vendor endpoint returns a partial result. A retry may be harmless in one phase and dangerous after a side effect has committed.
A spinner hides all of that behind animation.
A progress signal exposes enough state for the product, runtime, and operator to respond.
The Model Context Protocol already points in this direction with its progress utility for long-running requests. MCP progress notifications let a server report progress during an operation instead of waiting until the end to return success or failure. That is a small protocol feature with a large design lesson: once tools become part of agent workflows, the runtime needs in-flight visibility. The agent should not be a silent process that eventually produces a paragraph.
OpenAI’s Agents SDK makes a related point from the tracing side. Its tracing documentation treats agent activity as runs and spans: the kinds of execution units operators can inspect across generations, function calls, handoffs, and workflow steps. Traces are often discussed after the fact, for debugging or review. But progress signals should connect to the same structure while the work is happening. If a user sees “validating vendor response,” an operator should be able to find the span behind that claim.
OpenTelemetry’s generative AI semantic-convention work reinforces the broader direction. The details will keep evolving, but the signal is clear: AI operations need portable observability vocabulary. Agent progress should not live only in chat transcripts or frontend strings. It should be telemetry that can be queried, correlated, alerted on, and joined to logs, traces, tool receipts, approval records, and incident timelines.
The practical version does not need to be complicated. Start with a progress event that includes a run ID, session ID, phase, status, timestamp, last heartbeat, next expected transition, blocking dependency, cancelability, authority state, evidence pointer, and human-facing message. That is enough to separate “still working” from “stalled,” “waiting on user,” “waiting on vendor,” “retrying safely,” “retry would duplicate a side effect,” “approval expired,” and “ready for escalation.”
Those states are not interchangeable.
If an agent is gathering information, cancellation may be safe. If it has already submitted a payment, cancellation may mean starting a compensation workflow. If it is waiting on a human approval, retrying the tool will not help. If it is waiting on a vendor, the user needs an expectation and the operator needs a timeout. If the agent has stopped emitting heartbeats, the runtime should not wait forever just because the UI still has a spinner.
This is where lifecycle hooks become useful. Anthropic’s Claude Code hooks documentation describes deterministic commands that can run at events such as prompt submission, pre-tool-use, post-tool-use, notification, stop, and subagent stop. The broader architectural lesson applies beyond one product: agent systems need places where policy can observe the workflow without relying on the model to narrate itself correctly.
A pre-tool hook can emit “about to call external API.” A post-tool hook can emit “tool returned, evidence stored.” A stop hook can mark the run complete. A notification hook can distinguish “waiting for permission” from “still thinking.” A subagent-stop hook can record that delegated work returned control to the parent. These events are more trustworthy than asking the model to summarize its own progress every few seconds.
The user-facing layer should be simple. People do not need a trace dump. They need honest labels: collecting records, checking policy, waiting for approval, contacting vendor, retrying a transient failure, preparing final answer, or blocked and needs help. They also need to know whether they can safely cancel and what will happen if they do.
The operator-facing layer should be sharper. It should show phase names, run and span IDs, tool names, stale thresholds, elapsed time, retry counts, approval state, data scope, and the last evidence pointer. If the progress event says “waiting on external API,” the operator should know which API, when the wait began, what timeout applies, and whether another call would be idempotent.
That is the difference between cosmetic status and operational status.
NIST’s AI Risk Management Framework uses govern, map, measure, and manage as core functions. Those words can feel abstract until they become runtime artifacts. Progress signals are one of those artifacts. They help map what the agent is doing, measure whether it is behaving within expected bounds, manage timeouts and escalation, and govern which states are allowed to proceed without fresh human input.
They also improve trust in a more ordinary way: they make bad news visible sooner.
A good progress system should make stalled work look stalled. It should make ambiguous authority look ambiguous. It should make a missing evidence pointer obvious. It should make a repeated retry visible before the customer receives duplicate emails or the finance system receives duplicate requests. The goal is not to make agents seem more confident. The goal is to make uncertainty actionable.
For teams building production agents, the rollout can start small.
Pick one workflow where a spinner currently hides real complexity. Define five to seven phases with clear entry and exit conditions. Add structured progress events at tool boundaries, approval boundaries, handoffs, retries, and completion. Attach each event to the trace. Define stale thresholds for each phase. Decide which states are cancelable, which require escalation, and which must preserve an audit record before continuing.
Then test the unhappy paths. What does the user see when a vendor API hangs? What does the operator see when the agent stops emitting heartbeats? What happens when approval expires mid-run? What happens when a subagent completes but fails to return evidence? What happens when a retry would repeat a side effect?
If the answer is still “the spinner keeps spinning,” the system is not ready.
The next generation of agent reliability will not come only from better models or longer context windows. It will come from runtimes that expose the state of work clearly enough for humans and software to intervene. A spinner says, “trust me.” A progress signal says, “here is where the work is, here is why, here is what can happen next.”
Production agents need the second one.
Build AI Systems That Survive Contact With Real Work
We help teams turn AI research into practical automations, agent workflows, and operational systems that can be evaluated and improved.
Get the Field Guide — $10 →