At 2:14am your support agent issues a refund. Four thousand dollars, back to a customer, with nobody awake. In the morning someone asks the ordinary question: who approved that?
You open the ledger. The row says:
principal = svc-agent-prod
operation = refund
amount = 4000.00
customer = 78122
at = 02:14:07
svc-agent-prod is a service account: one set of long-lived credentials that the whole agent deployment uses for every action it takes, for every customer, on behalf of everyone. It is the name of a running process, not the name of a decision.
So the row establishes one fact. The production agent infrastructure caused this. It cannot tell you the thing you actually want to know.
Three different stories fit that row
At least three very different sequences produce that row, and it looks identical for all of them. A customer asked for a refund and the agent processed it. A customer asked about their bill and the agent decided a refund was the helpful answer. Or nobody asked, and an instruction buried in a support ticket told the agent to issue it.
That is the difference between a customer causing an action, a customer authorizing one, and an agent deciding alone. Your ledger cannot tell them apart, and neither can your on-call engineer at 8am.
Given a few hours you can usually still work it out from the chat transcript, the application logs and a trace. So be precise: the chain has not vanished, it has turned into archaeology. You reconstruct it afterwards across systems that were never designed to agree, and nothing you reconstruct is bound to the authorization decision itself. A trace proves two events belong to one request. It does not prove anyone was allowed to act.
This is not a rare configuration. In a January 2026 survey of 418 IT and security professionals by the Cloud Security Alliance, commissioned by Token Security, 68% said they had strong visibility into the agents in their environment and 82% had found at least one agent nobody had registered. Self-reported surveys measure belief, not systems, which is why the gap between those two numbers is the interesting part.
The chain was never there at hop one
The natural assumption is that identity gets lost somewhere deep in the call stack. Trace one tool call through the popular frameworks and you find something plainer.
In the OpenAI Agents SDK the client that talks to the model is built once from a key in a process-level global, not per request, and a tool call is your own Python function plus JSON arguments with no credential injected into it. Whatever that function uses to reach the outside world, it brought itself. LangGraph is the same shape: its tool node injects state, a store and a runtime object, and no authorization header. Off LangGraph’s hosted server the runtime’s user field is documented as None.
Now look at what crosses when one agent hands off to another: conversation items and an opaque application context. No field for the person who started this, none for what they authorized, none for how far that authority reaches. You can put a user ID in your own context object and nothing downstream will turn it into a credential.
So the frameworks do not lose the chain at hop three. They pass conversation, not authority, starting at hop one.
The protocols underneath do not rescue this. In the Model Context Protocol, which is how agents reach tools, authorization is explicitly optional, and over a standard input and output connection the spec says to take credentials from the environment instead; a tool call carries a name and arguments. The Agent2Agent protocol, which is how agents reach each other, is blunter: identity is handled at the protocol layer, not within its semantics, and credentials are obtained out of band. Both are defensible protocol design. Neither gives the payments API the sentence it needs.
Four things can cross a hop
There are exactly four options at each step, and the choice decides what an auditor can rebuild later.
You can forward the customer’s own token, which is honest about the human, silent about the agent, and hands over every permission the customer has rather than the one the agent needs. You can swap to the agent’s service account, which is what most stacks do, and lose the human entirely. You can exchange the credential for a short-lived one naming both parties and only the scope required. Or you can carry a workload credential plus a separately signed record of who delegated what, which works only if the receiving service checks that record instead of trusting whoever handed it over.
The third option has a specification, and it is not new. RFC 8693, OAuth 2.0 Token Exchange, has been published since 2020. It defines a subject_token for the party on whose behalf a request is made, an actor_token for the party doing the acting, and an act claim naming the delegate. It even says that a chain of delegation can be expressed by nesting one act claim inside another, which is precisely the refund problem written down years before anyone deployed a support agent.
The mental model that keeps this safe is authority reduction. An exchanged token should carry the intersection of what the customer may do, what the agent may do, and what the target service accepts. It is never an upgrade.
Rossoctl, an open source platform whose stated purpose is platform primitives for trustworthy AI agents, does implement token exchange. Immediately above the call that performs it:
// RFC 8693 Section 4.1 actor-token "act" claim chaining is not yet
// wired by any plugin; the wire-format support stays in
// exchange.ExchangeRequest.ActorToken for when a plugin needs it.
The field for the acting party is on the request structure. Nothing fills it in. That is not a knock on a project that is further along than most of its peers, and it is a fair picture of the state of the art: the delegation slot exists, and it is empty.
Four questions, four different mechanisms
The reason this stays unsolved is that “who authorized that?” is not one question. Workload identity answers what this running process is. Token exchange answers under whose authority it may act. An authorization policy answers whether this action is permitted. An audit record answers what happened. Teams buy a tool for one of those and expect the whole answer: cryptographic workload identity is excellent and will never tell you a customer approved a refund, and distributed tracing is excellent and is correlation, not authorization.
There is a fifth thing that no amount of identity plumbing replaces. Above some threshold, the action itself needs approving, with the amount and the account and the policy version written into the authorization record. You do not solve a $4,000 refund by adding another claim to a token.
The honest cost, and when a service account is right
Doing this properly is not free, and the frameworks that demand accountability are conspicuously silent on the bill. There is no published figure for what a token exchange costs you per hop, so do not believe one.
What you can say concretely is what you are signing up to run. The service that issues tokens is now in the path of live requests, so it is an availability dependency, and when it is down the exchange fails. Short-lived credentials shrink your revocation window and raise your issuance traffic. Every agent identity acquires a lifecycle, which means somebody has to answer who owns agent-8f7d when it needs rotating. The policy engine that decides whether the refund is allowed can now take the refund path down with it, so you must decide per resource whether a policy outage fails open or closed. A $4,000 refund fails closed. Fetching product documentation can degrade.
And the most realistic failure is not dramatic. The first two hops carry the delegation faithfully and the third one drops it, so the chain looks healthy in testing and is broken exactly where the money moves.
The cost scales with the number of trust boundaries you cross, not with the number of agents you run. which is what makes this defensible rather than lazy: a shared service account is fine when no individual person’s authority is being represented. A metrics collector, a nightly job clearing cache, a read-only summarizer inside one system with a small blast radius. Those do not need any of this, and wiring them up because a vendor said machine identities are the future is cargo cult work. Give an agent its own identity when you need to revoke it separately during an incident. Reach for token exchange when the authorization genuinely depends on which person caused the action.
Gartner’s Shiva Varma put the failure mode well in May 2026: enterprises treat agent governance as binary, either locked down or fully trusted, and that is the root cause of failure.
Not a standards vacuum. An integration vacuum.
It would be comforting to conclude that the industry is waiting on a standard. It is not. Token exchange with a delegation claim was specified in 2020. Workload attestation is mature. Policy engines are commodity. Audit pipelines are a solved problem in every other regulated system your company runs.
What is missing is the wiring between them, and you can watch that gap in a code comment where the delegation field sits on a request struct with nothing to fill it. Meanwhile the incidents keep arriving in a shape that makes the point better than any framework does. When researchers demonstrated the RovoBlast flaw in Atlassian’s AI assistant this July, the most instructive detail was not that a crafted link could make the assistant exfiltrate data using the victim’s own privileges. It was that chaining the steps inside a single agent run left behind an audit trail that looked like ordinary research activity. That is the failure: not an agent doing something it should not, but an agent doing something consequential and leaving a record that cannot be told apart from routine work.
So when you review your own stack, skip the question of whether your agents have identities. Ask the harder one: for the most consequential thing an agent in your company can do, what evidence reaches the other end, and would it survive somebody asking, nine months later, who said yes? Building that answer into the retrieval, authorization, and audit layers of production systems is the work we do at Gracient. Let’s talk.