Security, Governance and Platform · Principal
Two agent runs refresh the same OAuth credential. Why does one lose access?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Suppose both workers read refresh token R0 for the same connector account. Worker A exchanges R0 and receives R1. Worker B, still holding R0, tries to exchange it too. If the provider rotates refresh tokens and invalidates the old one, B can get a reuse error. A narrow grace period may allow both requests, but then their returned successors can race in our storage. The exact consequences are provider-specific. Amazon Cognito's rotation API explicitly describes issuing a new refresh token, invalidating the old one after an optional grace period, and a reuse exception.
The fix belongs in the credential service, not in a prompt telling agents to retry carefully. Give each connected account one refresh coordinator. On access-token expiry, a worker reads the current credential version, takes a short account-scoped lease or joins an in-flight refresh, and exchanges only that version. Store the new access and refresh tokens together with a compare-and-swap on the credential version. Other workers reread the stored result. A process-local mutex alone does not coordinate replicas.
The awkward case is a timeout after the provider rotated R0 but before we stored R1. Blindly replaying R0 can be wrong, and a local transaction cannot atomically commit the remote token exchange. Check whether the provider offers a grace or recovery mechanism. If it does not, mark the connection for explicit reauthorization rather than retrying indefinitely or claiming the agent completed its task. Keep the old credential encrypted only as long as the provider contract and recovery procedure require, and never log token values.
I would test two replicas refreshing simultaneously, a crash after provider success, a delayed stale writer and a revoked user grant. Metric-wise, distinguish temporary refresh contention from a permanently invalid connection and count affected runs by account. The identity provider rotated its signing key. Why did every agent tool call fail? covers verifying an identity provider's signing keys. This problem is about our ownership of a rotating secret used to obtain access tokens.
Continue reading
Related questions
Read beyond the question
Explore more security, governance and platform
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →