Security, Governance and Platform · Principal
The identity provider rotated its signing key. Why did every agent tool call fail?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
A tool service verifies bearer tokens using a cached set of public signing keys. The identity provider starts signing new tokens with a new key ID, kid. Some tool replicas still have a key set that contains only the old key. A signature check cannot even select the right key, so valid new tokens fail at those replicas. The failure looks random when the load balancer sends successive calls to different replicas.
OpenID Connect's signing-key rotation guidance describes publishing keys through a JWK set and coordinating rotation with cache duration. A safe planned rotation publishes the new public key before issuing tokens signed by it, keeps the old key available while old tokens and caches can still need it, and then retires it. The exact timing depends on cache lifetimes and token validity. In an emergency key compromise, the availability trade-off changes because accepting the old key may be unsafe.
On an unknown kid, a verifier can fetch the configured issuer's JWK set with bounded refresh and backoff. It must not fetch an arbitrary URL from the untrusted token or turn a network failure into “signature probably valid.” Validate issuer, audience, algorithm policy, time claims and signature using trusted configuration. Rate-limit refreshes so an attacker sending random key IDs cannot create a request storm against identity infrastructure.
Diagnose by correlating rejection rate with kid, issuer, verifier replica and JWK cache age. Check whether the new key was published before use and whether caches honor expected refresh rules. A canary token signed with the upcoming key, tested across all replicas before cutover, catches a planned-rotation failure. A rollback plan must account for tokens already issued under the new key.
The tempting shortcut is to disable signature verification until the key service recovers. That changes an availability incident into an authentication failure. Keep a verified, bounded cache of trusted keys and a rehearsed rotation process instead.
Continue reading
Related questions
Read beyond the question
Explore more security, governance and platform
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →