A tool service verifies bearer tokens using a cached set of public signing keys. The identity provider starts signing new tokens with a new key ID, kid. Some tool replicas still have a key set that contains only the old key. A signature check cannot even select the right key, so valid new tokens fail at those replicas. The failure looks random when the load balancer sends successive calls to different replicas.

OpenID Connect's signing-key rotation guidance describes publishing keys through a JWK set and coordinating rotation with cache duration. A safe planned rotation publishes the new public key before issuing tokens signed by it, keeps the old key available while old tokens and caches can still need it, and then retires it. The exact timing depends on cache lifetimes and token validity. In an emergency key compromise, the availability trade-off changes because accepting the old key may be unsafe.

On an unknown kid, a verifier can fetch the configured issuer's JWK set with bounded refresh and backoff. It must not fetch an arbitrary URL from the untrusted token or turn a network failure into “signature probably valid.” Validate issuer, audience, algorithm policy, time claims and signature using trusted configuration. Rate-limit refreshes so an attacker sending random key IDs cannot create a request storm against identity infrastructure.

Diagnose by correlating rejection rate with kid, issuer, verifier replica and JWK cache age. Check whether the new key was published before use and whether caches honor expected refresh rules. A canary token signed with the upcoming key, tested across all replicas before cutover, catches a planned-rotation failure. A rollback plan must account for tokens already issued under the new key.

The tempting shortcut is to disable signature verification until the key service recovers. That changes an availability incident into an authentication failure. Keep a verified, bounded cache of trusted keys and a rehearsed rotation process instead.