Resource serversenior8+ years

Your resource servers validate RS256 access tokens (15-minute lifetime) using `issuer-uri`. During an identity-provider key rotation, some pods reject valid tokens for a few minutes; a month later, a 20-minute identity-provider outage causes a partial outage of your API even though no tokens changed. Explain how the resource server finds keys, what went wrong in each incident, and what rotation procedure and settings you would agree with the identity team.

The resource server does not hold a configured key. It discovers them: issuer-uri → /.well-known/openid-configuration → jwks_uri → a JWK Set of public keys, each with a kid. Every token's header names the kid it was signed with, and the decoder verifies with that entry only, from a cached copy of the set. The rotation incident is a cache that did not know the new kid yet: the provider started signing with a new key, and pods whose cached set predated it rejected those tokens until they fetched the set again. The outage incident is the same cache from the other side: a pod needing to refresh the set (expiry, a new pod starting, an unknown kid) could not reach jwks_uri, and with no key it fails closed. The procedure: publish a new key before signing with it, keep a retired key published for at least the maximum token lifetime after signing stops, confirm the decoder re-fetches on an unknown kid, and size cache behaviour to ride out a short provider outage.

The lesson behind it →
More on Resource server