Errors and observabilitymedium3-5 years

A GraphQL API's uptime dashboard, built on HTTP status codes, reports 100% success for an hour during which half of all queries were actually failing. How is that possible, and what should the dashboard have measured instead?

A GraphQL response is { "data": ..., "errors": [...] }, and the HTTP status stays 200 as long as the query itself was parseable — a resolver throwing doesn't change the status code, it puts that field's value to null and adds an entry to the errors array instead. So an HTTP-status dashboard genuinely can't see resolver failures at all: every request that reaches a resolver and gets a response, successful or not, shows up as a 200. The metric that actually reflects health is a response with a non-empty errors array, which the lesson calls out directly as what the observability course's error metric should count here — not HTTP status, which was designed around REST's one-status-per-response model and doesn't map onto a protocol where partial failure is normal and still returns 200.

The lesson behind it →