gRPC reliabilitymedium3-5 years
A gRPC call sets a client-side timeout of 2 seconds and calls three downstream services in sequence. Two seconds after the call starts, the first downstream call is still in progress. What happens, and how is this different from each service independently having its own 2-second timeout?
A gRPC call carries a deadline — an absolute point in time set once by the original caller — not a per-hop timeout that resets at each service. That deadline is propagated to every downstream call the server makes on the request's behalf, so if the first downstream call has already used 1.5 of the 2 seconds, the second downstream call starts with only 0.5 seconds left, not a fresh 2. If each service independently ran its own 2-second timeout instead, the first hop alone could take 2 seconds, then the second another 2, then the third another 2 — six seconds total, with no service ever technically violating its own limit, while the caller who only wanted to wait 2 seconds has no idea any of that happened.
PreviousA `cancelOrder` mutation is designed to return `Order!` directly. A teammate wants to change it to return a `CancelOrderPayload` type with an `order` and an `error` field instead. Why is that a better design, and what does it actually fix?Next A gRPC service needs to signal 'the requested order doesn't exist' and 'you don't have permission to view this order' to its caller. What mechanism carries that, and how would you decide which status code to use for each?