gRPC reliabilitymedium3-5 years

A gRPC call sets a client-side timeout of 2 seconds and calls three downstream services in sequence. Two seconds after the call starts, the first downstream call is still in progress. What happens, and how is this different from each service independently having its own 2-second timeout?

A gRPC call carries a deadline — an absolute point in time set once by the original caller — not a per-hop timeout that resets at each service. That deadline is propagated to every downstream call the server makes on the request's behalf, so if the first downstream call has already used 1.5 of the 2 seconds, the second downstream call starts with only 0.5 seconds left, not a fresh 2. If each service independently ran its own 2-second timeout instead, the first hop alone could take 2 seconds, then the second another 2, then the third another 2 — six seconds total, with no service ever technically violating its own limit, while the caller who only wanted to wait 2 seconds has no idea any of that happened.

The lesson behind it →