A service replaces `Executors.newFixedThreadPool(20)` — which was silently queueing unboundedly in front of the pool — with a `ThreadPoolExecutor` backed by an `ArrayBlockingQueue` of 200. What's the actual difference between having the queue block the caller versus reject when full, and which do you pick for an HTTP request handler?
Executors.newFixedThreadPool(n) looks bounded because the pool size is fixed, but the queue in front of it is an unbounded LinkedBlockingQueue by default — the pool never grows past n threads, but the backlog of waiting work grows without limit, and measured, that backlog grows in a straight line (about 600 items a second in one trace) until either latency becomes unacceptable or the heap runs out. Bounding the queue to 200 forces a real choice at the moment it fills: block the producer (further submissions wait until there's room) or reject (return failure immediately). Blocking makes the producer run at the consumer's actual speed — true backpressure, the limit travelling upstream — but for an HTTP handler, the "producer" is a request thread, and blocking it just moves the unbounded wait into the server's own thread pool, where it's less visible, not gone. Rejecting immediately, returning a 503 with Retry-After, is usually the better answer for a request handler: the client learns immediately and can back off, instead of waiting five seconds for a response that was going to time out anyway.