Two Resilience4j breakers misbehave in the same incident: `payments` (about 20 calls/minute, `minimumNumberOfCalls: 100`) never opens during a five-minute outage, and `search` (`slidingWindowSize: 10`, `permittedNumberOfCallsInHalfOpenState: 1`, `waitDurationInOpenState: 60s`) stays open for an hour after a two-second blip. Explain both mechanically, and fix each.
payments: the breaker only evaluates the failure rate once its sliding window holds minimumNumberOfCalls — 100 — and at 20 calls a minute that takes five minutes to fill, so for the entire five-minute outage the breaker stayed CLOSED, every call timed out at full latency, and the fallback never engaged because the breaker never had enough evidence to judge. search: with only one trial call permitted in HALF_OPEN, a single unlucky trial (landing during, say, a GC pause) reopens the breaker for another 60 seconds — and it kept happening, so a dependency that actually recovered in two seconds stayed hidden behind a fallback for an hour, because one call was never a real sample of whether it had recovered. The fixes are symmetric: lower payments' minimumNumberOfCalls to something reachable within the time you're willing to be slow (say 10, judged within 30 seconds), and raise search's permittedNumberOfCallsInHalfOpenState to a genuine sample (say 10) with a shorter wait.