A checkout flow writes an order update, gets a `200` response from Elasticsearch, and immediately runs a search that's supposed to include it. The search doesn't find the just-written document. The write response says success and the cluster is green. What actually happened, and what's the one-line fix for this specific request?
A 200 response means the write was durably applied to the primary shard and replicated to its replicas — the document cannot be lost. It says nothing about whether the document is searchable yet, because searchability is a separate, later event: each shard copy periodically opens a new Lucene reader over its recently-written segments (the refresh, about once a second by default), and only from that moment do those segments' documents appear in search results on that copy. A search that runs before the next refresh simply won't find a document that is, in fact, already safely and durably stored. The fix for this one request is ?refresh=wait_for on the write, which holds the response until the next refresh has actually happened — at the cost of that specific write taking up to the refresh interval longer to return.