A nightly job bulk-indexes two million product documents one document at a time, and search latency on the same cluster degrades badly during the job, with some writes returning `429`. What's actually competing for resources, and what three changes would you make to the job itself?
Indexing and searching compete for the same nodes' CPU and I/O — a large bulk load fills the indexing queues with segment writes and merges, which slows down concurrent searches on those same nodes, and if the load is large enough the thread pool queues reject new requests with 429 es_rejected_execution_exception, which is backpressure the client is expected to handle, not an error to retry blindly against. Three fixes: batch documents into groups of a few thousand (or 5–15 MB) instead of one at a time, which amortises the per-request overhead across many documents; raise (or disable) refresh_interval for the duration of the load, since making new segments searchable every second is wasted work when nobody is searching this data yet; and treat any 429 as a signal to back off and retry with backoff, not a fatal error.