Under write latency pressure, a team proposes switching a Redis-backed cache from write-through to write-behind: acknowledge the write immediately, buffer it, and flush to the database on a schedule. What exactly is being traded away, and is it a bug to fix or a tradeoff to make deliberately?
Write-through writes the database synchronously before acknowledging, so a caller only ever gets a success response once the write is actually durable — if it fails, the caller finds out. Write-behind acknowledges the write immediately and buffers it, flushing to the database later in batches; between the acknowledgement and the flush, the only copy of that write anywhere is in the cache tier's memory, and if that process crashes, gets evicted, or its buffer is lost before the flush, the write is gone — silently, because the caller was already told it succeeded and nothing records that it didn't survive. That gap isn't a bug to patch; it's the entire reason write-behind is faster and able to batch writes at all, so the real engineering question is how big that window is and what's actually acceptable to lose during it, not how to eliminate it (eliminating it just turns write-behind back into write-through with extra steps).