Replication internalssenior8+ years

A team assumes the oplog is a literal record of the commands the application sent, and writes a tool that replays oplog entries against a second cluster for disaster recovery by re-issuing each entry's operation. Under what circumstance does this go wrong, and what should the tool rely on instead?

The oplog doesn't store the literal command the application issued — for most operators it stores an idempotent, per-document equivalent of it, one that's safe to apply more than once. Replication itself is at-least-once, not exactly-once, so a secondary (or a tool reading the oplog) can see the same entry twice — after a network retry, or a resync that overlaps already-applied entries — and idempotence is precisely what keeps that from silently double-applying something like a $inc. A tool that assumes the oplog holds the exact literal command and re-issues entries naively is fine as long as it never sees a duplicate; the moment it does (which at-least-once delivery makes a matter of when, not if, over a long enough run), replaying a naive literal reconstruction of the operation risks double-counting, where relying on the oplog's actual idempotent form does not.

The lesson behind it →