A leader holds a lease, checks it (valid), then pauses for eight seconds in a stop-the-world GC, resumes, and writes. During the pause, the lease expired and a new leader was elected and also wrote. Both writes land. Whose fault is this, and what actually fixes it?
Nobody made a mistake in the ordinary sense — the paused leader genuinely believed, by its own clock, that it still held a valid lease at the moment it checked. The bug is structural: a lease's expiry is measured by the grantor's clock, the holder checks it against its own clock, and between "I still hold it" and "it has expired" sits clock skew plus whatever pause (a GC, a scheduler delay, a slow network) separates the check from the actual write. Shortening the lease makes this more frequent; lengthening it makes failover slower; neither removes the gap. The real fix isn't a better clock — it's a fencing token: every lease grant carries a monotonically increasing number, every write to a downstream resource carries that number, and the resource itself rejects any write whose token is lower than the highest it has already seen. The paused leader isn't detected as stale — it's made unable to do harm, because its write, arriving with an old token, is refused by the resource regardless of what the leader itself still believes.