A pricing module has 92% line coverage, yet a boundary bug shipped: a refactor changed `amount > 5000` to `amount >= 5000` and no test failed. The team proposes raising the coverage gate to 95%. As the engineer responsible for test strategy, explain why that will not help, what mutation testing would have shown and how it works, and how you would introduce it without slowing every build.
Coverage says a line was executed, not that any test's outcome depended on it. The discount line ran in plenty of tests, just never at exactly 5000, so it was covered and unchecked; a higher coverage gate would be met by the same kind of tests. Mutation testing asks the question directly. A tool like pitest makes a small change to the compiled code (flip > to >=, replace a return value, remove a call), runs the tests that cover that line, and checks whether any fails. A change no test notices is a surviving mutant: a line that runs without being checked. The boundary mutation of amount > 5000 would have survived, pointing at the missing row 5000 -> 5000 before anyone refactored. It is expensive, so run it on the modules that carry business rules, periodically or incrementally, and treat the survivors as a review list rather than a CI gate.