Mutation testing evaluates whether tests detect deliberate code changes that alter the promised behavior.
Python mutation testing: prove a suite rejects selected wrong implementations
Operation contract
The quantity predicate accepts exact integers from zero through one million. Two owned mutants widen the type check to accept booleans or change the upper bound to exclude an accepted boundary. The contract cases detect both. The fixture reports the number killed instead of treating the number of executed lines as evidence that those behaviors are protected.
Failure and ownership boundary
Only these two defects are evaluated. Surviving mutants may be equivalent to the original or reveal missing cases; a killed score is not proof that every possible defect is detected. The code creates callable implementations directly rather than editing project files or running an external mutation service. Coverage.py branch coverage: execution paths are not correctness assertions, Python refactoring: preserve rejected inputs as well as accepted results and Hypothesis property tests: compare generated cases with an independent contract answer other testing questions.
Working program
def accepted_quantity(value):
return type(value) is int and 0 <= value <= 1000000
def boolean_mutant(value):
return isinstance(value, int) and 0 <= value <= 1000000
def boundary_mutant(value):
return type(value) is int and 0 <= value < 1000000
CASES = [(0, True), (125, True), (1000000, True), (True, False), (-1, False), (1000001, False), ("125", False)]
def passes_contract(operation):
return all(operation(value) is expected for value, expected in CASES)
print("original passes:", passes_contract(accepted_quantity))
mutants = [boolean_mutant, boundary_mutant]
print("selected mutants killed:", sum(not passes_contract(operation) for operation in mutants))
print("selected mutants:", len(mutants))Output
original passes: True
selected mutants killed: 2
selected mutants: 2Costs and limits
For m mutants and c cases, this fixture performs O(mc) predicate checks. Whole-project mutation campaigns can rebuild/run many tests and need timeouts, isolated state and a policy for equivalent changes. No whole-project mutation score is claimed.
Common Mistakes
- A killed score is meaningful only with its selected mutation set.
- Coverage and mutation detection are different evidence.
Connected lessons
Coverage.py branch coverage: execution paths are not correctness assertions, Python refactoring: preserve rejected inputs as well as accepted results, Hypothesis property tests: compare generated cases with an independent contract.
