Blog PROOF-9

Coverage says the line ran. Break the code on purpose to learn whether any test cares.

A suite written from its implementation asserts what the code does, so mutate one operator at a time and read every survivor before trusting green.

Ask an AI session for a function and then for its tests, and the tests will pass. They will usually reach full coverage as well. Neither fact says whether the suite would notice the function being wrong, because the expected values in those tests were read off the function they are testing. A test derived from the implementation agrees with the implementation by construction, including in the places where the implementation is mistaken.

PROOF-9. A test can pin a lie.

The method in one line: break the code on purpose, one operator at a time, rerun the suite after each break, and read every break the suite failed to notice. A test that stays green against broken code was never checking that line.

Coverage measures execution, not judgement

A coverage report counts a line as covered when the line runs. It does not ask whether any assertion depends on what the line computed. That gap is easy to state and easy to underestimate, so measure it directly. A test file that calls every branch of a module and asserts nothing at all reports 100 per cent line, branch and function coverage on that module. The report cannot tell it apart from a careful suite.

The measurement that can tell them apart is mutation testing. Change one thing in the source, such as >= to >, or 5 to 6, and run the tests. If the suite fails, the change was caught and the mutant is killed. If the suite passes, the mutant survived, and a survivor is a precise statement: this line can be wrong in this exact way and no test objects.

This matters more on AI-written code than on hand-written code, for a structural reason rather than a quality one. A person writing tests after the code at least carries an intention the code may have missed. A session that writes both works from the code it just produced, so the typical inputs it picks are the inputs the code already handles, and the boundaries are where nobody looked.

The check, in forty-two lines

Mature tools do this properly and belong in a real pipeline: Stryker for JavaScript and TypeScript, mutmut for Python, PIT for the JVM, cargo-mutants for Rust. They mutate the syntax tree rather than the text, run mutants in parallel, and most use coverage data to run only the tests that reach each mutant. The script below is the naive version, with no dependencies, and it is enough to run on one file tonight and see the shape of the answer.

// node mutate.mjs <source file> <test command...>
// Breaks one operator at a time, reruns the tests, reports what survived.
import { existsSync, readFileSync, writeFileSync } from 'node:fs'
import { spawnSync } from 'node:child_process'

const [file, ...cmd] = process.argv.slice(2)
const run = () => spawnSync(cmd[0], cmd.slice(1), { stdio: 'ignore', timeout: 60000 }).status === 0
const src = readFileSync(file, 'utf8')
const SWAPS = [
  [/>=/g, '>'], [/<=/g, '<'], [/===/g, '!=='], [/!==/g, '==='],
  [/(?<![=<>!-])>(?![=>])/g, '>='], [/(?<![<=])<(?![=<])/g, '<='],
  [/&&/g, '||'], [/\|\|/g, '&&'], [/ \+ /g, ' - '], [/ - /g, ' + '],
  [/ \* /g, ' / '], [/\b\d+\b/g, (n) => String(Number(n) + 1)],
]
// Reviewed equivalent mutants, one per line: the survivor line, then "# why".
const accepted = new Set(existsSync('.mutants-accepted')
  ? readFileSync('.mutants-accepted', 'utf8').split('\n').map((l) => l.split('#')[0].trim()).filter(Boolean)
  : [])
if (!run()) throw new Error('the suite fails before any mutation; fix that first')

let killed = 0
const survived = []
try {
  for (const [re, to] of SWAPS) {
    for (const m of src.matchAll(re)) {
      const before = src.slice(0, m.index).split('\n')
      const line = before.length, col = before.at(-1).length + 1
      if (/^\s*(\/\/|\*)/.test(src.split('\n')[line - 1])) continue // comments
      const rep = typeof to === 'function' ? to(m[0]) : to
      writeFileSync(file, src.slice(0, m.index) + rep + src.slice(m.index + m[0].length))
      if (run()) survived.push(`${file}:${line}:${col} ${m[0]} -> ${rep}`)
      else killed++
    }
  }
} finally {
  writeFileSync(file, src) // always restore the original
}
const open = survived.filter((s) => !accepted.has(s))
console.log(`[mutate] ${killed + survived.length} mutants, ${killed} killed, `
  + `${open.length} survived, ${survived.length - open.length} accepted as equivalent`)
for (const s of survived) console.log(`  ${accepted.has(s) ? 'ACCEPTED' : 'SURVIVED'} ${s}`)
process.exit(open.length ? 1 : 0)
mutate.mjs. Text substitution, one mutant at a time, the original restored in a finally block whatever happens.

Four details in it are load-bearing. The suite must pass before any mutation, or every mutant counts as killed and the score is meaningless. A mutant that makes the tests hang counts as killed, because the 60 second timeout fails the run. The original file is written back in a finally block, and it should still only be run on a clean working tree, so that a killed process leaves nothing that git cannot restore. And the acceptance key carries the column as well as the line, because two identical operators on one line are two different mutants, and accepting one must never silently accept the other.

What it reports on a suite at full coverage

The module below is three ordinary utilities of the kind every codebase carries: a retry delay with a cap and a budget, a batching helper and a page clamp. The suite beside it is written the way a suite is written when it is requested after the code: one typical input per branch, every expected value correct.

// Delay before retry `attempt` (0-based), or null once the budget is spent.
export function retryDelay(attempt, base = 200, cap = 10000, max = 5) {
  if (attempt >= max) return null
  const delay = base * 2 ** attempt
  return delay > cap ? cap : delay
}

// Split a list into batches of at most `size`.
export function chunk(items, size) {
  if (size < 1) throw new RangeError('size must be positive')
  const out = []
  for (let i = 0; i < items.length; i += size) out.push(items.slice(i, i + size))
  return out
}

// Clamp a requested page into 1..pages.
export function clampPage(page, pages) {
  if (pages === 0) return 1
  if (page < 1) return 1
  if (page > pages) return pages
  return page
}
limits.mjs. Nothing in it is wrong.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { retryDelay, chunk, clampPage } from './limits.mjs'

test('retryDelay', () => {
  assert.equal(retryDelay(0), 200)
  assert.equal(retryDelay(2), 800)
  assert.equal(retryDelay(9), null)
  assert.equal(retryDelay(6, 200, 1000, 10), 1000)
})

test('chunk', () => {
  assert.deepEqual(chunk([1, 2, 3, 4, 5], 2), [[1, 2], [3, 4], [5]])
  assert.throws(() => chunk([1], 0))
})

test('clampPage', () => {
  assert.equal(clampPage(3, 10), 3)
  assert.equal(clampPage(-2, 10), 1)
  assert.equal(clampPage(40, 10), 10)
  assert.equal(clampPage(5, 0), 1)
})
limits.test.mjs. Every assertion passes, and every line and branch of the module runs.

Node’s built-in coverage reports 100 per cent of lines, branches and functions. The mutation check generates 19 mutants from the module, and the suite kills 9 of them. Ten survive. The whole run is 20 test runs, one baseline and one per mutant, and takes about three seconds.

$ node --test --experimental-test-coverage limits.test.mjs
# file            | line % | branch % | funcs % | uncovered lines
# limits.mjs      | 100.00 |   100.00 |  100.00 |

$ node mutate.mjs limits.mjs node --test limits.test.mjs
[mutate] 19 mutants, 9 killed, 10 survived, 0 accepted as equivalent
  SURVIVED limits.mjs:3:15 >= -> >
  SURVIVED limits.mjs:5:16 > -> >=
  SURVIVED limits.mjs:20:12 > -> >=
  SURVIVED limits.mjs:10:12 < -> <=
  SURVIVED limits.mjs:12:21 < -> <=
  SURVIVED limits.mjs:19:12 < -> <=
  SURVIVED limits.mjs:2:55 10000 -> 10001
  SURVIVED limits.mjs:2:68 5 -> 6
  SURVIVED limits.mjs:10:14 1 -> 2
  SURVIVED limits.mjs:19:14 1 -> 2
Measured output, both commands, Node 22. The coverage table is trimmed to the module row.

Ten survivors is a count, and a count is not a finding until each item has been read. Reading them sorts them into three kinds.

Line:column and changeWhat it meansKind
3:15 >= -> >A budget of five would allow a sixth attemptMissing test
2:68 5 -> 6The default budget is never checkedMissing test
2:55 10000 -> 10001The default cap is never reachedMissing test
10:12 < -> <=A batch size of 1 would throwMissing test
10:14 1 -> 2The same boundary, from the other sideMissing test
12:21 < -> <=An exact multiple would gain an empty last batchMissing test
5:16 > -> >=A delay equal to the cap returns the cap either wayEquivalent
20:12 > -> >=A page equal to the last page returns it either wayEquivalent
19:12 < -> <=Page 1 returns 1 either wayEquivalent
19:14 1 -> 2Differs only for a fractional pageSpecification question
The ten survivors from the run above, each read against the source. Six are real gaps, three can never be told apart from the original, and one depends on a decision nobody wrote down.

All six real gaps sit on a boundary. That is not a coincidence of this fixture. Typical inputs are chosen to be comfortably inside the range a function handles, and off-by-one is exactly the class of defect that occurs at the edge of that range. A suite built from typical inputs is blind to the defect class most likely to be in the code.

Read every survivor, then decide

Each kind of survivor has one correct response, and none of them is raising a score for its own sake.

  1. A missing test is answered with the assertion that kills it. The survivor is already the specification: this input, at this boundary, must produce this result.
  2. An equivalent mutant is a change no input can distinguish from the original. It is proved by reasoning, not by running more tests, and then accepted by name with the reason written beside it.
  3. A specification question is a survivor that exposes behaviour nobody decided. Decide it. The answer becomes either a test or an accepted entry whose reason records the decision.
// retryDelay
assert.equal(retryDelay(5), null)                          // the last attempt is refused
assert.equal(retryDelay(7, 200, undefined, 10), 10000)     // the default cap holds
// chunk
assert.deepEqual(chunk([1, 2, 3], 1), [[1], [2], [3]])     // the smallest legal size
assert.deepEqual(chunk([1, 2, 3, 4], 2), [[1, 2], [3, 4]]) // an exact multiple
Four assertions added to limits.test.mjs, one per boundary.

Rerun, and the check reports 19 mutants, 15 killed and 4 surviving. Four assertions killed all six real survivors, because two of the boundaries each carried two mutants. The four that remain are the three equivalents and the fractional-page question, answered here by deciding that pages are whole numbers. They are accepted in a file the check reads, and the job goes green only once nothing unread remains.

limits.mjs:5:16 > -> >=    # a delay equal to the cap returns the cap either way
limits.mjs:20:12 > -> >=   # a page equal to pages returns pages either way
limits.mjs:19:12 < -> <=   # page 1 returns 1 either way
limits.mjs:19:14 1 -> 2    # differs only for fractional pages; pages are whole numbers

$ node mutate.mjs limits.mjs node --test limits.test.mjs
[mutate] 19 mutants, 15 killed, 0 survived, 4 accepted as equivalent
  ACCEPTED limits.mjs:5:16 > -> >=
  ACCEPTED limits.mjs:20:12 > -> >=
  ACCEPTED limits.mjs:19:12 < -> <=
  ACCEPTED limits.mjs:19:14 1 -> 2
.mutants-accepted, then the third run. Accepted survivors are still printed on every run; they are only excused from failing it.

The acceptance key carries the line and column on purpose. When the code around an accepted mutant changes, the key stops matching, the mutant reappears as a survivor, and the equivalence has to be argued again against the new code. Treat the file as code under review: an entry added to it is a claim that no input can tell two programs apart, and a reviewer should be able to check that claim from the reason alone.

The target is not a mutation score. It is zero survivors that nobody has read.

Put it where the tests are written

Every mutant costs one full run of the tests it is checked against, so mutating a whole repository on every push is the wrong shape. Scope it twice: to the source files the branch added or changed, and for each file, to its own test file rather than the whole suite. The fixture above costs 20 test runs for a 22 line module; the budget for a real change is roughly the number of operators and literals in the files it touches, multiplied by the time of one targeted test run.

# Mutate every source file this branch added or changed, against its own
# test file. Any unaccepted survivor fails the job.
git diff --name-only --diff-filter=AM origin/main...HEAD -- '*.mjs' ':!*.test.mjs' |
while read -r f; do
  node mutate.mjs "$f" node --test "${f%.mjs}.test.mjs" || exit 1
done
The whole CI step. It fails the job on the first file with an unaccepted survivor.

For AI sessions the check is also the best available brief. A survivor list is a specification more exact than any prompt: it names the file, the line, the column and the precise wrong behaviour that no test currently rejects. Give a session the survivor list and ask for the assertions that kill each one, and it is being asked for boundary tests with the boundaries already found. The gate in CI is what makes the result binding. The session can write the tests; only a failing job guarantees they exist.

What to do tonight

  1. Take a module your last AI session wrote together with its tests. Run the coverage report on it and write the number down.
  2. Save the script, and run it against that module with that module’s test file as the command. Work on a clean working tree.
  3. Read every survivor and sort it: missing test, equivalent mutant or specification question. Count each kind.
  4. Write the assertions that kill the missing-test survivors, decide the specification questions, and accept the equivalents by name with a reason each.
  5. Add the diff-scoped step to CI, so that a changed file with an unread survivor fails the build instead of passing it.

The general form is the oldest rule in testing, and coverage made it easy to forget: a test proves something only if it is able to fail. Coverage counts the lines a suite visited. Mutation counts the lines it would defend. When one hand wrote both the code and the tests, ask the only question that separates the two: which broken version of this code would this suite catch?

This post argues PROOF-9 from THE PLATFORM LAWS, the standing rules for anything Wavn, Inc. builds. Related reading: our principles, how we handle AI output and the record.

See it keep a real call.

Wavn takes the other side of your decision, keeps the reasoning, and holds the record.