The cheap gate did all the work
chdr ended with a conclusion: implementation isn't the bottleneck, validating the design is. Proof-Driven Development is what I built next. A method, a CLI and a gate chain for changing code nobody can verify by reading, ending in a pass or a fail with an exit code. PDD built itself: every change to it went through its own gates, so it had a lot of use behind it before anything measured it.
I built it, ran a controlled experiment against it, and the experiment is why it stopped.
The one thing worth taking from it
Validate your tests. Mutation testing is how.
That is the whole lesson, and its impact was out of all proportion to everything else in the system. On the target codebase the mutation gate surfaced 7,227 surviving mutants, of which roughly 6,782 were killable at 93% precision, sampled at n=385. Every finder in the experiment, across both arms, produced 15 test-adequacy findings between them.
On a different codebase, a public SDK with 226 stars, one module sat at 94% line coverage while about a third of the bugs I injected walked straight through its tests. Coverage counts lines the tests touched. Mutation counts bugs the tests would catch. Only the second one is the question you actually have.
Everything else PDD carried was scaffolding around that.
How I tested it
pdd-experiment-cave: three arms against the Cave engine, R a raw agent, C chunked, P full PDD. Prompts frozen, commits pinned, a kill-condition declared before the run, blind adjudication by a separate judge, and a results schema written before any results existed.
What it said
| R (raw) | P (full PDD) | |
|---|---|---|
| real issues | 37 | 150 |
| precision | 97% | 97% |
| high-severity real | 4 | 4 |
| cost | ~$125 | ~$484 |
| cost per real issue | ~$3.4 | ~$3.2 |
P found four times as many issues for four times the money, and every one of the extra findings was medium or low. Both arms found the same four high-severity ones.
Two of the static gates I wrote produced nothing usable: cover-the-mirror returned 0 real findings out of 36, pin-values 0 out of 4.
Why I stopped
The deterministic layer was doing the work, and the method wrapped around it was the expensive part. What the expense bought was more findings of the kind I did not need.
Finishing the experiment meant paying for more runs to sharpen a number whose direction was already clear, and the part worth keeping could be pulled out on its own.
The limit of the evidence
I wrote a kill-condition before the run and then never checked it. Arm C never ran. Runs 2 and 3 never ran. I stopped after one run, on one codebase, with one model.
So PDD was never killed by the rule I set for killing it. I stopped because the run pointed the same way and I did not want to pay for the rest. Read the results knowing the experiment is unfinished by its own design.
What survived
Mutation grading, the part that earned its cost, now runs as a gate on my current project.
The discipline catalog, stripped of project vocabulary, as engineering-discipline. Every rule there exists because something went wrong on a real project and the rule stopped that class of thing recurring. Each one names the class it prevents, carries the case that produced it, and cites the commits it was recovered from.
The claims gate. A sentence in the docs states a fact about the repository and says what would make it false. When that becomes true, the build fails.
Tiers. Every rule says whether anything checks it. chdr's rules said one thing while its build enforced another, and nobody noticed for a long time. Putting the answer next to the rule makes that gap visible.
Where it led
The method is gone and the enforcement survived. The skills and the gate chain now run against a real project rather than against themselves, which tests them harder than another arm would have.