PDD vs a raw agent, Cave engine

Run 1 of a three-arm experiment on a 9-subsystem brownfield target. Frozen prompts, pinned commits, a kill-condition declared before the run, blind adjudication by a separate judge. Arm C and runs 2 and 3 are not done, and the kill-condition has not been evaluated.

Method, results and the raw verdicts are in the repository, including the frozen prompts and the pinned commits that make run 1 reproducible.

The panels below are drawn by dod's dashkit, the same renderer the local dashboard uses.