You were measuring the reasoner
Three pre-registered campaigns, three voided or failed runs, one lesson. In the first, the control arm was handed a transcript and aced my consistency probes 288 out of 288. In the second, I hardened every probe and it scored 98.6 percent. In the third I removed the transcript entirely and flattened the memory records to remove urgency words, and the control still ranked the right thread first from content alone, 49 percent against a 40 percent gate.
The through-line: any information channel you leave within the control’s reach, a modern model will exploit to ceiling. Urgency does not need to be stated to be read. “Submitted the analysis job; results expected” carries its own priority, silently.
The design rule I now treat as law: the validity gate is not decoration. Before each run I wrote down the condition under which the run would declare itself uninterpretable, and twice that gate fired on my own experiment, exactly as written. Both times it was the best decision on the page, because without it I would have been holding results instead of the truth, which was that my instrument was grading reading comprehension.
A note on method: all scoring across these campaigns was mechanical. The first two used strict parsing of forced-choice answers; the third used a deterministic keyword matcher. No human and no model sits in any scoring path.
The numbers
- Campaign 1, briefed control on forks
- 288/288 (1.0000)
- Campaign 2, hardened forks
- 284/288 (0.9861), ceiling gate at 0.90 fired
- Campaign 3, flattened records
- stateless arm ranked the loaded thread first at 0.4931, leak gate at 0.40 fired
Primary documents
- Verdict reports M2, M3, M4 (campaigns 1 through 3) (available on request)
Reading
- Leakage and the reproducibility crisis in ML-based science (Kapoor & Narayanan, 2023)
- The preregistration revolution (Nosek et al., 2018)