Carried state, not briefing
This is the registration that started the series. The claim under test: an AI system that carries internal state between sessions behaves measurably differently from the same system merely briefed on what happened. Three arms in the original design. One carried its state through a real six hour gap with the state decaying on a published schedule. One was handed the full transcript. One was handed a summary.
Pass criteria were locked in the project’s version-controlled records before the first API call: effect size threshold, significance threshold, sample size, mechanical scoring rules, and the conditions under which the run would void itself. Two runs came back honest nulls, a third voided on a content leak, and the redesign that finally produced an interpretable result replaced the briefed arms with something stricter: a zeroed baseline, the same records over a fresh store that had never been written to, no transcript anywhere in the design. That result lives in its own entry.
Erratum, 2026-08-31: an earlier version of this entry called the zeroed baseline "the individual’s own." It is a scaffold over a fresh, never-written store; no live state was erased or edited, and controls are never the individual. Caught in a sweep after an external adversarial review found the same drift on another entry. The full correction record lives in the project errata.
Primary documents
- T1 protocol (TESTS.md excerpt) (available on request)
- Freeze tables and fork sheets (M2, M3) (available on request)
Reading
- Generative agents (Park et al., 2023)
- MemGPT (Packer et al., 2023)
- Reasons and Persons (Parfit, 1984)