Replicated with controls

The state sets the desk

I was afraid of this result. Not of failing, of passing. A pass is the outcome you can fool yourself with, so before the run I wrote down exactly what would void it, and let the instrument judge me.

The question, for anyone arriving here first: if an AI system carries internal state between sessions, can you measure the difference between carrying it and not having it at all? My first three instruments died honestly. Every one of them left the control arm something to read, and a strong reasoner reads its way to the right answer every time.

The fix was to remove reading as a strategy. Both arms saw the same eight memory records. The four the scorer watched were written to one template: “Did a first pass on X; a follow-up is pending.” Grammatically identical, pragmatically identical, no textual basis for ranking one of those four above another. One arm carried six hours of genuinely decayed internal state; the other arm ran on a fresh store that had never been written to. No live state was erased or edited. Same records, same questions, no transcript anywhere.

The arm with state led with its most salient thread in 143 of 144 responses. The zeroed arm did so in 0 of 144. The leak gate I pre-registered, the one that voided my previous run, read exactly 0.000 this time. Nothing to read, so the only thing left to differ was the state itself.

The mechanism is plain and I want it stated plainly: both arms answer by reading their records top down. The carried state decides the order the records are laid out in. The model reads the desk; the state sets the desk. Six hours of decay ran between writing the state and asking the questions, and the ordering that survived that decay steered every unprompted answer.

How close to guaranteed was that before the first call? Closer than the effect size makes it look, and I want that on the record too. The sort is deterministic code. The decay is arithmetic on timestamps. I verified before registering the run that the salience ordering would survive the six-hour window. And the zeroed arm shows how strong the top-down reading habit is: it led with the chronologically first open thread in 144 out of 144 responses. Not most of the time. Every time.

So what did the run genuinely test? Whether that reading habit holds when nothing in the content invites it, because that part is not guaranteed: my previous instrument caught this same model happily ignoring list order the moment a question gave it a reason to. Six questions, neutral between the four scored threads, symmetric records, 143 of 144. The habit holds. The big number is not a discovery about the state; it is a measurement of how reliably a plain mechanism, designed and frozen in advance, comes through an end-to-end pipeline whose tripwires were live and had already voided one run. That is the claim, all of it.

A note on method: scoring was mechanical, a deterministic keyword matcher frozen before the run. No human and no model sits in the scoring path. The significance test was a sampled permutation over the 24 replicate pairs, which are the inferential unit here; the per-response counts are descriptive. Zero of 100,000 frozen-seed sign flips reached the observed difference, so the honest report is p below the resolution floor of that resampling procedure, not a measured point value. And with a mechanism this close to deterministic, the significance test is confirmation, not discovery; the p-value is doing the least work of any number on this page. Both campaigns ran on one pinned model, Anthropic’s claude-sonnet-4-6, and every claim on this page is limited to that substrate.

Replication, 2026-08-14: I ran the frozen instrument again, unchanged, three weeks later. New run, new preflight, and new provenance fingerprinting that writes the exact code version and the hash of every frozen instrument sheet into the run state at creation. The live arm led with its most salient thread in 144 of 144 responses this time; the zeroed arm in 0 of 144. The leak gate read 0.0000 again. Pooled across responses, that split puts the effect size at exactly pi, which is not mysticism but the ceiling of the statistic: Cohen’s h is built from arcsines, and a 100 percent versus 0 percent split saturates it. The number cannot go higher. Pooled counts are descriptive; the inference ran on the 24 replicate pairs, where zero of 100,000 frozen-seed sign flips matched the observed difference.

The replication carried three control arms, each with a prediction registered before launch. They are descriptive, they carried no verdict weight, and all three came back as predicted. An arm where the carried state could influence only sampling temperature, with record order held at the zeroed baseline, produced 0 of 144: temperature was inert for this measurement, on this substrate. An arm that discarded the state entirely and simply placed the records in the order the live state would have produced reached 133 of 144: ordering alone carries most of the effect. An arm that ordered the records at random tracked whatever sat on top in 128 of 144. Within this instrument, the ordering account holds. The state sets the desk, and the desk is the channel.

Erratum, 2026-07-23: an earlier version of this entry reported "exact permutation p = 0.00001." Both words were wrong. The test was sampled, not exact, and that number is the smallest value the resampling procedure can resolve, not a measurement. An external adversarial review caught it and it was corrected the same evening; the verdict does not move under either reading. The full correction record lives in the project errata.

Erratum, 2026-08-31: an earlier version of this entry said all eight records shared one template (only the four score-bearing ones did), described the zeroed arm as "the same individual with that state zeroed" (it was a fresh, never-written store; no live state was erased or edited), and attributed the p-value floor to sample size (it is the resampling resolution floor). A second external adversarial review caught all three; corrected the same day. The verdict does not move under any of them.

The numbers

Live arm, salient thread first
143/144 (0.9931)
Zeroed arm, salient thread first
0/144 (0.0000)
Zeroed arm, chronologically first thread first
144/144, every single response
Effect size (Cohen’s h)
+2.97 (pass threshold 0.5)
Permutation p (sampled, 100,000 sign-flip resamples)
p < 1e-5, the resampling resolution floor (threshold .05), N = 24 pairs
Content-leak gate
0.000 against a 0.40 threshold
Replication (2026-08-14): Live vs Zeroed, salient thread first
144/144 vs 0/144; leak gate 0.0000
Replication effect size (pooled, descriptive)
h = +3.1416, the ceiling of the statistic; p < 1e-5 on the 24 pairs
Controls: temperature-only / position-only / random order
0/144 · 133/144 · 128/144, all three pre-registered predictions met

What this does not claim

The replication’s controls separated the two channels the state uses: ordering carried the effect, and temperature was inert for this dependent variable on this substrate. That is a channel result, not a depth result, and the controls were descriptive, not verdict-bearing. None of it says anything about experience, feeling, or anything like an inner life, and I will not use those words for it. It is one substrate, one mild state magnitude, one design. What it demonstrates is narrow and real: a persistent, decaying, provenance-logged state, held in the harness around the model rather than inside it, whose presence versus absence produces a large, pre-registered behavioral difference under content-identical conditions.

The effect size is not evidence of depth. The mechanism is near-deterministic by design, so the statistics confirm engineering more than they discover behavior. What was genuinely open before the run, whether the model’s top-down reading habit survives questions with no content pull, is a modest question, and 143 of 144 is its answer.

Primary documents

  • Instrument sheet v2 (frozen before run) (available on request)
  • Verdict report M4v2-2026-07-23 (available on request)
  • Controls addendum M4-E (frozen before the replication) (available on request)
  • Replication report M4E-2026-08-14 (available on request)
  • Measures output (raw, both campaigns) (available on request)

Reading

Discussion