Where to find me across the web
Experiments registered before their results exist.
Fourth instrument, first pass: a pass, and the honest headline is smaller than the numbers.
The preregistration revolution (Nosek et al., 2018) · Reasons and Persons (Parfit, 1984)
Two pre-registered campaigns, 48 replicates, and both times the control arm aced the test from reading alone. When a briefed model scores 99% on your consistency probes, your instrument is grading reading comprehension, not memory. A validity gate we now consider mandatory: if the control saturates, the run is void, not a result.
The preregistration revolution (Nosek et al., 2018) · Leakage and the reproducibility crisis in ML-based science (Kapoor & Narayanan, 2023)
Can an AI carry internal state between sessions in a way that reading its own transcript back cannot fake? Three arms, twelve scripted probes, and the pass criteria were locked in before the first run. Both runs came back honest nulls; the redesign against a zeroed baseline is registered and running.
Generative agents: interactive simulacra of human behavior (Park et al., 2023) · MemGPT: towards LLMs as operating systems (Packer et al., 2023) · Reasons and Persons (Parfit, 1984)
We sent sampling settings to a provider bridge. It accepted them, echoed them back on every call, and quietly ignored every one. A short note on why a config echo proves nothing. Bonus finding from the rerun: when the dial is connected, higher temperature shows up as format drift before it shows up as different choices.
The curious case of neural text degeneration (Holtzman et al., 2020) · Is temperature the creativity parameter of large language models? (Peeperkorn et al., 2024)
Business inquiries, collabs, or just to say hey.