Lab notes
What we measured, and what turned out to be wrong.
The research surface of Mnemosyne OS. What is deposited, what the numbers claim and what they do not, how they are produced, and the idea we cared most about and had to drop. There is no separate research site. This page is it.
Deposited work
Two artefacts are archived under permanent identifiers, released under CC BY 4.0 and attached to an ORCID. They can be cited, not just linked. A third party can point at them in ten years, whatever happens to this site.
Technical whitepaper, DOI 10.5281/zenodo.21728284Benchmark audit kit, DOI 10.5281/zenodo.21727140ORCID 0009-0009-1087-3917
The number, and what it does not claim
77.1% on LongMemEval-M, full-haystack variant, under the strict judge. That setting buries each question under roughly 480 distractor sessions, instead of the smaller slice most published numbers use.
It is a real 48-question run, replayed a second time and matching verdict for verdict, and the gain behind it was confirmed on a held-out set of questions. The same answers score 81.3% under the more flexible judge. July’s published floor, 72.9%, stays archived with its composition stated, under its own DOI. Two judges are two instruments, so none of these numbers is a progression of another.
It says nothing about systems evaluated on other benchmarks. Comparing a figure from here with one from elsewhere would be meaningless, in either direction.
How we measure
Protocols are pre-registered before the decisive runs, so a hypothesis cannot be quietly adjusted once the result is in. Every arm is compared against control arms, including deliberately meaningless ones. An effect a random control reproduces is not an effect.
Where the comparison is statistical, it is a seeded permutation test, not a difference judged by eye. Where a language model judges an answer, we measure the noise floor and state it. A gap smaller than that floor is reported as noise.
What we got wrong
The founder held for months that a particular mathematical structure would carry signal inside a memory engine. It was the idea he cared about most. We measured it across several independent tasks, in every framing we could build for it, including the version he considered the good one.
It carried no signal anywhere. Random controls reproduced every effect it appeared to have. A structural probe found nothing in the data for it to lock onto. The hypothesis was dropped.
That failure is the most useful thing on this page. Running it properly is what surfaced the design that did ship: consolidation that augments memory instead of replacing it. It keeps the raw record and adds to it, rather than summarising over it.
What this builds on
The lineage is set out with its references in the whitepaper, section 12. That covers the operating-system framing, the benchmark we submit to, the closest published architectures, and the work that argues an evaluation loop must be replayable rather than trusted. One canonical place beats two that drift.
Internal research program: blind pre-registration · control arms · 336 generations analyzed · negative results kept