Dossier 02 · Research labs
Research labs
The retrieval stack is measured in the open. 77.1% on LongMemEval-M, every question with its verdict, the grading protocol alongside. What we missed matters as much as what we scored.
If you work on long-term memory, retrieval or human–AI interaction and want to evaluate, replicate or collaborate, write to us.
Collaboration offer Independent researchers and domain labs in medicine, biology and beyond
We are looking for independent researchers and laboratories to build a memory engine for their field with us. Medicine, biology, chemistry, any discipline with a dense corpus. Your expertise and your corpus, our engine. It is a collaboration, not a product off the shelf.
The tool we want to build thinks like a researcher. It watches new publications on your exact themes, arXiv and beyond, and folds them into your memory. At night it dreams on those themes and surfaces the crossings between your corpus and your peers’ work. A lab assistant that lives on your machine.
It starts with a conversation, not a demo. How do you work? Where is the friction? What eats your weeks? The first researchers who test it shape the tool.
We know what honest science costs, because we applied it to ourselves. Protocols pre-registered before the decisive runs, control arms, seeded permutation tests, and the documented falsification of the hypothesis our founder cared about most.
Internal research program: blind pre-registration · control arms · 336 generations analyzed · falsification is part of the method
Cite this work Technical whitepaper, DOI 10.5281/zenodo.21728284Benchmark audit kit, DOI 10.5281/zenodo.21727140ORCID 0009-0009-1087-3917