engram: when does selection replace extraction?

Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough?

Paper: Rishabh Sharma and Rishika Lall, arXiv:2609.34227 (2026).

Within our study, the context budget decides. When the answer model reads only a few retrieved items, selecting raw turns with one request to Jev, a typed decision model, is non-inferior to an LLM-extraction memory, at a fraction of the cost to write. When it reads many, extraction is more accurate. The study was pre-registered and run on conversations never used for development.

Results

The paper labels every result as registered, exploratory or post-hoc. That the budget also explains why published studies disagree is our interpretation, not a tested claim.

The write path (top) and read path (bottom) of each system. Border colour shows what does the work: code (blue), an LLM call (amber), a Jev decision (purple), a store (green).
The write path (top) and read path (bottom) of each system. Border colour shows what does the work: code (blue), an LLM call (amber), a Jev decision (purple), a store (green).
The rerank's gain over cosine similarity falls as the budget grows, on both benchmarks.
The rerank's gain over cosine similarity falls as the budget grows, on both benchmarks.

Papers

Pre-registrations. v3 plan doi:10.5281/zenodo.22970745 and its amendment doi:10.5281/zenodo.22977848. The v2 plan doi:10.5281/zenodo.22948854 was paused before any held-out run.

How the v1 article's claims stand

Data, code and reproduction

Citation

@misc{sharma2026selection,
  author        = {Sharma, Rishabh and Lall, Rishika},
  title         = {When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model},
  year          = {2026},
  eprint        = {2609.34227},
  archivePrefix = {arXiv},
  doi           = {10.5281/zenodo.22985242},
  url           = {https://arxiv.org/abs/2609.34227}
}