Nearly every quantitative claim about cross-session memory traces back to LoCoMo. Its actual shape is worth knowing before you weigh a score on it: roughly 50 conversations, each up to 35 sessions and averaging around 300 turns, with about 200 question-answer pairs — and the conversations were generated by a human–LLM pipeline, not collected from real users.
That is a small and synthetic corpus to be carrying the evidential weight it currently carries.
It also has a documented methodological critique. LoCoMo-Plus (Li et al., February 2026) argues that the standard evaluation practice — disclosing the task type in the prompt, then scoring by string-matching — measures "models' adaptation to task prompts and generation styles rather than their ability to retain and apply conversational context," and that the benchmark targets "surface-level factual recall" rather than the implicit constraints that actually matter in extended work.
Read that carefully: the critique is not that scores are too low or too high. It is that the metric may not be measuring the construct. A leaderboard built on it can be internally consistent and still not be about memory.