Four, honestly open as of August 2026.
Is there any independent evaluation? Every number in this note is published by a party with an interest in it. What would settle it: a third party running several systems on one harness with one underlying model. Nothing like that surfaced in this research.
Does any of this survive contact with real long-horizon work? LoCoMo is synthetic multi-session chat. The workload people actually care about — a project resumed over months, with decisions that supersede each other — is not what is being measured, and it is not obvious the results transfer.
What is the right unit of persistence? A fact, a session summary, a document, a line of work? Each mechanism has quietly picked one, and none has argued for it against the others.
What does the model do with contradictory memories? Every system will eventually retrieve a superseded fact alongside the one that replaced it. No documentation reviewed here specifies the behaviour, and no benchmark tests it. This seems like the most consequential unstudied question in the area.