"Memory" is used for at least four distinct mechanisms with different failure modes. Conflating them is why comparisons in this space rarely mean anything.
- Context compaction / editing — dropping or summarising earlier turns so a long session does not exhaust its window. This solves running out of room inside one session. It is not cross-session persistence at all, and it is frequently counted as memory because it ships alongside it.
- Automatic extraction into a store — a layer watches the conversation, extracts facts, writes them to a vector or graph store, and retrieves by similarity at the start of the next session. Solves recall. Its failure mode is that nobody chose what was kept, so what comes back is whatever was similar.
- Agent-written files — the model decides what is worth writing and writes it to ordinary files it can read later. Solves recall with the selection decision moved to the model.
- Deliberately authored project context — a human or agent writes the durable state on purpose, and it is read at the start of the work. Solves intent and constraint transfer, which retrieval does not, because the thing you need next session is often not similar to anything you are about to say.
The distinction that matters: 1 is about capacity, 2 and 3 are about recall, and 4 is about intent. A tool that is excellent at one can be irrelevant to another, and the word does not distinguish them.