The intuitive model — more memory retrieved, better answer — does not hold, and this is the single most common surprise reported by teams adding a memory layer.
Retrieval introduces its own failure modes. A stale fact retrieved confidently is worse than no fact, because the agent has no signal that it is stale. Irrelevant-but-similar material consumes the window and displaces the material that mattered. Retrieval that misses an update will happily surface the superseded version. And every retrieval adds latency to the first token of every session.
The underlying reason is that similarity is not relevance. What you need at the start of a session is often the constraint you agreed three weeks ago, which is not textually similar to anything you are about to type — while the thing that is similar is a passing remark from yesterday.
This is why mechanism 4 above — deliberately authored context, read at the start rather than retrieved by similarity — behaves differently in practice, and why the two are complements rather than competitors.