The headline accuracy figures published by memory vendors are self-reported, and they are evaluated under different underlying models and different harness configurations. Two products quoting near-identical percentages on nominally the same benchmark are frequently not measuring the same thing, and the difference between them is smaller than the difference between their setups.
This is a structural property of the situation rather than an accusation: there is no neutral evaluator running these head to head, so each vendor runs its own, and each one — reasonably — configures the harness in the way that suits its architecture.
The practical consequence for anyone choosing: a percentage point difference between two self-reported scores carries no information. What does carry information is the mechanism (which of the four above), what happens when retrieval returns the wrong thing, and whether you can inspect what was stored. Those are answerable by reading the docs and are not on any leaderboard.