We used Claude Code agents running Opus 5 with each of the 12 memory systems. Each agent got the same work across multiple sessions and 150 questions.
Each memory tool runs through weeks of simulated work. Each agent chooses what to store and remember, guided by each tool’s own documentation.
Cognee and Karpathy’s Wiki tied for the highest accuracy score at 97.1 out of 100.
Using Cognee is significantly cheaper, at about half the wiki’s token cost at $256 per 1,000 questions vs. $503.
Every memory system has tradeoffs though, Cognee was the slowest system we tested, taking over five times as long per question as the wiki.
Across the tests, the biggest reason for failure was agents declining to answer questions they should have been able to answer. That accounted for 61% of failures.
We also tested whether each system leaked information it had been told to forget. 9/12 systems never repeated the facts they’d been told to forget, while three leaked data that should have been forgotten.
I just launched the Agentic Search Index, a public benchmark of the web-search tools AI agents use. Perplexity Sonar came last of the nine tools tested.
I ran the same Claude agent on 121 web-search tasks through every tool, three times for a total of 3,537 runs.
The site has the full results: the overall ranking, task-type breakdowns, all-in cost per successful answer, latency, failures, confidence intervals, and methodology.
Agentic Resource Radar puts every provider through the same test, so the results are directly comparable. Search is the first index; memory and second-brain tools are next, followed by lead enrichment tools.
as long as OpenAI and Anthropic keep subsidizing dirt cheap Codex or Claude Code usage, I'll just keep using them as evaluators. The trick is to have a fresh instance doing the reviewing, not the one that did the work.
> The trick is to have a fresh instance doing the reviewing, not the one that did the work.
In my experience that's not neccessary (some people even claim that you must use models from different vendors), and it's expensive since a fresh instance needs to rebuild all the context that's needed in order to properly and thoroughly review. LLMs have no problem throwing "them 5 minutes ago" under the bus when asked to review something "skeptically" and "with fresh eyes".
Doing it in the same session does save a ton of tokens but I find it's too biased towards its own implementation even if you tell it to use "fresh eyes" or to "act like a code reviewer in a bad mood." Including those strings in your prompt does show some improvement but not nearly as much as making it think from first principles in a fresh instance.
Each memory tool runs through weeks of simulated work. Each agent chooses what to store and remember, guided by each tool’s own documentation.
Cognee and Karpathy’s Wiki tied for the highest accuracy score at 97.1 out of 100.
Using Cognee is significantly cheaper, at about half the wiki’s token cost at $256 per 1,000 questions vs. $503.
Every memory system has tradeoffs though, Cognee was the slowest system we tested, taking over five times as long per question as the wiki.
Across the tests, the biggest reason for failure was agents declining to answer questions they should have been able to answer. That accounted for 61% of failures.
We also tested whether each system leaked information it had been told to forget. 9/12 systems never repeated the facts they’d been told to forget, while three leaked data that should have been forgotten.