Harness the Memory: A Holistic Benchmark of Agent Memory Substrates
Z. Huang et al. · 2026
What this means for traders
Retrieving more context is not automatically better: past a point it crowds out the information the model actually needs to act on.
Abstract
Evaluates eleven memory substrates across seven families, three backbones and four task suites under twenty-six metrics. No substrate dominates: the winner inverts by regime, with structured graphs leading on dialogue question answering and cheap flat retrieval leading on code. Widening retrieval helps question answering but measurably degrades sequential decision-making, with attention probes showing mass draining from the action context into the retrieved block. Production-grade memory systems are also shown to cost between 2,700 and 9,000 auxiliary LLM calls per long history, an overhead usually reported nowhere.