Aller au contenu principal
Retour à la recherche
Architecture système

Harness the Memory: A Holistic Benchmark of Agent Memory Substrates

Z. Huang et al. · 2026

Ce que cela signifie pour vous

Récupérer plus de contexte n'est pas automatiquement mieux : passé un certain point, cela chasse l'information dont le modèle a réellement besoin pour agir.

Résumé

Evaluates eleven memory substrates across seven families, three backbones and four task suites under twenty-six metrics. No substrate dominates: the winner inverts by regime, with structured graphs leading on dialogue question answering and cheap flat retrieval leading on code. Widening retrieval helps question answering but measurably degrades sequential decision-making, with attention probes showing mass draining from the action context into the retrieved block. Production-grade memory systems are also shown to cost between 2,700 and 9,000 auxiliary LLM calls per long history, an overhead usually reported nowhere.

Agent MemoryBenchmarkingRetrievalInference Cost
Lire l'article complet