MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
J. Jiang, D. Yuan, W. Li · 2026
What this means for traders
An assistant that remembers you should also know when a fact is out of date and forget it when you ask; a good average score says nothing about either.
Abstract
Argues that average scores hide the rare but serious failures of agents with long-term memory. Defines five risk types (stale facts after an update, unresolved conflicting facts, leakage across users, reuse of revoked memories, and gradual decay of standing constraints) and checks each one deterministically against the agent's execution trace instead of relying on an LLM judge. Across 120 scripted episodes and five small open models, overall pass rates conceal large per-risk gaps, with cross-user leakage the weakest category for every model tested. A subset that keeps only the episodes where models disagree reproduces the full ranking at a fifth of the cost.