Recursive Self-Improvement of AI Research Agents
D. Srikanth, B. Zhao, D. Xu, Y. Wu, Z. Jiang · 2026
What this means for traders
A system tuned and graded on the same examples ends up learning the exam; the only honest score comes from cases held back from every adjustment.
Abstract
Uses one AI agent to rewrite the code of another research agent in a two-level loop: the inner agent optimizes each task against a public score, while the outer loop keeps a rewrite only if it improves a private, held-out score the inner agent never sees. Over an eight-day autonomous run the system accepted seven improvements, including a bandit over drafting strategies and bounded summaries of long logs, and the resulting agent matched or beat a human-designed baseline on four held-out benchmarks. Its measured rate of reward hacking on a kernel-optimization test fell from 55% to 32%, on a small sample reported without error bars.