Skip to main content
Back to Research
Machine Learning

Recursive Self-Improvement of AI Research Agents

D. Srikanth, B. Zhao, D. Xu, Y. Wu, Z. Jiang · 2026

What this means for traders

A system tuned and graded on the same examples ends up learning the exam; the only honest score comes from cases held back from every adjustment.

Abstract

Uses one AI agent to rewrite the code of another research agent in a two-level loop: the inner agent optimizes each task against a public score, while the outer loop keeps a rewrite only if it improves a private, held-out score the inner agent never sees. Over an eight-day autonomous run the system accepted seven improvements, including a bandit over drafting strategies and bounded summaries of long logs, and the resulting agent matched or beat a human-designed baseline on four held-out benchmarks. Its measured rate of reward hacking on a kernel-optimization test fell from 55% to 32%, on a small sample reported without error bars.

Self-Improving AgentsHeld-Out EvaluationReward HackingAgent Design
Read the full paper