Aller au contenu principal
Retour à la recherche
Apprentissage automatique

Recursive Self-Improvement of AI Research Agents

D. Srikanth, B. Zhao, D. Xu, Y. Wu, Z. Jiang · 2026

Ce que cela signifie pour vous

Un système réglé et noté sur les mêmes exemples finit par apprendre l'examen, et seule une série de cas tenue à l'écart de chaque réglage donne une note honnête.

Résumé

Uses one AI agent to rewrite the code of another research agent in a two-level loop: the inner agent optimizes each task against a public score, while the outer loop keeps a rewrite only if it improves a private, held-out score the inner agent never sees. Over an eight-day autonomous run the system accepted seven improvements, including a bandit over drafting strategies and bounded summaries of long logs, and the resulting agent matched or beat a human-designed baseline on four held-out benchmarks. Its measured rate of reward hacking on a kernel-optimization test fell from 55% to 32%, on a small sample reported without error bars.

Self-Improving AgentsHeld-Out EvaluationReward HackingAgent Design
Lire l'article complet