Evaluation

Recursive self-improvement depends on scaling evals

Michael Siu · Sep 15, 2026

Agents are already accelerating research inside frontier labs. OpenAI reports that its researchers use coding agents throughout the day and are writing more code and running more experiments.[1] As agents take on more of this work, evaluation has to keep up: how do we tell which changes actually helped? I think evaluation is a major bottleneck for automated AI R&D.

I’ll skip the buzzword-maxxing here and try not to abuse the term the way people do on social media. By “RSI,” I mean automated AI R&D where the output of one research cycle improves the system’s ability to run the next. This is becoming a concrete research goal,[2] and frontier labs now discuss automated research and track related self-improvement capabilities.[1][3][4]

As models improve, benchmarks saturate and we need new ways to tell whether we’re making progress. By scaling evals, I mean both expanding what we can test and improving how we judge the results.

AI research today, especially post-training, is already eval-driven. We judge a training run by its benchmark scores, and we increasingly build training data around benchmark distributions. To a large extent, benchmarks set direction. What we want an agent to become is tightly coupled to how we design its benchmarks.

#1Evaluation is selection

In an improvement loop, the eval is the selection mechanism. Systems that rewrite their own agent code,[5] evolve programs through search,[6] or revise their harnesses all need a signal that decides which changes survive. That signal has to reflect the capability we want to improve. (I think a lot of harness-improvement work gets this wrong. It accumulates experience on the target benchmark in advance, and the result often does not generalize.)

A research system proposes and tests changes; evaluation determines whether to update the system or keep it unchanged before the next cycle
Figure 1. Evaluation determines which changes enter the research system that runs the next cycle. If no change holds up, the system stays as it is. Author’s illustration.

An eval used to choose changes becomes part of the optimization process. We still need independent checks on problems outside that search to see whether the gains carry over.

Evaluation constrains the loop in two ways:

  1. Coverage. The eval needs enough varied and difficult tasks to distinguish general progress from gains on a narrow benchmark.
  2. Direction. What the eval rewards determines which improvements survive and which capabilities the system develops.

As agents get better at proposing changes and carrying out experiments, I expect reliable selection to become a bigger part of the bottleneck. The balance still depends on the domain: with a clean verifier, compute or the quality of the proposed changes may matter more.

A benchmark provides tasks, but the score comes from everything around them: environment, harness, model configuration, compute budget, grader, aggregation rule. The same model on the same tasks can look better or worse depending on these choices. Together they form the evaluation contract.

The target capability guides an evaluation setup: what to test, how to run it, and how to score it. The model is evaluated under that setup, and the score informs a selection decision
Figure 2. What we want to improve guides the setup. A model’s score depends on the tasks, run conditions, and scoring rules; the score then informs selection. Author’s illustration.

The loop keeps changes that score well under this setup. If the score is noisy, it can pick winners by chance. If the contract is narrow, repeated selection compounds a benchmark-specific mistake. A one-off eval error is a one-off loss, but in a recursive loop the error feeds the next round of improvement.

#2Why evaluation is hard

Many eval failures are infrastructure problems and engineering defects. An agent eval measures the model and its harness together, and too many factors can pollute a result. The environment breaks, resources run out, or a tool call times out. Runs are also non-deterministic, so a task can pass one trial and fail the next.

The grader itself can also be the defect. For example, one model solved a flight-booking task through a policy loophole, giving the user a better answer while failing the eval as written.[7] OpenAI’s 2026 audit of SWE-Bench Pro estimated that roughly 30% of its tasks were broken.[8] Researchers must find and fix these defects before treating a score as evidence of capability.

Take one step back from the verifier. Before you can check the result, you have to come up with the task in the first place. A good engineer knows what a good coding task looks like and might reconstruct one from representative GitHub issues. A lawyer knows their own workflow and what an acceptable deliverable looks like. By judgment, I mean the decisions that make an eval meaningful: which capabilities matter, which proxies represent them, and which failures reveal a real model limitation. Research taste is one name for this kind of judgment. A broad goal like “be a better researcher” leaves all of those choices open.

Artificial Analysis’s Intelligence Index is a useful example of how these choices affect rankings. On September 3, the v4.1.1 index gave GPT-6 Astra the same score as GPT-5.6 Sol and placed it behind Muse Spark 1.3.[9] Over the next two versions, the index expanded its emphasis on agentic work and private tests. Astra rose to second place, then tied for first.[10]

Artificial Analysis Intelligence and Coding Agent leaderboards under index v4.1.1, with GPT-6 Astra tied with GPT-5.6 Sol in intelligence and behind Claude Fable 5.1 and Muse Spark 1.3 Artificial Analysis Intelligence and Coding Agent leaderboards under index v4.3, with GPT-6 Astra tied with Claude Fable 5.1 for first place in both
Figure 3. Artificial Analysis’s v4.1.1 and v4.3 leaderboards. Because the benchmark mix and normalization changed, compare the ordering rather than the raw scores. Click either image to enlarge.Charts: Artificial Analysis.[9][10]

These updates changed more than weights: they changed the task mix and, for Terminal-Bench, the run budgets and verification. The ranking changed with the evaluation setup, which is exactly why those choices matter. A lab makes similar choices when it builds its internal suite, and those choices shape which improvements it pursues.

The capabilities we care about also change. A suite built around knowledge questions and short reasoning says less once we want agents to carry out long-running work. Keeping an eval useful means revisiting which capabilities it should measure, as well as finding harder tasks.

#3The judgment gap

Building and maintaining evals this way is slow and expensive, so the natural thing is to ask agents to write the benchmarks themselves. They already help with parts of the process. Open benchmark efforts often gather a group of contributors to submit tasks, and models help standardize the format and verify the submissions. Inside companies, dedicated teams collect and clean the data and use models to help generate the tasks. In both cases, humans still decide what the benchmark should measure.

That still leaves a harder problem: models struggle to decide which work deserves an eval.

  • Relevance. A model needs context about real-world work to judge whether a task represents something useful and which failures matter.
  • Difficulty. A task is only hard relative to a moving frontier. Without testing against current models, it is hard to know whether a generated task is actually difficult.

Judgment enters at two points. Before a run, someone chooses what belongs in the suite. After the run, many failures only become visible in the traces, and someone has to decide which ones matter. A failure may come from the model, the task, or the harness. Telling the difference requires understanding the surrounding work. Agents can help with both decisions, but domain experts still provide much of the context for deciding what belongs in a benchmark and interpreting its failures.

This is why most auto-research work still starts from a metric a human has chosen, whether that is a nanoGPT loss to beat or a GPU kernel to speed up. One recent paper from ByteDance Seed tries the harder setting, letting the agent derive its own validation signal from a broad capability goal.[11] On their own checks, the agents appeared to improve. Those gains did not reliably carry over to the independently designed hidden evaluation. The paper attributes the gap to mismatched training data and self-evaluations that covered too narrow a slice of the target capability.

#4Learning judgment

I imagine an eval agent that keeps building and revising benchmarks as the model improves. It would look for gaps in coverage, propose harder tasks, check tasks and graders, and investigate failures. What it learns from one round would shape the next: which tasks to add, which tests to retire, and which apparent gains need another look.

In my own work with agents, this already works pretty well within a session. If you explain why a grader rejected a valid solution, an agent can often fix it and look for similar mistakes. The harder part is carrying that lesson into the next session, with a different task and a different failure.

Saving the lesson as a skill or a review guideline gives a future agent something to work with. But it still has to recognize when the lesson applies. “Don’t reject a valid solution over an implementation detail” is useful advice; deciding whether that detail is incidental or an actual requirement takes context.

The same problem comes up when choosing what to test. A task can be clearly written and correctly graded, yet tell us little about the work we want the model to do. Feedback from people who know that work needs to shape future task choices, too.

I don’t yet know how to make that judgment accumulate reliably across sessions and domains. Skills, memory, and training are ways to carry feedback forward. We would still need to check whether an agent uses it well on unfamiliar cases, and whether its tests stay useful as the model improves. Otherwise, we risk turning yesterday’s good advice into another fixed proxy.

Before an improvement enters the next research cycle, the system needs a good reason to believe it is real.


#References

  1. OpenAI. The AI Policy Window Is Open. We Need to Act.; Research Acceleration: The View Inside OpenAI; Our Updated Preparedness Framework. 2025–2026.
  2. Recursive. First Steps Toward Automated AI Research. 2026.
  3. Anthropic. Frontier Safety Roadmap. Updated 2026.
  4. Google DeepMind. Strengthening Our Frontier Safety Framework. Updated 2026.
  5. Jenny Zhang et al. Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv, 2025.
  6. AlphaEvolve Team. AlphaEvolve: A Gemini-Powered Coding Agent for Designing Advanced Algorithms. Google DeepMind, 2025.
  7. Anthropic. Demystifying Evals for AI Agents. Anthropic Engineering, 2026.
  8. OpenAI. Separating Signal from Noise in Coding Evaluations. 2026.
  9. Artificial Analysis. Initial GPT-6 Astra benchmark thread. September 3, 2026. See also Simon Willison’s contemporaneous summary.
  10. Artificial Analysis. Intelligence Benchmarking Methodology; Index v4.2 release notes; Index v4.3 release notes; updated GPT-6 Astra results. 2026.
  11. Yuhao Wu et al. Aspire: Can Models Self-Evolve from Vague Goals?. arXiv, 2026.

#Citation

Siu, Michael. “Recursive Self-Improvement Depends on Scaling Evals.” michaelsiu.me (Sep 2026). https://michaelsiu.me/blog/evals-and-rsi

@misc{siu2026evals,
  title  = {Recursive Self-Improvement Depends on Scaling Evals},
  author = {Siu, Michael},
  year   = {2026},
  month  = {September},
  url    = {https://michaelsiu.me/blog/evals-and-rsi}
}