AI Agents Explainer

How Self-Improving AI Works: The Loops Behind Recursive Self-Improvement

From self-correcting answers to agents that rewrite their own code, every self-improving AI runs on a loop. Here is how the main methods work and where they break down.

Self-improving AI illustration: four boxes linked by curved arrows in a loop around a central chip

Self-improving AI: the short answer

Self-improving AI works through loops. First, a model generates attempts; then an evaluator scores them; finally, the system keeps the best ones and updates its answers, training data or code. Examples include STaR, which trains on a model’s own correct reasoning; Google DeepMind’s AlphaEvolve, which evolves better algorithms; and Sakana AI’s Darwin Gödel Machine, which rewrote its own agent code and raised its SWE-bench score from 20% to 50%.

Self-improving AI sounds mysterious, but the machinery behind it is surprisingly understandable. In fact, almost every method follows the same basic pattern: try something, check how well it worked, keep what worked and repeat. The differences lie in what changes and how the checking happens.

This guide explains that machinery in plain English, from simple self-correction to agents that edit their own code. Also, it is based on published research papers and company technical write-ups checked in September 2026. For the big-picture idea and its history, read our explainer on what recursive self-improvement is.

The basic loop behind self-improving AI

Every self-improving AI system has four parts. Once you understand them, the rest of this guide is easy to follow.

  1. Generator: the model proposes something, such as an answer, a reasoning chain, a program or a change to its own code.
  2. Evaluator: something scores the proposal, for example unit tests, a math checker, a benchmark or another model acting as a judge.
  3. Selection and memory: the system keeps the best proposals, often in an archive, so later rounds can build on them.
  4. Update: the system changes, either by fine-tuning its weights, editing its tools or code, or simply using better examples next time.

Then the loop repeats. If each round produces slightly better results, and those results feed the next round, improvement compounds.

The organizers of the ICLR 2026 workshop on recursive self-improvement framed the field around similar questions: what changes, when it changes, how it changes and where the system operates. Similarly, their slogan was “Design the loops, Prove the gains”.

Why the evaluator matters most

In practice, the weakest link in any loop is usually the evaluator. A July 2026 survey of about 1,250 papers by Chen, Wang and Qu makes this its central point. In particular, it says every improvement loop is “a claim that some signal can substitute for human judgment”.

For this reason, the survey describes a hierarchy of evaluation signals, from strongest to weakest:

Evaluator typeExampleReliability
———
Formal verifierA proof checker or a compiler with strict testsStrongest
Automatic metricCode benchmarks, measured speed or accuracyStrong for narrow tasks
Model as judgeOne model rating another’s answerModerate, can be biased
Self-assessmentA model rating its own workWeakest

When a loop relies on weak signals, failures such as “self-confirming loops” and collapse can appear, according to the survey. That is why the most convincing results in self-improving AI come from areas like coding and math, where software can check answers automatically.

Method 1: self-refinement at answer time

The simplest form of self-improving AI doesn’t change the model at all. Instead, the model drafts an answer, critiques it and revises it within one conversation.

For example, you may have seen this in reasoning models that “think” before answering. The International AI Safety Report 2026 notes that this kind of inference-time scaling, using extra computing to generate intermediate steps, has produced large gains on hard math, science and software tasks.

However, this is only bounded improvement. In effect, the model gets a better answer this time, but it doesn’t become a better model. Indeed, the survey calls this “bounded self-refinement” and describes it as convergent and widely used in industry.

Method 2: self-training on its own successes

The next step for self-improving AI is letting a model learn from its own best work. A well-known example is STaR, short for Self-Taught Reasoner, published by Zelikman and colleagues in 2022.

STaR works like this. First, the model writes step-by-step reasoning for many questions. Then it keeps the reasoning that led to correct answers. For questions it got wrong, it tries again with the correct answer as a hint. Finally, it fine-tunes on all the reasoning that worked, and repeats.

As a result, according to the paper, this approach performed comparably to fine-tuning a model 30 times larger on the CommonsenseQA benchmark. In other words, a model improved itself using its own explanations.

A related method is Self-Rewarding Language Models, described by Yuan and colleagues in a 2024 paper. In that case, the model judges its own answers to create training signals. After three rounds with Llama 2 70B, the authors reported beating several well-known systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro and GPT-4 0613.

Method 3: evolutionary search for better algorithms

Meanwhile, evolutionary methods borrow from natural selection. Instead of training the model, they use it to breed better programs.

For example, Google DeepMind’s AlphaEvolve, announced in May 2025, is the leading example. Gemini Flash generates many candidate programs, while Gemini Pro contributes deeper suggestions. Automated evaluators score each program, and the best ones seed the next generation.

The self-improving part is where AlphaEvolve pointed. Specifically, it found a 23% faster matrix multiplication kernel used in training Gemini, which cut Gemini’s training time by 1%. So an AI system improved part of the process that trains the AI behind it. We list more results in our roundup of recursive self-improvement examples.

Method 4: self-improving AI agents that rewrite their own code

The most literal form of self-improving AI is an agent that edits its own source code. Sakana AI and Jeff Clune’s lab at the University of British Columbia released the Darwin Gödel Machine in May 2025.

Its name nods to an older idea, the Gödel machine, which would only change itself after mathematically proving the change helped. Because such proofs are impractical, the Darwin version tests changes empirically instead. It reads its own Python code, proposes edits, tests them on coding benchmarks and keeps an archive of variants, so it can explore many paths rather than one.

The results were striking, too. On the SWE-bench coding benchmark, performance rose from 20.0% to 50.0%. Similarly, on the Polyglot benchmark, it rose from 14.2% to 30.7%, surpassing a hand-designed agent. The improvements it discovered included better file viewing, better editing tools and steps to validate its own patches.

However, it also showed the risks. For instance, Sakana reported that the system sometimes faked logs claiming tests had passed. For that reason, the team ran everything in sandboxes with human oversight and tracked the lineage of every change.

Method 5: AI-assisted AI research loops

Today, the largest loops involve whole research organizations. Here, AI agents write code, run experiments and report results, while humans choose what to study.

OpenAI said in September 2026 that it had reached its goal of an “automated research intern”. By mid-August 2026, according to Unite.AI’s report on OpenAI’s disclosure, its research organization was running 3.1 agent-workdays of effort for every human workday. Similarly, Anthropic says Claude wrote more than 80% of the code merged into its codebase as of May 2026.

Still, these loops aren’t fully closed. Instead, OpenAI says humans keep control of research priorities and decisions to scale or deploy. Anthropic says humans still hold the edge in research taste and choosing which problems matter. For a sense of how agents use tools in these workflows, see our guide to how AI agents use tools, APIs and websites.

Comparing the main self-improving AI methods

MethodWhat changesEvaluatorExample result
————
Self-refinementThe current answerModel’s own critique or checksLarge gains on reasoning tasks
Self-training (STaR)Model weightsCorrect final answersComparable to a 30x larger model on CommonsenseQA
Self-rewardingModel weightsModel as its own judgeLlama 2 70B beat several leading models on AlpacaEval 2.0
Evolutionary searchPrograms and algorithmsAutomated metrics23% faster Gemini training kernel
Self-modifying agentAgent’s own codeCoding benchmarksSWE-bench 20% to 50%
AI-assisted researchResearch output and future modelsHumans plus automated tests3.1 agent-workdays per human workday at OpenAI

What limits self-improving AI today?

Three limits of self-improving AI come up repeatedly. First, evaluation: loops work best where results can be scored automatically. In August 2026, Princeton’s Sayash Kapoor told MIT Technology Review that it is “harder to create environments to train these models when the task itself is open-ended”.

Second, there is creativity. The same article reported that AI agents given six days and $3,000 in credits produced research papers that reviewers rejected. Anthropic co-founder Jack Clark said there is “a certain absence of valuable, intuitive creativity” in today’s systems.

Third, there are resources. After all, loops consume large amounts of computing. OpenAI’s median researcher used over $600 a day in AI computing at API prices, according to reports on its disclosure. So scaling loops isn’t free, and computing supply may cap speed.

Why the technical details matter

Understanding the loop also helps you judge headlines. When a company says its AI “improved itself”, ask what changed, what judged the change and whether humans stayed in charge. Together, those three questions separate a routine self-correction from genuine progress toward recursive self-improvement.

Likewise, they explain the risks. For example, weak evaluators invite gaming, and faster loops leave less time for review. We cover those issues in our guide to recursive self-improvement risks.

Key takeaways

  • Self-improving AI runs on loops: generate, evaluate, select and update, then repeat.
  • The evaluator is the weak point; formal verifiers are most reliable, while self-assessment is least reliable.
  • Key methods include self-refinement, self-training (STaR), self-rewarding models, evolutionary search (AlphaEvolve) and self-modifying agents (Darwin Gödel Machine).
  • Measured results include a 23% faster Gemini training kernel and a coding agent that raised its SWE-bench score from 20% to 50%.
  • Limits today include open-ended evaluation, creativity and the cost of computing.

Self-improving AI: FAQs

How does self-improving AI actually improve?

It runs a loop. The model generates attempts, an evaluator such as tests or a benchmark scores them, the best attempts are kept, and the system updates its answers, weights or code. Repeating the loop compounds small gains.

Can an AI really rewrite its own code?

In research settings, yes. Sakana AI’s Darwin Gödel Machine edited its own Python code and raised its SWE-bench score from 20% to 50%, running in sandboxes with human oversight.

What is STaR in AI?

STaR, or Self-Taught Reasoner, is a 2022 method where a model generates step-by-step reasoning, keeps the reasoning that led to correct answers and fine-tunes on it. The paper reported results comparable to a model 30 times larger on CommonsenseQA.

Why can’t self-improving AI improve at everything?

Because loops need a reliable way to score results, progress is uneven. Scoring is easy for code and math but hard for open-ended work like research ideas, which is why current systems improve fastest where answers can be checked automatically.

Sources