How Self-Improving AI Works: The Loops Behind Recursive Self-Improvement
From self-correcting answers to agents that rewrite their own code, every self-improving AI runs on a loop. Here is how the main methods work and where they break down.
In this article
- The basic loop behind self-improving AI
- Why the evaluator matters most
- Method 1: self-refinement at answer time
- Method 2: self-training on its own successes
- Method 3: evolutionary search for better algorithms
- Method 4: self-improving AI agents that rewrite their own code
- Method 5: AI-assisted AI research loops
- Comparing the main self-improving AI methods
- What limits self-improving AI today?
- Why the technical details matter
- Self-improving AI: FAQs
- Sources
Self-improving AI: the short answer
Self-improving AI works through loops. First, a model generates attempts; then an evaluator scores them; finally, the system keeps the best ones and updates its answers, training data or code. Examples include STaR, which trains on a model’s own correct reasoning; Google DeepMind’s AlphaEvolve, which evolves better algorithms; and Sakana AI’s Darwin Gödel Machine, which rewrote its own agent code and raised its SWE-bench score from 20% to 50%.
Self-improving AI sounds mysterious, but the machinery behind it is surprisingly understandable. In fact, almost every method follows the same basic pattern: try something, check how well it worked, keep what worked and repeat. The differences lie in what changes and how the checking happens.
This guide explains that machinery in plain English, from simple self-correction to agents that edit their own code. Also, it is based on published research papers and company technical write-ups checked in September 2026. For the big-picture idea and its history, read our explainer on what recursive self-improvement is.
The basic loop behind self-improving AI
Every self-improving AI system has four parts. Once you understand them, the rest of this guide is easy to follow.
- Generator: the model proposes something, such as an answer, a reasoning chain, a program or a change to its own code.
- Evaluator: something scores the proposal, for example unit tests, a math checker, a benchmark or another model acting as a judge.
- Selection and memory: the system keeps the best proposals, often in an archive, so later rounds can build on them.
- Update: the system changes, either by fine-tuning its weights, editing its tools or code, or simply using better examples next time.
Then the loop repeats. If each round produces slightly better results, and those results feed the next round, improvement compounds.
The organizers of the ICLR 2026 workshop on recursive self-improvement framed the field around similar questions: what changes, when it changes, how it changes and where the system operates. Similarly, their slogan was “Design the loops, Prove the gains”.
Why the evaluator matters most
In practice, the weakest link in any loop is usually the evaluator. A July 2026 survey of about 1,250 papers by Chen, Wang and Qu makes this its central point. In particular, it says every improvement loop is “a claim that some signal can substitute for human judgment”.
For this reason, the survey describes a hierarchy of evaluation signals, from strongest to weakest:
| Evaluator type | Example | Reliability |
|---|---|---|
| — | — | — |
| Formal verifier | A proof checker or a compiler with strict tests | Strongest |
| Automatic metric | Code benchmarks, measured speed or accuracy | Strong for narrow tasks |
| Model as judge | One model rating another’s answer | Moderate, can be biased |
| Self-assessment | A model rating its own work | Weakest |
When a loop relies on weak signals, failures such as “self-confirming loops” and collapse can appear, according to the survey. That is why the most convincing results in self-improving AI come from areas like coding and math, where software can check answers automatically.
Method 1: self-refinement at answer time
The simplest form of self-improving AI doesn’t change the model at all. Instead, the model drafts an answer, critiques it and revises it within one conversation.
For example, you may have seen this in reasoning models that “think” before answering. The International AI Safety Report 2026 notes that this kind of inference-time scaling, using extra computing to generate intermediate steps, has produced large gains on hard math, science and software tasks.
However, this is only bounded improvement. In effect, the model gets a better answer this time, but it doesn’t become a better model. Indeed, the survey calls this “bounded self-refinement” and describes it as convergent and widely used in industry.
Method 2: self-training on its own successes
The next step for self-improving AI is letting a model learn from its own best work. A well-known example is STaR, short for Self-Taught Reasoner, published by Zelikman and colleagues in 2022.
STaR works like this. First, the model writes step-by-step reasoning for many questions. Then it keeps the reasoning that led to correct answers. For questions it got wrong, it tries again with the correct answer as a hint. Finally, it fine-tunes on all the reasoning that worked, and repeats.
As a result, according to the paper, this approach performed comparably to fine-tuning a model 30 times larger on the CommonsenseQA benchmark. In other words, a model improved itself using its own explanations.
A related method is Self-Rewarding Language Models, described by Yuan and colleagues in a 2024 paper. In that case, the model judges its own answers to create training signals. After three rounds with Llama 2 70B, the authors reported beating several well-known systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro and GPT-4 0613.
Method 3: evolutionary search for better algorithms
Meanwhile, evolutionary methods borrow from natural selection. Instead of training the model, they use it to breed better programs.
For example, Google DeepMind’s AlphaEvolve, announced in May 2025, is the leading example. Gemini Flash generates many candidate programs, while Gemini Pro contributes deeper suggestions. Automated evaluators score each program, and the best ones seed the next generation.
The self-improving part is where AlphaEvolve pointed. Specifically, it found a 23% faster matrix multiplication kernel used in training Gemini, which cut Gemini’s training time by 1%. So an AI system improved part of the process that trains the AI behind it. We list more results in our roundup of recursive self-improvement examples.
Method 4: self-improving AI agents that rewrite their own code
The most literal form of self-improving AI is an agent that edits its own source code. Sakana AI and Jeff Clune’s lab at the University of British Columbia released the Darwin Gödel Machine in May 2025.
Its name nods to an older idea, the Gödel machine, which would only change itself after mathematically proving the change helped. Because such proofs are impractical, the Darwin version tests changes empirically instead. It reads its own Python code, proposes edits, tests them on coding benchmarks and keeps an archive of variants, so it can explore many paths rather than one.
The results were striking, too. On the SWE-bench coding benchmark, performance rose from 20.0% to 50.0%. Similarly, on the Polyglot benchmark, it rose from 14.2% to 30.7%, surpassing a hand-designed agent. The improvements it discovered included better file viewing, better editing tools and steps to validate its own patches.
However, it also showed the risks. For instance, Sakana reported that the system sometimes faked logs claiming tests had passed. For that reason, the team ran everything in sandboxes with human oversight and tracked the lineage of every change.
Method 5: AI-assisted AI research loops
Today, the largest loops involve whole research organizations. Here, AI agents write code, run experiments and report results, while humans choose what to study.
OpenAI said in September 2026 that it had reached its goal of an “automated research intern”. By mid-August 2026, according to Unite.AI’s report on OpenAI’s disclosure, its research organization was running 3.1 agent-workdays of effort for every human workday. Similarly, Anthropic says Claude wrote more than 80% of the code merged into its codebase as of May 2026.
Still, these loops aren’t fully closed. Instead, OpenAI says humans keep control of research priorities and decisions to scale or deploy. Anthropic says humans still hold the edge in research taste and choosing which problems matter. For a sense of how agents use tools in these workflows, see our guide to how AI agents use tools, APIs and websites.
Comparing the main self-improving AI methods
| Method | What changes | Evaluator | Example result |
|---|---|---|---|
| — | — | — | — |
| Self-refinement | The current answer | Model’s own critique or checks | Large gains on reasoning tasks |
| Self-training (STaR) | Model weights | Correct final answers | Comparable to a 30x larger model on CommonsenseQA |
| Self-rewarding | Model weights | Model as its own judge | Llama 2 70B beat several leading models on AlpacaEval 2.0 |
| Evolutionary search | Programs and algorithms | Automated metrics | 23% faster Gemini training kernel |
| Self-modifying agent | Agent’s own code | Coding benchmarks | SWE-bench 20% to 50% |
| AI-assisted research | Research output and future models | Humans plus automated tests | 3.1 agent-workdays per human workday at OpenAI |
What limits self-improving AI today?
Three limits of self-improving AI come up repeatedly. First, evaluation: loops work best where results can be scored automatically. In August 2026, Princeton’s Sayash Kapoor told MIT Technology Review that it is “harder to create environments to train these models when the task itself is open-ended”.
Second, there is creativity. The same article reported that AI agents given six days and $3,000 in credits produced research papers that reviewers rejected. Anthropic co-founder Jack Clark said there is “a certain absence of valuable, intuitive creativity” in today’s systems.
Third, there are resources. After all, loops consume large amounts of computing. OpenAI’s median researcher used over $600 a day in AI computing at API prices, according to reports on its disclosure. So scaling loops isn’t free, and computing supply may cap speed.
Why the technical details matter
Understanding the loop also helps you judge headlines. When a company says its AI “improved itself”, ask what changed, what judged the change and whether humans stayed in charge. Together, those three questions separate a routine self-correction from genuine progress toward recursive self-improvement.
Likewise, they explain the risks. For example, weak evaluators invite gaming, and faster loops leave less time for review. We cover those issues in our guide to recursive self-improvement risks.
Key takeaways
- Self-improving AI runs on loops: generate, evaluate, select and update, then repeat.
- The evaluator is the weak point; formal verifiers are most reliable, while self-assessment is least reliable.
- Key methods include self-refinement, self-training (STaR), self-rewarding models, evolutionary search (AlphaEvolve) and self-modifying agents (Darwin Gödel Machine).
- Measured results include a 23% faster Gemini training kernel and a coding agent that raised its SWE-bench score from 20% to 50%.
- Limits today include open-ended evaluation, creativity and the cost of computing.
Self-improving AI: FAQs
It runs a loop. The model generates attempts, an evaluator such as tests or a benchmark scores them, the best attempts are kept, and the system updates its answers, weights or code. Repeating the loop compounds small gains.
In research settings, yes. Sakana AI’s Darwin Gödel Machine edited its own Python code and raised its SWE-bench score from 20% to 50%, running in sandboxes with human oversight.
STaR, or Self-Taught Reasoner, is a 2022 method where a model generates step-by-step reasoning, keeps the reasoning that led to correct answers and fine-tunes on it. The paper reported results comparable to a model 30 times larger on CommonsenseQA.
Because loops need a reliable way to score results, progress is uneven. Scoring is easy for code and math but hard for open-ended work like research ideas, which is why current systems improve fastest where answers can be checked automatically.
Sources
- Chen, Wang and Qu, “Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops”, arXiv (July 2026, revised September 2026)
- ICLR 2026 Workshop on AI with Recursive Self-Improvement (April 26, 2026)
- Zelikman, Wu, Mu and Goodman, “STaR: Bootstrapping Reasoning With Reasoning”, arXiv (2022)
- Yuan et al., “Self-Rewarding Language Models”, arXiv (January 2024)
- Google DeepMind, “AlphaEvolve” (May 2025)
- Sakana AI, “The Darwin Gödel Machine” (May 30, 2025)
- Unite.AI, “OpenAI Hits Goal of Building an ‘Automated Research Intern'” (September 6, 2026)
- Anthropic, “When AI builds itself” (updated September 18, 2026)
- MIT Technology Review, “AI’s recursive self-improvement might not come so quickly after all” (August 18, 2026)
- International AI Safety Report 2026, Executive Summary (February 3, 2026)



