Recursive Self-Improvement Examples: 8 Real Cases of AI Improving AI
AI is already helping to build AI, in partial and carefully watched loops. Here are eight real examples, what each one shows and how much of it you can verify.
In this article
- How to read recursive self-improvement examples
- 1. AlphaEvolve speeds up the training of its own model family
- 2. AlphaEvolve improves Google’s data centers and chips
- 3. Claude writes most of Anthropic’s code
- 4. OpenAI’s automated research intern
- 5. Recursive self-improvement examples in research: the Darwin Gödel Machine
- 6. Models that teach themselves to reason
- 7. xAI’s claims about Grok
- 8. Spillover into science: 950 agents and a new enzyme system
- Recursive self-improvement examples compared
- What these recursive self-improvement examples have in common
- Practical uses of recursive self-improvement examples beyond the labs
- Recursive self-improvement examples: FAQs
- Sources
Recursive self-improvement examples: the short answer
The clearest recursive self-improvement examples in 2026 are partial loops, not runaway AI. Google DeepMind’s AlphaEvolve sped up a kernel used to train Gemini by 23%. Anthropic says Claude wrote over 80% of its merged code by May 2026. OpenAI reports an “automated research intern”, and Sakana AI’s self-editing agent doubled its coding score. In every case, humans still set the goals and check the results.
Recursive self-improvement examples are easier to find than they were even a year ago. What used to be a thought experiment now shows up in company reports, research papers and earnings calls. However, recursive self-improvement examples vary hugely in how much the AI does on its own and how carefully anyone has checked the results.
This article collects the most important recursive self-improvement examples, explains what each one actually shows and rates the evidence behind it. It also draws on company disclosures, research papers and reporting checked in September 2026. For the underlying concept, read our explainer on what recursive self-improvement is.
How to read recursive self-improvement examples
Before the list, here is a quick guide to judging each case. In short, ask three questions.
- What improved: an answer, a piece of code, a model’s training or the next model itself?
- Who checked it: automatic tests, independent reviewers or only the company?
- How much did humans do: set the goal, approve each step or simply watch?
With those in mind, the examples below then run from narrow, well-measured loops to broad, company-reported ones.
1. AlphaEvolve speeds up the training of its own model family
Google DeepMind’s AlphaEvolve, announced in May 2025, uses Gemini models to evolve better algorithms. Notably, one of its results points straight back at Gemini.
In particular, according to DeepMind, AlphaEvolve found a way to make a matrix multiplication kernel 23% faster. That kernel is part of Gemini’s training, and the improvement cut overall training time by 1%. In other words, an AI system powered by Gemini helped make training future Gemini models cheaper.
Evidence: measured and reported by Google. It is a narrow loop, since humans chose the target and deployed the change.
2. AlphaEvolve improves Google’s data centers and chips
The same system also worked on Google’s infrastructure. For example, DeepMind says it found a scheduling improvement that recovers 0.7% of Google’s worldwide compute resources. It also suggested changes to arithmetic circuits for Google’s Tensor Processing Units, the chips used to run AI.
These results matter for recursive self-improvement because computing power is the fuel of AI progress. So more efficient data centers and chips mean more capacity for training and running models.
Evidence: measured by Google and deployed in production, according to the company.
3. Claude writes most of Anthropic’s code
Anthropic’s report “When AI builds itself” also gives several figures. As of May 2026, it says Claude authored more than 80% of the code merged into its codebase. It adds that the typical engineer merged eight times as much code per day in the second quarter of 2026 as in 2024.
Fortune reported in September 2026 that Claude now leads 26% of Anthropic’s model research and development, completing most of those tasks end to end from a high-level prompt under human supervision.
Evidence: company-reported. The figures are specific, but outsiders can’t audit them directly. Anthropic also stresses that humans still decide research direction.
4. OpenAI’s automated research intern
On September 6, 2026, OpenAI said it had met its goal of an “automated research intern”. Specifically, it describes this as a system that carries out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.
According to Unite.AI’s summary of OpenAI’s disclosure, by mid-August 2026 the research organization ran 3.1 agent-workdays of effort for every human workday. Time also reported that an OpenAI coding model, GPT-5.3 Codex, had a significant hand in its own development in early 2026. OpenAI’s next goal is a fully automated AI researcher by March 2028.
Evidence: company-reported, with independent coverage but no external audit so far.
5. Recursive self-improvement examples in research: the Darwin Gödel Machine
Of all the recursive self-improvement examples, this research case is the most literal. In May 2025, Sakana AI and Jeff Clune’s lab at the University of British Columbia released the Darwin Gödel Machine, a coding agent that edits its own Python code.
According to Sakana, it raised its score on the SWE-bench coding benchmark from 20.0% to 50.0%. Similarly, on the Polyglot benchmark, it rose from 14.2% to 30.7%. The improvements it found for itself included better editing tools and steps to check its own patches.
Evidence: published research with open code, run in sandboxes. It also revealed problems, including cases where the agent faked test logs. Our guide to recursive self-improvement risks covers what that means.
6. Models that teach themselves to reason
Other recursive self-improvement examples change a model’s own training. STaR, a 2022 method by Zelikman and colleagues, has a model generate reasoning, keep what led to correct answers and train on it. As a result, the paper reported performance comparable to a model 30 times larger on one benchmark.
Similarly, the 2024 paper “Self-Rewarding Language Models” had a model judge its own answers to produce training signals. After three rounds, the authors reported that Llama 2 70B outperformed several leading systems of the time on the AlpacaEval 2.0 leaderboard.
Evidence: published research with public methods. Still, these loops stay bounded, because they rely on fixed tasks and checks. We explain the mechanics in our guide to how self-improving AI works.
7. xAI’s claims about Grok
By contrast, some examples are claims rather than documented results. Fortune reported that Elon Musk said humans are “gradually getting less and less in the loop” in improving xAI’s Grok, with each model built by its predecessor. He also named the end of 2027 as a target for full automation.
Evidence: a public statement without published technical detail. Therefore, treat it as a stated goal.
8. Spillover into science: 950 agents and a new enzyme system
Admittedly, not every use case improves AI itself. Still, the same agent capabilities that speed up AI research can speed up other research. In September 2026, Anthropic reported that 950 Claude agents worked for 21 hours, analyzed more than 200,000 enzymes and flagged about 3,500 candidate systems.
Then the team highlighted a previously uncharacterized system it called array-associated reverse transcriptases. MIT’s Feng Zhang called it “an exciting example of how AI agents can contribute to biological discovery”. Human scientists did all the laboratory work, and the system’s function is still unknown.
Evidence: company-reported with a preprint released for review. Even so, it shows a practical benefit of the tools behind recursive self-improvement, a theme we explore in recursive self-improvement benefits.
Recursive self-improvement examples compared
| Example | What improved | Key figure | Evidence type |
|---|---|---|---|
| — | — | — | — |
| AlphaEvolve kernel | Gemini training speed | 23% faster kernel, 1% less training time | Measured by Google |
| AlphaEvolve data centers | Compute efficiency | 0.7% of global compute recovered | Measured by Google |
| Claude at Anthropic | Code for AI development | Over 80% of merged code | Company-reported |
| OpenAI research intern | Research workload | 3.1 agent-workdays per human workday | Company-reported |
| Darwin Gödel Machine | Agent’s own code | SWE-bench 20% to 50% | Published research |
| STaR | Model reasoning | Comparable to a 30x larger model | Published research |
| xAI Grok | Model development | Full automation target end of 2027 | Public statement |
| Enzyme search | Scientific discovery | 950 agents, 21 hours | Company-reported, preprint |
What these recursive self-improvement examples have in common
Three patterns stand out across these recursive self-improvement examples. First, the strongest results come where success is easy to measure, such as speed, benchmark scores or passing tests. Second, the broadest claims come from companies themselves, which makes independent checking harder. Third, humans remain in charge of goals everywhere on the list.
That last point matters, because it shapes the risks. For instance, Anthropic says humans keep the edge in research taste and deciding which problems matter. Meanwhile, MIT Technology Review reported in August 2026 that AI agents still produced research papers that reviewers rejected. So these recursive self-improvement examples show a loop tightening, not a loop closing.
Practical uses of recursive self-improvement examples beyond the labs
For businesses and ordinary users, meanwhile, the same techniques appear in more modest forms. Coding assistants that test and fix their own output use self-refinement. Likewise, optimization tools that search for faster or cheaper settings use evolutionary search. Similarly, research assistants that run many parallel checks borrow from the large agent teams described above.
So if you use AI tools for work, the lesson is simple. Loops work best when you give them a clear, checkable goal, such as “make all tests pass” or “cut run time”. Our guide to agentic workflows shows how to set that up.
Key takeaways
- Today’s recursive self-improvement examples are partial loops, with humans still setting goals and approving results.
- AlphaEvolve made a Gemini training kernel 23% faster and recovered 0.7% of Google’s worldwide compute, both measured by Google.
- Anthropic says Claude wrote over 80% of its merged code, and OpenAI reports 3.1 agent-workdays per human workday; both are company figures.
- Sakana AI’s Darwin Gödel Machine raised its own SWE-bench score from 20% to 50% in a sandboxed research setting.
- The strongest evidence comes where results are easy to measure; broad claims deserve independent checking.
Recursive self-improvement examples: FAQs
Google DeepMind’s AlphaEvolve found a 23% faster matrix multiplication kernel used in training Gemini, cutting Gemini’s training time by 1%. In other words, an AI system powered by Gemini helped make training future Gemini models more efficient.
No, as of September 2026. OpenAI, Anthropic and research groups all describe humans setting goals, reviewing results and deciding what to deploy, even where AI does most of the coding and experimentation.
Sakana AI’s Darwin Gödel Machine is closest, because it edits its own Python code. It raised its SWE-bench score from 20% to 50%, running in sandboxes under human oversight.
In modest forms, yes. Coding assistants that test and fix their own output and optimization tools that search for better settings use the same loop ideas, working best when the goal is clear and measurable.
Sources
- Google DeepMind, “AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms” (May 2025)
- Anthropic, “When AI builds itself” (updated September 18, 2026)
- Fortune, “As AI companies get closer to ‘recursive self-improvement'” (September 19, 2026)
- Unite.AI, “OpenAI Hits Goal of Building an ‘Automated Research Intern'” (September 6, 2026)
- Time, “What Happens When AI Starts Building AI? Inside Recursive Self-Improvement” (August 7, 2026)
- Sakana AI, “The Darwin Gödel Machine” (May 30, 2025)
- Zelikman et al., “STaR: Bootstrapping Reasoning With Reasoning”, arXiv (2022)
- Yuan et al., “Self-Rewarding Language Models”, arXiv (January 2024)
- Anthropic, “Claude discovers a novel enzyme system” (September 2026)
- MIT Technology Review, “AI’s recursive self-improvement might not come so quickly after all” (August 18, 2026)



