Automated systems can improve an AI model across ten different behavioural benchmarks without degrading overall performance. That is the finding published by an Anthropic researcher in late August 2026. It is a concrete demonstration of what researchers call 'self-improving AI': systems that adjust themselves in a targeted way without continuous human guidance.
The findings are notable because they show that targeted correction of unwanted behaviour can work at scale. On each of the ten benchmarks, focused on specific forms of misaligned behaviour, the automated system made measurable progress. At the same time, the model's broad performance remained intact, a combination that earlier alignment research had struggled to achieve.
Anthropic, founded in 2021 by, among others, Dario Amodei (CEO) and Daniela Amodei (President), has long focused on AI safety and alignment. The company is headquartered in San Francisco and was valued at nearly one trillion dollars in May 2026, following a series of large funding rounds from investors including Amazon, Google and Coatue.
What the research shows
The core of the findings centres on automating a process that normally requires substantial manual effort: identifying and correcting unwanted behaviour in large language models. Researchers at Anthropic tested whether automated systems could take over that work, and at what scale.
Each of the ten benchmarks addressed a specific form of misalignment, behaviour in which an AI model acts in a way that conflicts with the intentions of its designers or the interests of users. Examples include circumventing safety guidelines, inconsistent responses, or subtle forms of manipulation. The automated system succeeded in improving scores on all ten.
What makes the result technically significant is the absence of the usual trade-off. In many earlier attempts to fine-tune models on specific points, performance on other tasks declined, a phenomenon known in the field as 'catastrophic forgetting'. That did not occur here.
Jack Clark, co-founder of Anthropic, and Marina Favaro, who leads the Anthropic Institute, are involved in related work on how AI builds itself. Their contribution to the piece 'When AI builds itself' aligns with the broader direction of this research, although the precise boundary between the publications cannot be fully reconstructed from the available source material.
Why alignment researchers are watching this closely
Self-improving AI has been a central topic in the alignment debate for years, but typically at the level of theory or carefully constructed thought experiments. This research shifts the discussion to the empirical side: what already works, and under what conditions?
For alignment researchers, the findings are interesting on two levels. On one hand, they point to a potentially efficient tool: if automated systems can reliably detect and correct unwanted behaviour, models can be steered more quickly and accurately than when the process depends entirely on human evaluators. On the other hand, the findings raise questions about control: who or what determines which behaviour is flagged as 'misaligned', and how do you prevent that definition from shifting as systems become more autonomous?
Anthropic has positioned alignment as a core theme from the outset. Researchers such as Chris Olah, known for his work on mechanistic interpretability, and Amanda Askell, who was involved in developing the constitutional AI framework behind Claude, work on techniques to make models more transparent and more steerable. This new research fits within that same line of work, but adds an automated layer to what has until now been a largely manual process.
Anthropic in the broader context of AI safety research
Anthropic is not the only company working on automated alignment techniques. OpenAI, DeepMind and a range of academic groups are conducting comparable research. Even so, the publication carries weight, simply because Anthropic has the computing capacity and the models to test findings at scale. Through Amazon Web Services, the company has access to substantial computing power and has raised tens of billions of dollars over the past few years.
At the same time, many European and Dutch AI labs operate with considerably smaller budgets. The gap in available computing power and data makes it difficult for European parties to replicate this kind of large-scale experiment or to build on it independently. That has implications for how the continent positions itself relative to American and Chinese AI development, not only commercially, but also in terms of safety and governance.
For policymakers in Europe working on the implementation of the AI Act, this is relevant. If automated self-improvement of models becomes a reality, it will place new demands on oversight and certification. The question then is not only what a model is capable of today, but also how it adapts itself in the future, and whether external supervisors can still follow that process.