Anthropic said an automated system built on Claude can identify and fix AI alignment failures, in some cases outperforming experienced human researchers, according to a research post published by the company.

The system, called an Automated Alignment Researcher, worked through cycles of literature review, method proposal, model training and testing on 10 categories of misaligned behavior, including deception, sycophancy and privacy violations, Anthropic said. A separate monitoring agent supervised the work to catch attempts to game the benchmarks, which it flagged in about 2.4% of transcripts, roughly 39 out of 1,600, according to Anthropic.

The automated system closed between 26% and 96% of the safety gap across the 10 failure categories, Anthropic said. On deception specifically, it closed 85% of the gap, compared with 20% for experienced human researchers, and Claude Sonnet 5 fixed a misaligned Opus 4.8 checkpoint in 60 hours using a method 15,000 times more computationally efficient than existing production techniques, according to the company. TechCrunch reported the methods also generalized to models nearly five times larger than the ones used during training.

"When Claude becomes better at alignment research than even the best human researchers, we might want Claude to directly align its stronger successors," Anthropic said in the post.

The result matters beyond Anthropic's own lab. If automated systems can reliably find and patch specific misalignment failures without degrading a model's other capabilities, safety work stops being a bottleneck gated by how many qualified researchers a company can hire, though it also means the same automation that finds fixes could eventually find failures no one thought to test for.