3 days ago
Anthropic Study Finds AI Models Improving Other Models
Anthropic tested whether one AI system could help train another AI system.
It created automated alignment researchers using Claude Opus 4.8.
These systems searched research papers, suggested training methods, and trained another model.
They worked on 10 tests designed to detect unsafe or misaligned behavior.
The automated systems improved every test without making the model worse overall.
In some cases, they performed better than methods suggested by human researchers.
They also cost less to operate than human researchers.
However, people are still needed to design good tests and provide useful research information.
Anthropic developed automated alignment researchers powered by Claude Opus 4.8.
The systems improved performance across all 10 tested alignment benchmarks.
The best automated method outperformed methods proposed by human researchers within six hours.
Anthropic estimated automated researchers cost about $4 per hour, compared with $150 for humans.
Researchers cautioned that automated systems still depend on human-designed benchmarks and research literature.
- Who
- Anthropic researchers, including Chen Yueh-Han, conducted the study using Claude Opus 4.8.
- What
- The study tested automated alignment researchers that train AI models to reduce specific alignment failures.
- Where
- The articles do not specify a physical location; the experiments used a Nvidia H200 GPU.
- When
- The paper was published on Friday, August 28.
- Why
- The research examined whether automated systems could improve AI alignment and potentially contribute to recursive self-improvement.
Evidence of Progress
Important Limitations
Path toward recursive self-improvement
Evidence of Progress
The automated systems improved every tested alignment benchmark without degrading overall performance, suggesting that AI-assisted model improvement may become practical soon.
Important Limitations
The findings are early evidence rather than proof that artificial general intelligence or full recursive self-improvement has been achieved.
Automation versus human researchers
Evidence of Progress
The best automated method beat methods proposed by experienced human researchers on average within six hours and operated at a much lower stated hourly cost.
Important Limitations
Automated researchers remain dependent on humans to define meaningful alignment benchmarks and expand the research literature they use.
Key facts
- Study title
- Automated Researchers Can Reliably Mitigate Alignment Failures
- Developer
- Anthropic
- AI system used
- Claude Opus 4.8
- Benchmarks tested
- 10 alignment benchmarks
- Training hardware
- Nvidia H200 GPU
- Automated researcher cost
- About $4 per hour in API inference
- Human researcher cost
- About $150 per hour
Quotes
Anthropic study authors
Researchers who conducted and reported Anthropic’s automated alignment study
“Overall, these results provide early evidence that automated alignment post-training could become practical in the near term.”
indianexpress.com
“cost roughly $4 per hour in API inference against the $150 per hour we pay our human researchers.”
indianexpress.com
Mark Chen
OpenAI’s chief research officer, as reportedly quoted regarding progress toward AGI
“80 per cent of the way”
indianexpress.com






