3 days ago

Anthropic Study Finds AI Models Improving Other Models

Anthropic Study Finds AI Models Improving Other Models
AI models are getting better at training other models, Anthropic study finds · indianexpress.com

Anthropic tested whether one AI system could help train another AI system.

It created automated alignment researchers using Claude Opus 4.8.

These systems searched research papers, suggested training methods, and trained another model.

They worked on 10 tests designed to detect unsafe or misaligned behavior.

The automated systems improved every test without making the model worse overall.

In some cases, they performed better than methods suggested by human researchers.

They also cost less to operate than human researchers.

However, people are still needed to design good tests and provide useful research information.

Key facts

Study title
Automated Researchers Can Reliably Mitigate Alignment Failures
Developer
Anthropic
AI system used
Claude Opus 4.8
Benchmarks tested
10 alignment benchmarks
Training hardware
Nvidia H200 GPU
Automated researcher cost
About $4 per hour in API inference
Human researcher cost
About $150 per hour

Quotes

Anthropic study authors

Researchers who conducted and reported Anthropic’s automated alignment study

“Overall, these results provide early evidence that automated alignment post-training could become practical in the near term.”
indianexpress.com
“cost roughly $4 per hour in API inference against the $150 per hour we pay our human researchers.”
indianexpress.com

Mark Chen

OpenAI’s chief research officer, as reportedly quoted regarding progress toward AGI

“80 per cent of the way”
indianexpress.com

Sources

Related news