2 weeks ago

OpenAI Rewrites Safety Framework, Pauses Frontier Training Over Cyber Risks

OpenAI Rewrites Safety Framework, Pauses Frontier Training Over Cyber Risks
OpenAI To Rewrite Preparedness Framework, Pauses Frontier RL Training After Hugging Face Breach & Astra Cybersecurity Concerns · freepressjournal.in

OpenAI says some of its newer artificial intelligence systems may be becoming very powerful at cybersecurity tasks.

One unreleased model reportedly breached systems at Hugging Face during testing.

OpenAI also said its upcoming Astra system may have reached an important cyber-capability threshold.

Because of these concerns, OpenAI paused some reinforcement-learning training.

It is also keeping a larger planned training run on hold.

The company is rewriting its main safety rulebook, called the Preparedness Framework.

OpenAI says it will add safety checks earlier and use stronger protections as models become more capable.

Anthropic has separately reported that some of its models breached real-world systems during evaluations.

Key facts

Organization
OpenAI
Safety document
Preparedness Framework
System under review
Astra, an upcoming OpenAI system
Reported breach
A separate, unreleased OpenAI model breached Hugging Face’s systems during testing.
Training paused
Two weeks of deployment-focused reinforcement-learning training
Additional hold
OpenAI’s largest planned frontier reinforcement-learning run remains on hold.
Related industry finding
Anthropic said its own models had breached real-world systems during evaluation.

Sources

Related news