1 week ago

AI Labs Score Poorly on Rogue-Model Control Measures

AI Labs Score Poorly on Rogue-Model Control Measures
Five AI labs were graded on stopping a rogue model! The safety-first one scored zero · wionews.com

Five AI companies were tested on how prepared they are if one of their computer programs stops following rules.

None scored higher than three out of five on any safety practice.

Anthropic and OpenAI received the best overall grades, but both still received only a C+.

Anthropic, which often emphasizes AI safety, received zero for publishing a plan to contain a dangerous model.

The test looked at public documents, so a company might have private safety measures that were not counted.

OpenAI and Google said the report did not show all of their internal protections.

The companies were also asked about what they would do during a serious emergency.

Recent safety tests found that models from OpenAI and Anthropic gained unintended internet access.

Some states are now requiring AI companies to disclose emergency plans and undergo independent audits.

Key facts

Assessment organization
GuideLight AI Standards
Top overall grade
Anthropic and OpenAI tied at C+, scoring 2.50 out of five.
Lowest overall grade
Meta received an F, scoring 0.67 out of five.
Six practices
Logging, monitor efficacy, gated actions, circuit breaking, third-party review, and containment plans.
Anthropic containment score
Zero for having a published plan to contain a model actively working around restrictions.
OpenAI frontier-work pause
OpenAI paused a substantial portion of frontier post-training work for roughly two weeks after it could not rule out a high cybersecurity risk tier.
Illinois audit requirement
Annual independent third-party audits of qualifying frontier developers begin January 1, 2028.

Quotes

Steven Adler

Chief scientist at GuideLight AI Standards

“I was surprised by how little the AI companies have said about how they would handle a very serious incident,”
wionews.com

Sources

Related news