3 weeks ago
OpenAI delays Astra launch over possible critical cyber capabilities
OpenAI made a new computer program called Astra that is incredibly smart.
Astra can solve math problems that experts have been stuck on for a very long time.
But OpenAI is worried that Astra might also be able to do dangerous things with computers.
That is why OpenAI is slowing down before letting people use Astra.
OpenAI is doing more safety tests and stopping some work that is not safe enough.
Astra recently solved ten super-hard math problems.
The answers were checked by a special computer program called Lean, which makes sure every step is correct.
OpenAI even told the White House about its plan to slow down.
Another company called Anthropic thinks pausing too much could be a problem too.
In the end, everyone wants to make sure powerful AI is used safely.
OpenAI told Axios it 'cannot rule out' that its upcoming Astra model has 'critical' cyber capabilities.
The company widened its safety testing regime and halted internal work that does not meet stricter safeguards.
Astra solved ten mathematics problems that had stood open for at least a decade, each verified with machine-checkable Lean certificates published on GitHub.
The 249-page results, published on 1 August, cost roughly $2,000 in API charges for the successful runs.
OpenAI proactively briefed the White House on its release delay, while Anthropic has argued unilateral pauses could backfire.
- Who
- OpenAI and its forthcoming Astra AI model, with Anthropic and the White House/Trump administration, plus mathematicians including Tim Gowers and Thomas Bloom.
- What
- OpenAI is slowing Astra's release after internal evaluations could not rule out 'critical cyber capabilities', expanding safety testing and delaying development until protections are in place.
- Where
- United States, where OpenAI briefed the White House; verification files were published publicly on GitHub.
- When
- As reported in early August; OpenAI released Astra's verified maths results on 1 August.
- Why
- To ensure adequate safety protections are in place before a potential public release of a model with possible critical cyber capabilities.
Caution: delay high-risk models
Progress: avoid unilateral pauses
Delaying powerful AI releases
Caution: delay high-risk models
OpenAI is slowing Astra's development until adequate protections exist, expanding security testing and halting internal work that fails stricter safeguards.
Progress: avoid unilateral pauses
Anthropic's updated Responsible Scaling Policy argues unilateral pauses could backfire, potentially leaving the AI landscape less safe if rivals keep shipping without comparable safeguards.
Pausing self-improving AI models
Caution: delay high-risk models
Anthropic's June blog post called for an industry-wide pause on models capable of improving themselves.
Progress: avoid unilateral pauses
Anthropic scaled back its own blanket pause pledge in February and instead released a more heavily safeguarded version of Mythos, its most cyber-capable model, in June.
Key facts
- AI model
- Astra (OpenAI)
- Risk finding
- OpenAI cannot rule out 'critical cyber capabilities'
- Safety response
- Expanded security testing; halted internal work failing stricter safeguards
- Math achievement
- Solved 10 problems open for at least a decade, including three from Paul Erdős's catalogue
- Verification
- Machine-checkable certificates formalised in Lean and published on GitHub
- API cost
- Roughly $2,000 for the successful runs
- White House briefing
- OpenAI informed the administration of its plan to delay the release
- Related release
- Anthropic released heavily safeguarded Claude Mythos 5 in June
Quotes
OpenAI spokesperson
Official representing OpenAI
“"The firm has also signalled it intends to slow Astra's development until adequate protections are in place, a step that could push back any eventual public release."”
livemint.com
“"OpenAI has widened its security testing regime and halted internal work that does not meet tighter safeguard requirements."”
livemint.com




