10 hrs ago
AI Giants Launch Advanced Models Amid Intensifying Safety Debate
Several big technology companies released new artificial-intelligence models.
OpenAI’s Astra is especially good at finding computer-security weaknesses.
That power could help defenders, but it could also help someone attack protected systems.
Some researchers worry that Astra’s hidden reasoning may be harder to inspect.
OpenAI says the concerns exaggerate how much its architecture changed and that it still values chain-of-thought monitoring.
The company also says it added protections and restricted access to some features.
Anthropic released models aimed at coding, science, and other difficult tasks.
Google released new models for reasoning, cybersecurity, video understanding, and image editing.
The debate is about how to gain the benefits of stronger AI while keeping it safe and understandable.
OpenAI says its Astra model reached the “Critical” cybersecurity capability tier and found two genuine zero-day vulnerabilities in testing.
Researchers warn Astra’s reported recurrent-depth architecture could make its reasoning harder to monitor through chain-of-thought analysis.
OpenAI says it delayed parts of Astra’s release, strengthened protections, and will initially limit advanced cybersecurity access to vetted users.
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, highlighting coding and scientific-research performance while reducing typical-workload prices by about 25 percent.
Google introduced Gemini 3.8 Flash, Gemini 3.8 Flash Cyber, agentic video understanding, and the Google Pics image creation and editing tool.
- Who
- OpenAI, Anthropic, and Google, along with AI safety researchers and other commentators.
- What
- The companies announced advanced AI models and tools, prompting debate over cybersecurity risks, transparency, and responsible deployment.
- Where
- The announcements involved the companies’ AI products; OpenAI is identified as being based in San Francisco.
- When
- The articles do not provide a specific announcement date.
- Why
- The companies said the releases offer stronger capabilities in cybersecurity, coding, research, reasoning, video understanding, and image editing, while critics questioned whether the systems can be adequately monitored.
Safety Critics
AI Companies and Defenders
Monitorability of Astra
Safety Critics
Researchers argue that recurrent depth moves reasoning into unreadable internal computations, making it harder to verify the model’s behavior and potentially weakening chain-of-thought monitoring.
AI Companies and Defenders
OpenAI Chief Scientist Jakub Pachocki called the concerns based on confused reporting, saying the architectural change is more limited than some reports suggest and that OpenAI has worked to preserve chain-of-thought monitoring.
Speed of capability development
Safety Critics
Critics say increasingly capable models may be advancing faster than methods for deploying them responsibly and warn that the shift could undermine existing safety practices.
AI Companies and Defenders
OpenAI says Astra represents progress in both capabilities and alignment, while also saying it delayed parts of the release and strengthened safeguards before deployment.
Cybersecurity access
Safety Critics
Safety concerns focus on Astra’s ability, with the right tools and access, to discover and exploit previously unknown flaws across well-protected systems.
AI Companies and Defenders
OpenAI says access to the most advanced cybersecurity features will initially be restricted to vetted users, with broader defensive use offered through Daybreak Blue; Google similarly limits Gemini 3.8 Flash Cyber through the Fairwind programme.
Key facts
- OpenAI model
- Astra
- Astra cybersecurity result
- It achieved a perfect score on the public ExploitBench benchmark for known vulnerabilities and found two genuine zero-day vulnerabilities in a separate internal evaluation.
- Astra access
- Its most advanced cybersecurity features will initially be limited to vetted users; defensive access is also planned through the Daybreak Blue programme.
- Anthropic models
- Claude Fable 5.1 and Claude Mythos 5.1, with Mythos restricted to vetted users for sensitive cybersecurity and biology work.
- Anthropic pricing
- Anthropic said typical-workload prices would fall by around 25 percent.
- Google video capability
- Agentic video understanding can dynamically search and scan video, with Google reporting up to 88 percent lower token consumption, 66 percent lower costs, and 7 percent higher accuracy.
- Google image tool
- Google Pics is an image creation and editing tool rolling out to Google AI Pro and Ultra subscribers and most Workspace business customers.
Quotes
Ryan Greenblatt
Chief scientist at Redwood Research
“With the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.”
rediff.com
“the single worst development for AI security and safety to date”
rediff.com








