10 hrs ago

AI Giants Launch Advanced Models Amid Intensifying Safety Debate

AI Giants Launch Advanced Models Amid Intensifying Safety Debate
AI Giants Unveil Advanced Models, Raising Critical Safety Questions · rediff.com

Several big technology companies released new artificial-intelligence models.

OpenAI’s Astra is especially good at finding computer-security weaknesses.

That power could help defenders, but it could also help someone attack protected systems.

Some researchers worry that Astra’s hidden reasoning may be harder to inspect.

OpenAI says the concerns exaggerate how much its architecture changed and that it still values chain-of-thought monitoring.

The company also says it added protections and restricted access to some features.

Anthropic released models aimed at coding, science, and other difficult tasks.

Google released new models for reasoning, cybersecurity, video understanding, and image editing.

The debate is about how to gain the benefits of stronger AI while keeping it safe and understandable.

Key facts

OpenAI model
Astra
Astra cybersecurity result
It achieved a perfect score on the public ExploitBench benchmark for known vulnerabilities and found two genuine zero-day vulnerabilities in a separate internal evaluation.
Astra access
Its most advanced cybersecurity features will initially be limited to vetted users; defensive access is also planned through the Daybreak Blue programme.
Anthropic models
Claude Fable 5.1 and Claude Mythos 5.1, with Mythos restricted to vetted users for sensitive cybersecurity and biology work.
Anthropic pricing
Anthropic said typical-workload prices would fall by around 25 percent.
Google video capability
Agentic video understanding can dynamically search and scan video, with Google reporting up to 88 percent lower token consumption, 66 percent lower costs, and 7 percent higher accuracy.
Google image tool
Google Pics is an image creation and editing tool rolling out to Google AI Pro and Ultra subscribers and most Workspace business customers.

Quotes

Ryan Greenblatt

Chief scientist at Redwood Research

“With the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.”
rediff.com
“the single worst development for AI security and safety to date”
rediff.com

Sources

Related news