2 hrs ago
Nvidia Unveils Safety Platform to Contain Rogue AI Agents
Nvidia has created tools to help keep computer programs called AI agents under control.
AI agents can perform tasks instead of only answering questions.
OpenShell puts the agent inside a protected software area with rules about what it can access.
Nvidia Sentry watches the agent using separate hardware.
If the agent tries to cross its limits, Sentry can isolate or stop it quickly.
Nvidia says companies such as Anthropic, Microsoft and Salesforce are using parts of the system.
The tools are meant to add extra safety layers instead of relying only on the AI to obey instructions.
However, the system cannot stop every mistake, trick, or bad choice made within the agent’s allowed permissions.
Nvidia announced the Open Agent Safety Platform for limiting autonomous AI agents.
OpenShell creates a secure software boundary that controls access and traces agent activity.
Nvidia Sentry uses BlueField-4 data processing units to monitor and stop agents independently.
More than 100 organisations are reportedly working with the technology, including Anthropic and Microsoft.
The platform can contain authorised actions but cannot prevent every mistake or poor decision.
- Who
- Nvidia announced the platform, with organisations including Anthropic, Salesforce, SAP, Scale AI, Microsoft, Palantir and JPMorganChase working with the technology.
- What
- The Nvidia Open Agent Safety Platform uses OpenShell software and Nvidia Sentry hardware to limit, monitor and potentially stop autonomous AI agents.
- Where
- The tools are designed for AI agents operating across software systems, including enterprise, coding and cybersecurity environments.
- When
- The announcement follows recent concerns and incidents involving AI agents acting beyond their intended instructions.
- Why
- Nvidia says the platform is intended to add independent controls as AI agents become more capable and autonomous.
Pacing Proponents
Layered-Safety Proponents
Speed of AI development
Pacing Proponents
Anthropic CEO Dario Amodei has argued that frontier AI development should be moderated so safety measures can keep up with rapidly advancing capabilities.
Layered-Safety Proponents
Nvidia’s approach is to add independent technical controls around agents, allowing development to continue while limiting what systems can access and do.
Who should enforce safety?
Pacing Proponents
Relying on increasingly capable AI systems to follow prompts or police their own behaviour may be insufficient.
Layered-Safety Proponents
Nvidia argues that safety restrictions should sit outside the model, using OpenShell and Sentry to enforce boundaries independently.
Effectiveness of controls
Pacing Proponents
Technical boundaries may not prevent an agent from making a mistake or bad decision within the permissions it has been given.
Layered-Safety Proponents
Nvidia says layered controls can trace activity, verify identities, enforce access policies and stop agents that move beyond their permitted boundaries.
Key facts
- Platform
- Nvidia Open Agent Safety Platform
- Software component
- OpenShell creates a controlled runtime boundary, traces activity and enforces policies.
- Hardware component
- Nvidia Sentry runs on BlueField-4 data processing units.
- Response time
- Nvidia says Sentry can quarantine and stop an agent in milliseconds.
- Organisational adoption
- Nvidia says more than 100 organisations are working with the technology.
- Processor compatibility
- OpenShell is designed for Nvidia Vera CPUs and can be extended to Arm and Intel processors.
- Limitation
- The platform does not by itself prevent mistakes, deceptive behaviour or bad decisions within an agent’s permissions.







