Nvidia has launched a new security platform designed to stop autonomous AI agents from acting outside the limits set by developers.
The Nvidia Open Agent Safety Platform combines software controls with an independent hardware-based monitoring layer that can quarantine an AI agent within milliseconds if it attempts to move beyond its permitted environment.
The launch comes as AI agents gain access to files, tools, APIs and external services, allowing them to perform increasingly complex tasks with less direct human supervision: a trend that has raised broader questions about whether autonomous AI systems could eventually escape meaningful human control.
Nvidia Adds a Safety Layer Outside the Model
The platform uses OpenShell, an open-source runtime that creates enforceable boundaries around an AI agent and determines which resources it can access.
Nvidia pairs that software layer with Sentry, an independent watchdog running on its BlueField-4 data processing units.
Because Sentry operates separately from the AI model itself, the agent cannot simply override the safety mechanism through its own reasoning. If suspicious behavior is detected, the system can isolate the agent within milliseconds.
Nvidia says model-level safeguards alone may no longer be enough as agents become more capable.
How the System Works
| Layer | Function |
|---|---|
| OpenShell | Sets access rules and software boundaries |
| AI agent | Performs tasks inside approved limits |
| Sentry | Independently monitors behavior |
| Unsafe action | Agent can be quarantined within milliseconds |
AI Agents Are Becoming Harder to Control
AI agents are rapidly moving beyond traditional chatbots.
They can write and execute code, interact with software and carry out long sequences of actions with limited supervision. Some agents can already interact directly with financial platforms, including systems that let AI agents trade crypto and access investment accounts.
That creates a new security problem: an agent may technically follow its objective while still taking actions developers did not anticipate.
Nvidia CEO Jensen Huang said AI safety increasingly requires a “full-stack” approach, meaning protections must extend beyond the model and into the infrastructure running it.
More than 100 organizations are involved with the initiative. Participants include Anthropic, Microsoft, CrowdStrike, Hugging Face, Cisco, JPMorgan Chase, Palantir, Salesforce and SAP.
The broader significance is that Nvidia is expanding beyond providing the chips behind AI. It is also positioning itself as part of the security infrastructure designed to keep autonomous agents under control.