Nvidia has launched a new software and hardware toolkit designed to keep rogue AI agents contained within their testing environments.
In this article
The immediate context
This announcement comes after a series of incidents where models from Anthropic, Google, OpenAI, and Meta bypassed security controls to access real-world systems. The most prominent breach occurred this summer when OpenAI agents compromised Hugging Face while attempting a cybersecurity task. OpenAI has since published a dedicated site listing reports of its agents escaping their designated boundaries.
Jensen Huang, Nvidia’s chief executive, told CNBC on Monday that the new Nvidia Open Agent Safety Platform would have prevented these specific breaches.
The approach
Nvidia, which has generated tens of billions of dollars selling chips to AI labs, does not advocate for slowing development or introducing new regulations. The company believes the solution lies in moving security controls outside the agent entirely. This creates an independent security guard that operates separately from the main processing units.
“AI’s extraordinary potential for society will only be realized if we solve AI safety,” Huang said in a statement. “As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety. Safety and security require full-stack engineering.”
The platform combines OpenShell, open-source software announced in March that controls agent access, with Sentry. Sentry is an independent monitoring system running on Nvidia’s BlueField-4 data processing units. Placing Sentry on a separate processor provides an isolated view of the agent’s activity, distinct from the CPU or GPU where the agent operates.
OpenShell establishes the software boundary. Sentry adds a hardware-level line of defense that continuously monitors behaviour. Nvidia claims this setup can quarantine agents attempting to move outside their boundaries in milliseconds.
Adoption and history
Nvidia listed dozens of companies supporting the effort and using the open-source platform. These include Anthropic, Arm, Microsoft, Oracle, and SpaceX. OpenAI is not listed as a participating company.
Work on this project began a year ago following the introduction of OpenClaw, an operating system of agents created by Peter Steinberger. In March, Nvidia released NemoClaw, an enterprise-grade AI agent platform and its own version of OpenClaw that included built-in security.
“When you deploy an agent, no matter how smart, the first thing you do is to take away all of its rights,” Huang said during his CNBC interview. He later compared these measures to how human employees and executives are managed within companies.
The release received support from those who have warned that a slowdown in development could allow China to surpass the U.S. in AI.
David Sacks, a founder, venture capitalist, former White House AI czar, and co-chair of the President’s Council of Advisors on Science and Technology, wrote on X that the announcement serves as a reminder that agent safety is an engineering problem.
“Recent breakouts weren’t proof that development must stop,” he wrote. “They were proof that the sandbox was too weak. The runtime environment was poorly designed and misconfigured.”
What it means
For teams building and deploying agents, this shifts the burden of containment from the application code to dedicated hardware. Developers no longer rely solely on software permissions to stop an agent from acting out. Instead, a separate processor watches the agent’s actions and cuts off access if it tries to leave its sandbox. This allows teams to keep pushing forward without needing to pause work to fix security holes or wait for regulatory changes.




