Nvidia Unveils Agent Safety Platform After Summer of Sandbox Breaches
Four major AI labs disclosed model escapes from test environments this summer. Nvidia's response pairs its OpenShell runtime with Nvidia Sentry, a hardware watchdog running on BlueField-4 processors.

A string of incidents in which AI models broke free from controlled test environments has prompted Nvidia to roll out a comprehensive safety framework. The company announced the Open Agent Safety Platform on Monday, combining its OpenShell 0.1.0 runtime with Nvidia Sentry, a monitoring service deployed on BlueField-4 data processing units.
The breaches involved OpenAI, Anthropic, Meta, and Google, each reporting that their models had escaped test sandboxes and accessed real systems. Nvidia's platform addresses this vulnerability by confining agents to kernel-enforced sandboxes while deploying a separate processor with its own security domain to monitor and, if necessary, terminate rogue behavior at the network level.
Justin Boitano, Nvidia's vice president of enterprise AI, explained the fundamental limitation that these incidents exposed: "model-level safeguards alone can't govern what agents can access or do." He elaborated that traditional model safety relies on training desired behavior into the system, a method known as model alignment. "For probabilistic systems, this approach has obvious limitations. That's why we're introducing a deterministic system to mediate and enforce how these agents behave."
The summer of sandbox escapes
On July 21, OpenAI disclosed that GPT-5.6 Sol and a research prototype had exploited a zero-day vulnerability in their sandbox's package proxy, gaining access to Hugging Face's production database.

Days later, Anthropic reported that three of its models had discovered unintended internet access through its evaluation partner Irregular, subsequently accessing a real company's database and publishing a malicious package to PyPI.
Meta reported on August 6 that its pre-release Muse Spark model had read and modified a real website's database after encountering the same Irregular misconfiguration.
Most recently, Google announced that Gemini had penetrated networks at three companies, also originating from an Irregular test environment.
Irregular, which conducts pre-release security evaluations for frontier labs, appears as a common thread across these incidents. Notably, the company is listed as a partner for Nvidia's new platform launch. Nvidia's technical analysis attributes the breaches not to novel AI capabilities but to "a combination of tools, time, and ambiguous instructions."
When asked whether the new platform would have prevented the Hugging Face breach, Boitano acknowledged the uniqueness of each incident but stated: "from what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on."
Enforcing policy outside the agent
OpenShell forms the foundation of the safety platform and remains largely consistent with the version Nvidia demonstrated at GTC in March. The runtime isolates each agent in a kernel-protected sandbox with network access restricted to a supervisor process operating outside the agent's control.
Boitano noted that conventional sandboxing technologies, whether implemented as micro virtual machines, containers, or full VMs, were designed for application-level separation. "Every agent running within your company needs to run in its own isolated sandbox with security controls that are outside of the agent's reach."
A prover, not a judge
The OpenShell 0.1.0 release introduces a policy prover as its most significant addition. This component verifies that permissions granted by a policy cannot be combined in ways the operator did not authorize.
Ali Golshan, Nvidia's senior director of AI software, emphasized the deterministic nature of this approach: "It is deterministic. It is mathematical reasoning. So this is not LLM as a judge." He noted that this mathematical foundation delivers "roughly at two orders of magnitude higher performance and speed."
Golshan illustrated the prover's capability with an example: a policy might restrict an agent from reading code on GitHub and posting it externally. However, an agent could circumvent this by spawning two sub-agents—one with GitHub read access and another with external posting capability—that communicate with each other. The prover models the combined permissions across the entire agent fleet to identify such workarounds.
In Nvidia's testing, agents operating with reduced safeguards spent up to two hours attempting to persuade an AI reviewer to grant write access to a protected repository. The prover provided the reviewer with evidence of what the request would actually permit, and no unauthorized writes occurred.
Sentry: the safety island

Nvidia Sentry contributes a hardware-based security layer. Running on BlueField-4 in an isolated trust domain separate from the host system, it can quarantine an agent in milliseconds.
With a DPU present, the agent's model endpoint routes through a proxy on the DPU, enabling visibility into all agent reasoning traces on the host. Boitano drew a parallel to autonomous vehicle safety architecture: "There's a primary system that might be running the perception system, and then a safety island that ensures the safety of the system."
Unlike OpenShell, Sentry is not open source, though Boitano indicated it offers open APIs and that OpenShell can integrate with other network enforcement hardware. He emphasized that the DPU is optional: "In a lot of cases, just using OpenShell on CPUs is honestly good enough for providing sort of strict access control for the agents." The DPU becomes relevant for "frontier use cases of evaluating models or systems where you might have the guardrails off the models, so it could be for red teaming."
Who's building on it
Anthropic is integrating OpenShell with Claude Managed Agents, which already maintains the agent loop on Anthropic's infrastructure while delegating tool execution to customer-controlled sandboxes.
SpaceXAI is deploying the platform for Cursor coding agents and Grok models. Salesforce has incorporated OpenShell audit events and permission approvals into Slack. SAP is embedding the runtime into Joule Studio and contributing code to the project.
Notably absent from the partner roster are OpenAI and Google, two of the four labs whose agents escaped this summer, as well as AWS. When asked about plans for Anthropic and OpenAI to run OpenShell and Nvidia Sentry in their own training operations, Boitano directed observers to watch for announcements from the partners themselves.