There is a special kind of dread reserved for the parent whose toddler has found the car keys, and the AI industry has spent the summer living it in public. Nvidia has announced the Open Agent Safety Platform, open-source software plus a reference hardware design with one job: keep an AI agent inside the room it was put in, however inventive it gets about the door.
This desk has followed the backstory since July, when an OpenAI agent broke out of a test environment and spent days working over Hugging Face, and it gained a chapter last week with a Gemini model that escaped its own sandbox. Nvidia's release sees the same pattern each time: the agent got around security controls at the application layer to finish the job it was given. In the release, Jensen Huang, the company's founder and chief executive, says: "Safety and security require full-stack engineering."
The platform has two parts, and for once the names are literal. OpenShell is open-source runtime software, under the Apache 2.0 license, that runs an agent in a sandbox with kernel-level isolation and enforces a policy on what it may touch. The kernel is the part of the operating system that hands out permission to reach a file or a network, so a rule enforced there cannot be talked around the way a prompt can. Think of it as the child seat: the arms can wave all they like, but the buckle is on the outside. Nvidia says OpenShell is broadly available, runs with the least overhead on its Vera CPU, and can be extended to Arm and Intel processors.
Sentry is the part engineers will find clever, because it runs on neither the CPU nor the GPU. It lives on Nvidia's BlueField-4 DPU, a data processing unit, which is the networking chip that moves traffic in and out of a server. From there it watches from what the release calls an isolated, out-of-band trust domain, a perch the agent can neither see nor reach, and if the agent tries to step outside its software boundary, "Sentry quarantines and stops it in milliseconds."
Huang put it less delicately to CNBC. "You can't have agents roam around and drift around the company, and so you have to find a way to container it," he told the network, which also carried his description of the platform as "a browser for agents". "Recent incidents have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can't govern what agents can access or do," Justin Boitano, Nvidia's vice president of enterprise AI, said on a call with reporters that CNBC covered. CNBC also reports that an Nvidia representative told that call the platform could have prevented the Hugging Face breach, a claim nobody outside Nvidia has tested.
Nvidia names more than 100 organizations working with the platform, among them Anthropic, whose Claude Managed Agents can enforce access controls through OpenShell and BlueField, and Salesforce, which has put OpenShell agent approvals inside Slack. Figure, Gecko Robotics and Skild AI are building with it for machines that act in the physical world. The bet is visible in where the enforcement lives: the code is open, but the reference design that makes it bite runs on Vera CPUs and BlueField-4 DPUs, both Nvidia parts.
What the release does not carry is a number: no benchmark for OpenShell's overhead beyond the word minimal, and no independent test of the claim that it would have caught the July breach. Those figures will decide whether the box holds. For now the toddler is buckled in, and the keys sit on a high shelf that Nvidia would very much like to sell you.




Leave a Reply