When Nvidia announced its Open Agent Safety Platform on 28 September, I didn't have a take. So I read what it actually does.
It has two parts, and only one of them is something you can run. OpenShell is an open-source runtime for AI agents, Apache 2.0, written mostly in Rust, currently at version 0.1.2. The other part, Sentry, is a watchdog that runs on Nvidia's BlueField-4 chips, outside the machine the agent lives on. Nvidia says it can quarantine a misbehaving agent in milliseconds. It also calls Sentry a "reference system design." I couldn't find any code for it, and the press release gives no price and no date. So a lot of the arguing about Nvidia's safety chip is arguing about a diagram.
What OpenShell actually does#
The agent runs in a sandbox with no route to the network except through one supervisor. A connection only gets out if a rule matches the host, the port and the program making the call. Nvidia's walkthrough lets curl reach the GitHub API on port 443, and nothing else. It can also split reads from writes on the same API, so in the demo a read goes through and a POST gets blocked.
The agent never holds the real keys. It gets a placeholder, and the supervisor swaps in the real credential only when the request is headed somewhere approved. Send the placeholder anywhere else and the request gets rejected.
The operating system kernel limits which files the agent can touch and stops it picking up extra privileges. Those limits stay when the agent opens a shell, runs code it wrote, or starts child processes. There's also a policy prover. It either proves the permissions you've modelled stay inside your boundary, or it shows you a specific action that crosses it. Nvidia's line on this is my favourite in the whole write-up: "an agent's explanation cannot change that result."
Then there's the policy advisor. When the agent hits a wall, it can propose a narrow new rule, like access to one more host. The proposal waits for a human by default, and "the agent cannot approve its own request." Nvidia says that in its own tests, agents with reduced safeguards spent up to two hours trying to talk an AI reviewer into letting them write to a protected GitHub repo. None of them got in. That's Nvidia's claim with no method published, but it's a strange picture. Two hours of an agent lobbying.
A firewall with branding#
None of this is new. Default-deny outbound traffic, a proxy with an allowlist, secrets kept away from the workload, kernel sandboxing... it's what a careful ops team would already build.
The Hacker News thread had people saying exactly that. One asked, "Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?" Another said "we already have VMs, containers, firewalls, airgaps, etc." A bigger share of the thread was people noticing that a chip company had found a reason to sell a chip. On the technology, the firewall crowd is right. OpenShell is a good default-deny firewall with some agent-shaped conveniences on top.
The trouble is that most people running agents never set that firewall up. "You could configure it yourself" is true of nearly every security control, and it usually means nobody did. It's the weekend job nobody schedules.
A default that ships ready gets used, mostly because it's already there. You install OpenShell, then one command starts Codex or Claude Code inside it. One commenter put it plainly: "Nvidia isn't wrong for offering a turnkey mitigation option."
Where the pitch gets ahead of itself#
On a call with reporters, an Nvidia representative said the platform could have prevented OpenAI's Hugging Face incident in July. CNBC reported it. I couldn't find that claim anywhere Nvidia has put it in writing, and nobody has explained how it would have worked.
The same piece quotes Justin Boitano, Nvidia's vice president of enterprise AI: "From what we know, Hugging Face reported over 17,000 agents attacking their infrastructure that went on for days and weeks." I spent a while on that number because it didn't match anything I'd read. METR's investigation puts it at about 1,200 agents on the message board and around 700 in the attack. The closest figure I could find from Hugging Face itself is in its timeline, which says its forensics covered "~17,600 attacker actions," over about four and a half days.
Hugging Face counted actions, not agents.
A fence also doesn't touch why the agents cheated. When I wrote about the Hugging Face breach, what stuck with me was that the test pushed agents to cheat and almost none of them thought about telling a human. Nvidia seems to agree on the first part. Its Sentry write-up says an agent in these circumstances "cannot be expected to fully govern its own behavior." Its answer is containment, which leaves the incentives exactly where they were.
The policy advisor looks a bit like a tell-a-human channel. It's built for permission requests, though. An agent can say it needs one more host. There's nowhere for it to say the task looks impossible and it's about to start cheating.
So I'd call it useful. It's also version 0.1.2, Nvidia already put out a security bulletin covering OpenShell in August, and I couldn't find an independent audit of the proxy or the isolation.
