NVIDIA has introduced the Open Agent Safety Platform — an open reference security stack for autonomous AI agents consisting of two layers. The OpenShell runtime under the Apache 2.0 license runs the agent in a sandbox and passes its actions through kernel-level restrictions, while the optional NVIDIA Sentry hardware layer on the BlueField-4 DPU operates outside the reach of the agent and host and, according to the press release, is capable of quarantining an out-of-bounds agent in milliseconds. OpenShell 0.1.0 was released on September 25, 2026, the current release v0.1.2 is dated September 28, and the NVIDIA/OpenShell repository has already gathered 12.4k stars. The point of the release is that agent security no longer depends on the model's promises: execution is mechanically restricted by the infrastructure.

What Happened
The company released the platform as an open reference implementation of continuous in-silicon agent monitoring and published the OpenShell runtime in a public GitHub repository under the Apache 2.0 license. Agent operating rules are described in a YAML policy: it lists permitted files, system calls, network requests down to the HTTP method and specific path, processes, and external providers. The enforcement logic is executed by an instrumented kernel, and every policy change before agent launch undergoes formal verification in a prover component. The agent does not see real keys: the OpenShell gateway substitutes them inside requests to policy-approved addresses, so secrets remain outside the sandbox. The NVIDIA Sentry hardware layer operates on the BlueField-4 DPU through the DOCA environment, lives in a trust domain separate from the host, identifies agents, and isolates those that have crossed established boundaries.
Context
Until now, agentic AI security has mainly relied on prompts and "model self-restrictions": the developer asked the agent not to go out of bounds and hoped the model would hold. NVIDIA is moving control from the realm of trust in the model to the realm of a verifiable environment — the kernel and hardware, where promises do not matter. The company explains the logic with its own analogy: the browser made the internet safe not by the promises of page creators, but by ceasing to trust their code. The Sentry hardware layer addresses the known research problem of drift: during long operation or when hitting a restriction, the agent deviates from the original task and cannot restrain itself, so control must be external — hence the idea of a host-independent trust domain on the DPU. The bet on openness — Apache 2.0, SDKs in four languages, extensibility on Arm and Intel architectures — is supported by industry engagement: more than 100 organizations have joined at launch, and Anthropic, Salesforce, SAP, Red Hat, and Scale AI are integrating the platform into their products such as Claude Managed Agents, Slack, and Joule Studio. Against this backdrop, the release looks not like a proprietary NVIDIA feature, but like a contender for a general runtime standard for agentic AI.
Why This Matters for the Industry
For the industry, the main point is not the fact of the release itself, but the appearance of a ready-made building block. OpenShell turns agent security into policy-as-code: the YAML policy can be audited, changes are formally verified, the gateway hides keys from the agent, and the appearance of prover stages in CI/CD and libraries of ready-made policies for typical scenarios looks like an expected continuation. This commoditization of sandboxes reduces the cost for agentic startups to enter the enterprise market: instead of developing their own, they can wrap already used agents in a controlled environment and pass corporate security checks faster — a quick win in time-to-market with minimal entry cost. There is also a downside: the most powerful mechanism — quarantine outside the host — is tied to the BlueField-4 DPU, which could split the ecosystem into those deploying NVIDIA hardware and those limiting themselves to the software runtime. If the open runtime with verifiable policies takes hold, the market will have a measurable baseline for comparing agentic runtimes — from the percentage of successful escapes to the latency of response to a boundary violation — but this is a prospect on the horizon of a couple of years, not an established fact.
Why This Matters for Users
If you work with Claude Code, Codex, OpenCode, or GitHub Copilot CLI, you can run these agents in a sandbox with a single command: installation is done via the install.sh script, then the openshell sandbox create command is used. You describe the policy for files, network, and processes in YAML, keys remain outside the sandbox, and network rules can be edited on the fly — for example, you can allow the agent to only read the GitHub API without giving write permissions. This works on Linux and macOS with Apple Silicon; for Windows, a path via WSL 2 is provided, currently in experimental mode. Release 0.1.x is already marked as stable, and the Apache 2.0 license makes the tool free — a working option for enthusiasts and small teams, not just for corporate security departments. For researchers, the project provides open code that can be audited and used as a basis for their own agent security assessments.
What Is Still Unknown / Limitations
Key impressive characteristics remain the company's own claims: Sentry's capabilities — quarantine in milliseconds and independence from the host — are taken from the press release and have not been verified by an independent party. The sources contain no data on the sandbox's resistance to escapes or red-team tests, and the repository's popularity is a metric of audience attention, not code quality. The rapid replacement of early 0.1.x versions every few days indicates an unstable API, and Windows support via WSL 2 is positioned as experimental. The only thing known about the prover is that it formally verifies policy changes; what formal model stands behind it and what specific properties are proven is not disclosed — "verified" without such methodology creates a risk of false confidence. Hiding keys closes one class of risks — secret leakage — but does not answer the main question about the sufficiency of the policy against agent drift from the original task. Finally, there are no independent measurements of sandbox overhead, which determines real production deployment, in the sources.
Sources
- NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment | NVIDIA Newsroom
- NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring | NVIDIA Technical Blog
- NVIDIA/OpenShell — safe, private runtime for autonomous AI agents (GitHub, Apache 2.0, v0.1.0 released 2026-09-25)
Author
Look at AI, editorial team
