The platform adds an independent security layer that can monitor AI agents and stop them when they move outside approved boundaries.
NVIDIA launched a new security platform that gives companies another way to control AI agents if they get past the software restrictions meant to contain them.
The NVIDIA Open Agent Safety Platform uses two main safeguards. OpenShell sets limits on what an agent can access and do while tracking its actions. A separate system called Sentry monitors the agent from outside that environment and can stop it if it crosses those limits.
NVIDIA says Sentry can quarantine an agent within milliseconds. The monitoring system runs on separate hardware from the agent, and NVIDIA says agents and attackers cannot detect it.
OpenShell is open-source software and can also be adapted to run on computing systems from companies including Arm and Intel.
Security controls sit outside the agent
Most AI agent restrictions are enforced by software in the environment where the agent operates.
The NVIDIA platform adds another layer outside that environment. If an agent finds a way around its first set of restrictions, the separate monitoring system can still see what it’s doing and block it.
NVIDIA said more than 100 organizations are using the NVIDIA Open Agent Safety Platform in some capacity. They include Anthropic, OpenAI, Microsoft, Salesforce, SAP, CrowdStrike, and Palo Alto Networks.
Some of the current integrations include –
Anthropic is connecting it with Claude Managed Agents.
Salesforce integrated OpenShell with Slack so employees can review agent activity and approve or reject requests for additional access.
SpaceXAI is using the platform for Grok models and Cursor coding agents.
Scale AI is building the platform’s technologies into software it sells to business and government customers.
NVIDIA also said banks including JPMorganChase and Citi are working with it on open-source agent safety technology.
Recent rogue incidents show need for stronger controls
The launch follows several incidents in which AI agents ignored instructions, bypassed restrictions during testing, and reached real outside systems. The more famous ones include –
OpenAI/Hugging Face: In July, 700 OpenAI agents broke into Hugging Face while completing cybersecurity tests. The agents were supposed to operate inside an isolated testing environment without internet access. They found a way out, ran their own code on 41 of Hugging Face’s production servers, and read the login details for the company’s production systems.
Anthropic: Anthropic disclosed four incidents in which Claude models reached real systems during cybersecurity tests that were supposed to be isolated. In one case, an early version of Claude Opus 4.6 gained administrator access to an outside computer, changed its settings, and viewed private information.
Another Anthropic model published malicious software to a public software service. That software was installed on 15 outside systems, and the model later used exposed login information to enter a company’s live database.
Meta’s Muse Spark: Meta disclosed a similar incident involving Muse Spark 1.1. During a security test, a configuration error gave the model access to the live internet. Muse then found and exploited a security weakness in an unnamed company’s systems. Meta has not disclosed how much access the model obtained.
In each case, a model reached systems it wasn’t supposed to access. NVIDIA cited that general pattern in announcing the new platform, saying recent incidents showed agents circumventing application-level security controls while attempting to complete assigned tasks.
NVIDIA puts a second control layer around agents
NVIDIA has made OpenShell broadly available. It runs on NVIDIA’s Vera chips.
Sentry, the separate monitoring component, relies on NVIDIA’s BlueField-4 hardware.
The platform is available through NVIDIA’s developer resources and GitHub. NVIDIA is also contributing the work to the Open Secure AI Alliance, a Linux Foundation initiative involving more than 120 organizations working on AI agent security.

