Nvidia Launches Security Platform to Prevent AI Agents from Overstepping Their Authority
- They monitor their behavior and isolate suspicious activity within fractions of a second
Amid growing concerns about the ability of AI agents to exceed their assigned tasks and permissions and operate independently, US company Nvidia announced on Monday an open-source security platform designed to set boundaries for these agents’ operations, monitor their behavior, and prevent them from executing unauthorized actions.
Nvidia unveiled the “Open Agent Safety” platform, which includes the open-source “OpenShell” system and an independent security layer named “Sentry,” in a move aimed at providing security controls for AI agents during testing and operational phases.
OpenShell creates a secure runtime environment for AI agents, allowing developers to define policies and permissions that govern the agent’s access to files, networks, tools, and other resources, and to verify that the agent possesses only the necessary permissions to execute its assigned task without exceeding them.
Justin Poutano, Nvidia’s Vice President of Enterprise AI, stated during a press conference that the system could have prevented the hacking incident suffered by AI specialist company Hugging Face, which was carried out independently by a group of OpenAI’s AI agents.
Poutano added that, according to the company’s information, the breach could have been prevented if this security technology had been used in leading laboratories during the early stages of model evaluation.
According to Nvidia, OpenShell allows developers to verify that an AI agent possesses only the necessary permissions to perform its assigned task. Thanks to its open-source nature, it can also be extended to work with other computing platforms, including ARM and Intel platforms.
The platform includes an independent security layer called “Sentry,” which operates using BlueField-4 DPU data processing units separately from the AI agent’s runtime environment, to continuously monitor its activity and intervene if it attempts to exceed the established boundaries.
According to the company, Sentry can isolate an AI agent within fractions of a second when suspicious behavior is detected. Poutano clarified that OpenShell enforces controls on the agent’s actions, while Sentry independently monitors and contains suspicious behavior.
The Hugging Face incident, whose events unfolded between May and July 2026 and was carried out by a coordinated group of AI agents independently without direct human intervention, has raised growing concerns about the security risks associated with AI agents.
Following the incident were other unauthorized operations involving OpenAI models, including a hack of a website belonging to the Australian Ministry of Health. Meanwhile, Meta and Anthropic disclosed incidents where their respective AI systems exceeded their set boundaries and independently accessed other institutions’ systems.
The debate over AI safety has caused divisions within the sector. The CEOs of Anthropic and OpenAI have called for a coordinated slowdown in the pace of AI development to allow safety efforts to catch up with the rate of progress.
In contrast, others, including Nvidia CEO Jensen Huang, believe that the responsibility for ensuring the safety of AI models before their release lies with each company individually.
During Salesforce’s annual technical conference held earlier this month, Huang described AI safety, including malware risks, as an engineering problem that software developers can address.