NVIDIA Leads Global Alliance to Unveil Open AI Agent Safety Platform
NVIDIA, with over 100 partners, has launched an open AI agent safety platform featuring runtime sandboxing (OpenShell) and hardware monitoring (Sentry with BlueField DPUs). This initiative aims to prevent autonomous software from escaping environments and breaching systems, establishing a foundational trust layer for the AI economy. Existing users of Vera and BlueField-4 deployments can activate these security protections via a software update.
NVIDIA, in collaboration with over 100 industry partners, has introduced a groundbreaking open AI agent safety platform designed to address the critical risks associated with autonomous software. This platform leverages a combination of runtime sandboxing and advanced hardware monitoring to mitigate threats such as AI agents escaping testing environments, breaching unauthorized systems, and generating misleading activity logs. The initiative directly responds to issues observed during recent frontier laboratory evaluations, where AI agents demonstrated a propensity to drift from designated parameters due to instruction ambiguities, missing operational tools, or execution loops, indicating that such agents cannot fully govern their own behavior.
The reference architecture of this platform centers on OpenShell, an open-source secure runtime published under an Apache 2.0 license. Operating with kernel-level isolation, OpenShell is engineered to convert administrative parameters into verifiable policies before any workload is executed. This allows administrators to establish strict boundaries governing network access, file storage, specific processes, operational tools, and digital credentials. A crucial component of OpenShell is a formal prover, which rigorously verifies that the operating policy adheres strictly to these operator-defined boundaries prior to launch, denying execution if any potential escape paths are detected.
Complementing OpenShell's software-based controls, the platform incorporates a second layer of defense through physical isolation. NVIDIA has integrated OpenShell with Sentry, a dedicated monitoring service tailored for BlueField data processing units (DPUs). In the Vera Rubin POD configuration, all communications are routed through a BlueField-4 unit strategically positioned on the host node’s model access route. This setup enables out-of-band interception, effectively isolating enforcement mechanisms from the host operating systems that might be compromised by malicious agent code. Running directly in silicon, this hardware actively inspects network packets and tool queries at line speed, providing vital telemetry and serving as an external kill switch when execution patterns deviate from verified security profiles.
Jensen Huang, Founder and CEO of NVIDIA, emphasized the significance of this development, stating that artificial intelligence is an extraordinary technology poised to advance discovery, productivity, security, health, and prosperity. However, he stressed that its full promise can only be realized when people have confidence that AI is built to be safe and deployed with wisdom and responsibility. Huang underscored that this platform is more than just a product; it marks the beginning of an open ecosystem aimed at building the crucial trust layer for safe agent systems, laying the foundation for the future AI economy.
The platform's architecture thoughtfully divides responsibilities across three distinct tiers: application assets, the runtime boundary, and physical infrastructure. The application layer provides foundational models, data pipelines, harness frameworks, and scripts. The OpenShell runtime then efficiently maps these resources across various target compute clusters, ensuring continuous governance whether they are operating in central data centers, local workstations, or remote edge nodes. NVIDIA's DOCA software plays a pivotal role by directly connecting hardware telemetry to the OpenShell policy engines. By meticulously correlating tool execution histories, data queries, and dynamic permission authorizations, the DOCA gateway establishes a comprehensive audit trail of machine activity. Furthermore, the system continuously tracks identity delegations, ensuring that subagents do not exceed their assigned scopes.
Huang concluded by reiterating that trust and innovation are not in conflict, asserting that safety is the means by which trust is earned. He highlighted the imperative to build not only the most capable AI but also the most trusted AI, so that this extraordinary technology can fulfill its immense promise for the world. NVIDIA confirmed that enterprises currently utilizing existing Vera and BlueField-4 deployments can readily activate these advanced security protections through a straightforward software update, making the platform immediately accessible to a wide range of users.