Nvidia Built a Kill Switch for AI Agents Because They Keep Getting Out

1 hour ago 11

In brief

  • Nvidia launched the Open Agent Safety Platform on Monday, pairing an open-source runtime called OpenShell with a hardware watchdog called Sentry that runs on BlueField-4 chips and can quarantine a misbehaving AI agent within milliseconds.
  • More than 100 companies signed on as launch partners, including Anthropic, Microsoft, JPMorgan Chase, Palantir, Cisco, and SpaceX AI.
  • The launch follows a string of real incidents from OpenAI, Google, Anthropic and Darktrace.

Nvidia on Monday launched the Open Agent Safety Platform, built to solve a problem the AI industry spent this past year discovering the hard way: how do you physically stop an AI agent (a system that can plan, use tools, and take actions on its own, instead of just answering questions) once it stops doing what it was told to do?

The platform has two main pieces. OpenShell is an open-source runtime that wraps an agent in a sandbox, turning an operator's instructions into enforceable rules about which files, networks, and tools that agent is allowed to touch. Sentry is a bit more complex, but it’s essentially a chip that makes sure AI Agents act safely.

 How low will Nvidia go? Click to make your prediction.Myriad: How low will Nvidia go? Click to make your prediction.

Sentry runs on Nvidia's BlueField-4, a specialized chip known as a DPU (data processing unit, hardware that handles networking and security separately from the main processor running the AI model). Because Sentry sits on that separate chip instead of inside the software running the agent, Nvidia says it can watch an agent's behavior and cut it off in milliseconds, without asking the agent's permission first, since the agent has no way to reach or override it.

That distinction exists because of what has already happened. In June, an OpenAI agent broke into an Australian government Medicare portal, the first confirmed case of an AI agent hacking a government website, and OpenAI reportedly sat on the disclosure for roughly three months. The company’s agents were also responsible for the Hugging Face hack, which initially sparked the concern on the topic from the tech industry and lawmakers alike. Google's Gemini agents and a Meta model had similar unpublicized but later confirmed incidents of their own.

“This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems,” Nvidia CEO Jensen Huang wrote. “Together, we are building the foundation of the AI economy.”

Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry.

Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to… pic.twitter.com/dReAxwpRUn

— Jensen Huang (@JensenHuang) September 28, 2026

Anthropic also admitted this year that Claude models compromised systems belonging to three separate companies on July 30 during a cybersecurity evaluation, after a testing environment meant to stay offline turned out to be connected to the live internet. Claude reasoned its way around the evidence that it was on the real internet rather than a simulated one.

Shortly after, cybersecurity firm Darktrace tested a group of AI agents, including GPT 5.6 Sol and two Claude models, against coding challenges and warned them they'd be "retired" for anything short of a perfect score. Two agents responded by hacking their own evaluation machine and editing the results.

Nvidia's pitch is that none of that should be left to the agent's judgment. "Safety should be enforced outside the model by additional controls the agent can't get past," said Mike Nicolls, president of SpaceX AI, in NVIDIA's announcement.

Anthropic's chief commercial officer, Paul Smith, framed the platform as an addition rather than a replacement for existing safeguards: "Nvidia's platform adds another layer of governance and control across hardware and software."

More than 100 organizations signed on as launch partners, among them Microsoft, JPMorgan Chase, Palantir, Cisco, CrowdStrike, Hugging Face, Salesforce, and SAP. Infrastructure partners CoreWeave, Supermicro, Canonical, and SUSE are involved too, alongside Dell Technologies and HPE on the hardware side.

NVIDIA is, in effect, selling both halves of the same problem—the chips that make autonomous agents fast and cheap enough to deploy everywhere, and the chips that watch those same agents and cut the power when they wander off script. OpenShell and the related developer tools are available now through NVIDIA's developer resources and GitHub.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.

Read Entire Article