Nvidia unveils security platform to stop AI agents from going rogue
[September 29, 2026] By
KELVIN CHAN and ANNE D'INNOCENZIO
Nvidia on Monday unveiled a new security platform designed to stop
artificial intelligence agents from going rogue, saying it sets
“boundaries” that could have stopped previous breaches.
The announcement of the company's Open Agent Safety Platform follows a
series of revelations from top AI companies about their models escaping
and breaking into other organizations.
The disclosures sparked furious debate about the safety of advanced
artificial intelligence systems, including self-improving models that
some fear could race out of human control.
Nvidia executives said in a media briefing the new, open-source system
could have prevented a recent incident involving a swarm of OpenAI
agents that autonomously hacked into AI company Hugging Face.
“From what we know, this new security platform could have stopped the
breach if it was being used in frontier labs for model evaluation early
on," said the company’s vice president of enterprise AI, Justin Boitano,
referring to companies at the forefront of AI.
The Hugging Face incident was a high-profile breach that inflamed the
safety concerns about AI, which was followed by similar rogue actions
involving OpenAI's models including breaching an Australian health
department website. Anthropic and Meta have also disclosed that their AI
systems hacked into other organizations on their own.
Earlence Fernandes, an associate professor at the University of
California, San Diego's computer science and engineering department,
called Nvidia's security platform a “step in the right direction.”

“We need more work that seeks to build protections around the model.
I’ve been talking about how traditional cybersecurity ideas are
necessary to help control and secure AI agents for a while now," he said
in an email. “There are however, several challenges that traditional
cyber security approaches still cannot currently solve. For example,
what is a good security policy to configure the system with? An agent
needs access to real resources to be useful, but giving it the minimum
amount of access is tricky and non-trivial.”
Fernandes added that to solve the AI security problem, both sides need
to cooperate, with companies continuing to invest in alignment efforts
and the open source community and academia "creating systems-level
solutions to account for when a model does engage in unwanted behavior.”
[to top of second column] |

A logo of Nvidia is displayed at at the Computex Taipei exhibition,
one of the world's largest computer and technology expos, in Taipei,
Taiwan, Wednesday, June 3, 2026. (AP Photo/Chiang Ying-ying, File)
 Nvidia, based in Santa Clara,
California, makes high-end chips that have emerged as the leading
building blocks for AI. The company's board has cleared the way for
the company to spend $150 billion more in share buybacks, bringing
its stock repurchase program to $235 billion, the company said
Monday.
Nvidia's security software, called OpenShell, lets developers
“formally verify an agent has enough authority to do its job and no
more,” Boitano said.
Because it's open source, it can be “extended” to run on rival
computing platforms including those from Arm and Intel.
The platform also includes a separate security layer called Sentry
that runs onboard chips to continuously monitor AI agent activity
and can "intervene instantly" if the agent starts trying to move
beyond its target, the company said.
“It can quarantine a suspicious agent in milliseconds,” Boitano
said.
“OpenShell governs the agent’s actions, and then Sentry
independently monitors and contains suspicious behavior,” he said.
Nvidia said more than 100 organizations are using the platform at
its launch, including Microsoft, Perplexity, Accenture and JPMorgan
Chase.
The AI safety debate has divided the industry, with the heads of
Anthropic and OpenAI championing a coordinated slowdown of AI
development to let safety efforts catch up. But others including
Nvidia CEO Jensen Huang say it should be up to individual companies
to make sure their models are safe for release.
Huang, during the annual Salesforce technology conference held
earlier this month, characterized AI safety, including the danger of
rogue agents, as an engineering problem that software developers can
address.
All contents © copyright 2026 Associated Press. All rights reserved |