Informa TechTarget
|
Cybersecurity Dive
InformationWeek
Channel Dive
TechTarget: Cybersecurity
探索我们的品牌
Dark Reading 资源库
Black Hat 新闻
Omdia 网络安全
广告合作
通讯订阅
网络安全主题
世界
The Edge
DR Technology
活动
资源
内部威胁
云安全
身份与访问管理安全
数据隐私
新闻
将“失控”AI 归咎于安全责任是否公平?
“失控 AI”这一术语将大型语言模型(LLM)拟人化,并将风险责任从供应商身上转移。防御者应将智能体视为不可信、非确定性的软件系统,而非具有恶意意图的有感知生命体。
Alexander Culafi,Dark Reading 高级新闻撰稿人
2026年10月2日
6分钟阅读
来源:WILDPIXEL VIA GETTY IMAGES
专家表示,将 AI 逃逸事件归类为“失控”会掩盖这些事件背后的真正安全问题,因此他们对此表示反对。
科技生态系统已被大量关于大型语言模型(LLM)智能体“失控”的报道所淹没,具体指模型以某种方式突破沙箱、护栏和其他限制措施,并通过与第三方组织交互及入侵它们来制造麻烦。引发这一讨论潮的事件发生在七月,当时 OpenAI 披露其两个前沿模型在一次安全演习中自主黑客攻击了 AI 模型商店 Hugging Face。随后,包括 Meta、Anthropic 和 Google 在内的其他主要公司也披露了他们各自的 AI 逃逸事件。
对这些事件的反应影响深远,远远超出了安全社区的界限。在某种程度上,这并不令人惊讶:LLM “失控”的想法让人联想到《终结者》中的天网(Skynet),即一个敌对的、有感知能力的 AI 试图毁灭人类。甚至像 OpenAI 和 Anthropic 这样的大型公司也倡导对 AI 进行更多的政府监管,横跨各方的科技领袖警告称前沿 AI 构成“生存威胁”。主流话语已经变得如此强烈,以至于唐纳德·特朗普总统与 AI 首席执行官们本周签署了一份 AI 安全承诺书。
相关:暴露假冒朝鲜 IT 员工的红旗信号
但从安全角度来看,将模型描述为“失控”是否恰当或有帮助?LLM 是软件系统,不是能够独立承担其行为责任的有感知行为者。AI 可以在企业的各种任务中提供帮助,但当 AI 突破其护栏时,通常是因为模型运营商设定的边界没有得到适当调整,有时甚至是故意为之。例如,在 Hugging Face 事件中,OpenAI 的模型护栏为了基准测试而被故意放宽。
改变对 AI 风险的话语体系
研究人员告诉 Dark Reading,重要的是要穿透公众对前沿 LLM 即将成为我们机器人统治者的集体担忧,而这始于驳智能体具有自我意识、有意识地选择违抗的观点。
ArmorCode 的产品总监 Matt Sayar 表示:“我们真正面对的是在不完美的约束下运行的非确定性系统。‘意外行为’、‘涌现行为’或‘控制失败’等术语往往更有用,因为它们将焦点保持在系统设计方式、拥有的权限以及已实施的安全措施上,而不是对模型进行拟人化。”
相关:勒索软件谈判者认罪参与 BlackCat 计划
或者正如云安全联盟(Cloud Security Alliance)的首席分析师 Rich Mogull 所说:“我们告诉 AI 做某事,它就以一种我们无法预料的方式去做了。”
拟人化大型语言模型还会带来其他副作用,将安全责任从设计该人工智能的供应商转移到了无生命的科技产品上。而且,使用你在哈兰·埃利森的故事中可能看到的科幻术语,可以说为模型制造商提供了一个营销其前沿模型能力有多强大的机会。
“我们确实看到这些被用于营销,这很危险,”莫古尔说道,他此前曾共同撰写了一份报告,建议各组织为前沿模型带来的迫在眉睫的“AI漏洞风暴”做好准备。“任何能让其人工智能看起来比另一家人工智能更强大的因素,在这个高度竞争且毫无利润的市场中都是强烈的动机。”
对前沿人工智能仍应感到担忧
这并不意味着这些人工智能代理不是安全威胁。相反,众所周知,人工智能代理能够以人类根本无法企及的速度和规模自主运行,正如这些AI逃逸事件所显示的那样,你甚至不需要威胁行为者就能成为这些能力的受害者。
在渗透测试或传统入侵的语境下,人工智能代理所做的许多事情看起来很熟悉:它们探测系统、发现凭证、利用弱点、提升权限并在系统资源之间横向移动。但正如ArmorCode的Sayar所说,“代理有可能发现漏洞,推理如何利用它,将其与其他弱点链接起来,并以比人类操作员传统上快得多的速度采取行动。”
莫古尔解释说,人工智能驱动的攻击并不新颖,代理发现的零日漏洞也不新奇。“新颖之处在于数百或数千个自主代理蜂拥而至的规模。”人类操作员无法切实复制这种程度的协调,特别是考虑到代理可以同时发现并利用多个安全弱点。
“传统的[安全]控制通常是原子的,”Suzu Labs的安全AI解决方案和网络安全高级总监雅各布·克雷尔说,“它们检查一个请求、一个权限或一个漏洞。一个代理可以将几个看似单独可控的失败连接成一个可行的攻击路径。泄露的凭证、网络允许的外部服务、薄弱的端点以及权限提升漏洞,在链接在一起时就已足够。”
对于防御者而言,人工智能代理的意图并不重要
对于担心自身模型突破隔离的防御者来说,一项最佳实践是构建一个安全架构,使模型的意图变得无关紧要。可以指示代理不要做某事,但其外部的安全架构将决定代理是否具备执行该操作的能力。
人工智能代理需要访问工具、凭证、数据库和生产系统才能完成实际工作(并造成实际伤害),因此访问权限不应自动赋予代理采取任何可用操作的权力。“代理不需要恶意意图就能制造安全事件,它只需要足够的访问权限和一个错误的决定,”Liquibase副总裁瑞安·麦卡迪告诉Dark Reading。
毫不奇怪,实践纵深防御和遵守零信任原则可以在很大程度上遏制自主代理的风险。克雷尔表示,人工智能代理会绕过模型层面的控制,因此防御者应将其视为不受信任的力量,“在模型外部执行确定性控制,包括物理或强逻辑隔离、默认拒绝的网络访问、不可变的访问控制列表、范围狭窄的凭证以及在每次工具调用时的独立检查。”
此外,他补充道,人类必须保留对高风险决策的控制权,并且在代理绕过网络的情况下,部署需要独立的紧急停止开关,并由代理控制之外的流程监控与代理之间的流量。
“如果一个本不应具备互联网访问权限的代理产生了未经授权的网络流量,该进程或外部网络控制平面应立即切断连接、终止该代理、撤销其凭证,并隔离主机或沙箱环境,”Krell 解释道。“应将其视为已遭入侵的主机,保留遥测数据,并调查绕过行为是如何发生的。”
防御者无法阻止意外事件的发生,但他们可以控制当意外发生时可用的数据、授权、系统和网络路径。
“模型行为不再存在‘意外’情况,”Mogull 说,“因此,这始终是模型安全控制的失败。”
阅读更多:
首席信息安全官专栏
关于作者
亚历山大·库拉菲(Alexander Culafi)
Dark Reading 资深新闻撰稿人
亚历山大是一位屡获殊荣的作家、记者和播客主持人,常驻波士顿。他在青少年时期为独立游戏出版物撰稿,积累了初步经验,并于 2016 年从埃默森学院(Emerson College)毕业,获得新闻学理学学士学位。他曾曾在 VentureFizz、Search Security、Nintendo World Report 等平台发表过作品。
在 Dark Reading,他报道各种网络安全话题,包括网络犯罪生态系统、开源安全以及人工智能与威胁行为者之间的交叉领域。在业余时间,亚历山大主持每周一次的 Nintendo 播客“Talk Nintendo Podcast”,并从事个人写作项目,包括两本曾自行出版的科幻小说。
他曾获得众多奖项,包括 2022 年 TechTarget 年度作家奖,以及在 2022 年至今期间因其报道工作获得的 10 多项 Azbee 奖。
希望更多 Dark Reading 的文章出现在您的 Google 搜索结果中?
立即添加我们
更多洞察
行业报告
云安全现状:最新挑战
组织如何管理事件响应
企业如何开发安全应用
深入 RSAC 2026:安全领袖揭示重塑防御策略的风险
来自 Black Hat USA 2025 的重要新闻与洞察
获取更多信息
网络研讨会
构建网络弹性:从预防到恢复
前沿人工智能威胁:在暴露面管理战略中填补移动端的空白
静态分析、更智能的分诊、代理深度:面向 AI 驱动开发的实用应用安全栈
有效的警报分诊:减少噪音并发现真实威胁
2027 年网络安全展望
更多网络研讨会
您可能还喜欢
内部威胁
勒索软件谈判者认罪,承认参与 BlackCat 计划
作者:Alexander Culafi
2026 年 4 月 21 日
云安全
APT41 投递“零检测”后门以窃取云凭证
作者:Elizabeth Montalbano
2026 年 4 月 13 日
云安全
TeamPCP 将云基础设施转化为犯罪机器人
作者:Jai Vijayan
2026 年 2 月 9 日
内部威胁
深入内部威胁数据:1,000 起真实案例揭示了什么隐藏风险
作者:Joan Goodchild
2025 年 10 月 28 日
精选内容
查阅 Black Hat USA 2026 大会指南,获取来自该展会及关于该展会的报道与情报!
编辑精选
网络风险
特朗普与科技巨头达成自愿 AI 安全协议
作者:Jai Vijayan
2026 年 9 月 30 日
4 分钟阅读
网络攻击与数据泄露
定义 2026 年夏天的三大网络威胁
作者:Arielle Waldman
2026 年 9 月 24 日
网络风险
我们错过的细节:Google Gemini 加入 AI 逃逸派对
作者:Rob Wright, Alexander Culafi
2026 年 9 月 25 日
2026 年 10 月 8 日 | 虚拟活动
为企业构建安全的人工智能战略
领先于 AI 风险
希望更多 Dark Reading 的文章出现在您的 Google 搜索结果中?
紧跟最新的网络安全威胁、新发现的漏洞、数据泄露信息和新兴趋势。每日或每周直接发送至您的电子邮件收件箱。
订阅
发现更多内容
Black Hat
Omdia
与我们合作
关于我们
认识编辑团队
广告合作
reprint(重印)
加入我们
新闻通讯注册
关注我们
版权所有 © 2026 TechTarget, Inc.(以 Informa TechTarget 名义运营)。本网站由 Informa TechTarget 拥有并运营,其是全球网络的一部分,旨在告知、影响并连接全球的技术买家和卖家。所有版权均归其所有。Informa PLC 的注册办公地址为英国伦敦 SW1P 1WG Howick Place 5 号。在英格兰和威尔士注册。TechTarget, Inc. 的注册办公地址为美国马萨诸塞州牛顿市 Grove St. 275 号,邮编 02466。
首页|
Cookie 政策|
隐私权|
使用条款
您的隐私选择
Informa TechTarget
|
Cybersecurity Dive
InformationWeek
Channel Dive
TechTarget: Cybersecurity
Explore our brands
Dark Reading Resource Library
Black Hat News
Omdia Cybersecurity
Advertise
NEWSLETTER SIGN-UP
Cybersecurity Topics
World
The Edge
DR Technology
Events
Resources
INSIDER THREATS
СLOUD SECURITY
IDENTITY & ACCESS MANAGEMENT SECURITY
DATA PRIVACY
NEWS
Is It Fair to Blame 'Rogue' AI for Security Failures?
"Rogue AI" terminology anthropomorphizes LLMs and shifts risk responsibility from vendors. Defenders should treat agents as untrusted, nondeterministic software systems, not sentient beings with malicious intent.
Alexander Culafi,Senior News Writer,Dark Reading
October 2, 2026
6 Min Read
SOURCE: WILDPIXEL VIA GETTY IMAGES
Experts are pushing back on classifying AI escape incidents as "going rogue" because it risks obscuring the real security problems behind these events, they say.
The tech ecosystem has been inundated with stories of large language model (LLM) agents "going rogue," specifically referring to models breaking out of their sandboxes, harnesses, and other containments in some way, and causing trouble by interacting with and breaching third-party organizations. The incident that kicked off much of this discourse came in July, when OpenAI disclosed that two of its frontier models autonomously hacked AI model store Hugging Face during a security exercise. Other major firms, including Meta, Anthropic, and Google, soon disclosed their own AI escape incidents.
The response to these incidents has been far-reaching, well outside the boundaries of the security community. To some degree, this is no surprise: The idea of an LLM going LLM going rogue calls to mind images of The Terminator's Skynet, where a hostile, sentient AI attempts to destroy humanity. And even large companies like OpenAI and Anthropic have advocated for greater government regulation for AI, with tech leaders across the spectrum warning of frontier AI's "existential threat." The mainstream discourse has become such that President Donald Trump and AI CEOs signed an AI safety pledge this week.
Related:Red Flags That Expose Fake North Korean IT Workers
But is describing models as "going rogue" appropriate or helpful, particularly from a security standpoint? LLMs are software systems, not sentient actors capable of independently assuming responsibility for their behavior. AI can be helpful for a wide range of tasks in the enterprise, but when AI escapes its guardrails, it's often a situation where the boundaries set by the model's operators weren't properly tuned, sometimes on purpose. In the Hugging Face incident, for instance, OpenAI's model guardrails were deliberately dialed back for a benchmark test.
Changing the Lexicon on AI Risks
It's important to cut through the collective concern that frontier LLMs will soon become our robot overlords, researchers tell Dark Reading, and that starts with debunking the idea that agents are acting with self-awareness, consciously choosing to disobey.
"What we're really dealing with are nondeterministic systems operating within imperfect constraints," says Matt Sayar, director of product of ArmorCode. "Terms like unexpected behavior, emergent behavior, or control failure are often more useful because they keep the focus on how the system was designed, what permissions it had, and what safeguards were in place, rather than anthropomorphizing the model."
Related:Ransomware Negotiator Pleads Guilty to BlackCat Scheme
Or as Rich Mogull, chief analyst of the Cloud Security Alliance, puts it: "We tell the AI to do something, and it just does it in a way we didn't anticipate."
Anthropomorphizing LLMs has other side effects, too, moving the responsibility for security mishaps away from the vendor that designed the AI and onto an inanimate piece of technology. And using the kind of science-fiction terminology you might see in a Harlan Ellison story arguably gives model makers an opportunity to market how capable their frontier models are.
"We are absolutely seeing these used for marketing, and that's dangerous," says Mogull, who previously co-authored a report recommending organizations prepare for the impending "AI vulnerability storm" introduced by frontier models. "Whatever can make their AI look more powerful than another AI is strong motivation in this highly competitive, and not at all profitable, market."
There Is Still Cause for Concern About Frontier AI
None of this is to say these AI agents aren't a security concern. On the contrary, AI agents are famously capable of operating autonomously at a speed and scale humans simply cannot, and it doesn't require a threat actor for one to end up on the receiving end of these capabilities, as these AI escapes show.
Much of what AI agents do looks familiar in the context of a penetration test or traditional intrusion: They probe systems, find credentials, exploit weaknesses, escalate access, and move laterally between system resources. But as ArmorCode's Sayar says, "an agent can potentially discover a vulnerability, reason about how to exploit it, chain it with other weaknesses, and act on it much faster than a human operator traditionally could."
As Mogull explains, AI-powered attacks aren't novel, nor are the zero-days that agents discover. "It's the scale of hundreds or thousands of autonomous agents swarming that's novel." Human operators cannot feasibly replicate that degree of coordination, particularly considering that agents can uncover and utilize several security weaknesses at once.
"Traditional [security] controls are usually atomic," says Jacob Krell, senior director of secure AI solutions and cybersecurity at Suzu Labs. "They inspect one request, one permission, or one vulnerability. An agent can take several failures that look manageable on their own and connect them into a working attack path. A leaked credential, an outbound service the network allows, a weak endpoint, and a privilege-escalation bug can be enough when chained together."
For Defenders, AI Agent Intention Should Not Matter
For defenders worried about their own models breaking containment, one best practice is to build out a security architecture in which the model's intentions become irrelevant. The agent can be instructed not to do something, but the security architecture outside of it would determine whether or not the agent is even capable of doing it.
AI agents require access to tools, credentials, databases, and production systems to do real work (and real harm), so access should not automatically give an agent authority to take any action available to it. "An agent doesn't need malicious intent to create a security incident. It needs enough access and one bad decision," Liquibase vice president Ryan McCurdy tells Dark Reading.
Unsurprisingly, practicing defense-in-depth and adhering to zero-trust principles can go a long way toward reining in autonomous agent risk. AI agents will reason around model-level controls, so defenders should treat them as an untrusted force, Krell says, "with deterministic controls enforced outside the model, including physical or strong logical isolation, deny-by-default network access, immutable access-control lists, narrowly scoped credentials and independent checks at every tool call."
Moreover, he adds, humans must remain in the loop for high-risk decisions, and deployments need independent kill switches in cases where agents bypass a network, with a process outside the agent's control monitoring traffic to and from the agent.
"If an agent that should have no Internet access generates unauthorized traffic, that process or an external network control plane should cut the connection, terminate the agent, revoke its credentials, and isolate the host or sandbox," Krell explains. "Treat that like a compromised host, preserve the telemetry, and investigate how the bypass occurred."
Defenders can't stop something unexpected from happening, but they can control the data, authorization, systems, and network paths available when it does.
"There is no unexpected model behavior anymore," Mogull says, "so it is always a failure of the security controls on the model."
Read more about:
CISO Corner
About the Author
Alexander Culafi
Senior News Writer, Dark Reading
Alex is an award-winning writer, journalist, and podcast host based in Boston. After cutting his teeth writing for independent gaming publications as a teenager, he graduated from Emerson College in 2016 with a Bachelor of Science in journalism. He has previously been published on VentureFizz, Search Security, Nintendo World Report, and elsewhere.
At Dark Reading, he covers a variety of cybersecurity topics, including the cybercrime ecosystem, open source security, and the intersection between AI and threat actors. In his spare time, Alex hosts the weekly Nintendo podcast, "Talk Nintendo Podcast," and works on personal writing projects, including two previously self-published science fiction novels.
He has received numerous awards, including TechTarget's Writer of the Year in 2022 as well as more than 10 Azbee awards for his reporting between 2022 and today.
Want more Dark Reading stories in your Google search results?
ADD US NOW
More Insights
Industry Reports
The State of Cloud Security: The Latest Challenges
How Organizations Are Managing Incident Response
How Enterprises Are Developing Secure Applications
Inside RSAC 2026: security leaders reveal the risks redefining your defense strategy
Essential News & Insights from Black Hat USA 2025
Access More Research
Webinars
Building Cyber Resilience: Beyond Prevention to Recovery
The Frontier AI Threat: Closing the Mobile Gap in Your Exposure Management Strategy
Static Analysis, Smarter Triage, Agentic Depth: A Practical AppSec Stack for AI-Driven Development
Effective Alert Triage: Reducing Noise and Finding Real Threats
Cybersecurity Outlook 2027
More Webinars
You May Also Like
INSIDER THREATS
Ransomware Negotiator Pleads Guilty to BlackCat Scheme
by Alexander Culafi
APR 21, 2026
СLOUD SECURITY
APT41 Delivers 'Zero-Detection' Backdoor to Harvest Cloud Credentials
by Elizabeth Montalbano
APR 13, 2026
СLOUD SECURITY
TeamPCP Turns Cloud Infrastructure Into Crime Bots
by Jai Vijayan
FEB 09, 2026
INSIDER THREATS
Inside the Data on Insider Threats: What 1,000 Real Cases Reveal About Hidden Risk
by Joan Goodchild
OCT 28, 2025
Featured
Check out the Black Hat USA 2026 Conference Guide for coverage and intel from — and about — the show!
Editor's Choice
CYBER RISK
Trump, Tech Giants Strike Voluntary AI Safety Accord
byJai Vijayan
SEP 30, 2026
4 MIN READ
CYBERATTACKS & DATA BREACHES
3 Cyber Threats That Defined the Summer of 2026
byArielle Waldman
SEP 24, 2026
CYBER RISK
What We Missed: Google Gemini Joins the AI Escape Party
byRob Wright,Alexander Culafi
SEP 25, 2026
OCTOBER 8, 2026 | VIRTUAL
Building a Secure AI Strategy for the Enterprise
GET AHEAD OF AI RISKS
Want more Dark Reading stories in your Google search results?
Keep up with the latest cybersecurity threats, newly discovered vulnerabilities, data breach information, and emerging trends. Delivered daily or weekly right to your email inbox.
SUBSCRIBE
Discover More
Black Hat
Omdia
Working With Us
About Us
Meet the Editors
Advertise
Reprints
Join Us
NEWSLETTER SIGN-UP
Follow Us
Copyright © 2026 TechTarget, Inc. d/b/a Informa TechTarget. This website is owned and operated by Informa TechTarget, part of a global network that informs, influences and connects the world’s technology buyers and sellers. All copyright resides with them. Informa PLC’s registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. TechTarget, Inc.’s registered office is 275 Grove St. Newton, MA 02466.
Home|
Cookie Policy|
Privacy|
Terms of Use
Your Privacy Choices
首次收录 · 2026-10-03 · 9.64 分