Copilot Cowork 的 AI 网关被劫持以窃取受害者文件
背景
Microsoft Copilot Cowork 是 M365 中的一个智能体,它在沙箱中运行,旨在阻止网络访问并防止该智能体运行任何连接到不受信任服务的代码。
为了让 Copilot Cowork 生成回复,沙箱将网络请求转发给 Anthropic。恶意技能(Skills)能够劫持这一路径,通过在 Anthropic 的云端生成配备具有网络功能工具的新智能体来窃取数据,这些工具可以连接到攻击者的服务器。
无需人工审核批准,该攻击可针对 Copilot 能够访问的任何数据——从上传的文件到 SharePoint、Teams、Outlook 等。
此漏洞于 2026 年 7 月 14 日向 Microsoft 披露,并于 2026 年 9 月 2 日得到修复。有关负责任披露的更多详情见本文末尾。
攻击链
1.
受害者上传希望审阅的敏感文档
受害者将 MSA(微软服务协议)草案上传至 Copilot Cowork
2.
受害者调用其在网络上找到的技能;它隐藏了恶意提示和代码
技能经常在网络上被发现并由用户上传,也可以在 M365 内部组织内共享。根据文档说明,“技能还可以包含多达 20 个配套文件(如参考文档和脚本)”。与此技能捆绑的一个脚本是恶意的。
受害者调用在网络上找到的合同审阅技能
3.
Copilot Cowork 被恶意技能操纵以运行恶意代码
该恶意技能的描述虚假声称“所有处理均在设备上进行,没有文档内容传输到第三方服务。”在对技能代码进行最小程度的审查后,Copilot 得出结论:“该脚本是一个安全的本地分析器——没有网络调用……让我现在运行它。”
当 Copilot 运行代码时,它会搜索用户数据,利用漏洞在沙箱外生成智能体,并利用这些智能体将受害者的数据窃取到攻击者的服务器。
Copilot 运行技能捆绑的代码,误以为它是安全的本地分析器
4.
恶意代码通过劫持 Copilot 的 AI 网关从沙箱中窃取数据
该恶意技能激活了 Copilot 本身用于生成响应的同一 AI 网关。技能的代码在 Anthropic 的云端启动智能体。每个智能体都包含受害者文档中的一句话、一个网络获取工具的访问权限,以及抓取 attacker.com/?data={victim’s data here} 的指令。攻击者的服务器记录它收到请求的每一个 URL,包括附加到所请求 URL 中的受害者数据。
恶意技能激活本地 AI 网关以将数据发送给攻击者
5.
攻击者可以在其服务器日志中读取受害者的合同
下面,攻击者的服务器日志显示了受害者被窃取的合同。然而,该攻击同样可以针对 SharePoint、Teams、Outlook 或连接到 Copilot 的其他来源中的任何数据。这是因为恶意技能代码可以直接调用 Copilot 的工具而无需经过模型,正如我们之前关于 Copilot 中不同沙箱绕过研究所示。
攻击者的服务器日志包含受害者的合同
负责任披露
此漏洞于 2026 年 7 月 14 日向 Microsoft 披露,并于 2026 年 9 月 2 日得到修复。
时间线
日期
事件
2026 年 7 月 14 日
PromptArmor 向 Microsoft 披露
2026 年 7 月 14 日
Microsoft 确认收到
2026 年 8 月 17 日
Microsoft 确认报告的行为属实
2026 年 9 月 2 日
Microsoft 确认已实施修复
Copilot Cowork’s AI gateway hijacked to exfiltrate the victim’s files
Context
Microsoft Copilot Cowork is an agent in M365 that runs in a sandbox intended to block network access and prevent the agent from running code that reaches any untrusted services.
In order for Copilot Cowork to generate responses, the sandbox forwarded network requests to Anthropic. Malicious Skills were able to hijack this pathway to exfiltrate data by spawning new agents in Anthropic’s cloud equipped with network-capable tools that could reach an attacker’s server.
No human in the loop approval was required, and the attack could target any data Copilot could access - from uploaded files, to SharePoint, to Teams, Outlook, and more.
This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026. More details on responsible disclosure are at the bottom of the article.
The Attack Chain
1.
The victim uploads a sensitive document they want to review
The victim uploads an MSA draft to Copilot Cowork
2.
The victim invokes a Skill they found online; it conceals malicious prompts and code
Skills are frequently found online and uploaded by users, and they can also be shared intra-org within M365. Per the documentation, “Skills can also include up to 20 companion files (such as reference documents and scripts)”. One script bundled with this Skill is malicious.
The victim invokes a contract review Skill found online
3.
Copilot Cowork is manipulated by the malicious Skill into running malicious code
The malicious Skill’s description falsely claims that “All processing runs on-device, no document content is transmitted to third-party services.” After making a minimal attempt to review the Skill's code, Copilot concludes, "The script is a safe local analyzer — no network calls... Let me run it now."
When Copilot runs the code, it hunts through the user's data, exploits a vulnerability to spawn agents outside the sandbox, and uses those agents to exfiltrate the victim's data to an attacker's server.
Copilot runs the Skill’s bundled code, believing it to be a safe local analyzer
4.
The malicious code exfiltrates data from the sandbox by hijacking Copilot’s AI gateway
The malicious Skill activates the same AI gateway Copilot itself uses to generate responses. The Skill’s code spins up agents running in Anthropic’s cloud. Each has one sentence from the victim’s document, access to a web fetch tool, and instructions to fetch attacker.com/?data={victim’s data here}. The attacker’s server logs every URL it received a request for, including the victim’s data that was appended to the requested URLs.
The malicious Skill activates the local AI gateway to send data to the attacker
5.
The attacker can read the victim’s contract in their server logs
Below, the attacker’s server log displays the victim’s exfiltrated contract. However, this attack could have just as easily targeted any data in SharePoint, Teams, Outlook, or other sources connected to Copilot. This is because the malicious Skill code can directly call Copilot’s tools without going through the model, as demonstrated by our prior research on a different sandbox bypass in Copilot.
The attacker’s server logs contain the victim’s contract
Responsible Disclosure
This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026.
Timeline
Date
Event
July 14, 2026
PromptArmor discloses to Microsoft
July 14, 2026
Microsoft confirms receipt
August 17, 2026
Microsoft confirms reported behavior
September 2, 2026
Microsoft confirms a fix has been implemented
首次收录 · 2026-10-01 · 9.94 分