间接提示注入(IPI)是一种不断演变的威胁向量,针对拥有多个数据源的复杂 AI 应用程序用户,例如 Workspace with Gemini。该技术使攻击者能够通过向 LLM 在响应用户查询时所使用的数据或工具中注入恶意指令,从而影响大语言模型的行为。甚至可能在没有任何来自用户的直接输入的情况下实现这一点。
IPI 并非那种可以“解决”后就置之不理的技术问题。随着智能体自动化使用的增加以及内容来源的多样化,先进的大语言模型为对抗性攻击提供了一个极度动态且不断演变的演练场。这就是为什么 Google 采取复杂且全面的方法来应对这些攻击。我们持续改进大语言模型对 IPI 攻击的抵抗力,并推出具有日益增强防御能力的 AI 应用程序功能。在间接提示注入攻击的最新态势中保持领先,对于我们实现 Workspace with Gemini 安全化的使命至关重要。
在上一篇博客《通过分层防御策略缓解提示注入攻击》中,我们回顾了 IPI 防御的分层架构。在本篇博客中,我们将分享更多关于我们如何持续改进这些防御措施并应对新攻击的细节。
新型攻击发现
通过内部和外部项目主动发现和编目新的攻击向量,我们可以识别漏洞并在对抗性活动发生之前部署强大的防御措施。
人工红队测试
人工红队测试利用对抗性模拟来揭示安全和功能漏洞。专业团队基于逼真的用户画像执行攻击以利用弱点,并与产品团队协调以解决已发现的问题。
自动化红队测试
自动化红队测试通过动态的、由机器学习驱动的基础设施来对环境进行压力测试。通过算法生成并迭代攻击载荷,我们可以大规模地模拟复杂威胁的行为。这使得我们能够绘制复杂的攻击路径,并在比人工测试本身所能实现的更广泛的边缘情况下验证我们安全控制的有效性。
Google AI 漏洞奖励计划(VRP)
Google AI 漏洞奖励计划(VRP)是促进 Google 与发现利用 IPI 的新攻击的外部安全研究人员之间协作的关键工具。通过该 VRP,我们认可并奖励贡献者的研究工作。我们还定期举办现场黑客松活动,为受邀研究人员提供访问预发布功能的权限,主动发现新型漏洞。这些合作伙伴关系使 Google 能够快速验证、复现和解决外部发现的问题。
公开披露的 AI 攻击
Google 利用开源情报源来掌握最新公开披露的 IPI 攻击,涵盖社交媒体、新闻稿、博客等。在此基础上,我们在内部对新的 AI 漏洞进行来源收集、复现和编目,以确保我们的产品不受影响。
漏洞目录
所有新发现的漏洞都要经过 Google 信任、安全与团队执行的全面分析流程。每个新漏洞都会被复现、检查是否重复、映射到攻击技术/影响类别,并分配给相关的所有者。新型攻击发现来源与漏洞目录流程的结合,帮助 Google 以可操作的方式掌握最新的攻击态势。
Indirect prompt injection (IPI) is an evolving threat vector targeting users of complex AI applications with multiple data sources, such as Workspace with Gemini. This technique enables the attacker to influence the behavior of an LLM by injecting malicious instructions into the data or tools used by the LLM as it completes the user’s query. This may even be possible without any input directly from the user.
IPI is not the kind of technical problem you “solve” and move on. Sophisticated LLMs with increasing use of agentic automation combined with a wide range of content create an ultra-dynamic and evolving playground for adversarial attacks. That’s why Google takes a sophisticated and comprehensive approach to these attacks. We’re continuously improving LLM resistance to IPI attacks and launching AI application capabilities with ever-improving defenses. Staying ahead of the latest indirect prompt injection attacks is critical to our mission of securing Workspace with Gemini.
In our previous blog “ Mitigating prompt injection attacks with a layered defense strategy ”, we reviewed the layered architecture of our IPI defenses. In this blog, we’ll share more detail on the continuous approach we take to improve these defenses and to solve for new attacks.
New attack discovery
By proactively discovering and cataloging new attack vectors through internal and external programs, we can identify vulnerabilities and deploy robust defenses ahead of adversarial activity.
Human Red-Teaming
Human Red-Teaming uses adversarial simulations to uncover security and safety vulnerabilities. Specialized teams execute attacks based on realistic user profiles to exploit weaknesses, coordinating with product teams to resolve identified issues.
Automated Red-Teaming
Automated Red-Teaming is done via dynamic, machine-learning-driven frameworks to stress-test environments. By algorithmically generating and iterating on attack payloads, we can mimic the behavior of sophisticated threats at scale. This allows us to map complex attack paths and validate the effectiveness of our security controls across a much wider range of edge cases than manual testing could achieve on its own.
Google AI Vulnerability Rewards Program (VRP)
The Google AI Vulnerability Rewards Program (VRP) is a critical tool for enabling collaboration between Google and external security researchers who discover new attacks leveraging IPI. Through this VRP, we recognize and reward contributors for their research. We also host regular, live hacking events where we provide invited researchers access to pre-release features, proactively uncovering novel vulnerabilities. These partnerships enable Google to quickly validate, reproduce, and resolve externally-discovered issues.
Publicly disclosed AI attacks
Google utilizes open-source intelligence feeds to stay on top of the latest publicly disclosed IPI attacks, across social media, press releases, blogs, and more. From there, new AI vulnerabilities are sourced, reproduced, and catalogued internally to ensure our products are not impacted.
Vulnerability catalog
All newly discovered vulnerabilities go through a comprehensive analysis process performed by the Google Trust, Security, & Safety teams. Each new vulnerability is reproduced, checked for duplications, mapped into attack technique / impact category, and assigned to relevant owners. The combination of new attack discovery sources and vulnerability catalog process helps Google stay on top of the latest attacks in an actionable manner.
首次收录 · 2026-09-24 · 6.4 分