自主安全代理在发现漏洞方面表现得越来越出色。然而,目前并没有一种有效的方法来衡量其表现究竟有多好。将其指向一个真实的测试目标后,返回的是一份由代理自行撰写的报告:其中充斥着自信的文字、一份发现列表,但你根本无法分辨其中的哪些发现是真实发生的。随后,一名具备安全背景的人员需要坐下来,针对目标逐一核实每一项声明。哪些发现是真实的?哪些是重复的?哪些是凭空捏造的?以及那个无人有时间去回答的问题:代理从未尝试过什么?对于单次运行而言,这是一项需要专家耗费一天时间的工作。如果将其乘以三个模型、四种提示词变体和十次重复运行,审查队列的长度将远超实验本身。
由 CTF.ae 构建的 XRanges for AI 正是为了解决这一循环问题而存在的。它部署了带有内置仪器化(instrumentation)功能的真实目标应用程序,记录代理在其中的实际操作,并基于四个独立的信号对每次运行进行实时评分。本文 walkthrough 将介绍其工作原理、从部署到比较的一次运行流程,以及其在压力测试中的应用情况。
反馈循环难题
构建自主渗透测试或漏洞赏金代理的团队往往共享同一种工作流程:构建一个看起来像真实公司的目标,运行代理,然后阅读输出结果。问题就出在输出环节。
一个声称利用了访问控制缺陷的代理,可能确实利用了该缺陷,可能只是擦边而过地触及了相关线索,也可能完全是凭空捏造。在这三种情况下,报告的措辞完全相同。此外,报告仅列出代理发现的内容。对于代理从未打开的四十个功能、从未枚举的 API,以及在其发现第一个漏洞的那个端点中存在的第二个漏洞,报告对此保持沉默。还有那些在寻找漏洞的过程中删除数据表或吊销所有 API 密钥的代理。没有任何客户会接受这样的结果,而发现列表中也没有任何内容能记录这一点。
人工审查可以应对单次运行。但它无法应对 AI 工程团队真正需要的实验矩阵——即对许多模型、许多配置和许多重复运行进行诚实的比较。
XRanges for AI 是什么
XRanges for AI 是一个用于整个评估的工作空间。它包含两个部分。
第一部分是基准目标库。每个目标都是一个完整的应用程序,而不仅仅是一组谜题挑战:它是一个拥有自身业务逻辑、种子数据、后台任务和模拟用户流量的多服务公司,跨多种语言和框架构建,因为这才是真实软件的开发方式。每个目标都包含 20 多个注入的漏洞,从单步缺陷到跨越服务边界的链式漏洞,其中包括 CTF.ae 自己的研究人员发现的零日漏洞。这些内容均未存在于公共训练语料库中,这一点的重要性每月都在增加。
第二部分是仪器化层。目标中的每个服务都通过 OpenTelemetry 发出结构化遥测数据。这些仪器化代码由应用安全工程师和软件工程师手工编写,专为特定目标设计。通用的 HTTP 日志记录会遗漏大多数关键信息。该平台在 ai.xranges.com 上摄取每次部署的遥测数据,并将其转化为四个分数,这些分数会在代理仍在运行时不断更新。
四个信号
每次部署都基于四个相互独立的测量指标进行评分。代理无法通过操纵其中一个指标来改善另一个,这种独立性正是其核心所在。
覆盖率回答了代理是否探索了目标的问题。每个面向用户的功能都是一个覆盖点,以业务操作而非 URL 的形式表述:注册账户、浏览职位发布、打开共享对话、在评估中运行代码。覆盖点只能通过正常使用到达,绝不能通过利用漏洞到达,因此该分数是衡量代理对合法表面工作彻底程度的纯净指标。未命中的点是代理的盲点,按名称列出。大多数团队发现这个列表比分数本身更有用。
覆盖点是业务操作,每个操作都有命中时间线。未命中的点是代理的盲点。
边界条件用于回答代理是否遵守了交战规则。每个目标都附带守卫规则,例如“不得删除招聘内容”或“不得撤销 API 密钥”。违规行为会在发生的瞬间被记录,并包含容器和 timestamp(时间戳)。零违规是预期结果。任何其他情况都是关于代理的发现,而非目标,且通常更为紧急。
针对每个目标的交战规则,在运行期间持续追踪违规行为。
利用情况回答了代理实际利用了哪些漏洞。每个漏洞都被定义为一系列有序阶段的杀伤链,从接触易受攻击的表面开始,直到仅在成功时触发的利用信号。由于每个阶段都是从目标内部检测到的,平台知道代理完成了哪一步以及在哪一步停滞,无论代理在其报告中写了什么。一个三步访问控制链如果在第二步停止,就会如实显示为:三步中的两步,并附带时间戳。
完整性回答了目标是否存活下来。完整性检查每分钟运行一次,确认应用程序在功能上仍然正确:种子数据存在、服务以正确的内容响应、跨服务信任保持完整。无论原因如何,检查失败都会导致扣分。它能捕捉到那些通过破坏其周围环境来发现漏洞的代理。
这四项汇总为一个分数,而该分数是页面上最无趣的数字。细分数据才是工作的重点。
一次端到端的运行
目标在约九十秒内部署为一个隔离的多容器环境。该平台可同时运行多达一千个部署,因此 AI 工程师、软件工程师和基础设施团队可以各自运行自己的实验,无需互相排队等待。
代理随后独立针对部署端点运行。平台从不介入代理与目标之间。它从内部进行监视。
在代理工作时,时间线记录其实际执行的操作,并将其映射到业务功能:预览职位发布、提交企业请求、生成 API 密钥。想要原始材料的工程师可以直接读取 OpenTelemetry 流,并使用支持正则表达式和属性过滤器的日志查询语言对其进行查询。当代理报告的内容不在目标的漏洞目录中时,时间线会对此进行裁定。有时这是误报。有时代理确实发现了无人植入的真实漏洞,这种情况已发生过多次。
重新测试并不意味着第二个实验室。可以在运行的部署中就地切换或修补漏洞。有些补丁在运行时生效,而另一些则需要一两分钟的重启。无论哪种方式,代理都会针对具有相同状态的同一环境进行重新测试。
每次部署还附带自定义元数据:模型名称、代理版本、提示变体、执行该操作的工程师。部署被分组,一个组显示其运行期间的平均分数和最佳分数,以及按漏洞划分的视图,展示哪次运行完成了哪个杀伤链。并排重复运行是将方差与改进区分开来的方法。单次运行证明的东西非常有限,而该平台正是基于这一假设构建的。
两次针对同一目标的并排运行,按信号细分。
自动化
控制台中的所有功能也通过 API 和 Model Context Protocol(模型上下文协议)服务器提供,需使用承载令牌(bearer token)。部署一批目标、启动代理、拉取覆盖范围和杀伤链进度,并在最后收集比较结果,都可以从 CI 管道或无人值守的聊天助手运行。控制台供阅读结果的人员使用,而 API 则用于实验矩阵。
实地验证:DEF CON 34 期间的 48 小时
Bug Bounty Village 每年在 DEF CON 期间为漏洞挖掘社区举办一场 CTF(夺旗赛),这是该活动中最精彩的活动之一。在定于 2026 年 8 月举行的 DEF CON 34 届大会上,CTF.ae 构建了名为 Xenoptic 的靶场:一家虚构的 AI 公司,其测试范围复杂且细致。545 名注册选手每人获得了一个独立隔离的公司副本,XRanges for AI 在整整 48 小时内对每一个副本进行了全程监控。
这样做的目的是为了确保公平性。此类竞赛的评判依据是提交的报告,而仅凭一份报告无法说明选手是如何达成目标的。有人可能会触发一个非预期的漏洞,一次性暴露所有旗帜;或者利用外部已知的 CVE(通用漏洞披露)绕过容器,在无需接触应用程序的情况下收集旗帜。在 DEF CON 的语境下,这是一种合理的假设。CTF.ae 需要观察每位选手及其代理人在环境中真实的行动路径,以便将提交的报告与实际发生的情况进行比对验证。
这正是该平台所提供的功能。在 850 多个部署实例中,它流式传输了目前用于评估代理程序的四种相同信号:
完整性确认(Integrity confirmed)确保每个环境保持健康状态。
边界记录(Boundaries recorded)捕捉任何超出交战规则的行为。
覆盖率(Coverage)显示每位选手实际处理了多少目标范围。
漏洞利用记录(Exploited)记录了哪些漏洞被真正利用,以及在哪个步骤被利用,从而使得每一份提交都能与真实路径进行核对。
所有数据均来自靶场内部,而非选手的本地机器。在持续遭受熟练研究人员攻击的情况下运行数百个并发部署实例,比测试单个实验室中的单一代理程序更具挑战性。目前用于评估代理程序的正是同样的监控机制。
适用对象
XRanges for AI 面向开发自主安全代理程序的团队,帮助团队了解其代理程序实际执行的操作,而非其声称执行的操作。该服务以托管云服务形式提供,也可在客户自有基础设施上自托管,确保数据完全保留在客户环境内部。如果团队希望将代理程序投入从未见过的目标进行测试,欢迎与我们联系!
觉得这篇文章有趣吗?
本文由我们的一位尊贵合作伙伴供稿。请通过以下平台关注我们,阅读更多独家内容:
Google News ,
Twitter 和
Autonomous security agents are getting good at finding bugs. Nobody has a good way to measure how good. Point one at a realistic target and what comes back is a report the agent wrote about itself: confident prose, a list of findings, and no way to tell which of them happened. Someone with a security background then sits down and checks every claim against the target. Which findings are real, which are duplicates, which are inventions, and, the question nobody has time for, what did the agent never try? That is a day of expert work for one run. Multiply it by three models, four prompt variants and ten repetitions, and the review queue is longer than the experiment.
XRanges for AI, built by CTF.ae, exists for that loop. It deploys realistic target applications with instrumentation baked into every service, records what an agent actually does inside them, and scores each run live on four independent signals. This walkthrough covers how it works, what a run looks like from deployment to comparison, and where it has been stress-tested.
The feedback loop problem
Teams building autonomous pentesting or bug bounty agents tend to share a workflow. Build a target that looks like a real company, run the agent, read the output. The output is where it goes wrong.
An agent that says it exploited an access control flaw may have exploited it, may have brushed past a hint of it, or may have made it up. The report reads identically in all three cases. A report also only lists what the agent found. It is silent on the forty features the agent never opened, the API it never enumerated, and the second bug sitting in the very endpoint where it found the first. Then there is the agent that deletes a table or revokes every API key on its way to a finding. No client would accept that result, and nothing in a findings list records it.
Manual review copes with one run. It does not cope with the experiment matrix an AI engineering team actually needs, which is many models, many configurations, many repetitions, compared honestly.
What XRanges for AI is
XRanges for AI is one workspace for the whole evaluation. It has two halves.
The first is a library of benchmark targets. Each is a complete application, not a set of puzzle challenges: a multi-service company with its own business logic, seeded data, background jobs and simulated user traffic, built across several languages and frameworks because that is how real software gets built. Each target carries 20 or more injected vulnerabilities, from single-step flaws to chains that cross service boundaries, including zero-days found by CTF.ae's own researchers. None of it exists in public training corpora, which matters more every month.
The second half is the instrumentation layer. Every service in every target emits structured telemetry through OpenTelemetry. The instrumentation is written by hand, by application security and software engineers, for that specific target. Generic HTTP logging would miss most of what matters. The platform, at ai.xranges.com , ingests the telemetry per deployment and turns it into four scores that update while the agent is still working.
The four signals
Each deployment is scored on four measurements chosen to be independent of each other. An agent cannot improve one by gaming another, and that independence is the whole point.
Coverage answers whether the agent explored the target. Every user-facing feature is a coverage point, phrased as a business action rather than a URL: registered an account, browsed job postings, opened a shared conversation, ran code in an assessment. A coverage point is only reachable through normal use, never through an exploit, so the score is a clean measure of how thoroughly the agent worked the legitimate surface. The unhit points are the agent's blind spots, listed by name. Most teams find that list more useful than the score.
Coverage points are business actions, each with a hit timeline. Unhit points are the agent's blind spots.
Boundaries answer whether the agent respected the rules of engagement. Each target ships with guard rules such as "must not delete hiring content" or "must not revoke API keys". A violation is recorded the moment it happens, with the container and timestamp. Zero violations is the expectation. Anything else is a finding about the agent, not the target, and usually a more urgent one.
Rules of engagement per target, each tracked for violations across the run.
Exploited answers which vulnerabilities the agent actually exploited. Every vulnerability is defined as a kill chain of ordered phases, from first contact with the vulnerable surface to an exploitation signal that only fires on success. Because each phase is detected from inside the target, the platform knows which step the agent completed and where it stalled, regardless of what the agent wrote in its report. A three-step access control chain that stopped at step two shows up as exactly that: two of three, with timestamps.
Integrity answers whether the target survived. Integrity checks run every minute and confirm the application is still functionally correct: seed data present, services answering with the right content, cross-service trust intact. A failed check is a penalty regardless of cause. It catches the agent that found a bug by breaking the environment around it.
The four roll up into one score, and the score is the least interesting number on the page. The breakdown is where the work is.
A run, end to end
A target deploys as an isolated multi-container environment in about ninety seconds. The platform runs up to a thousand deployments at once, so the AI engineers, the software engineers and the infrastructure team can each run their own experiments without queueing behind each other.
The agent then runs against the deployment endpoint on its own. The platform never sits between the agent and the target. It watches from the inside.
While the agent works, the timeline records what it actually did, mapped to business functionality: previewed a job posting, submitted an enterprise request, minted an API key. Engineers who want the raw material can read the OpenTelemetry stream directly and query it with a log query language that handles regular expressions and attribute filters. When the agent reports something that is not in the target's vulnerability catalogue, the timeline settles it. Sometimes it is a false positive. Sometimes the agent found a real bug nobody planted, which has happened more than once.
Retesting does not mean a second lab. Vulnerabilities can be toggled or patched in place on a running deployment. Some patches apply at runtime. Others need a restart of a minute or two. Either way the agent retests against the same environment with the same state.
Every deployment also carries custom metadata: model name, agent version, prompt variant, the engineer who ran it. Deployments are grouped, and a group shows the average and best score across its runs plus a per-vulnerability view of which run completed which chain. Repeated runs side by side are how variance gets separated from improvement. A single run proves very little, and the platform is built on that assumption.
Two runs of the same target side by side, broken down by signal.
Automation
Everything in the console is also available through an API and a Model Context Protocol server, with a bearer token. Deploying a batch of targets, launching the agent, pulling coverage and kill-chain progress and collecting the comparison at the end can run from a CI pipeline or from a chat assistant with nobody watching. The console is for people reading results. The API is for the experiment matrix.
Field proof: 48 hours at DEF CON 34
Bug Bounty Village runs a CTF at DEF CON every year for the bug hunting community, and it is one of the better things that happens there. For the DEF CON 34 edition in August 2026, CTF.ae built the target: Xenoptic, a fictional AI company with a sophisticated scope. Each of the 545 registered players got their own isolated copy of the whole company, and XRanges for AI watched every one of them for the full 48 hours.
The reason was fairness. A contest like this is judged on submitted reports, and a report on its own says nothing about how the player got there. Someone might hit an unintended bug that exposes every flag at once, or bring a CVE from outside, escape the container and collect the flags without touching the application. At DEF CON that is a fair assumption. CTF.ae needed to see how each player, and each player's agent, was really moving through the environment, so that what was submitted could be checked against what actually happened in that player's deployment.
That is what the platform delivered. Across 850+ deployments it streamed the same four signals it now uses on agents:
Integrity confirmed each environment stayed healthy.
Boundaries recorded anyone stepping outside the rules of engagement.
Coverage showed how much of the target each player had actually worked through.
Exploited recorded which vulnerabilities were really exploited, and at which step, so every submission could be checked against a real path.
All of it came from inside the target, never from the player's machine. Hundreds of concurrent deployments under sustained attack from skilled researchers is a harder test than one agent on one lab. The same instrumentation now evaluates agents.
Who it is for
XRanges for AI is for teams developing autonomous security agents who need to know what their agent did rather than what it said it did. It runs as a managed cloud service or self-hosted on the customer's own infrastructure, where nothing leaves their environment. Teams that want to put an agent against a target it has never seen can talk to us !
Found this article interesting?
This article is a contributed piece from one of our valued partners. Follow us on
Google News ,
Twitter and
LinkedIn to read more exclusive content we post.
首次收录 · 2026-09-24 · 9.44 分