使用工具的LLM智能体不再仅从单一检索到的段落中读取信息。通过模型上下文协议(Model Context Protocol, MCP),智能体可以调用搜索工具、检查结构化的患者或账户记录、查询数据库并提取元数据,然后将所有这些内容整合成一个答案。这使得关于事实性的常见讨论变得比表面看起来更为微妙。大多数用于检查LLM答案的系统——从RAGAS的忠实度指标到MiniCheck、AlignScore和SummaC等细粒度检查器——在证据汇集后,只询问某个主张是否得到了可用证据的支持。在它们通常的形式中,它们并未告诉我们哪个MCP工具输出支持了每个主张,也未指出该输出是否是答案所引用的来源。
我们最新的论文《ProvenanceGuard:面向基于MCP的LLM智能体的源感知事实性验证》(可在此阅读 Hugging Face 版本,或暂阅 arXiv 版本)旨在填补这一空白。我们所关注的故障模式是我们所称的跨源混淆:一个主张在证据中的某处为真,却被归因于错误的来源。一个不感知来源的验证器可能会通过它,因为该事实确实存在于证据池中;而一个感知来源的验证器则不应通过。
问题在于:“在某处得到支持”并不等同于“由正确的来源得到支持”。
考虑一个客户服务智能体回答:“根据账户记录,此计划包含30天的退款窗口。”这个退款窗口可能完全真实,但它陈述于政策文档中,而非答案所指向的账户记录中。将两者合并在一起时,该主张看起来得到了支持;但若将它们分开,归因就是错误的,而在数据敏感的环境中,错误的归因可能与错误的事实一样具有破坏性。同样的模式也出现在临床智能体中:从患者历史工具中提取的特定于患者的药物细节,一旦答案将其呈现为医学文献的发现,就会变得具有误导性。
一个主张可能由某个MCP来源支持,但答案却将其归因于另一个来源。不感知来源的评分在汇集的证据中看到支持性并通过;ProvenanceGuard则单独检查支持来源是否与答案所述或暗示的来源匹配。来源:论文图1。
这就是为什么尽管忠实度分数很有用,但对于MCP智能体来说却不够的原因。一个答案带有来源信息,有时是显式的(“根据账户记录”),有时是隐式的。ProvenanceGuard保留了主张与来源之间的连接,以便进行检查。
ProvenanceGuard的功能
ProvenageGuard是一个位于黑盒MCP智能体之上的生成后验证层。它在智能体生成答案后运行,且从不将证据合并为一个匿名上下文。相反,它将来源身份一直保留在管道中。它读取捕获的MCP轨迹,包括工具输出及其来源ID,而无需重新训练智能体。然后,它按顺序执行五项操作:将答案分解为具体主张,找到与每个主张最相关的来源,检查该来源是否实际支持该主张,将该来源与答案所述或暗示的来源进行比较,最后输出每个主张的来源判定以及全局的、答案级别的允许或阻止决策。
验证流程。来源身份在分解、路由、支持评分、归因检查和修复过程中得以保留,而非被汇集。被阻止的答案可以通过RARR风格的修复并进行重新验证。来源:论文图2。
有几个设计选择值得特别指出。在论文的实验部分,我们使用了本地模型,以便在受控的离线环境中处理捕获的追踪数据:MiniLM 用于查找相关来源,DeBERTa NLI(自然语言推理)验证器模型检查该来源是否支持该主张,而本地语言模型则有助于将答案拆解为各个主张。验证器还会仔细检查字面值:如果数字、日期或标识符在来源中不存在,仅凭句子听起来合理是无法通过验证的。一个经过校准的决策步骤会综合这些信号。如果某个答案被阻断,RARR 风格的修复步骤可以尝试基于来源的修订或安全的回退方案,随后由验证器再次进行检查。
上述提到的模型是我们评估的配置,而非 ProvenanceGuard 的要求。相同的声明、来源和决策步骤可以适配到托管模型中,如果团队更倾向于使用云服务;新的配置则需要其自身的测试和校准。我们报告的结果来自本地配置。其保守的决策策略适合对数据敏感的场景,在这些场景中,确保来源准确比生成尽可能快的答案更为重要。
结果
我们在一个使用了患者记录、研究文章和其他工具的医疗代理生成的答案上测试了 ProvenanceGuard。这为我们提供了 281 条真实追踪数据供研究。医学是一个有用的测试领域,因为来自患者记录的常识和来自一般研究的常识不能被视为相同的来源。当代理保留其工具输出和来源 ID 的记录时,该方法也可用于其他领域。在主要测试中,人类专家检查了从用于开发系统的数据中分离出来的 40 个答案中的 361 个主张。
最直接的结果是:专家称有 139 个主张不应通过,而 ProvenanceGuard 捕获了其中的 138 个。它让一个主张通过了。它还扣留了 67 个专家认为是得到支持的声明,并将它们发送回审查或修复。这反映了我们测试的谨慎设置:它倾向于对某些得到支持的主张进行二次检查,而不是让未得到支持的主张通过。对于有可识别来源的主张,在此测试中,它也大约 86% 的时间选对了来源。
我们在相同的主张上运行了其他四种支持检查器。ProvenanceGuard 在论文衡量系统阻止应被阻止的主张并避免不必要阻止的效果方面得分最高。此比较中的其他检查器并未告诉我们哪个工具输出了支持每个声明。ProvenanceGuard 记录了这种关联,因此审查者可以看到为每个主张检查的来源及其产生的决策。
| 验证器 | 拒绝/阻断 F1 | 发出声明到来源 ID |
|---|---|---|
| ProvenanceGuard( ours) | 0.802 | 是 |
| MiniCheck | 0.783 | 否 |
| RAGAS Faithfulness | 0.758 | 否 |
| AlignScore | 0.662 | 否 |
| SummaC-ZS | 0.436 | 否 |
同一保留声明数据包上的二元支持指标。ProvenanceGuard 在阻断方面匹配或优于源盲基线,同时还为每个声明生成来源裁决。来源:论文摘要和表 III。
检查来源看似相似的主张
在另一项涉及多个相似来源的更困难测试中,ProvenanceGuard 在决定阻止哪些主张方面的 F1 得分为 0.846,但在 50.3% 的主张中正确识别了确切来源。区分相似的来源仍然是改进的重要领域。
我们还运行了一个专注于错误归因的控制测试:我们在 50 个案例中更改了命名来源,同时保留了支持证据不变。ProvenanceGuard 捕获了所有 50 次交换。这表明它可以检测出明显的来源错误,而更困难的测试则显示了在众多看似合理的来源中进行选择的挑战。
修复被阻断的答案
只有当针对被拦截的答案有所作为时,拦截机制才具有实际效用。在连接至 RARR 风格的修复循环后,全轨迹运行解决了全部 173 个被拦截的答案,尽管其中 144 个以回退文本告终,而非实质性的重写内容;这表明系统选择避免提供无法验证的答案,而不是凭空捏造一个。在重建的多源测试轨迹上,一次新的修复运行解决了所有 59 个最初被拦截的答案,仅出现两次终端回退。作为离线门控机制,其开销适中,在报告的本地配置下,每个答案约需半秒,而 NLI(自然语言推理)和路由调用本身仅需数十毫秒。
为何这契合 Multiverse Computing
随着智能体从单段落 RAG 转向多工具 MCP 设置,事实来源的问题不再只是附注,而是成为事实性定义的一部分。ProvenanceGuard 使这种来源连接在每一个声明中变得可见。对于 Multiverse Computing 而言,这意味着一种检查现有智能体的方法,同时在需要时将敏感轨迹保存在受控环境中。医学研究是其中一个用例;同样的方法可以适应于任何智能体轨迹保留其工具和来源的场景。
这种适应性已在 NVIDIA NVFlow 中显现,后者为其金融智能体合并了一个可选的接地验证阶段。它将完成的答案与智能体检索到的 SEC 摘录进行比对,并保存独立的决策,而不改变原始的回放或训练数据。NVFlow 的贡献采用了 ProvenanceGuard 的源感知验证方法;上述修复循环属于更广泛的研究系统。
ProvenanceGuard 也在加州大学伯克利分校举行的 Agentic AI Summit 2026 上作为海报进行了展示。
想要获取完整的技术细节,包括路由和 NLI 推导、校准消融实验、多源压力切片以及完整的结果表?请在 Hugging Face 上阅读全文,或联系我们的团队探讨将源感知验证应用于您自己的智能体。
Tool-using LLM agents no longer read from a single retrieved passage. Through the
Model Context Protocol (MCP) , an agent can call a search tool, inspect a structured patient or account record, query a database, and pull metadata, then weave all of it into one answer. That m aakes the usual question of factuality more subtle than it looks. Most of the systems built to check LLM answers, from
RAGAS faithfulness to fine-grained checkers like MiniCheck, AlignScore, and SummaC, ask whether a claim is supported by the available evidence once that evidence has been pooled together. In their usual form, they do not tell us which MCP tool output supports each claim, or whether that is the source the answer names.
Our latest paper, ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (read it on Hugging Face , or on arXiv in the meantime), targets that gap. The failure mode we care about is one we call cross-source conflation: a claim that is true somewhere in the evidence, but attributed to the wrong source. A source-blind verifier may pass it, because the fact does exist in the pool. A source-aware verifier should not.
The problem: supported somewhere is not the same as supported by the right source
Consider a customer support agent that answers, "According to the account record, this plan includes a 30-day refund window." The refund window may be perfectly real, but stated in a policy document, not in the account record the answer points to. Pool the two together and the claim looks supported. Keep them separate and the attribution is wrong, and in a data-sensitive setting a wrong attribution can be as damaging as a wrong fact. The same pattern shows up in a clinical agent, where a patient-specific medication detail taken from a patient-history tool becomes misleading the moment the answer presents it as a finding from the medical literature.
A claim can be supported by one MCP source while the answer attributes it to another. Source-blind scoring sees support in the pooled evidence and passes it; ProvenanceGuard separately checks whether the supporting source matches the one the answer states or implies. Source: paper Figure 1.
This is why faithfulness scores, useful as they are, are not enough for MCP agents. An answer carries provenance, sometimes explicitly ("according to the account record") and sometimes implicitly. ProvenanceGuard keeps that connection between claim and source available for inspection.
What ProvenanceGuard does
ProvenanceGuard is a post-generation verification layer that sits on top of a black-box MCP agent. It runs after an agent produces an answer, and never collapses the evidence into one anonymous context. Instead it carries the source identity all the way through the pipeline. It reads the captured MCP trace, including the tool outputs and their source IDs, without retraining the agent. Then it does five things in sequence: it breaks the answer into specific claims, finds the source most relevant to each one, checks whether that source actually supports it, compares the source with the one the answer names or implies, and finally emits both a per-claim source verdict and a global, answer-level allow or block decision.
The verification flow. Source identity is preserved through decomposition, routing, support scoring, attribution checking, and repair, rather than being pooled. Blocked answers can go through RARR-style repair and be re-verified. Source: paper Figure 2.
A few of the design choices are worth calling out. For the experiments in our paper, we used local models so the captured traces could be processed in a controlled, offline setup: MiniLM helps find the relevant source, a DeBERTa NLI verifier model checks whether that source supports the claim, and a local language model helps break answers into claims. The verifier also checks literal values closely: a number, date, or identifier absent from the source cannot pass merely because the sentence sounds plausible. A calibrated decision step combines these signals. If an answer is blocked, a RARR -style repair step can try a source-grounded revision or a safe fallback, which the verifier then checks again.
Those named models are the setup we evaluated, not a requirement of ProvenanceGuard. The same claim, source, and decision steps can be adapted to hosted models where a team prefers cloud services; a new setup would need its own testing and calibration. Our reported results come from the local configuration. Its conservative decision policy suits data-sensitive review, where getting the source right matters more than producing the fastest possible answer.
Results
We tested ProvenanceGuard on answers from a medical agent that had used patient records, research articles, and other tools. This gave us 281 real traces to study. Medicine is a useful test because a fact from a patient's record and a fact from general research cannot be treated as the same source. The method can also be used in other fields when an agent keeps a record of its tool outputs and source IDs. For the main test, human experts checked 361 claims from 40 answers set aside from the data used to develop the system.
The most direct result is this: experts said 139 claims should not pass, and ProvenanceGuard caught 138 of them. It let one through. It also held 67 claims that the experts considered supported, sending them for review or repair. This reflects the cautious setting we tested: it favors a second look at some supported claims over letting unsupported ones through. For claims with an identifiable source, it also picked the right source about 86% of the time in this test.
We ran four other support checkers on the same claims. ProvenanceGuard scored highest on the paper's measure of how well a system catches claims that should be blocked while avoiding unnecessary blocks. The other checkers in this comparison did not tell us which tool output supported each claim. ProvenanceGuard records that connection, so a reviewer can see the source checked for each claim and the decision it produced.
Verifier
Reject/block F1
Emits claim-to-source ID
ProvenanceGuard (ours)
0.802
Yes
MiniCheck
0.783
No
RAGAS Faithfulness
0.758
No
AlignScore
0.662
No
SummaC-ZS
0.436
No
Binary support metrics on the same held-out claim packet. ProvenanceGuard matches or beats the source-blind baselines on blocking while also producing per-claim source verdicts. Source: paper abstract and Table III.
Checking claims when sources look similar
In a separate, harder test with several similar sources, ProvenanceGuard scored 0.846 F1 for deciding which claims to block, but identified the exact source correctly in 50.3% of claims. Telling similar sources apart remains an important area for improvement.
We also ran a controlled test focused on wrong attribution: we changed the named source in 50 cases while leaving the supporting evidence intact. ProvenanceGuard caught all 50 swaps. This shows it can detect a clear source error, while the harder test shows the challenge of choosing among many plausible sources.
Repairing blocked answers
Blocking is only useful if there is something to do with a blocked answer. Wired to the RARR-style repair loop, the full-trace run resolved all 173 blocked answers, though 144 of them ended in fallback text rather than a substantive rewrite, which is the system choosing to avoid an unverifiable answer rather than manufacture one. On reconstructed multi-source test traces, a fresh repair run resolved all 59 initially blocked answers with only two terminal fallbacks. As an offline gate the overhead is modest, roughly half a second per answer on the reported local configuration, with the NLI and routing calls themselves in the tens of milliseconds.
Why this fits Multiverse Computing
As agents move from single-passage RAG to multi-tool MCP setups, the question of which source a fact actually came from stops being a footnote and becomes part of what factuality means. ProvenanceGuard makes that source connection visible claim by claim. For Multiverse Computing, that means a way to check existing agents while keeping sensitive traces in a controlled environment when needed. The medical study is one use case; the same approach can be adapted wherever an agent's trace preserves its tools and sources.
That adaptation is already visible in NVIDIA NVFlow , which merged an optional grounding-verification stage for its finance agent. It checks completed answers against the SEC excerpts the agent retrieved and saves separate decisions without changing the original rollout or training data. The NVFlow contribution uses ProvenanceGuard's source-aware verification approach; the repair loop discussed above belongs to the broader research system.
ProvenanceGuard was also presented as a poster at the Agentic AI Summit 2026 at UC Berkeley .
Want the full technical details, including the routing and NLI derivations, the calibration ablations, the multi-source stress slices, and the complete results tables? Read the full paper on Hugging Face , or get in touch with our team to talk about applying source-aware verification to your own agents.
首次收录 · 2026-09-30 · 10.5 分