EvalEval 联盟很高兴地宣布,英国人工智能安全研究所(AISI)正在使用 EvalEval 的基础设施来公开共享评估结果,以支持更具可复现性和可验证性的评估科学。
AISI 和 EvalEval 此前曾在 NeurIPS 2025 期间联合举办的研讨会中开展了合作研究,来自该研究所的反馈有助于塑造“Every Eval Ever”(EEE)模式。此次合作的下一阶段旨在将这一共享基础设施付诸实践。
为何可复现的评估报告至关重要
随着人工智能部署的加速,评估正成为关于模型和系统性能的证据的重要来源。然而,结果往往以多种格式、平台和渠道发布,且通常缺乏足够的信息以供复现。重新运行这些评估本身可能成本高昂到令人望而却步。
EvalEval 的使命是通过共享的报告模式“Every Eval Ever”以及开放平台“Evaluation Cards”,改善这一生态系统,将评估结果及解释所需的信息整合为统一的结构。
这自然地建立在 AISI 通过 OptStop 提高评估效率、通过 HiBayES 增强统计严谨性,以及在包括转录分析和能力激发在内的领域实现标准化的工作基础之上。AISI 和 EvalEval 正共同努力诊断评估报告中的差距,并构建共享基础设施以弥补这些不足。
AISI 正在共享的内容
转录级别的透明度不仅对可复现性至关重要,也对分析和诊断具有重要意义。在此次合作的新阶段,AISI 通过“Evaluation Cards”公开了适当的已发布评估方法和发现。此次发布包括论文主要实验中五个基准测试的验证结果、上下文和配置信息:
HealthBench
FrontierMath
Humanity's Last Exam
SWE-Bench Pro
Terminal-Bench 2.0
这些结果涵盖了六个前沿模型:Claude Opus 4、Claude Opus 4.5、Claude Opus 4.6、GPT-5、GPT-5.2 和 GPT-5.4。此次发布还包括来自两项相关网络评估——Cyber CTFs 和 The Last Ones——的结果,这些评估使用了另一组部分重叠的模型。这些数据伴随 AISI 的论文《推理计算如何塑造前沿大语言模型评估》(How Inference Compute Shapes Frontier LLM Evaluation)一同发布,该论文研究了基准表现如何依赖于推理时的计算量和评估协议。
在“Humanity's Last Exam”上的表现随评估协议和推理计算量的变化而变化。每条曲线显示了在给定令牌数量内解决的尝试任务的累积比例,使用每个任务观察到的最早成功次数。当模型在每次尝试后从预言机获得正确性反馈时,随着令牌使用量的增加,它们继续解决额外的任务。
当结果连同设置信息一起公开发布时,研究人员和从业者可以更仔细地研究单个研究,并在更广泛的生态系统中比较发现。在其他报告缺乏这些细节的情况下,AISI 的发布提供了经过验证的参考点,用于在上下文中解释评估——例如,帮助研究人员理解设置选择如何影响报告的表现。随着更多评估者采用 EEE,此类开放比较可以支持更广泛、更可靠的元研究。
AISI 的 Terminal-Bench 2.0 结果与其他针对相同模型在不同评估设置下报告的评估结果并列。
我们对这一采纳感到兴奋,并期待与 AISI 及其他人工智能评估组织进一步标准化和共享评估。
为共同使命贡献力量
关于 EvalEval 联盟
EvalEval联盟是一个研究社区,致力于开发具有科学依据的研究成果以及稳健的部署基础设施,以支持评估生态系统。其目标是提升评估科学的水平,解决在记录评估适用性和效用方面缺乏共识的问题,并扩大对影响科学研究和政策分析的关键影响的覆盖范围。
该联盟的旗舰项目包括“Ever Eval”(Every Eval Ever),这是一个用于评估结果的共享模式和存储库;以及“Evaluation Cards”,它将基准元数据、评估运行数据和模型元数据整合为可解释的记录。二者结合使用,使得人们更容易理解看似相似的分数是否是在实质不同的条件下产生的。
关于英国AI安全研究所
英国AI安全研究所(AISI)是英国科学与技术创新部下属的一个研究机构。其使命是为政府提供对先进人工智能所带来风险的科学理解。AISI开展研究并构建基础设施,以理解先进人工智能的能力与影响,开发并测试缓解措施,并为政策制定提供信息支持。
延伸阅读
The
EvalEval Coalition is thrilled to share that the
UK AI Security Institute (AISI) is using EvalEval's infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.
AISI and EvalEval have previously collaborated on research that began at a joint workshop alongside NeurIPS 2025 , and feedback from the Institute has helped shape the Every Eval Ever (EEE) schema . This next phase of the collaboration puts that shared infrastructure into practice.
Why reproducible evaluation reporting matters
As AI deployment accelerates, evaluations are becoming increasingly important sources of evidence about model and system performance. Yet results are reported across many formats, platforms, and outlets, often without enough information to reproduce them. Running the evaluations again may itself be prohibitively expensive.
EvalEval's mission is to improve this ecosystem through a shared reporting schema, Every Eval Ever , and an open platform, Evaluation Cards , that brings evaluation results and the information needed to interpret them into a common structure.
This builds naturally on AISI's work to make evaluation more efficient through OptStop , more statistically rigorous through HiBayES , and more standardised in areas including transcript analysis and capability elicitation. Together, AISI and EvalEval are working to diagnose gaps in evaluation reporting and build shared infrastructure to close them.
What AISI is sharing
Transcript-level transparency matters not only for reproducibility, but also for analysis and diagnosis. In this new phase of the collaboration, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate. The release includes verified results, context, and configuration information for the five benchmarks in the paper's main experiment:
HealthBench
FrontierMath
Humanity's Last Exam
SWE-Bench Pro
Terminal-Bench 2.0
These results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. The release also includes results from two related cyber evaluations—Cyber CTFs and The Last Ones—which use a different, partially overlapping set of models. The data accompany AISI's paper, How Inference Compute Shapes Frontier LLM Evaluation , which studies how benchmark performance depends on inference-time compute and evaluation protocol.
Performance on Humanity's Last Exam changes with evaluation protocol and inference compute. Each curve shows the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task. When models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased.
When results are openly released with setup information, researchers and practitioners can examine individual studies more closely and compare findings across the wider ecosystem. Where other reports lack these details, releases like AISI's provide verified reference points for interpreting evaluations in context—for example, by helping researchers understand how setup choices may influence reported performance. As more evaluators adopt EEE, open comparisons like these can support broader and more reliable meta-research.
AISI's Terminal-Bench 2.0 results alongside other reported evaluations for the same models, under different evaluation setups.
We are excited about this adoption and look forward to further standardising and sharing evaluations with AISI and other AI evaluation organisations.
Contribute to the shared mission
About the EvalEval Coalition
The EvalEval Coalition is a research community developing scientifically grounded research and robust deployment infrastructure for the evaluation ecosystem. Its goal is to improve evaluation science, address the lack of consensus around documenting evaluation applicability and utility, and broaden coverage of the impacts that matter for scientific research and policy analysis.
The coalition's flagship projects include Every Eval Ever , a shared schema and repository for evaluation results, and Evaluation Cards , which combines benchmark metadata, evaluation-run data, and model metadata into interpretable records. Together, they make it easier to understand when apparently similar scores were produced under meaningfully different conditions.
About the UK AI Security Institute
The UK AI Security Institute is a research organisation within the UK government's Department for Science, Innovation and Technology. Its mission is to equip governments with a scientific understanding of the risks posed by advanced AI. AISI conducts research and builds infrastructure to understand advanced AI capabilities and impacts, develop and test mitigations, and inform policy.
Further reading
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-27 | 8.4 | 24 | 入选 |
| 2026-09-26 | 8.49 | 39 | 未入选 |
| 2026-09-25 | 8.77 | 42 | 未入选 |
| 2026-09-24 | 9.22 | 45 | 未入选 |
| 2026-09-23 | 9.95 | 31 | 未入选 |