🤗
模型 | 📊
数据
语言模型已经能够帮助研究人员搜索文献、综合证据并解决复杂问题。但科研工作对这些模型提出了特殊要求——答案必须立足于证据,模型需要保留证据实际支持的内容,而不是悄悄扩大研究的结论范围,并且研究人员需要能够验证最终输出结果。
我们在科学家使用 Asta(我们的科研代理平台)的方式中看到了这一点。用户通常不是进行简单的关键词搜索,而是提供大量的上下文和许多约束条件——例如,要求 Asta 在考虑特定方法、人群或研究环境的情况下,比较一系列文献中的不同方法。许多人之后还会回头查看生成的报告,将其视为可编辑的研究工作成果,而非一次性答案。
我们希望帮助科学家更快地生成带有引用的报告,并提供一个他们可以自行下载和运行的模型。为此,我们测试了专门针对科学报告生成训练的小型开源模型,看其能否在减少生成时间和服务成本的同时,达到与我们正在使用的专有模型相当的报告质量。
我们构建了 AstaBrief 8B,这是一个能将研究问题和检索到的文献片段转化为带引用报告的模型。AstaBrief 现已作为“快速模式”(Fast mode)与基于 Claude 的“思考模式”(Thinking mode)一同提供在 Asta 的“生成报告”功能中。我们还开源了该模型及其训练数据,以便其他人能够研究、复现并在此基础上构建。
开发 AstaBrief 需要数万个真实的研究查询、专注于引用的过滤、偏好数据,以及一个重新设计的报告生成流水线,该流水线一次性写出完整报告,而非逐节生成。与我们要追踪的专有模型相比,结果使报告生成时间减少了近一个数量级——在整个 Asta 流水线中,“快速模式”平均每份报告耗时 51.1 秒,而“思考模式”为 178.5 秒,速度约为 3.5 倍。
这些效率提升共同使得 AstaBrief 成为一个有用的测试案例,用于实现更广泛的目标:构建能够适应科研工作特定需求的开源语言模型。
开放权重还将使机构能够在其自有基础设施上运行 AstaBrief,这在研究问题涉及敏感或未发表的工作时是必要的。除了模型权重外,我们还发布了一个示例工作流,研究人员可以据此调整,从自己的 PDF 文件中创建报告,为本地报告生成提供起点。
本文介绍了我们如何训练 AstaBrief、我们在将其扎根于科学证据方面学到了什么,以及我们认为哪些方法部分可以延续到未来的科学模型中。文中描述的大部分训练和评估工作在 2025 年完成,因此用于生成训练数据和作为比较基准的专有模型反映了当时的前沿水平。我们没有针对当今的前沿模型重新运行完整的评估;以下结果最好被解读为关于我们所测试的特定训练和系统设计选择的证据。
模型训练
我们的目标是通过 AstaBrief 构建一个开源权重模型,具备对长篇科学综合最重要的所有特质:答案质量、相关性、结构和引用扎根性。我们从 Qwen3-8B 出发,将大部分精力集中在后训练数据、评估以及周围的报告生成支撑结构上。
将通用模型适应科学工作——以及从头训练新的科学模型——是我们正在 Ai2 广泛探索的方向。通过由 Ai2 领导的美国国家倡议 NSF OMAI(旨在为科学发现构建完全开放的 AI 基础设施和模型),我们的研究人员正直接与科学界合作,了解他们对未来开放模型的需求,以及当前通用模型的不足之处。这包括研究不同科学领域和工作流程中的需求差异,该研究的更多成果将在未来分享。
最近的工作,包括我们的 DR Tulu,表明基于强化学习(RL)的方法可以提高开放权重模型的长文报告生成质量,特别是在评审模型参与训练循环时。我们曾考虑过为 AstaBrief 采用这一路径,但最终专注于一种更简单的方案,该方案围绕监督微调(SFT)和直接偏好优化(DPO)构建。
基于 RL 的训练可能不稳定且成本高昂。我们希望看到在更便宜、更具操作管理性的设置下,报告生成质量能提升多少——这种设置也更易于调试和迭代。
这使得训练数据的质量变得尤为重要。我们没有依赖更复杂的优化方法来弥补噪声样本的不足,而是花了大量时间研究如何生成、选择和过滤真正展示我们所需写作行为的样本。
我们还希望 AstaBrief 运行得更快,以便用户能够快速获得初步报告,并在随后的交互中进行迭代。为了提升速度,我们决定让 AstaBrief 在给定用户查询和相关检索片段的情况下,直接一次性生成最终报告,从而绕过基于 Claude 的 Thinking 模式所使用的昂贵片段摘要和聚类阶段,也不分段写出答案部分。有趣的是,我们发现这样做并不会牺牲性能。
收集 SFT 训练数据
训练流程始于通过我们论文《Synthesizing scientific literature with retrieval-augmented LMs》中描述的系统以及 ScholarQA(即支撑 Asta “生成报告”功能的基础框架)提交的真实用户查询。我们不想仅针对合成提示或基准类任务进行训练,而是希望 AstaBrief 能从真实科学家的真实查询中学习。
我们的研究表明,科学家对语言模型的要求与用户对通用聊天机器人或传统搜索工具的要求往往不同。在对数十万条 Asta 查询的分析中,专家研究人员经常提供大量的上下文、多个约束条件以及概念之间的关系,而不是依赖简短的关键词式提示。
最近对 Asta 用户的研究也揭示了研究人员希望 AI 参与其工作的差异——有些人乐于使用模型进行构思或实验,而另一些人则倾向于让 AI 在综合、文献监控或模式发现中扮演更狭窄的角色。尽管存在这些差异,参与者都要求更清晰的来源可追溯性、对模型运作过程更高的可见性,以及对所用上下文更大的控制权。
我们对收集的用户日志进行了质量、相关性和隐私方面的过滤,剔除了测试版用户和机器人的流量,删除了过于简短而无意义的查询,并使用基于 LLM 的过滤步骤来捕获非英语查询、非科学请求以及包含个人信息的提示。最终留下了 90,000 个专注于研究的查询。
对于监督微调(SFT),我们利用 Asta 报告生成背后的多步 ScholarQA 管道,从过滤后的查询中生成完整的报告目标输出。该管道检索相关文献,将材料组织成各个章节,并使用一个支持性的报告生成模型将证据综合为一份带有引用的报告。我们采用了多种专有系统:Claude 3.5 Sonnet、Claude 3.7 Sonnet、o3、o4-mini 和 GPT-4.1。经过质量过滤后,这产生了 47,000 个可用的训练样本。
创建 DPO 对
DPO 需要不同类型的训练数据。与每个查询对应单一目标报告不同,我们需要成对的报告,其中一对中的一个优于另一个。
我们基于 SFT 数据生成过程中未使用的一个独立查询子集构建了这些配对。每个查询的一份报告来自现有的 ScholarQA 管道,通常由 Claude 3.5 Sonnet 或 3.7 Sonnet 提供支持。竞争性报告则是通过将 ScholarQA 检索到的文献摘录输入到不同的模型中生成的:根据具体样本的不同,使用 o3、o4-mini、DeepSeek-V3 或 DeepSeek-R1。
两个评判模型——GPT-4.1 和 DeepSeek-R1——对每对报告进行比较并选出优胜者。我们确保大语言模型(LLM)评判者与人类偏好保持一致(一致率达 95%),并且仅保留两位评判者意见一致的配对,从而获得了更纯净的偏好数据集,并大幅减少了通常在大规模生成的偏好数据中出现的噪声。
经过质量过滤后,最终的 DPO 数据集包含约 6,000 个样本。
使用多个生成器并要求两位评判者达成一致,为我们提供了一种相对简单的方法来构建偏好数据,而无需将任何单一模型的输出或判断视为绝对真理。
过滤数据以改善归因
我们的主要评估目标是 SQABench-CS2,这是一组由用户撰写的 200 个计算机科学研究问题。在 AstaBrief 的开发过程中,我们跟踪了四项指标:
评分量表得分(Rubric score),用于衡量报告覆盖了多少必要内容。
答案精确度(Answer precision),用于衡量每个段落是否与问题相关。
引用精确度(Citation precision),用于衡量每条引用是否支持其所附着的论点。
引用召回率(Citation recall),用于衡量报告的论点是否完全由所提供的引用所支持。
对于我们的最终模型,我们还进行了次要评估:DeepScholarBench,这是一个基于近期 ArXiv 论文构建的、包含 63 个查询的长文研究综合基准;以及针对由 Claude 支持的管道生成的报告进行的两次独立配对评估——一次是在 SQABench-CS2 上通过 LLM 评判的比较,另一次是小型的人类研究。
一份报告可能听起来经过精心打磨且完整无缺,但却偏离了问题核心,或将引用附着在缺乏底层证据支持的论点上。对于科学综合而言,我们需要分别衡量这些行为。然而,引用支持只是科学忠实性的一部分——模型可以引用正确的研究,但仍提出比该研究本身所支持更强的论点。这种情况可能以微妙的方式发生,例如,将对特定样本的发现转化为关于整个人群的通用主张,将过去时态报告的结果转变为听起来更具普遍真理性的现在时态陈述,或将描述性发现转化为对临床医生、政策制定者或研究人员应采取行动的推荐。
这类概括对于科学报告生成尤为重要,因为每一步都可能扩大证据的表观范围,而不引入明显错误的陈述。因此,被引用的句子在技术上可能与来源相关,但仍可能夸大研究人员实际确立的内容。我们的开发指标主要侧重于相关性、覆盖范围和引用依据;对科学报告撰写者的更丰富评估还应测试其是否保留了来源中论点的范围和强度。
我们的首轮监督微调(SFT)运行提升了整体内容质量,但在答案精确度和引用质量方面,仍落后于我们基于 Claude 的报告生成流水线。换句话说,模型在撰写报告方面变得更好了,但它仍然没有像我们在科学合成中所需的那样始终如一地以证据为基础。
这促使我们投入更多时间提升数据质量。我们测试了四种基于统计的过滤器,以识别较弱的合成训练样本:
输出与输入令牌比率。比率极高的答案通常噪声较大,因为它们用过少的证据生成了大量文本。
引用相关性。对于训练集中的每个合成报告,我们计算了其引用论文检索相关性的平均值。较低的平均值表明该报告过于依赖排名较低的证据。
引用密度。我们测量了至少有一条引用的陈述所占的比例。低密度的报告通常包含大段缺乏支持的文本。
引用多样性:我们衡量了在答案中引用的论文比例,前提是这些论文来自基于 Claude 的报告检索流水线返回的集合。较低的分值表明该报告过度依赖少数几篇论文。
最强的增益来自于过滤掉引用密度较低的合成报告;更激进的过滤、过滤器组合以及学习率扫描并未带来有意义的提升。
这是该项目中最清晰的教训之一:更复杂的过滤并不一定更好。一个相对简单的信号——即合成报告是否始终如一地引用其主张——比我们要尝试的几种更复杂的组合更有用。换句话说,科学专业化并不一定是向预训练中添加更多科学文本的问题;后训练数据的构成和质量,以及它是否展示了如 grounding(接地/依据)和 attribution(归因)等行为,可以实质性地改变最终模型的性能表现。
这种对基于证据且有用的输出的关注,也与我们在 Asta 用户研究中听到的反馈一致。参与者指出,生成更多文本并不一定更有帮助;他们希望获得简洁的合成结果,并拥有足够的来源可追溯性,以便在不浏览不必要输出的情况下审查和验证结果。
一旦我们获得了更强的 SFT 检查点,我们就在其基础上进行了 DPO(直接偏好优化)训练。这一阶段进一步提升了性能,使 AstaBrief 在报告生成方面达到了与 Asta 中的 Claude 驱动报告流水线以及 DR Tulu 相当的水平。
验证方法
由于该模型旨在作为我们 Agentic Asta 报告生成框架的一部分工作(不一定是作为独立模型),我们的主要问题是 AstaBrief 能否在保持我们所关注的报告质量的同时,实现显著更快且更便宜的报告生成流水线。换句话说,我们不仅询问模型是否能在个别基准测试中匹敌更强的专有模型;我们还想知道,在使用简单得多的系统时,我们能保留多少这样的质量。
每一行均按最佳结果优先排序;所有指标越高越好。Qwen3-8B 仅在 SQABench-CS2 测试集上进行了评估。SQABench-CS2 是一组用户编写的计算机科学研究问题;DeepScholarBench 使用其自身指标对长篇研究合成进行评分,这些指标与 SQABench-CS2 的指标不可比。
在我们开发过程中使用的评估中,AstaBrief 在答案和引用质量的几项衡量标准上与 Claude 驱动的流水线及 DR Tulu 具有竞争力。下图展示了由 LLM 判定的比较结果;在另一项包含 14 个问题的独立人类研究中,三位科研人员各自贡献了 4-5 个问题,并对来自三个系统的报告在整体偏好、完整性、相关性、组织结构和引用准确性方面进行了排名(允许并列)。在整体偏好方面,DR-Tulu 胜出,但三位研究人员中有两位在引用准确性指标上更偏爱 AstaBrief 而非其他系统,这证明了我们 SFT 数据质量过滤器的效用。
图表显示了在相同问题上,各系统在与“思维模式”进行的由大语言模型评判的报告对比中获胜的比例。人类评判是单独评估的,未包含在内。“思维模式”作为比较基准,没有显示柱状图。与 DR-Tulu 不同,Asta Brief 在 DPO(直接偏好优化)阶段针对这种成对报告排序进行了优化。
这些数字最好被解读为对我们开发当时所采用的工程方法的验证,而不是关于该特定基础模型相对于当今前沿模型所处位置的断言。模型生态系统发展迅速——我们期望泛化的部分在于数据构建、归因过滤和部署服务方面的经验教训。
在 Asta 中验证 AstaBrief 的实用性方面,“快速模式”已显示出令人鼓舞的早期使用情况。在尝试过该功能的 374 名 Asta 用户中,有 29.1% 的用户连续使用了两天或更长时间,且用户平均生成了 3.67 个报告线程。
🤗
Model | 📊
Data
Language models can already help researchers search the literature, synthesize evidence, and work through complex questions. But scientific work places particular demands on these models—answers need to stay grounded in evidence, the models need to preserve what the evidence actually supports rather than quietly broadening a study’s conclusions, and researchers need to be able to verify the final outputs.
We see that in how scientists use Asta , our agentic platform for scientific work. Instead of simple keyword searches, users often bring substantial context and many constraints—for example, asking Asta to compare approaches across a body of literature while accounting for a particular method, population, or setting. Many also return to generated reports later, treating them as working research artifacts rather than one-off answers.
We wanted to help scientists generate cited reports faster, with a model they could download and run themselves. To do that, we tested whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models we were using, while reducing generation time and serving costs.
We built AstaBrief 8B , a model that turns a research question and retrieved literature excerpts into a cited report. AstaBrief is available in Asta’s Generate a report feature today as Fast mode alongside Claude-powered Thinking mode, and we’re also open-sourcing it and the training data so others can study, reproduce, and build on our approach.
Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section. The result is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked—across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5× faster.
Together, those efficiency gains made AstaBrief a useful test case for a broader goal: building open language models that can be adapted to the specific demands of scientific work.
Open weights will also let institutions run AstaBrief on their own infrastructure, which is necessary when research questions reveal sensitive or unpublished work. Alongside the model weights, we’re releasing an example workflow that researchers can adapt to create reports from their own PDFs , providing a starting point for local report generation
This post covers how we trained AstaBrief, what we learned about grounding it in scientific evidence, and which parts of our approach we think can carry forward to future models for science. Most of the training and evaluation described was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time. We haven’t rerun the full evaluation against today’s frontier models; the results below are best read as evidence about the particular training and system design choices we tested.
Training the model
Our goal with AstaBrief was to build an open-weights model with all the qualities that matter most for long-form scientific synthesis: answer quality, relevance, structure, and citation grounding. We started from Qwen3-8B and focused most of our effort on the post-training data, evaluation, and surrounding report-generation scaffolding.
Adapting general-purpose models for scientific work – and training new scientific models from scratch – is something we're exploring broadly across Ai2. Through NSF OMAI , a U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, our researchers are working directly with scientific communities to understand what they need from future open models and where today's general-purpose models fall short. That includes studying how needs differ across scientific fields and workflows, with more findings from that research to share in the future.
Recent work, including our DR Tulu , has shown that reinforcement-learning-based (RL) methods can improve long-form report generation for open-weights models, especially when judge models are involved in the training loop. We considered that path for AstaBrief, but ultimately focused on a simpler recipe built around supervised fine-tuning (SFT) and direct preference optimization (DPO).
RL-based training can be unstable and expensive. We wanted to see how far we could push report generation quality with a cheaper, more operationally manageable setup—one that's also easier to debug and iterate on.
That made the quality of the training data especially important. Rather than relying on a more complex optimization method to compensate for noisy examples, we spent much of the project figuring out how to generate, select, and filter examples that actually demonstrated the report-writing behavior we wanted.
We also wanted AstaBrief to be faster so that users could get preliminary reports quickly that they could then iterate over in subsequent turns. For speed improvements, we decided to train AstaBrief to directly generate the final report in one pass given a user query and relevant retrieved snippets, bypassing the expensive snippet summarization and clustering stages our Claude-based Thinking mode uses and not writing out the answer section-by-section. Interestingly, we found it was possible to do so without sacrificing performance.
Collecting SFT training data
The training pipeline began with real user queries submitted through the system described in our paper “ Synthesizing scientific literature with retrieval-augmented LMs ” and ScholarQA , the framework that now underpins Asta’s Generate a report feature. Rather than training only on synthetic prompts or benchmark-style tasks, we wanted AstaBrief to learn from real queries from real scientists.
Our research suggests that scientists often ask different things of language models than users do of general-purpose chatbots or traditional search tools. In our analysis of hundreds of thousands of Asta queries , expert researchers frequently supplied substantial context, multiple constraints, and relationships between concepts rather than relying on short, keyword-style prompts.
More recent Asta user studies have also surfaced differences in how researchers want AI involved in their work—some are comfortable using models for ideation or experimentation, while others prefer a narrower role in synthesis, literature surveillance, or pattern-finding. Across those differences, participants want clearer source traceability, more visibility into what a model is doing, and greater control over the context it uses.
We filtered the user logs we collected for quality, relevance, and privacy, stripping out beta-tester and bot traffic, dropping queries that were too short to be meaningful, and using an LLM-based filtering pass to catch non-English queries, non-scientific requests, and prompts containing personal information. That left a pool of 90K research-focused queries.
For SFT, we generated full-report target outputs from the filtered queries using the multi-step ScholarQA pipeline behind Asta's report generation. The pipeline retrieved relevant literature, organized the material into sections, and used a backing report-generating model to synthesize the evidence into a cited report. We drew on a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, this yielded 47K usable training examples.
Creating DPO pairs
DPO required a different kind of training data. Instead of a single target report per query, we needed pairs of reports with one preferred over the other.
We built those pairs from a separate subset of queries not used during SFT data generation. One report per query came from the existing ScholarQA pipeline, typically backed by Claude 3.5 Sonnet or 3.7 Sonnet. The competing report was generated by feeding ScholarQA's retrieved literature excerpts to a different model: o3, o4-mini, DeepSeek-V3, or DeepSeek-R1, depending on the example.
Two judge models – GPT-4.1 and DeepSeek-R1 – compared each pair and picked a winner. We ensured that LLM judges were aligned with human preferences (95% agreement) and only kept pairs where both judges agreed, which gave us a cleaner preference set and cut much of the noise that typically shows up in preference data generated at scale.
After quality filtering, the final DPO dataset came to about 6K examples.
Using multiple generators and requiring agreement between two judges gave us a relatively simple way to construct preference data without treating any single model’s output or judgment as ground truth.
Filtering data for better attribution
Our main evaluation target was SQABench-CS2 , a set of 200 user-written computer science research questions. We tracked four metrics throughout the development of AstaBrief:
Rubric score , which measures how much necessary content is covered by the report.
Answer precision , which measures whether each paragraph is relevant to the question.
Citation precision , which measures whether each citation supports the claim it's attached to.
Citation recall , which measures whether the report's claims are fully supported by the citations provided.
For our final model, we also ran secondary evaluations: DeepScholarBench , a 63-query benchmark for long-form research synthesis built from recent ArXiv papers, and two separate pairwise evaluations against reports generated by the Claude-powered pipeline—an LLM-judged comparison on SQABench-CS2 and a small human study.
A report can sound polished and complete while meandering from the question or attaching citations to claims from which the underlying evidence doesn't follow. For scientific synthesis, we needed to measure those behaviors separately. But citation support is only part of scientific faithfulness—a model can cite the right study and still make a stronger claim than the study itself supports. This can happen in subtle ways , for example, turning a finding about a particular sample into a generic claim about an entire population, shifting a result reported in the past tense into a present-tense statement that sounds more universally true, or turning a descriptive finding into a recommendation for what clinicians, policymakers, or researchers should do.
Those kinds of generalizations are especially important for scientific report generation because each step can broaden the apparent scope of the evidence without introducing an obviously false statement. A cited sentence may therefore be technically related to its source while still overstating what researchers actually established. Our development metrics focused primarily on relevance, coverage, and citation grounding; a richer evaluation of scientific report writers should also test whether they preserve the scope and strength of the claims in their sources.
Our first SFT runs improved overall content quality, but they still lagged behind our Claude-powered report generation pipeline on answer precision and citation quality. In other words, the model got better at writing reports, but it still wasn’t grounded in evidence as consistently as we needed for scientific synthesis.
That pushed us to spend more time on data quality. We tested four statistics-based filters to identify weaker synthetic training examples:
Output-to-input token ratio . Answers with very high ratios were often noisy because they were generating a lot of text from too little evidence.
Citation relevance . For each synthetic report in the training set, we averaged the retrieval relevance scores of its cited papers. Low averages suggested the report was relying too heavily on lower-ranked evidence.
Citation density . We measured the share of statements that had at least one citation. Low-density reports often had large stretches of unsupported text.
Citation diversity: We measured the share of papers cited in the answer, given the set returned by the Claude-powered report retrieval pipeline. Low scores suggested the report was overly reliant on a few papers.
The strongest gains came from filtering out synthetic reports with low citation density; more aggressive filtering, filter combinations, and learning-rate sweeps didn't add meaningful gains.
That was one of the clearest lessons from the project: more elaborate filtering wasn’t necessarily better. A relatively simple signal – whether the synthetic reports consistently cited their claims – was more useful than several more complicated combinations we tried. Scientific specialization, in other words, isn't necessarily a matter of adding more scientific text to pretraining; the composition and quality of post-training data and whether it demonstrates behaviors like grounding and attribution can materially change how the resulting model performs.
That focus on grounded, useful output also lines up with what we’ve heard in Asta user research. Participants note that generating more text isn't necessarily more helpful; they want concise synthesis and enough source traceability to review and verify results without wading through unnecessary outputs.
Once we had a stronger SFT checkpoint, we ran DPO training on top of it. That stage pushed performance further, bringing AstaBrief within range of the Claude-powered report pipeline in Asta and DR Tulu on report generation.
Validating the approach
Because this model was intended to work as part of our agentic Asta report generation framework (not necessarily as a standalone model), our main question was whether AstaBrief could preserve the report qualities we cared about while enabling a substantially faster and cheaper report-generation pipeline. In other words, we weren’t only asking whether the model could match a stronger proprietary model on individual benchmarks; we wanted to know how much of that quality we could retain with a much simpler system.
Each row is ordered best first; higher is better on every metric. Qwen3-8B was evaluated on SQABench-CS2 test only. SQABench-CS2 is a set of user-written computer science research questions; DeepScholarBench scores long-form research synthesis with its own metrics, which are not comparable with SQABench-CS2's.
In the evaluations we used during development, AstaBrief was competitive with the Claude-powered pipeline and DR Tulu across several measures of answer and citation quality. The chart below shows the LLM-judged comparison—in a separate 14-question human study, three scientific researchers each contributed 4-5 questions and ranked reports from the three systems on overall preference, completeness, relevance, organization, and citation accuracy (with ties allowed). On overall preference, DR-Tulu wins, but two of the three researchers prefer AstaBrief over other systems on citation accuracy metrics, demonstrating the utility of our SFT data quality filters.
Bars show the share of LLM-judged report comparisons each system won against Thinking mode on the same questions. Human judgments were evaluated separately and are not included. Thinking mode is the comparison reference and has no bar. Unlike DR-Tulu, Asta Brief was optimized for this pairwise report ranking during the DPO stage.
These numbers are best read as validation of the engineering approach at the time we developed it, rather than as a claim about where this particular base model sits relative to today’s frontier. The model ecosystem moves quickly—the data construction, attribution filtering, and serving lessons are the pieces we expect to generalize.
Validating the usefulness of AstaBrief in Asta, Fast mode has shown encouraging early usage. Among 374 Asta users who’ve tried it, 29.1% have used it for two or more days, and users on average generate 3.67 report threads with it. Tw
首次收录 · 2026-10-03 · 10.61 分