该项目致力于开发一种名为 Atria Dawn Preview 的智能体语言模型,该模型基于拥有 7440 亿参数的混合专家架构构建,专为研究和工程任务而设计。
该模型通过一个将每个任务与实际执行环境绑定在一起的流程进行训练。它调用工具、生成中间结果,并根据测试、指标或源代码证据等外部信号进行检查。团队表示,它在 16 项基准测试中的五项领先,包括网络搜索和网络安全领域,尽管它并未在整体上超越竞争对手。
Atria Dawn Preview 在 AutomationBench、CyberGym 和 MLE-Bench Lite 中处于领先地位,但在 GDPval 和 SWE-Bench Pro 中落后于其他竞争者。| 图片来源:Atria 团队
三分之一的已完成 AI 辅助任务如果没有 AI 协助本不会被尝试
在被审查的任务中,有 96.5% 使用了 AI。在整个项目过程中,参与者将越来越多的工作移交给智能体。四周内,智能体操作与人类输入的中位数比率从 11 上升至 28.5。团队警告不要将此解读为自主性的增强。每一次人类决策都导致了更多的智能体步骤,但这并不意味着智能体本身做出了更多决策。
在四周的时间里,每个人类输入对应的智能体操作中位数从 11 次增加到 28.5 次。| 图片来源:Atria 团队
研究还询问参与者是否能在没有 AI 的情况下,以相同的范围和完成同样的任务份额。在 455 个已完成的 AI 辅助任务中,有 151 个被评定为在没有 AI 的情况下不可行,约占三分之一。这些任务分布在 56 名参与者中的 27 人身上,因此并非仅来自少数高级用户。在这些情况下,AI 并没有加速现有工作,而是使原本根本不会开始的工作成为可能。
参与者表示,约三分之一的 AI 辅助任务在没有 AI 的情况下无法完成。| 图片来源:Atria 团队
AI 提议,人类选择
在方法和参数方面,最常见的模式是“AI 提议,人类选择”,占比为 55.4%。总体而言,人类做出了 85.5% 的方法和参数决策,而 AI 仅占 9.2%。在 93.4% 的情况下,人类对目标和范围拥有最终决定权。
AI 提出方案的份额因决策类型而异,介于 17% 到 55% 之间。其最终决策的份额始终是个位数。谁提出选项差异很大,但人类始终做出了大多数最终选择。即使在被评定为没有 AI 不可行的 151 个任务中,人类在 95.4% 的情况下选择了目标。
智能体通常提供方法提议,但在超过 80% 的情况下,人类做出最终决定。| 图片来源:Atria 团队
人类提供上下文,而非体力劳动
当出现问题时,同样的模式也会出现。在 588 个记录了难度的任务中,76% 通过人类干预得以推进,23% 由智能体自行解决问题。人类的帮助几乎总是以信息的形式出现,要么是通过添加上下文或澄清要求(35.2%),要么是通过诊断问题并切换方法(34.7%)。人类很少亲自做这些工作。部分编辑占案例的 3.2%,完全接管仅占 0.7%。
在四分之三的问题案例中,人类干预推动了工作进展,主要通过提供上下文或诊断,而非接管工作。| 图片来源:Atria 团队
当 AI 输出需要修订时,AI 在收到人类反馈后,有 75.4% 的时间自行处理更改。人类判断而非执行成为了瓶颈。
团队描述了 AI 角色的三个阶段,从研究的对象转变为单个任务的工具,现在是项目合作伙伴。在当前角色中,AI 在人类设定的目标内起草和调整计划。推测的第四阶段将涉及递归自我改进,更强的模型将产生更强大的后继者。
Atria团队将AI的演变描述为从研究对象到项目合作伙伴的转变。下一阶段——递归式自我改进——仍然是一个未解之谜。| 图片:Atria团队
作者表示,一个模型可以在其训练任务上取得进步,而无需在开发其继任者方面变得更好。AI如何能够提出多样化的研究方向并在结果出现之前评估其价值,这仍然是一个悬而未决的问题。
橡皮图章风险
当每一个决策都依赖于人类无法审查的更长的代理工作链时,监督就变得困难了。在最坏的情况下,人类变成了只能对他们所看到的进行“橡皮图章”式批准的审查者,团队写道。许多参与者还以自主模式运行代理,以避免因持续批准而打断长时间运行的任务。这一界限的划定是出于便利,而非关于AI应拥有多少权力的任何深思熟虑的选择。
这篇论文出现在关于递归式自我改进的辩论中心。Anthropic认为,开发其自身继任者的AI可能比预期更早成为现实,因此CEO达里奥·阿莫代伊呼吁为该行业设定速度限制。据Anthropic称,人类现在在公司关于研究方向的决策中仅占个位数百分比。OpenAI在其整个开发周期中使用GPT-5.6 Sol,而Google和DeepMind则通过Dream-RSI让AI代理通过记录的搜索轨迹探索替代策略,尽管它们只改进搜索策略,而非模型本身。
最近,领先AI公司的一千多名员工发出警告,称他们的组织可能正处于自动化AI研究的边缘。普林斯顿大学和英国AI安全研究所的一项独立研究得出了与Atria团队发现更一致的结论,表明前沿模型能够处理研究工程,但在真正重要的判断方面却失败了。
The project centered on developing an agentic language model called Atria Dawn Preview , built on a mixture-of-experts architecture with 744 billion parameters and designed for research and engineering tasks.
The model was trained through a pipeline that ties each task to a real execution environment. It calls tools, generates intermediate results, and gets checked against external signals like tests, metrics, or source evidence. The team says it leads on five of 16 benchmarks, including web search and cybersecurity, though it doesn't hold an overall edge over competitors.
Atria Dawn Preview leads in AutomationBench, CyberGym, and MLE-Bench Lite but trails the field in GDPval and SWE-Bench Pro. | Image: Atria Team
A third of completed AI-assisted tasks wouldn't have been attempted without AI
AI was used in 96.5 percent of the tasks reviewed. Over the course of the project, participants handed off more and more to agents. The median ratio of agent actions to human inputs rose from 11 to 28.5 over four weeks. The team cautions against reading this as growing autonomy. Each human decision led to more agent steps, which didn't mean the agents were making more decisions themselves.
Over four weeks, the median number of agent actions per human input rose from 11 to 28.5. | Image: Atria Team
Participants were also asked whether they could have completed their share of a task without AI, at the same scope and quality. Of 455 completed AI-assisted tasks, 151 were rated infeasible without AI, roughly a third. These tasks were spread across 27 of the 56 participants, so they didn't come from just a handful of power users. AI didn't speed up existing work in these cases. It made work possible that would never have been started otherwise.
Participants said about a third of AI-assisted tasks couldn't have been completed without AI. | Image: Atria Team
AI proposes, humans choose
For methods and parameters, the most common pattern was "AI proposes, human selects" at 55.4 percent. Overall, humans made 85.5 percent of decisions about methods and parameters, while AI made just 9.2 percent. Humans made the final decision on goals and scope in 93.4 percent of cases.
AI's share of proposals ranged from 17 to 55 percent depending on the decision type. Its share of final decisions stayed in the single digits. Who proposed the options varied widely, but humans consistently made most of the final choices. Even among the 151 tasks rated infeasible without AI, humans chose the goal 95.4 percent of the time.
Agents often supply the method proposals, but humans make the final call in over 80 percent of cases. | Image: Atria Team
Humans supply context, not manual labor
The same pattern shows up when things go wrong. Of 588 tasks with a recorded difficulty, 76 percent moved forward through human intervention, and in 23 percent the agent solved the problem on its own. Human help almost always came in the form of information, either by adding context or clarifying requirements (35.2 percent) or by diagnosing issues and switching methods (34.7 percent). Humans rarely did the work themselves. Partial edits accounted for 3.2 percent of cases, and full takeovers just 0.7 percent.
In three quarters of problem cases, human intervention moved work forward, mostly through context or diagnosis rather than taking over. | Image: Atria Team
When AI outputs needed revision, the AI handled the changes itself 75.4 percent of the time after receiving human feedback. Human judgment, rather than execution, was the bottleneck.
The team describes three phases in AI's role, from a subject of research to a tool for individual tasks and now a project partner. In that current role, AI drafts and adjusts plans within goals set by humans. A speculative fourth phase would involve recursive self-improvement, with stronger models producing stronger successors.
The Atria team describes AI's evolution from research object to project partner. The next stage, recursive self-improvement, remains an open question. | Image: Atria Team
The authors say a model can improve at its training tasks without getting better at developing its successor. How AI could propose varied research directions and assess their value before results are available remains an open question.
The rubber-stamp risk
When every decision rests on a longer chain of agent work than any human can review, oversight gets hard. In the worst case, humans become reviewers who can only rubber-stamp what they see, the team writes. Many participants also ran agents in autonomous modes to avoid interrupting long runs with constant approvals. That boundary was drawn out of convenience, not from any deliberate choice about how much authority AI should have.
The paper lands in the middle of a debate about recursive self-improvement. Anthropic considers an AI that develops its own successor possible sooner than expected , and CEO Dario Amodei is calling for a speed limit for the industry as a result. According to Anthropic, humans now make only a single-digit percentage of decisions about research direction at the company. OpenAI uses GPT-5.6 Sol across its entire development cycle , while Google and DeepMind let AI agents explore alternative strategies through recorded search trajectories with Dream-RSI, though they only improve the search strategy, not the model itself.
Over a thousand employees at leading AI companies recently warned that their organizations may be on the verge of automating AI research . A separate study from Princeton and the UK AI Security Institute reached a conclusion more in line with the Atria team's findings, showing that frontier models can handle research engineering but fail at the judgment calls that actually matter.
首次收录 · 2026-09-28 · 10.61 分