GPT-6.1
Sol
在各项任务中更强大的 Sol
编码
专业工作
计算机使用
科学研究
事实准确性
安全部署 GPT-6.1 Sol
定价与可用性
以五分之一价格实现接近 Astra 的智能
我们推出 GPT‑6.1 Sol,这是对 GPT‑6 Sol 的升级,它在智能代理编码、计算机使用和专业工作方面的能力几乎可与 GPT‑6 Astra 媲美,而标准输入和输出令牌的价格仅为 Astra 的五分之一。缓存输入成本仅为每百万令牌 0.10 美元——比标准输入定价低 95%,比 GPT‑6 Sol 的缓存输入定价低 50%——为开发者提供了更多空间来构建和运行能够在请求间复用上下文的强大智能体。
在各项任务中更强大的 Sol
GPT‑6.1 Sol 为重要的日常工作提供了能力与成本的新平衡。它在复杂的专业任务方面相比 GPT‑6 Sol 取得了显著改进,从编写和调试代码到理解文档以及执行多步骤业务流程。在这些评估中的几项上,它以显著更低的成本接近 GPT‑6 Astra 的表现。
编码
在 DeepSWE v1.1(在真实代码库中评估复杂软件工程任务)上,GPT‑6.1 Sol 以约五分之一的成本与 GPT‑6 Astra 持平,同时以更低的推理努力和成本超越了 GPT‑6 Sol 的最高得分 6.4 个百分点。
DeepSWE
GPT-6.1 Sol
GPT-6 Sol
GPT-6 Astra
$0
$2
$4
$6
每个任务的成本
0%
20%
40%
60%
80%
得分
GPT-6.1 Sol
GPT-6 Sol
GPT-6 Astra
在 DeepSWE 1.1(在新窗口中打开)中,AI 智能体解决原创的、长周期的软件工程任务。
专业工作
在 GDP.pdf(衡量模型使用包含表格、图表、示意图和细则细节的复杂 PDF 文档回答专业问题的准确性)上,GPT‑6.1 Sol 在低于一半的任务成本下,以超过 Opus 5.5 的得分领先,涵盖了所有测试的推理设置。它也以约五分之一的任务成本接近 GPT‑6 Astra 的最先进表现。
在 GDP.pdf(在新窗口中打开)中,模型必须回答来自金融、医疗、法律和其他七个专业领域工作流程中提取的真实世界复杂 PDF 提示。
在 AutomationBench(衡量智能体是否正确完成多步骤业务流程)上,GPT‑6.1 Sol 在中度推理努力下比 Opus 5.5 高出 2.2 个百分点,成本约为其三分之一。该得分也比相同设置下的 GPT‑6 Sol 高出 4.8 个百分点。
在 AutomationBench 1.0.6(在新窗口中打开)中,AI 智能体在销售、营销、运营、支持、财务和人力资源方面使用 47 种工具进行端到端工作流程测试。Claude Fable 5.1 的数据点低估了其实际成本,因为它省略了回退的成本,而回退发生在约 40% 的任务中。
计算机使用
GPT‑6.1 Sol 在需要与计算机应用程序交互的任务上也取得了显著进展。在 OSWorld 2.0 的离线集(评估智能体应对高难度计算机使用工作流程的能力)上,GPT‑6.1 Sol 以低于一半的成本,在最大推理努力下比 GPT‑6 Sol 高出七个百分点。它以约七分之一的任务成本,在最大推理努力下距离 Astra 的得分仅差 2.1 个百分点。
在 OSWorld 2.0(在新窗口中打开)中,AI 智能体尝试涵盖日常和专业任务的长周期计算机使用工作流程。我们报告的是 v2026.08.08 版本离线集上的部分奖励。
科学研究
在 Terminal-Bench Science 0.1(评估包括数据分析、仿真和定理证明在内的科学工作流程)上,GPT‑6.1 Sol 以低于一半的任务成本,在最大推理努力下将 GPT‑6 Sol 的得分翻倍以上。在最大努力下,GPT‑6.1 Sol 每个任务的平均成本为 5.47 美元,而 Opus 5.5 为 23.21 美元,Astra 为 23.80 美元,以比这两种模型低超过 75% 的成本提供显著的科学能力。
在测试的模型中,GPT‑6 Astra 仍以 68.1% 的最高得分位居榜首,适用于最具挑战性的科学研究任务。
在 Terminal-Bench Science 0.1(在新窗口中打开)中,智能体使用代码和终端工具完成科学研究工作流,包括数据分析、运行模拟和模型拟合。
事实准确性
GPT‑6.1 Sol 也在困难提示下提高了事实准确性。与 GPT‑6 Sol 相比,其最大的事实准确性提升出现在低推理努力场景下,其中包含事实错误的回答比例从 11.4% 降至 7.7%,降幅约为 32%。在测试的各种推理设置中,其错误率始终保持在低于 GPT‑6 Astra 1.9 个百分点的范围内,且每项任务的成本不足其五分之一。
该评估衡量的是在用户标记了早期模型错误的去标识化对话中,包含至少一个事实错误的回答所占的比例。这些刻意设计的困难提示并不能代表典型的使用场景。
我们在去标识化的 ChatGPT 对话中对事实准确性进行了评估,这些对话中的用户曾指出先前模型的事实错误。这些诱发错误的对话并不代表典型使用情况,因为在典型使用中事实错误更为罕见。
安全部署 GPT‑6.1 Sol
在我们的对齐评估中,GPT‑6.1 Sol 相比 GPT‑6 Sol 取得了显著改进,使其更接近 GPT‑6 Astra。
GPT‑6.1 Sol 对其局限性更加透明,在尊重用户意图和安全约束方面也更加可靠。在具有挑战性的评估中,它在关于损坏搜索工具的透明度、尊重明确限制以及在智能体任务中避免未经授权的结果方面,失败率低于 GPT‑6 Sol。我们未观察到任何试图绕过自动安全审查者的尝试,这与 GPT‑6 Astra 和 GPT‑6 Sol 的表现一致。详细信息请参阅 GPT‑6.1 Sol 系统卡附录(在新窗口中打开)。
以下评估刻意测试了具有挑战性的情境,并未衡量典型使用情况下的失败率。
损坏的搜索工具
审查者绕过
警告规避
计算机使用安全
该评估测试智能体是否在搜索工具损坏时告知用户,而不是给出其最佳猜测。GPT‑6.1 Sol 在 2.1% 的情况下未能披露问题,而 GPT‑6 Sol 为 4.9%,GPT‑6 Astra 为 1.5%,GPT‑6 Luna 为 28.7%。所选任务旨在诱发失败,并不代表典型使用情况。推理努力设置为最大值。
定价与可用性
GPT‑6.1 Sol 今日起向 ChatGPT Work 和 Codex 中的所有 Plus、Pro、Business、Enterprise 和 Edu 用户开放。GPT‑6.1 Sol 目前尚未在 Chat 中提供。开发者也可通过 OpenAI API 以 gpt-6.1-sol 的形式访问它。其标准 API 价格为每百万输入令牌 2 美元,每百万缓存输入令牌 0.10 美元,每百万输出令牌 10 美元。在未来几天内,我们还将提供 GPT‑6.1 Sol Ultrafast(在新窗口中打开),其在 Codex 中的令牌生成速度最高可达标准速度的 8 倍。
2026
GPT
作者
OpenAI
对 GPT 的评估是在我们的研究环境中或通过我们的 API 进行的,由于系统提示、可用工具、推理努力等方面的差异,其输出可能与生产环境的 ChatGPT 略有不同。竞争对手模型的评估数据来自公开可用的报告。
继续阅读
查看全部
推出 GPT-6 Sol 和 Luna
产品
2026年9月22日
GPT-6 Astra:新一代智能
研究
2026年9月3日
DevDay 2026 回顾
公司
2026年9月29日
GPT-6.1
Sol
A more capable Sol across tasks
Coding
Professional work
Computer use
Scientific research
Factuality
Deploying GPT-6.1 Sol safely
Pricing and availability
Near-Astra intelligence for a fifth of the price
We’re introducing GPT‑6.1 Sol, an upgrade to GPT‑6 Sol that nearly matches GPT‑6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices. Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing—giving developers more room to build and run capable agents that reuse context across requests.
A more capable Sol across tasks
GPT‑6.1 Sol offers a new balance of capability and cost for important everyday work. It delivers substantial improvements over GPT‑6 Sol across complex professional tasks, from writing and debugging code to understanding documents and executing multi-step business workflows. On several of these evaluations, it approaches GPT‑6 Astra’s performance at substantially lower cost.
Coding
On DeepSWE v1.1, which evaluates complex software-engineering tasks in real codebases, GPT‑6.1 Sol matches GPT‑6 Astra at roughly one-fifth of the cost, while eclipsing GPT‑6 Sol’s best score by 6.4 percentage points at a lower reasoning effort and cost.
DeepSWE
GPT-6.1 Sol
GPT-6 Sol
GPT-6 Astra
$0
$2
$4
$6
$8
Cost per task
0%
20%
40%
60%
80%
Score
GPT-6.1 Sol
GPT-6 Sol
GPT-6 Astra
In DeepSWE 1.1
(opens in a new window)
, AI agents solve original, long-horizon software engineering tasks.
Professional work
On GDP.pdf, which measures how accurately models answer professional questions using complex PDF documents, including tables, charts, diagrams, and fine-print details, GPT‑6.1 Sol scores higher than Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings. It also approaches GPT‑6 Astra’s state-of-the-art performance at roughly one-fifth the cost per task.
In GDP.pdf
(opens in a new window)
, models must answer real-world prompts about complex PDFs pulled from professional workflows in finance, healthcare, legal, and seven other professional domains.
On AutomationBench, which measures whether agents correctly complete multi-step business workflows, GPT‑6.1 Sol scores 2.2 percentage points above Opus 5.5 at medium reasoning effort, at roughly a third of the cost. That score is also up 4.8 percentage points from GPT‑6 Sol at the same setting.
In AutomationBench 1.0.6
(opens in a new window)
, AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance, and HR. The datapoint for Claude Fable 5.1 understates its actual cost, as it omits the cost of fallbacks, which occurred on ~40% of tasks.
Computer use
GPT‑6.1 Sol also makes substantial progress on tasks that require interacting with computer applications. On OSWorld 2.0’s offline set, which evaluates agents on demanding computer-use workflows, GPT‑6.1 Sol outperforms GPT‑6 Sol by seven percentage points at maximum reasoning effort at less than half the cost. It comes within 2.1 percentage points of Astra’s score at maximum reasoning effort at roughly one-seventh the cost per task.
In OSWorld 2.0
(opens in a new window)
, AI agents attempt long-horizon computer-use workflows spanning everyday and professional tasks. We report the partial reward on the offline set from the v2026.08.08 release.
Scientific research
On Terminal-Bench Science 0.1, which evaluates scientific workflows including data analysis, simulation, and theorem proving, GPT‑6.1 Sol more than doubles GPT‑6 Sol’s score at maximum reasoning effort at less than half the cost per task. At maximum effort, GPT‑6.1 Sol costs $5.47 per task on average, compared with $23.21 for Opus 5.5 and $23.80 for Astra, delivering substantial scientific capability at over 75% lower cost than either model.
GPT‑6 Astra still achieves the highest score among the models tested at 68.1%, and should be used for the most difficult scientific research tasks.
In Terminal-Bench Science 0.1
(opens in a new window)
, agents complete scientific research workflows using code and terminal tools, including analyzing data, running simulations, and fitting models.
Factuality
GPT‑6.1 Sol also improves factual accuracy on difficult prompts. Its largest factuality improvement over GPT‑6 Sol comes at low reasoning effort, where it reduces the share of responses containing a factual error from 11.4% to 7.7%—a reduction of approximately 32%. Across the tested reasoning settings, its error rate remains within 1.9 percentage points of GPT‑6 Astra’s, at less than one-fifth the cost per task.
This evaluation measures the share of answers containing at least one factual error on de-identified conversations where users flagged an earlier model’s error. These deliberately difficult prompts are not representative of typical usage.
We evaluate factuality on de-identified ChatGPT conversations where users had flagged a factual error from a prior model. These error-inducing conversations are not representative of typical usage, where factual errors are more rare.
Deploying GPT‑6.1 Sol safely
GPT‑6.1 Sol shows substantial improvements over GPT‑6 Sol in our alignment evaluations, bringing it closer to GPT‑6 Astra.
GPT‑6.1 Sol is more transparent about its limitations and more reliable at respecting user intent and safety constraints. In challenging evaluations, it shows lower failure rates than GPT‑6 Sol on transparency about broken search tools, respecting explicit restrictions, and avoiding unauthorized outcomes during agentic tasks. We observed no attempts to bypass an automated safety reviewer, matching GPT‑6 Astra and GPT‑6 Sol. Full details can be found in the GPT‑6.1 Sol system card addendum
(opens in a new window)
.
The evaluations below deliberately test challenging situations and do not measure failure rates in typical use.
Broken search tool
Reviewer bypass
Warning circumvention
Computer-use safety
This evaluation tests whether agents tell the user when their search tool is broken instead of giving their best guess. GPT‑6.1 Sol fails to disclose the problem in 2.1% of cases, compared with 4.9% for GPT‑6 Sol, 1.5% for GPT‑6 Astra, and 28.7% for GPT‑6 Luna. Tasks are selected to elicit failures and do not represent typical usage. Effort was set to maximum.
Pricing and availability
GPT‑6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. GPT‑6.1 Sol is not yet available in Chat. Developers can also access it through the OpenAI API as gpt-6.1-sol. Its standard API prices are $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast, with up to 8x faster token generation compared to its standard speed in Codex.
2026
GPT
Author
OpenAI
Evaluations of GPT were performed in our research environment or via our API, which may provide slightly different output from production ChatGPT due to differences in the system prompts, tools available, efforts, etc. Evaluations of competitor models were taken from publicly available reports.
Keep reading
View all
Introducing GPT-6 Sol and Luna
Product
Sep 22, 2026
GPT-6 Astra: A new generation of intelligence
Research
Sep 3, 2026
DevDay 2026 Recap
Company
Sep 29, 2026
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-30 | 9.7 | 7 | 入选 |
| 2026-09-30 | 12.36 | 7 | 入选 |