随着 GPT-6 Sol 和 Luna 的推出,OpenAI 在其产品系列中增加了两款更便宜的模型,它们在性能上与前代产品相当,但代币价格减半,旨在与 Anthropic 部分更昂贵的模型相抗衡。
最大的变化是相比 GPT-5.6 Sol 和 Luna,价格降低了 50%。GPT-6 Sol 的输入代币价格为每百万个 2 美元,输出代币价格为每百万个 10 美元;而 Luna 的输入价格为每百万个 0.10 美元,输出价格为每百万个 0.50 美元。
OpenAI 将较低的价格归因于缓存和推理方面的改进,并表示将这些节省直接传递给用户。这使得其定价与较便宜的开源权重模型处于同一范围。此前系列中最便宜的模型 Terra 已不再可用。
模型
输入
输出
GPT-6 Sol 对比 GPT-5.6 Sol
$4 → $2
$20 → $10
GPT-6 Luna 对比 GPT-5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
据 OpenAI 称,Sol 旨在处理重复性的复杂任务,如构建新功能、审查代码、调试和分析数据。Luna 则旨在以低成本处理大量定义明确的任务,例如总结文档、提取信息和回答简短问题。
除了降低代币价格外,OpenAI 表示已改进 GPT-6 的提示缓存功能,对缓存的输入代币提供 90% 的折扣。新的提示缓存仪表板和诊断工具旨在帮助开发者优化缓存使用,且开发者现在可以在不使缓存失效的情况下更改推理力度和工具可用性。
在发布时,这两款模型均面向 Plus、Pro、Business、Enterprise 和 Edu 订阅用户,在 ChatGPT Work 和 Codex 中可用。Free 和 Go 用户可通过桌面应用访问 Luna,但这两款模型最初并未在普通聊天中提供。API 将它们作为 gpt-6-sol 和 gpt-6-luna 提供,而 ChatGPT 中的访问权限正在逐步推出。
GPT-6 Sol 和 Luna 的基准优势在于成本
OpenAI 主要将新模型定位为与 Anthropic 的 Claude 系列相比的性价比。在测试计算机使用的 OSWorld 2.0 上,OpenAI 表示 GPT-6 提供的结果与 Claude Opus 5 相似,但成本降低了约 80%,尽管 Astra 在此类别中仍领先。
在测试跨 47 种工具的业务工作流的 AutomationBench 上,据报道,GPT-6 Sol 在其最高力度下击败了处于最大力度的 Claude Opus 5。OpenAI 指出,Sol 每项任务的成本仅为 Opus 的 9%。与此同时,Luna 相比其前代产品提升了 5.4 个百分点,而成本降低了 58%。
在编程方面,OpenAI 提供了来自两个基准测试的结果。FrontierCode 1.1 测试 AI 代理生成的代码是否真正可以集成到现有代码库中,检查测试质量、代码风格以及对需求的合规性。GPT-6 Sol 在最大力度下得分为 49.3%,每项任务成本为 2.14 美元,这与 Claude Fable 5.1 大致持平,后者在最大力度下得分为 50.3%,但成本是前者的六倍,为 12.83 美元。Claude Opus 5 在中等力度下达到 53.4% 的得分,该项测试中该设置产生了其最佳结果,每项任务成本为 4.31 美元。
OpenAI 的推理力度正在失控
在 DeepSWE v1.1(一项针对真实代码库中长期进行的复杂软件工程任务的基准测试)上,OpenAI 报告称 GPT-6 Sol 在最大力度下的得分为 68.8%。这与 Claude Fable 5 在“xhigh”设置下的最佳得分 69.9% 仅相差 1.1 个百分点,而 Fable 在“max”设置下每项任务成本为 21.63 美元时得分为 69.7%。
当然,Sol 并非在此方面追逐前沿。Claude Opus 5 在最大力度下达到 73.7% 的得分,每项任务成本为 11.84 美元,而 OpenAI 自己的 GPT-5.6 Sol 得分为 72.7%,每项任务成本为 6.46 美元。GPT-6 Sol 的目标是以极低的成本接近这些数字。在“xhigh”设置下,它每项任务成本为 1.00 美元时提供 66.6% 的得分。Luna 更加便宜,在最大力度下达到相同的得分,仅需 0.22 美元。OpenAI 表示,Luna 的结果与 Claude Opus 5 和 Claude Fable 5 在中等力度下的结果相当,而成本比 Opus 低 93%,比 Fable 低 96%。
DeepSWE 的结果也让 OpenAI 自有模型之间的选择变得令人头疼。Luna 在最大努力模式下与 Sol 在“xhigh”模式下的表现持平,但成本却低了 78%。将 Sol 提升至最大努力模式仅能获得 68.8% 的成绩,比 Luna 高出区区 2.2 个百分点。在实际使用中,究竟该由谁来梳理这一切呢?
OpenAI 是自身最强劲的竞争对手。Luna 的性能与 Sol 相当,但成本却低得多。| 图片来源:OpenAI
总体而言,OpenAI 的基准测试选择看起来像是经过精心挑选的。诸如用于知识工作的 GDPval 或用于智能体编码的 Terminal-Bench 4.0 等指标缺失了,尽管它们通常是常规榜单的一部分。OpenAI 似乎还错过了 Opus 5.5 的发布,该模型的性能显著优于 Opus 5,且价格可能低多达 40%。
独立分析显示实际进展甚微
根据 Artificial Analysis 的分析,GPT-6 Sol 和 Luna 相比其前身将每项任务的成本减半,但智能评分仍保持在 GPT-5.6 的水平,在某些评估中有所提升,而在其他评估中则出现倒退。在编码智能体指数方面,Sol 提升了 2 分,而 Luna 下降了 2 分。
在 Artificial Analysis 智能指数上,GPT-6 Sol 从其前身的 47 分上升至 48 分,而 Luna 保持在 37 分。真正的差异在于每项任务的成本大幅降低,正如底部图表所示。| 图片来源:Artificial Analysis
Artificial Analysis 还在两个关键知识工作基准测试中发现了倒退现象。在测试跨 44 个专业领域的计算机化知识工作的 GDPval-AA v2.1 上,Sol 损失了约 100 个 Elo 分,Luna 下降了约 75 分。人工检查发现,这些倒退主要归因于呈现质量较低和结果不完整。
不过,Artificial Analysis 此前也曾因基准测试过时而给 OpenAI 的 Astra 模型评分过低而受到批评。随后进行了两次基准测试更新,Astra 才重新夺回榜首位置。我们来看看这次会发生什么。
无论如何,很明显 OpenAI 正在 GPT-6 Sol 和 Luna 上重金押注价格。现在,这些模型需要在日常工作中证明自己的实力,因为基准测试只能讲述故事的一部分。这也可能解释了为什么 OpenAI 遗漏了其中一些测试。结合不同的推理水平以及有时看起来是为了迎合基准测试而调整的结果(无论有意还是无意),随着每次新模型的发布,整个基准测试游戏似乎变得越来越荒谬。
没有炒作成分的 AI 新闻 – 由人工策划
订阅 THE DECODER,享受无广告阅读、每周 AI 通讯、每年六次的独家“AI Radar”前沿报告、完整档案访问权限以及评论板块的参与权。
立即订阅
With GPT-6 Sol and Luna, OpenAI adds two cheaper models to its lineup that match their predecessors' performance at half the token price and aim to rival some of Anthropic's more expensive models.
The biggest change is a 50 percent price cut compared with GPT-5.6 Sol and Luna. GPT-6 Sol now costs $2 per million input tokens and $10 per million output tokens, while Luna comes in at $0.10 for input and $0.50 for output.
OpenAI attributes the lower prices to improvements in caching and inference, saying it's passing those savings directly to users. That puts its pricing in the same range as cheaper open-weight models. Terra, previously the cheapest model in the lineup, is no longer available.
Model
Input
Output
GPT-6 Sol vs. GPT-5.6 Sol
$4 → $2
$20 → $10
GPT-6 Luna vs. GPT-5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
According to OpenAI, Sol is designed for recurring complex tasks such as building new features, reviewing code, debugging, and analyzing data. Luna is meant to handle large volumes of well-defined tasks at low cost, like summarizing documents, extracting information, and answering short questions.
Along with cutting token prices, OpenAI says it has improved prompt caching for GPT-6, offering a 90 percent discount on cached input tokens. A new prompt caching dashboard and diagnostics tool are meant to help developers optimize cache usage, and developers can now change reasoning effort and tool availability without invalidating the cache.
At launch, both models are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu subscribers. Free and Go users get access to Luna through the desktop app, but neither model is initially available in regular chat. The API offers them as gpt-6-sol and gpt-6-luna , while access in ChatGPT is rolling out gradually.
GPT-6 Sol and Luna's benchmark advantage comes down to cost
OpenAI is positioning the new models primarily on price-to-performance compared with Anthropic's Claude lineup. On OSWorld 2.0, which tests computer use, OpenAI says GPT-6 delivers results similar to Claude Opus 5 at roughly 80 percent lower cost, though Astra still leads in this category.
On AutomationBench, which tests business workflows across 47 tools, GPT-6 Sol at its highest effort level reportedly beats Claude Opus 5 at maximum effort. OpenAI puts Sol's cost per task at just 9 percent of Opus's. Luna, meanwhile, improves on its predecessor by 5.4 percentage points while costing 58 percent less.
For coding, OpenAI provides results from two benchmarks. FrontierCode 1.1 tests whether AI agents produce code that can actually be integrated into an existing codebase, checking test quality, code style, and compliance with requirements. GPT-6 Sol scores 49.3 percent at maximum effort for $2.14 per task, putting it roughly on par with Claude Fable 5.1, which scores 50.3 percent at maximum effort but costs six times as much at $12.83. Claude Opus 5 reaches 53.4 percent for $4.31 at medium effort, the setting that produced its best result on this test.
OpenAI's reasoning levels are getting out of hand
On DeepSWE v1.1, a benchmark for demanding software engineering tasks over long stretches in real codebases, OpenAI reports 68.8 percent for GPT-6 Sol at maximum effort. That's within 1.1 percentage points of Claude Fable 5's best score of 69.9 percent at "xhigh," while Fable at "max" hits 69.7 percent for $21.63 per task.
Of course, Sol isn't chasing the frontier here. Claude Opus 5 reaches 73.7 percent at maximum effort for $11.84 per task, and OpenAI's own GPT-5.6 Sol scores 72.7 percent for $6.46. The point of GPT-6 Sol is to land close to those numbers for a fraction of the price. At "xhigh," it delivers 66.6 percent for $1.00 per task. Luna is cheaper still, matching that score at maximum effort for just $0.22. OpenAI says Luna's result is comparable to Claude Opus 5 and Claude Fable 5 at medium effort, while costing 93 percent less than Opus and 96 percent less than Fable.
The DeepSWE results also turn the choice between OpenAI's own models into a headache. Luna at maximum effort matches Sol at "xhigh" while costing 78 percent less. Cranking Sol up to maximum effort only gets you to 68.8 percent, a mere 2.2 percentage points above Luna. Who's supposed to sort through all this in real-world use?
OpenAI is its own toughest competitor. Luna matches Sol's performance but costs far less. | Image: OpenAI
Overall, OpenAI's benchmark selection looks cherry-picked. Metrics like GDPval for knowledge work or Terminal-Bench 4.0 for agentic coding are missing, even though they're part of the usual lineup. OpenAI also appears to have missed the launch of Opus 5.5 , which is potentially up to 40 percent cheaper than Opus 5 with significantly better performance.
Independent analysis sees little actual progress
According to Artificial Analysis , GPT-6 Sol and Luna cut per-task costs in half compared to their predecessors, but intelligence scores stay at GPT-5.6 levels, with gains in some evaluations and regressions in others. On the coding agent index, Sol improves by 2 points while Luna drops by 2, according to the analysis.
On the Artificial Analysis Intelligence Index, GPT-6 Sol climbs from 47 to 48 points over its predecessor while Luna stays at 37. The real difference is the much lower cost per task, as the bottom chart shows. | Image: Artificial Analysis
Artificial Analysis also found regressions on two key knowledge-work benchmarks. On GDPval-AA v2.1, which tests computer-based knowledge work across 44 professional fields, Sol loses about 100 Elo points and Luna drops about 75. Manual inspection traced the regressions mainly to lower presentation quality and incomplete results.
Artificial Analysis has faced criticism before , though, when it rated OpenAI's Astra model too low because of outdated benchmarks. Two benchmark updates followed, after which Astra was back on top. We'll see what happens this time.
Either way, it's very much clear that OpenAI is betting heavily on price with GPT-6 Sol and Luna. Now the models need to prove themselves in day-to-day work, since benchmarks only tell part of the story. That may also explain why OpenAI left some of them out. Combined with the different reasoning levels and results that sometimes look tuned for benchmarks, whether intentionally or not, the whole benchmarking game seems more ridiculous with each new model launch.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-09-23 · 10.87 分