xAI 以优惠价格推出 Grok 4.7,但基准测试显示其与 Claude 和 GPT-6 之间存在巨大差距
埃隆·马斯克(Elon Musk)旗下的 xAI 推出了 Grok 4.7,这是其迄今为止在编码和知识工作方面能力最强的模型。据该公司称,该模型基于更大的基础模型构建,经过更长时间的强化学习训练,并被设计用于更好地验证其自身的输出。其定价为每百万输入令牌 2 美元,每百万输出令牌 6 美元。这些费率更接近中国模型,而非西方前沿模型,这背后可能有其原因。在结合十项基准测试的独立 Artificial Analysis Intelligence Index(v4.3.2)中,Grok 4.7 得分 46,位列中游。Claude Fable 5.1 和 GPT-6 以各 53 分领先。
Grok 4.7(黑色)总体得分为 46,明显落后于各得 53 分的 Claude Fable 5.1 和 GPT-6。其两个最高推理水平的表现似乎大致相同。| 图片来源:Artificial Analysis
在智能体编码方面,差距进一步扩大。在 Terminal-Bench 4.0 中,Grok 4.7 仅达到 26%,而 GPT-6 Astra 为 60%,Claude Fable 5.1 为 55%。甚至更便宜的 DeepSeek V4.1 Flash 也以 27% 的成绩超越它。该模型可通过 Grok API、Cursor 和 Grok Build 获取。
Grok 4.7 在 Terminal-Bench 4.0 的智能体编码测试中得分为 26%,远远落后于 GPT-6 Astra(60%)和 Claude Fable 5.1(55%)。| 图片来源:Artificial Analysis
没有炒作成分的 AI 新闻 – 由人工策划
订阅 THE DECODER,享受无广告阅读、每周 AI 通讯、每年六次的独家“AI Radar”前沿报告、完整档案访问权限以及评论版块访问权。
立即订阅
xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
Elon Musk's xAI has introduced Grok 4.7, its most capable model yet for coding and knowledge work. It's built on a larger base model, trained with longer reinforcement learning, and designed to better verify its own output, according to the company. Pricing sits at $2 per million input tokens and $6 per million output tokens. Those rates are closer to Chinese models than Western frontier models, probably for good reason. On the independent Artificial Analysis Intelligence Index (v4.3.2), which combines ten benchmarks, Grok 4.7 scores 46 and lands mid-pack. Claude Fable 5.1 and GPT-6 lead with 53 each.
Grok 4.7 (black) scores 46 overall, well behind Claude Fable 5.1 and GPT-6 at 53 each. Its two highest reasoning levels appear to perform about the same. | Image: Artificial Analysis
The gap grows wider in agentic coding. On Terminal-Bench 4.0, Grok 4.7 hits just 26 percent, versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1. Even the cheaper DeepSeek V4.1 Flash edges past it at 27 percent. The model is available through the Grok API , Cursor , and Grok Build .
Grok 4.7 scores 26 percent on Terminal-Bench 4.0's agentic coding test, far behind GPT-6 Astra (60 percent) and Claude Fable 5.1 (55 percent). | Image: Artificial Analysis
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-09-22 · 10.7 分