Qwen 3.8 是阿里巴巴于 2026 年 8 月发布的模型系列。该系列包含四个共享名称的模型,它们在关键方面存在差异,这在你选择其中一个之前至关重要。你可以以 Apache 2.0 许可证下载 Qwen3.8 27B,并以 Qwen3.8-Max 许可证下载 Qwen3.8 2.4T A95B。Qwen3.8 Flash 没有自己的仓库,但其基础模型 Qwen3.8-Flash-Next 可以在 Qwen Community License 1.0 下下载。Qwen3.8 Max 没有公开权重。
阿里巴巴宣布 Qwen3.8-Max 为该系列的托管模型。QwenLM/Qwen3.8 发行说明记录了 Qwen3.8-2.4T-A95B 权重于 2026 年 8 月 12 日登陆 Hugging Face,以及 27B 权重于 2026 年 8 月 14 日上线。2.4T 模型卡将此描述为首个开源发布的 Qwen-Max 级模型。
你可以下载的模型与你可调用的模型并不相同。阿里巴巴的模型卡将 Qwen3.8-Max 描述为 Qwen3.8-2.4T-A95B 的托管版本,并增加了额外功能,包括视觉输入、非思考模式、1,000,000 token 的默认上下文以及内置工具。Flash 也遵循相同的关系。阿里巴巴将 Qwen3.8-Flash 描述为 Qwen3.8-Flash-Next 的托管版本,具有 1,000,000 token 的默认上下文和内置工具。
我们将所有四个模型进行路由。本文介绍每个模型的性质、哪些模型可供下载、许可证要求以及运行每个模型的成本。
四个 Qwen 3.8 模型
价格为美元每百万 token(提示词随后生成),数据来源于我们 2026 年 9 月 11 日的模型目录。27B 的价格是一个范围,因为有 14 个提供商以不同费率提供服务。七家 2.4T 提供商中有六家收费为 $2 / $6,Venice 收费为 $2.50 / $7.50。
模型 | 权重许可证 | 输入上下文(原生/最大托管) | 价格
qwen/qwen3.8-max-0902 | 仅限 API | 专有协议 | 文本、图像、视频 | 未公布 / 1,000,000 | $2 / $6
qwen/qwen3.8-2.4t-a95b | Hugging Face | Qwen3.8-Max 许可证 | 文本 | 262,144 / 1,048,576 | $2 / $6
qwen/qwen3.8-27b | Hugging Face | Apache 2.0 | 文本、图像、视频 | 262,144 / 1,000,000 | $0.15 至 $0.45 / $2.00 至 $3.20
qwen/qwen3.8-flash | Hugging Face(作为 Qwen3.8-Flash-Next) | Qwen Community License 1.0(在 Qwen3.8-Flash-Next 上) | 文本、图像、视频 | Qwen3.8-Flash-Next 上的 262,144 / 1,000,000 | $0.15 / $0.47
模型 ID qwen/qwen3.8-max 是一个别名。它当前解析为 qwen/qwen3.8-max-0902,即 dated 2026 年 9 月 2 日的快照。如果你希望请求在阿里巴巴发布新版本后仍持续命中同一个快照,请使用带日期的 ID。
上下文列显示两个数字,因为它们衡量的是不同的东西。第一个是模型训练的窗口,阿里巴巴为开源权重模型公布此数据,但未为 Max 公布。第二个是我们目录中任何提供商提供的最大窗口,它更大是因为窗口可以扩展到训练长度之外。Qwen3.8-27B 模型卡描述了使用 YaRN(一种 RoPE 缩放技术)将 27B 扩展到 1,000,000 token,并指出静态 YaRN 可能会降低对较短输入的 performance。提供商是否扩展窗口以及如何扩展是提供商的决定,因此你能获得的上限取决于我们将你的请求路由到哪里。
如果你使用提示词缓存(prompt caching),缓存费率因模型和提供商而异。Max 读取缓存的费用为每百万 token $0.25,写入费用为 $2.50。Flash 读取费用为 $0.016,在阿里巴巴上写入费用为 $0.20。2.4T 在大多数提供商上读取费用为 $0.25,仅在阿里巴巴列出写入费率,为 $2.50。27B 的缓存读取费率在提供该功能的提供商中范围为 $0.032 至 $0.18,仅阿里巴巴列出写入费率,为 $0.53。
Qwen 3.8 是开源的吗?
通过订阅,你同意接收 OpenRouter 通讯:模型使用数据、产品更新和研究报道,大约每周一封电子邮件。随时通过每封电子邮件中的链接取消订阅。请参阅我们的隐私政策。
部分是,答案取决于你指的是哪个模型。
你无法下载 Max。它没有 Hugging Face 仓库。Max 运行在阿里巴巴的服务器上,你通过 API 访问它。
你可以下载 2.4T 版本,但其许可协议在大规模应用时附加了条件。Qwen/Qwen3.8-2.4T-A95B 是一个经过指令微调的模型,拥有 2.4 万亿个总参数和每个 token 950 亿个活跃参数。它不使用 Apache 2.0 许可协议。Qwen 为其制定了 Qwen3.8-Max 许可协议。Hugging Face 元数据将该许可记录为名称为 qwen3.8-max 的其他类型。
该许可授予使用、复制、修改、分发、再许可、销售、部署、托管、微调以及创建衍生作品的权利,但需满足两个条件。
归属要求。如果基于该模型构建的商业产品或服务拥有超过 1 亿月活跃用户或月收入超过 2,000 万美元,则必须在其用户界面中显著展示模型名称。
针对大型模型即服务(Model-as-a-Service)业务的单独许可。如果你或你的关联公司运营模型即服务或 AI 工作助手业务,且在任何连续十二个月内总收入超过 5,000 万美元,则在任何商业使用前需要获得 Qwen 的单独许可。此条件不适用于内部使用,即不向第三方提供该模型、其输出或其能力的情况。
该许可将“模型即服务”定义为以允许第三方控制输入、参数或训练数据的方式,向第三方提供模型推理或微调访问权限。它不包括转发对由他人托管的模型的请求。它将“AI 工作助手”定义为主要设计用于 AI 辅助编码或办公生产力的独立产品。
你可以下载 27B 版本,其采用 Apache 2.0 许可协议。Qwen/Qwen3.8-27B 的 Hugging Face 元数据中记录为 apache-2.0。Apache 2.0 列于开源促进会(Open Source Initiative)批准的列表中,且不附带任何收入或用户数量条件。
Flash 版本以不同的名称和第三种许可协议提供开放权重。目前不存在 Qwen3.8 Flash 的代码库。存在一个名为 Qwen3.8-Flash-Next 的代码库,阿里巴巴将其描述为将支撑 Qwen4 的架构的实验性预览。其许可协议为 Qwen Community License 1.0。它与 Qwen3.8-Max 许可协议具有相同的归属要求,并且对于任何模型即服务或 AI 工作助手业务,在商业使用前需要获得 Qwen 的单独许可,且没有收入门槛。
因此,三个开放权重模型中的两个均在由 Qwen 编写且未获开源促进会(OSI)批准的许可协议下发布。开放权重意味着你可以下载并检查检查点(checkpoint)。但这并不意味着该许可是开源的。
Qwen 3 是较早的一代
Qwen 3 并非 Qwen 3.8 的前一个版本。Qwen 3 于 2025 年 4 月发布,而 Qwen 3.5、Qwen 3.6 和 Qwen 3.7 在这两者之间发布。Qwen 3.8 紧随 Qwen 3.7 之后。适用于 Qwen 3 模型的许可条款并不能告诉你关于本文提到的四个模型的任何信息。
你可以在本地运行 Qwen 3.8 吗?
你可以在自己的硬件上运行 27B 版本。2.4T 版本需要多 GPU 基础设施。
图 1. 针对不同输入模态的可下载权重。2.4T 是该系列中唯一仅读取文本的模型。没有哪个 Qwen 3.8 模型是仅限 API 且仅处理文本的。
2.4T 是一个混合专家(Mixture-of-Experts)模型。在任何单个 token 上只有 950 亿个参数处于活跃状态,但必须加载全部 2.4 万亿个参数,因为模型为每个 token 选择不同的专家集。你的硬件需要容纳的参数数量是 2.4 万亿。QwenLM/Qwen3.8 代码库中的 README 仅显示了针对 27B 版本的部署命令。
27B 是一个拥有 270 亿个参数的密集模型。在 bf16 精度下,每个参数占用两个字节,权重约占 54 GB,而在 fp8 精度下约占其一半。阿里巴巴的部署示例使用 SGLang、vLLM 或 TokenSpeed 在四块 GPU 上运行它,并充分利用 262,144 个 token 的原生窗口。
vllm serve Qwen/Qwen3.8-27B --port 8000 --tensor-parallel-size 4 --max-model-len 262144 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder
以每秒 1,000,000 个 token 的速度提供服务需要模型卡片中描述的 YaRN 覆盖,且模型卡片建议仅在需要更长上下文时才更改 RoPE 参数。
大型模型的开源权重仅支持文本读取。托管版 Max 接受图像和视频,开源权重的 27B 版本也支持。但开源权重的 2.4T 版本不支持。当我们于 2026 年 9 月 11 日向 qwen/qwen3.8-2.4t-a95b 发送图像时,请求失败,返回 HTTP 404 错误,消息为“未找到支持图像输入的端点”。向 qwen/qwen3.8-27b 和 qwen/qwen3.8-flash 发出的相同请求则成功。
在文本方面,开源权重的得分与 Max 接近。
| 模型智能指数 | 编码指数 | 智能体指数 |
|---|---|---|
| Qwen3.8 Max (0902) | 40.3 | 71.8 |
| Qwen3.8 2.4T A95B | 40.0 | 71.9 |
| Qwen3.8 27B (xhigh) | 33.9 | 68.1 |
这些分数来自独立基准测试机构 Artificial Analysis,我们将其发布在每个模型的页面上。我们于 2026 年 9 月 11 日获取了这些数据。Artificial Analysis 测量的是每个模型的过时快照,因此 Max 行是 2026 年 9 月 2 日的快照,而 27B 行是在 xhigh 推理强度下测量的。如果 Artificial Analysis 重新测量模型,或者您在不同的推理设置下运行模型,分数将会发生变化。
在智能指数上,2.4T 落后 Max 0.3 分,在编码指数上领先 0.1 分,在智能体指数上领先 0.8 分。27B 分别落后 Max 6.4、3.7 和 3.1 分。
如果您希望自托管能够读取截图的 Max 级别权重的负载,那么 27B 是支持图像读取的模型。
选择模型
利用这些事实来缩小选择范围。
如果您需要图像或视频输入并希望持有权重,27B 是我们目录中唯一接受它们的开源 Qwen 3.8 模型。它采用 Apache 2.0 许可证。
如果您需要图像或视频输入并希望获得该系列中最高的 Artificial Analysis 分数,Max 在智能指数上得分最高。
如果您希望最低的每 token 价格,Flash 的提示 token 价格为每百万 0.15 美元,补全 token 价格为每百万 0.47 美元,并且它接受图像和视频。您无法下载 Qwen3.8 Flash 本身,只能下载 Qwen3.8-Flash-Next。
如果您处理文本并希望持有权重,2.4T 是 Max 的开源变体。请根据上述许可阈值与您的具体数据进行对比。
如果您处理文本且不希望托管任何内容,Max 和 2.4T 在我们提供的 2.4T 七个提供商中的六个上成本相同,因此区别在于您是否需要图像输入。
如何调用 Qwen 3.8 模型
更改模型 ID,请求的其余部分保持不变。此示例使用 TypeScript SDK 。
import { OpenRouter } from '@openrouter/sdk' ;
const openRouter = new OpenRouter ({
apiKey: process.env. OPENROUTER_API_KEY ?? '' ,
});
const result = await openRouter.chat. send ({
chatRequest: {
model: 'qwen/qwen3.8-27b' ,
messages: [
{
role: 'user' ,
content: 'Summarize the trade-offs between fp8 and bf16 inference in two sentences.' ,
},
],
stream: false ,
},
});
if (result instanceof ReadableStream ) {
throw new Error ( 'Expected a non-streaming response' );
}
console. log (result.choices[ 0 ].message.content);
要发送图像,请传递内容部分数组而不是字符串。有关请求形状,请参阅图像输入。
const result = await openRouter.chat. send ({
chatRequest: {
model: 'qwen/qwen3.8-27b' ,
messages: [
{
role: 'user' ,
content: [
{ type: 'text' , text: 'What changed in this screenshot?' },
{ type: 'image_url' , imageUrl: { url: screenshotUrl } },
],
},
],
stream: false ,
},
});
if (result instanceof ReadableStream ) {
throw new Error ( 'Expected a non-streaming response' );
}
console. log (result.choices[ 0 ].message.content);
向 qwen/qwen3.8-2.4t-a95b 发出相同请求会失败,因为其提供商均不接受图像输入。
影响成本的因素
有两点因素会使你的账单超出所列费率。模型在回答之前会进行推理,而你需为推理令牌付费。为你提供路由的服务商设定价格。
推理默认开启
所有四个模型默认都会在回答前进行推理。推理令牌的计费采用完成费率,即两个费率中较高的一个。请参阅“推理令牌”部分了解推理参数的工作原理。
你的输出中推理所占的比例可能很大。我们向通过 Alibaba 接入的 27B 模型三次提问,要求用两句话解释 fp8 量化如何改变模型的输出。在每次运行中,推理令牌占完成令牌的比例均为 74%,具体数值分别为 128/174、171/231 和 181/243。在对同一模型发出图像描述请求时,121 个完成令牌中有 87 个是推理令牌。
能否关闭推理取决于模型和端点。在 2026 年 9 月 11 日,我们向每个模型发送了 reasoning: { effort: 'none' }。
模型 effort: 'none' 的结果
qwen/qwen3.8-max-0902 HTTP 400,此端点强制要求推理,无法禁用。
qwen/qwen3.8-2.4t-a95b HTTP 400,相同提示
qwen/qwen3.8-27b HTTP 200,零推理令牌
qwen/qwen3.8-flash HTTP 200,零推理令牌
Alibaba 的模型卡将非思考模式列为托管 Max 的一项功能,但我们路由到的 Alibaba 端点拒绝了该请求。27B 的结果涵盖了我们测试过的路由,即 Alibaba 和 Chutes。Flash 的结果涵盖 Makora。其他服务商的行为可能不同,因此在依赖推理关闭时,请检查响应中的 usage 字段。
要在 27B 上禁用推理,只需添加一个字段。
const result = await openRouter.chat.send({
chatRequest: {
model: 'qwen/qwen3.8-27b',
messages: [{ role: 'user', content: 'Reply with the single word OK.' }],
reasoning: { effort: 'none' },
stream: false,
},
});
我们分别通过 Alibaba 三次开启推理和三次关闭推理运行了 fp8 问题。开启推理时,每次调用的平均完成令牌数为 216,费用为 $0.00058。关闭推理时,平均完成令牌数为 61,费用为 $0.00017。无论按令牌数还是金额计算,这都减少了约 70%。
Max 和 2.4T 仍接受其他 effort 值。在同一测试运行中,对这四个模型设置 effort 为 max、high 和 minimal 的请求均返回 HTTP 200,因此你可以控制它们的推理程度。但在当前端点上无法完全停止它们。
服务商定价与上下文限制
Max 在我们的目录中只有一个服务商 Alibaba。Flash 有两个:Alibaba 和 Makora。2.4T 有七个。27B 有十四个,且它们的收费并不相同。在 2026 年 9 月 11 日,27B 的提示价格范围为每百万令牌 $0.15 至 $0.45,完成价格范围为 $2.00 至 $3.20。当前列表请参阅模型页面,且该列表会发生变化,因此请将此处数据视为快照。
27B 的服务商提供的上下文长度也不尽相同。十四个中有两个提供 1,000,000 令牌的窗口,一个在 65,536 处停止,其余的在 262,144 处停止。携带 400,000 令牌输入的请求必须路由到接受该长度的服务商。服务商的量化的方式也不同。十四个中有大多数列出 fp8,一个列出 fp4,还有几个将其量化列为未知。同一模型 ID 背后的权重在不同服务商之间并不完全相同,上述 Artificial Analysis 评分描述的是由 Artificial Analysis 测量的模型,而非任何特定服务商的版本。
没有一个单一数字能准确描述 27B 的成本。如果价格对你很重要,请指定一个服务商。你可以使用 order 字段选择服务商并关闭回退机制。服务商路由具备全套控制选项。
推理令牌并未在所有地方单独列示。在2026年9月11日的测试中,Chutes返回了模型的推理文本,但在用量报告中报告的推理令牌数为零,三次运行均如此。而Alibaba、CoreWeave和Novita在每次运行中都报告了非零的推理令牌数量。无论哪种方式你都需要为这些令牌付费,因此一个跨提供商对推理令牌数进行求和的控制台将会低估实际用量。
直接对接27B参数的提供商意味着需要管理14个账户和14次集成。通过我们,只需在请求中使用一个模型ID即可;当某个提供商响应变慢或返回错误时,我们会将请求路由到其他提供商。
常见问题解答
Qwen 3.8是开源的吗?
部分是。Qwen3.8 27B版本采用Apache 2.0许可证发布。Qwen3.8 2.4T A95B版本已发布
Qwen 3.8 is Alibaba’s model generation released in August 2026. Four models share the name, and they differ in ways that matter before you pick one. You can download Qwen3.8 27B under Apache 2.0 and Qwen3.8 2.4T A95B under the Qwen3.8-Max License. Qwen3.8 Flash has no repository of its own, but Qwen3.8-Flash-Next , the model it is based on, is downloadable under the Qwen Community License 1.0. Qwen3.8 Max has no public weights.
Alibaba announced Qwen3.8-Max as the hosted model of the generation. The QwenLM/Qwen3.8 release notes record the Qwen3.8-2.4T-A95B weights arriving on Hugging Face on 12 August 2026 and the 27B weights on 14 August 2026. The 2.4T model card describes this as the first open release of a Qwen-Max-class model.
The model you can download is not the same as the model you can call. Alibaba’s model card describes Qwen3.8-Max as the hosted version of Qwen3.8-2.4T-A95B with added features, including vision input, non-thinking mode, a 1,000,000-token default context, and built-in tools. The same relationship holds for Flash. Alibaba describes Qwen3.8-Flash as the hosted version of Qwen3.8-Flash-Next with a 1,000,000-token default context and built-in tools.
We route all four. This post covers what each model is, which ones you can download, what the licenses require, and what each one costs to run.
The four Qwen 3.8 models
Prices are US dollars per million tokens, prompt then completion, from our model catalog on 11 September 2026. The 27B price is a range because 14 providers serve it at different rates. Six of the seven 2.4T providers charge $2 / $6, and Venice charges $2.50 / $7.50.
Model Weights License Input Context, native / largest hosted Price
qwen/qwen3.8-max-0902 API only Proprietary Text, image, video Not published / 1,000,000 $2 / $6
qwen/qwen3.8-2.4t-a95b Hugging Face Qwen3.8-Max License Text 262,144 / 1,048,576 $2 / $6
qwen/qwen3.8-27b Hugging Face Apache 2.0 Text, image, video 262,144 / 1,000,000 $0.15 to $0.45 / $2.00 to $3.20
qwen/qwen3.8-flash Hugging Face , as Qwen3.8-Flash-Next Qwen Community License 1.0, on Qwen3.8-Flash-Next Text, image, video 262,144 on Qwen3.8-Flash-Next / 1,000,000 $0.15 / $0.47
The model ID qwen/qwen3.8-max is an alias. It currently resolves to qwen/qwen3.8-max-0902 , the snapshot dated 2 September 2026. Use the dated ID if you want the request to keep hitting the same snapshot after Alibaba publishes a new one.
The context column shows two numbers because they measure different things. The first is the window the model was trained on, which Alibaba publishes for the open-weight models and does not publish for Max. The second is the largest window any provider in our catalog offers, and it is larger because the window can be extended past the trained length. The Qwen3.8-27B model card describes extending the 27B to 1,000,000 tokens with YaRN, a RoPE scaling technique, and notes that static YaRN can reduce performance on shorter inputs. Whether a provider extends the window, and how, is that provider’s decision, so the ceiling you get depends on where we route your request.
If you use prompt caching , the cache rates differ by model and provider. Max charges $0.25 per million tokens to read from the cache and $2.50 to write to it. Flash charges $0.016 to read and, on Alibaba, $0.20 to write. The 2.4T charges $0.25 to read on most providers and lists a write rate only on Alibaba, at $2.50. The 27B cache read rate ranges from $0.032 to $0.18 across the providers that offer it, and only Alibaba lists a write rate, at $0.53.
Is Qwen 3.8 open source?
By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy .
Partly, and the answer depends on which model you mean.
You cannot download Max. It has no Hugging Face repository. Max runs on Alibaba’s servers and you reach it through an API.
You can download the 2.4T, and its license adds conditions at scale. Qwen/Qwen3.8-2.4T-A95B is the instruction-tuned model with 2.4 trillion total parameters and 95 billion active per token. It does not use Apache 2.0. Qwen wrote the Qwen3.8-Max License for it. The Hugging Face metadata records the license as other with the name qwen3.8-max .
The license grants the rights to use, copy, modify, distribute, sublicense, sell, deploy, host, fine-tune, and create derivative works, subject to two conditions.
Attribution. If a commercial product or service built on the model has more than 100,000,000 monthly active users or more than US$20,000,000 in monthly revenue, the model name must be prominently displayed in its user interface.
Separate license for large model-as-a-service businesses. If you or your affiliates run a Model as a Service or AI Work Assistant business, and your aggregate revenue exceeds US$50,000,000 in any consecutive twelve months, you need a separate license from Qwen before any commercial use. This condition does not apply to internal use that does not make the model, its outputs, or its capabilities available to a third party.
The license defines Model as a Service as giving a third party access to model inference or fine-tuning in a way that lets them control the inputs, parameters, or training data. It excludes relaying requests to models hosted by others. It defines AI Work Assistant as an independent product primarily designed for AI-assisted coding or office productivity.
You can download the 27B under Apache 2.0. Qwen/Qwen3.8-27B records apache-2.0 in its Hugging Face metadata. Apache 2.0 is on the Open Source Initiative approved list and carries no revenue or user-count conditions.
Flash has open weights under a different name and a third license. There is no repository for Qwen3.8 Flash. There is one for Qwen3.8-Flash-Next , which Alibaba describes as an experimental preview of the architecture that will underpin Qwen4. Its license is the Qwen Community License 1.0. It has the same attribution condition as the Qwen3.8-Max License, and it requires a separate license from Qwen for any Model as a Service or AI Work Assistant business before commercial use, with no revenue threshold.
Two of the three open-weight models therefore ship under licenses that Qwen wrote and that are not OSI-approved. Open weights means you can download and inspect the checkpoint. It does not mean the license is open source.
Qwen 3 is an earlier generation
Qwen 3 is not the version before Qwen 3.8. Qwen 3 was released in April 2025, and Qwen 3.5, Qwen 3.6, and Qwen 3.7 shipped between the two. Qwen 3.8 follows Qwen 3.7. A license term that applied to a Qwen 3 model does not tell you anything about the four models in this post.
Can you run Qwen 3.8 locally?
You can run the 27B on your own hardware. The 2.4T requires multi-GPU infrastructure.
Figure 1. Downloadable weights against input modality. The 2.4T is the only model in the family that reads text alone. No Qwen 3.8 model is API only and text only.
The 2.4T is a mixture-of-experts model. Only 95 billion parameters are active on any one token, but all 2.4 trillion parameters have to be loaded, because the model selects a different set of experts for each token. The number your hardware has to hold is 2.4 trillion. The README in the QwenLM/Qwen3.8 repository shows serving commands for the 27B only.
The 27B is a dense model with 27 billion parameters. At bf16, two bytes per parameter, the weights occupy about 54 GB, and about half that at fp8. Alibaba’s serving examples run it across four GPUs at the full 262,144-token native window using SGLang, vLLM, or TokenSpeed.
vllm serve Qwen/Qwen3.8-27B --port 8000 --tensor-parallel-size 4 --max-model-len 262144 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder
Serving it at 1,000,000 tokens requires the YaRN override described in the model card, and the model card advises changing the RoPE parameters only when you need the longer context.
The open weights of the large model only read text. The hosted Max accepts images and video, and so does the open-weight 27B. The open-weight 2.4T does not. When we sent an image to qwen/qwen3.8-2.4t-a95b on 11 September 2026, the request failed with HTTP 404 and the message No endpoints found that support image input . The same request to qwen/qwen3.8-27b and qwen/qwen3.8-flash succeeded.
On text, the open weights score close to Max.
Model Intelligence Index Coding Index Agentic Index
Qwen3.8 Max (0902) 40.3 71.8 49.6
Qwen3.8 2.4T A95B 40.0 71.9 50.4
Qwen3.8 27B (xhigh) 33.9 68.1 46.5
These scores come from Artificial Analysis , an independent benchmarking organization, and we publish them on each model’s page. We retrieved them on 11 September 2026. Artificial Analysis measures a dated snapshot of each model, so the Max row is the 2 September 2026 snapshot and the 27B row was measured at xhigh reasoning effort. The scores will move if Artificial Analysis remeasures a model or if you run the model at a different reasoning setting.
The 2.4T is 0.3 points behind Max on the Intelligence Index, 0.1 ahead on the Coding Index, and 0.8 ahead on the Agentic Index. The 27B is behind Max by 6.4, 3.7, and 3.1 points respectively.
If you want Max-class weights to self-host a workload that reads screenshots, the 27B is the model that reads images.
Choosing a model
Use these facts to narrow the choice.
If you need image or video input and want to hold the weights, the 27B is the only open-weight Qwen 3.8 model in our catalog that accepts them. It is Apache 2.0.
If you need image or video input and want the highest Artificial Analysis scores in the family, Max scores highest on the Intelligence Index.
If you want the lowest per-token price, Flash is $0.15 per million prompt tokens and $0.47 per million completion tokens, and it accepts images and video. You cannot download Qwen3.8 Flash itself, only Qwen3.8-Flash-Next.
If you work with text and want to hold the weights, the 2.4T is the open-weight variant of Max. Check the license thresholds above against your own numbers.
If you work with text and do not want to host anything, Max and the 2.4T cost the same through us on six of the 2.4T’s seven providers, so the difference is whether you need image input.
How to call the Qwen 3.8 models
Change the model ID and the rest of the request stays the same. This example uses the TypeScript SDK .
import { OpenRouter } from '@openrouter/sdk' ;
const openRouter = new OpenRouter ({
apiKey: process.env. OPENROUTER_API_KEY ?? '' ,
});
const result = await openRouter.chat. send ({
chatRequest: {
model: 'qwen/qwen3.8-27b' ,
messages: [
{
role: 'user' ,
content: 'Summarize the trade-offs between fp8 and bf16 inference in two sentences.' ,
},
],
stream: false ,
},
});
if (result instanceof ReadableStream ) {
throw new Error ( 'Expected a non-streaming response' );
}
console. log (result.choices[ 0 ].message.content);
To send an image, pass an array of content parts instead of a string. See image inputs for the request shape.
const result = await openRouter.chat. send ({
chatRequest: {
model: 'qwen/qwen3.8-27b' ,
messages: [
{
role: 'user' ,
content: [
{ type: 'text' , text: 'What changed in this screenshot?' },
{ type: 'image_url' , imageUrl: { url: screenshotUrl } },
],
},
],
stream: false ,
},
});
if (result instanceof ReadableStream ) {
throw new Error ( 'Expected a non-streaming response' );
}
console. log (result.choices[ 0 ].message.content);
The same request to qwen/qwen3.8-2.4t-a95b fails, because none of its providers accept image input.
What affects the cost
Two things change your bill beyond the listed rates. The models reason before they answer, and you pay for the reasoning tokens. The provider we route you to sets the price.
Reasoning is on by default
All four models reason before answering by default. Reasoning tokens are billed at the completion rate, which is the higher of the two rates. See reasoning tokens for how the reasoning parameter works.
The share of your output that is reasoning can be large. We asked the 27B through Alibaba, three times, to explain in two sentences why fp8 quantization changes a model’s output. Reasoning tokens were 74% of the completion tokens on each run, at 128 of 174, 171 of 231, and 181 of 243. On an image description request to the same model, 87 of 121 completion tokens were reasoning.
Whether you can turn reasoning off depends on the model and the endpoint. On 11 September 2026 we sent reasoning: { effort: 'none' } to each model.
Model Result of effort: 'none'
qwen/qwen3.8-max-0902 HTTP 400, Reasoning is mandatory for this endpoint and cannot be disabled.
qwen/qwen3.8-2.4t-a95b HTTP 400, same message
qwen/qwen3.8-27b HTTP 200, zero reasoning tokens
qwen/qwen3.8-flash HTTP 200, zero reasoning tokens
Alibaba’s model card lists non-thinking mode as a feature of the hosted Max, but the Alibaba endpoint we route to rejects the request. The 27B result covers the routes we tested, which were Alibaba and Chutes. The Flash result covers Makora. Other providers can behave differently, so check the response usage field when you rely on reasoning being off.
To disable reasoning on the 27B, add one field.
const result = await openRouter.chat. send ({
chatRequest: {
model: 'qwen/qwen3.8-27b' ,
messages: [{ role: 'user' , content: 'Reply with the single word OK.' }],
reasoning: { effort: 'none' },
stream: false ,
},
});
We ran the fp8 question three times with reasoning on and three times with it off, both through Alibaba. With reasoning on, the runs averaged 216 completion tokens and $0.00058 per call. With it off, they averaged 61 completion tokens and $0.00017 per call. That is roughly 70% less, whether you count tokens or dollars.
Max and the 2.4T still accept the other effort values. Requests with effort set to max , high , and minimal returned HTTP 200 on all four models in the same test run, so you can control how much they reason. You cannot stop them on the current endpoints.
Provider pricing and context limits
Max has one provider in our catalog, Alibaba. Flash has two, Alibaba and Makora. The 2.4T has seven. The 27B has 14, and they do not charge the same. On 11 September 2026, 27B prompt prices ran from $0.15 to $0.45 per million tokens and completion prices from $2.00 to $3.20. The current list is on the model page, and it changes, so treat the figures here as a snapshot.
The 27B providers do not all serve the same context either. Two of the 14 offer the 1,000,000-token window, one stops at 65,536, and the rest stop at 262,144. A request carrying 400,000 tokens of input has to be routed to a provider that accepts it. Providers also quantize differently. Most of the 14 list fp8, one lists fp4, and several list their quantization as unknown. The weights behind one model ID are not identical from provider to provider, and the Artificial Analysis scores above describe the model as measured by Artificial Analysis, not any particular provider’s copy.
No single figure describes what the 27B costs. Pin a provider if the number matters to you. You can pick a provider with the order field and switch off fallbacks. Provider routing has the full set of controls.
Reasoning tokens are not itemized everywhere. In our test on 11 September 2026, Chutes returned the model’s reasoning text but reported zero reasoning tokens in usage , three runs out of three. Alibaba, CoreWeave, and Novita reported non-zero reasoning token counts on every run. You pay for those tokens either way, so a dashboard that sums reasoning_tokens across providers will undercount.
Going direct to the 27B providers would mean 14 accounts and 14 integrations. Through us it is one model ID in the request, and when a provider slows down or returns an error, we route the request to another one.
FAQ
Is Qwen 3.8 open source?
Partly. Qwen3.8 27B is published under Apache 2.0. Qwen3.8 2.4T A95B is publis
首次收录 · 2026-09-29 · 9.95 分