智能体在规划任务时可能需要做出一次艰难的决策,而在执行该任务时则可能涉及数十次调用。这些后续的调用负责读取文件、选择工具、验证结果并决定下一步的操作。NVIDIA 为此工作流中高频调用的部分构建了 Nemotron 3.5 Lightning。
Nemotron 3.5 Lightning 是一个拥有 300 亿参数的混合专家(MoE)模型,每个 token 激活约 30 亿个参数。它支持工具调用、编码任务、指令遵循以及其他范围明确的智能体步骤。
该名称容易与 Nemotron 3 Ultra 混淆,但这两个模型是为智能体工作流的不同部分而构建的。在它们之间进行选择会影响整个运行过程的成本、延迟和可靠性。
Nemotron 3.5 Lightning 概览
规格:Nemotron 3.5 Lightning
模型类型:混合 Mamba-2、注意力机制和混合专家语言模型
模型大小:总共 300 亿参数,激活 30 亿
输入与输出:文本到文本
模型上下文限制:根据 NVIDIA 的模型卡片,最多支持 100 万个 token
OpenRouter 标准模型:所有当前提供商均提供 262,144-token 上下文。完成限制因提供商而异,从 32,768 到 235,929 个 token 不等
OpenRouter 免费模型:1,000,000-token 上下文,最多 65,536 个完成 token
工具调用:两个模型均列出。当前四个标准提供商中的两个列出了该功能
结构化输出:在所有四个当前提供商的标准模型中列出。未在免费模型中列出
OpenRouter 模型 ID:nvidia/nemotron-3.5-lightning 和 nvidia/nemotron-3.5-lightning:free
权重:BF16 参考权重,以及优化的 NVFP4 和 GGUF 版本
许可证:OpenMDW-1.1
发布日期:2026 年 8 月 11 日
本表中的端点限制和功能数据来源于我们于 2026 年 9 月 11 日的模型页面。在设计特定限制时,请先查看标准模型页面或免费模型页面。
30B MoE 且 3B 激活参数意味着什么?
通过订阅,您同意接收 OpenRouter 通讯:模型使用数据、产品更新和研究报道,大约每周一封电子邮件。您可以通过每封电子邮件中的链接随时取消订阅。请参阅我们的隐私政策。
Nemotron 3.5 Lightning 是一个稀疏混合专家模型。它总共包含 300 亿个参数,但并非对每个 token 都使用所有专家。路由器会在模型处理每个 token 时选择一个较小的专家子集,从而将激活参数数量降至约 30 亿。您也会看到这种写法为 30B-A3B,意为总共 300 亿参数和约 30 亿激活参数。
这种设计使模型拥有比 3B 稠密模型更大的总容量,而无需为每个 token 运行全部 300 亿个参数。这就是 Lightning 如何将相对较大的检查点与频繁智能体调用所需的吞吐量相结合的方式。
激活数量并非完整的计算或内存规格。整个模型仍然需要存储,并且在推理期间共享层也会运行。硬件要求取决于精度、上下文长度、推理引擎、批处理策略和缓存配置。
Lightning 还使用混合架构,而非纯粹的 Transformer 堆栈。NVIDIA 的模型卡片描述了交错排列的 Mamba-2 和 MoE 层以及选定的注意力层。它还支持可配置的推理,并附带多 token 预测功能以及两种推测解码草稿器 DSpark 和 DFlash,以实现更快的生成,同时提供针对推理优化的 NVFP4 检查点。
Nemotron 3.5 Lightning 旨在用于什么场景?
长期运行的智能体会进行大量模型调用。其中一些调用需要复杂的规划。大部分工作则较为狭窄,例如选择工具、生成参数、检查结果、编辑文件或决定下一步操作。
NVIDIA 将 Lightning 定位为这一执行层。当您的工作负载具有以下特征时,该模型是一个候选方案:
智能体进行多次调用,且延迟或成本在运行过程中累积。
每次调用都有明确的目标以及完成该目标所需的足够上下文。
模型需要使用工具、遵循指令或返回结构化数据。
您需要可自定义的开放权重模型,以便针对特定领域或部署目标进行调整。
例如,编码代理可以使用更大的推理模型来检查代码库并制定实施计划。Lightning 随后可以处理单个步骤,如定位符号、编辑文件、运行工具以及解释结果。
Nemotron 3.5 Lightning 与 Nemotron 3 Ultra 对比
NVIDIA 从 Nemotron 3 Ultra 中提炼出了 Lightning,但这两个模型并不 interchangeable(可互换)。Lightning 是用于高吞吐量代理执行的小型模型。Ultra 是用于复杂推理和编排的大型模型。
Nemotron 3.5 Lightning | Nemotron 3 Ultra
OpenRouter 模型 ID | nvidia/nemotron-3.5-lightning | nvidia/nemotron-3-ultra-550b-a55b
总参数量 | 30B | 550B
激活参数量 | 3B | 55B
主要用途 | 高吞吐量执行和专门任务 | 复杂推理和编排
OpenRouter 上的上下文支持 | 当前所有提供商均支持 262,144 tokens | 根据提供商不同,支持 202,800 至 262,144 tokens
工具调用支持 | 在四个当前提供商中的两个上列出 | 在所有当前提供商上列出
结构化输出支持 | 在所有当前提供商上列出 | 在四个当前提供商中的一个上列出
2026年9月11日的提供商价格 | 输入每百万 token $0.065 至 $0.10,输出每百万 token $0.18 至 $0.25 | 输入每百万 token $0.50 至 $0.625,输出每百万 token $2.20 至 $3.125
NVIDIA 报告称,Lightning 的吞吐量可达同类开放模型的四倍。在 NVIDIA 的 PinchBench 测试中,它在保持相当准确度的情况下,完成 10,000 个代理任务的速度提高了多达 30%。这些是厂商报告的测试结果,因此请根据您的自身工作负载测试吞吐量和任务完成情况。
区别在于调用的难度和后果。当某个步骤需要对模糊问题进行深度推理、为长工作流程制定计划或做出难以逆转的决策时,Ultra 是更合适的选择。当任务范围明确且重复频率足够高,使得延迟和成本变得重要时,Lightning 是更合适的选择。
您无需将一个模型分配给整个代理。将复杂调用路由到 Ultra,将频繁执行调用路由到 Lightning。然后衡量这种组合是否在不降低可靠性的情况下提高了每个完成任务的成本效益。
在 OpenRouter 上在这两个模型之间进行路由
如果您更愿意将模型选择交给我们,请使用 Auto Router(自动路由器)。将 openrouter/auto 作为模型 ID 发送,Auto Router 会在根据任务类型、模型能力、工具支持和成本选择模型之前对提示进行分类。cost_tier(成本层级)设置选择一个从低到高的成本区间。它不会将从较便宜模型的对话升级到较昂贵的模型。如果您希望 Auto Router 仅在 Lightning 和 Ultra 之间进行选择,请使用 allowed_models(允许使用的模型)限制其选择范围。
在 OpenRouter 之外在这两个模型之间进行路由
NVIDIA 的 NeMo Switchyard 将此路由模式实现为开源编排层。其升级路由器以较低成本的模型开始每次对话,当检测到持续困难时,LLM 裁判会将会话移动到更具能力的模型。在 LangChain 对 145 个多步骤代理任务的评估中,在 Lightning 和 Claude Opus 4.8 之间进行路由,与仅使用前沿模型的基线相比,成本降低了 74%,同时将约 7% 的调用发送到前沿模型。准确度降低了约六个百分点。这一结果展示了该基准测试上的路由权衡。它并非针对 Lightning 和 Ultra 组合的基准测试。在依赖任何节省之前,请在您自己的任务上测试任何路由设置。
上下文窗口有多长?
Nemotron 3.5 Lightning 在模型级别支持多达 100 万个 tokens。您的应用程序可用的上下文取决于为其提供服务的端点。
2026年9月11日,所有提供标准模型ID的提供商均开放了262,144个token的上下文窗口。不同提供商的生成限制各不相同,范围从32,768到235,929个token不等。:free模型则暴露出完整的100万token上下文窗口,并将生成量限制在65,536个token。
如果请求需要超过262,144个token,目前只有免费端点在OpenRouter上暴露该模型完整的100万token上下文。NVIDIA在该端点上的通知指出,使用情况会被记录用于安全目的以及改进NVIDIA的产品和服务,并建议用户不要上传机密信息或个人数据。对于敏感或生产环境下的长上下文负载,请选择其他端点或模型。
在两个模型ID之间切换时,这种差异至关重要。在免费端点上成功的请求可能会超出标准端点的上下文限制。免费端点也列出了工具调用功能,但未列出response_format或结构化输出。
请查看模型页面,而不要将架构的最大上下文视为每个提供商的承诺。我们在模型页面上显示了每个端点的当前限制和支持的参数。
开放权重与本地部署
NVIDIA在OpenMDW-1.1许可下发布了BF16参考检查点。模型卡片将BF16版本描述为后训练、领域适配、研究以及生产优化变体的起点。NVIDIA还发布了用于定制和评估的训练数据和配方,模型卡片指出该版本已准备好用于商业用途。
对于直接推理,NVIDIA推荐其NVFP4版本。它还提供了支持本地系统的GGUF检查点。BF16指南列出了单GPU部署所需的H100 80GB或A100 80GB硬件。优化后的版本针对RTX 5090、RTX PRO 6000和DGX Spark等硬件进行了目标优化。
在本地运行模型并不意味着每台机器都能提供完整的100万token上下文窗口。在NVIDIA当前的Ollama指南中,默认上下文大小随可用VRAM缩放:低于24GB时为4K,24GB至低于48GB时为32K,48GB及以上时为256K。llama.cpp示例使用约40K。您可以提高这些限制,但这可能需要更多的内存或CPU卸载。
权重已公开,“开放权重”是最准确的描述。在重新分发修改后的模型或围绕其条款做出部署决策之前,请阅读OpenMDW-1.1许可协议。
如何在OpenRouter上调用Nemotron 3.5 Lightning
如果您不想配置GPU或管理推理引擎,可以通过OpenRouter调用Lightning。标准模型ID保持相同的API,同时我们将请求路由到该模型的可用提供商。
安装OpenRouter TypeScript SDK。
npm install @openrouter/sdk
在环境中设置OPENROUTER_API_KEY,然后使用Lightning模型ID发送请求。标准模型列出了结构化输出,因此您可以要求JSON schema并让响应符合该schema。
import { OpenRouter } from "@openrouter/sdk" ;
const openRouter = new OpenRouter ({
apiKey: process.env. OPENROUTER_API_KEY ,
});
const result = await openRouter.chat. send ({
chatRequest: {
model: "nvidia/nemotron-3.5-lightning" ,
messages: [
{
role: "user" ,
content:
'Ticket: "Checkout returns a 500 after I click Pay. Started this morning, three customers affected." Return its category and priority.' ,
},
],
responseFormat: {
type: "json_schema" ,
jsonSchema: {
name: "triage" ,
strict: true ,
schema: {
type: "object" ,
properties: {
category: { type: "string" },
priority: { type: "string" , enum: [ "low" , "medium" , "high" ] },
},
required: [ "category" , "priority" ],
additionalProperties: false ,
},
},
},
provider: {
requireParameters: true ,
},
},
});
if ( ! ( "choices" in result)) {
throw new Error ( "Expected a non-streaming response" );
}
console.log(result.choices[0]?.message.content);
当我们在2026年9月11日运行此请求时,模型返回了{"category": "Bug (Payment Checkout)", "priority": "medium"}。采样输出在不同运行之间会有所差异,而请求保证的是模式(schema)而非具体值。
同一模型对 response_format 和结构化输出的支持因提供商而异。默认情况下,我们优先选择支持您发送的 tools 和 response_format 参数的提供商;不支持某个参数的提供商将忽略该参数。如示例所示,将 require_parameters 设置为 true 会将请求限制为支持其中所有参数的提供商。有关完整行为,请参阅提供商路由文档。
要尝试免费端点,请将模型 ID 更改为 nvidia/nemotron-3.5-lightning:free。免费端点未列出 response_format,因此请删除该字段并自行验证 JSON。它适用于评估和低容量实验,其限制和可用性也与标准端点不同。
不要通过免费端点提交机密信息或个人数据。其模型页面载有 NVIDIA 的通知,指出使用记录将用于安全目的以及改进 NVIDIA 的产品和服务。在决定发送哪些提示词之前,请查阅该通知。
默认情况下,我们按价格顺序对标准模型的提供商进行负载均衡。有两种变体可以改变这种排序。nvidia/nemotron-3.5-lightning:nitro 按吞吐量对提供商进行排序,而 nvidia/nemotron-3.5-lightning:exacto 则偏好具有更强工具调用质量信号的提供商。对于选择用于延迟和工具使用的模型,Nitro 和 Exacto 是两种值得了解的变体。
Lightning 接受 tools 和 tool_choice,因此你也可以在智能体循环中使用它。请参阅我们的工具调用智能体循环指南,以获取完整的请求、工具执行和迭代流程。
你应该使用 Nemotron 3.5 Lightning 吗?
当你需要一个快速、开放权重的模型用于频繁且定义明确的智能体步骤时,Nemotron 3.5 Lightning 值得评估。其 30B-A3B 架构、工具支持和较低的列示令牌价格使其成为那些原本会将每次调用发送给更大推理模型的智能体的候选执行模型。
将模型名称作为起点,而非决策依据。从你的智能体执行的工具调用和任务中构建评估集。测量整个运行过程中的任务成功率、重试次数、延迟和总成本。如果智能体重复执行该调用或频繁升级结果以至于抵消了节省的成本,那么更便宜的调用并无帮助。
从 nvidia/nemotron-3.5-lightning 开始,或使用 nvidia/nemotron-3.5-lightning:free 在添加付费流量之前测试该模型。
常见问题解答
什么是 Nemotron 3.5 Lightning?
Nemotron 3.5 Lightning 是 NVIDIA 的开放权重 30B 混合专家语言模型,每个令牌约有 3B 个激活参数。NVIDIA 将其定位为用于高容量智能体执行、工具使用、编码、指令遵循和专门任务。
Nemotron 3.5 Lightning 与 Nemotron 3 Ultra 相同吗?
不。Nemotron 3.5 Lightning 有 30B 总参数和 3B 激活参数。Nemotron 3 Ultra 有 550B 总参数和 55B 激活参数。Lightning 针对频繁的执行步骤,而 Ultra 针对复杂的推理和编排。
30B-A3B 是什么意思?
30B-A3B 意味着该模型总共有 300 亿个参数,每个令牌激活约 30 亿个参数。混合专家路由器为每个令牌选择模型专家的一个子集,而不是运行所有专家。
标准 OpenRouter 模型 nvidia/nemotron-3.5-lightning 将其支持的参数列为 tools、tool_choice、response_format 和结构化输出。支持情况因提供商而异,因此如果你的请求依赖于其中任何一个参数,请将 require_parameters 设置为 true。免费模型列出了 tools 和 tool_choice,但未列出 response_format 或结构化输出。
是否有免费的 Nemotron 3.5 Lightning API?
是的。我们列出了由 NVIDIA 免费提供的 nvidia/nemotron-3.5-lightning:free 模型,该模型由 NVIDIA 提供服务且不收取任何令牌费用。免费模型的速率限制和可用性不同于付费模型,且此端点附带 NVIDIA 数据通知。请勿通过该端点发送机密信息或个人数据。
我可以在本地运行 Nemotron 3.5 Lightning 吗?
可以。NVIDIA 在 OpenMDW-1.1 许可证下发布了 BF16、NVFP4 和 GGUF 权重。BF16 参考检查点针对单个 80GB H100 或 A100 GPU。
An agent may make one difficult call to plan a task and dozens more to carry it out. Those later calls read files, choose tools, validate results, and decide what to do next. NVIDIA built Nemotron 3.5 Lightning for that high-volume part of the workflow.
Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model that activates about 3 billion parameters for each token. It supports tool calls, coding tasks, instruction following, and other well-scoped agent steps.
The name is easy to confuse with Nemotron 3 Ultra , but the two models are built for different parts of an agent workflow. Choosing between them affects the cost, latency, and reliability of the complete run.
Nemotron 3.5 Lightning at a glance
Specification Nemotron 3.5 Lightning
Model type Hybrid Mamba-2, attention, and mixture-of-experts language model
Model size 30B total parameters, 3B active
Input and output Text to text
Model context limit Up to 1 million tokens, per NVIDIA’s model card
OpenRouter standard model 262,144-token context on every current provider. Completion limits vary by provider, from 32,768 to 235,929 tokens
OpenRouter free model 1,000,000-token context, up to 65,536 completion tokens
Tool calling Listed for both models. Two of the four current standard providers list it
Structured outputs Listed for the standard model on all four current providers. Not listed for the free model
OpenRouter model IDs nvidia/nemotron-3.5-lightning and nvidia/nemotron-3.5-lightning:free
Weights BF16 reference weights, plus optimized NVFP4 and GGUF releases
License OpenMDW-1.1
Released August 11, 2026
The endpoint limits and features in this table come from our model pages on September 11, 2026. Check the standard model page or the free model page before designing around a specific limit.
What does 30B MoE with 3B active parameters mean?
By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy .
Nemotron 3.5 Lightning is a sparse mixture-of-experts model. It contains 30 billion parameters in total, but it does not use every expert for every token. A router selects a smaller subset of experts as the model processes each token, which brings the active parameter count to about 3 billion. You will also see this written as 30B-A3B, meaning 30 billion total parameters and about 3 billion active parameters.
This design gives the model more total capacity than a 3B dense model without running all 30 billion parameters for every token. That is how Lightning combines a relatively large checkpoint with the throughput needed for frequent agent calls.
The active count is not a complete compute or memory specification. The full model still has to be stored, and shared layers also run during inference. Hardware requirements depend on the precision, context length, inference engine, batching strategy, and cache configuration.
Lightning also uses a hybrid architecture rather than a pure Transformer stack. The NVIDIA model card describes interleaved Mamba-2 and MoE layers with selected attention layers. It also supports configurable reasoning and ships with multi-token prediction and two speculative-decoding drafters, DSpark and DFlash, for faster generation, along with an NVFP4 checkpoint tuned for inference.
What is Nemotron 3.5 Lightning designed for?
Long-running agents make many model calls. Some calls require difficult planning. Much of the work is narrower, such as choosing a tool, generating arguments, checking a result, editing a file, or deciding what to do next.
NVIDIA positions Lightning for that execution layer. The model is a candidate when your workload has the following characteristics:
The agent makes many calls, and latency or cost compounds across the run.
Each call has a clear objective and enough context to complete it.
The model needs to use tools, follow instructions, or return structured data.
You want open weights that you can customize for a domain or deployment target.
For example, a coding agent could use a larger reasoning model to inspect a repository and decide on an implementation plan. Lightning could then handle individual steps such as locating a symbol, editing one file, running a tool, and interpreting the result.
Nemotron 3.5 Lightning compared with Nemotron 3 Ultra
NVIDIA distilled Lightning from Nemotron 3 Ultra , but the two models are not interchangeable. Lightning is the smaller model for high-volume agent execution. Ultra is the larger model for complex reasoning and orchestration.
Nemotron 3.5 Lightning Nemotron 3 Ultra
OpenRouter model ID nvidia/nemotron-3.5-lightning nvidia/nemotron-3-ultra-550b-a55b
Total parameters 30B 550B
Active parameters 3B 55B
Primary role High-volume execution and specialized tasks Complex reasoning and orchestration
Context on OpenRouter 262,144 tokens on every current provider 202,800 to 262,144 tokens, depending on provider
Tool calling Listed on two of four current providers Listed on every current provider
Structured outputs Listed on every current provider Listed on one of four current providers
Provider prices on September 11, 2026 $0.065 to $0.10/M input, $0.18 to $0.25/M output $0.50 to $0.625/M input, $2.20 to $3.125/M output
NVIDIA reports that Lightning reaches up to four times the throughput of similarly sized open models. In NVIDIA’s PinchBench testing, it completed 10,000 agent tasks up to 30% faster at comparable accuracy. These are vendor-reported results, so test throughput and task completion on your own workload.
The distinction is the difficulty and consequence of the call. Ultra is the stronger fit when a step requires deep reasoning across an ambiguous problem, produces the plan for a long workflow, or makes a decision that is expensive to reverse. Lightning is the stronger fit when the task is well-scoped and repeated often enough for latency and cost to matter.
You do not need to assign one model to the entire agent. Route complex calls to Ultra and frequent execution calls to Lightning. Then measure whether the mix improves cost per completed task without reducing reliability.
Routing between the two models on OpenRouter
If you would rather hand model selection to us, use Auto Router . Send openrouter/auto as the model ID, and Auto Router classifies the prompt before selecting a model based on the task type, model capabilities, tool support, and cost. The cost_tier setting selects a cost band, from low to max . It does not escalate a conversation from a cheaper model to a more expensive one. If you want Auto Router to choose only between Lightning and Ultra, restrict its choices with allowed_models .
Routing between the two models outside OpenRouter
NVIDIA’s NeMo Switchyard implements this routing pattern as an open-source orchestration layer. Its escalation router starts each conversation with a lower-cost model, and an LLM judge moves the session to a more capable model when it detects sustained difficulty. In LangChain’s evaluation of 145 multi-step agent tasks, routing between Lightning and Claude Opus 4.8 reduced cost by 74% compared with the frontier-only baseline while sending about 7% of calls to the frontier model. Accuracy was about six points lower. This result shows the routing trade-off on that benchmark. It is not a benchmark of a Lightning-and-Ultra combination. Test any routing setup on your own tasks before you rely on the savings.
How long is the context window?
Nemotron 3.5 Lightning supports up to 1 million tokens at the model level. The context available to your application depends on the endpoint that serves it.
On September 11, 2026, every provider behind the standard model ID exposes a 262,144-token context window. Completion limits differ by provider, from 32,768 to 235,929 tokens. The :free model exposes the full 1-million-token context window and caps completions at 65,536 tokens.
If a request needs more than 262,144 tokens, only the free endpoint currently exposes the model’s full 1-million-token context on OpenRouter. NVIDIA’s notice on that endpoint states that use is logged for security purposes and to improve NVIDIA products and services, and asks you not to upload confidential information or personal data. For sensitive or production long-context workloads, choose another endpoint or model.
This difference matters when you switch between the two model IDs. A request that succeeds on the free endpoint may exceed the context limit of the standard endpoint. The free endpoint also lists tool calling but does not list response_format or structured outputs.
Check the model page instead of treating the architecture’s maximum context as a promise from every provider. We show the current limits and supported parameters for each endpoint on the model page.
Open weights and local deployment
NVIDIA publishes the BF16 reference checkpoint under the OpenMDW-1.1 license. The model card describes the BF16 release as a starting point for post-training, domain adaptation, research, and producing optimized variants. NVIDIA also publishes training data and recipes for customization and evaluation, and the model card says the release is ready for commercial use.
For direct inference, NVIDIA recommends its NVFP4 release. It also provides a GGUF checkpoint for supported local systems. The BF16 guidance lists an H100 80GB or A100 80GB for single-GPU deployment. The optimized releases target hardware such as the RTX 5090, RTX PRO 6000, and DGX Spark.
Running the model locally does not mean every machine can serve the full 1-million-token context window. In NVIDIA’s current Ollama guidance, the default context scales with available VRAM. It is 4K below 24GB, 32K from 24GB to below 48GB, and 256K with 48GB or more. The llama.cpp example uses about 40K. You can raise those limits, and doing so may require more memory or CPU offloading.
The weights are publicly available, and “open-weight” is the clearest description. Read the OpenMDW-1.1 license before you redistribute a modified model or make deployment decisions around its terms.
How to call Nemotron 3.5 Lightning on OpenRouter
If you do not want to provision GPUs or manage an inference engine, call Lightning through OpenRouter. The standard model ID keeps the same API while we route requests across the available providers for that model.
Install the OpenRouter TypeScript SDK.
npm install @openrouter/sdk
Set OPENROUTER_API_KEY in your environment, then send a request with the Lightning model ID. The standard model lists structured outputs, so you can ask for a JSON schema and have the response conform to it.
import { OpenRouter } from "@openrouter/sdk" ;
const openRouter = new OpenRouter ({
apiKey: process.env. OPENROUTER_API_KEY ,
});
const result = await openRouter.chat. send ({
chatRequest: {
model: "nvidia/nemotron-3.5-lightning" ,
messages: [
{
role: "user" ,
content:
'Ticket: "Checkout returns a 500 after I click Pay. Started this morning, three customers affected." Return its category and priority.' ,
},
],
responseFormat: {
type: "json_schema" ,
jsonSchema: {
name: "triage" ,
strict: true ,
schema: {
type: "object" ,
properties: {
category: { type: "string" },
priority: { type: "string" , enum: [ "low" , "medium" , "high" ] },
},
required: [ "category" , "priority" ],
additionalProperties: false ,
},
},
},
provider: {
requireParameters: true ,
},
},
});
if ( ! ( "choices" in result)) {
throw new Error ( "Expected a non-streaming response" );
}
console. log (result.choices[ 0 ]?.message.content);
When we ran this request on September 11, 2026, the model returned {"category": "Bug (Payment Checkout)", "priority": "medium"} . Sampled outputs vary between runs, and the schema, not the values, is what the request guarantees.
Provider support for response_format and structured outputs differs within the same model. By default we prefer providers that support the tools and response_format parameters you send, and a provider that does not support a parameter ignores it. Setting require_parameters to true , as the example does, restricts the request to providers that support every parameter in it. See the provider routing docs for the full behavior.
To try the free endpoint, change the model ID to nvidia/nemotron-3.5-lightning:free . The free endpoint does not list response_format , so drop that field there and validate the JSON yourself. It is useful for evaluation and low-volume experiments, and it has different limits and availability from the standard endpoint.
Do not submit confidential information or personal data through the free endpoint. Its model page carries NVIDIA’s notice that use is logged for security purposes and to improve NVIDIA products and services. Review that notice before deciding which prompts to send.
By default we load-balance the standard model across its providers, ordered by price. Two variants change that ordering. nvidia/nemotron-3.5-lightning:nitro sorts providers by throughput, and nvidia/nemotron-3.5-lightning:exacto prefers providers with stronger tool-calling quality signals. For a model chosen for latency and tool use, Nitro and Exacto are the two variants worth knowing.
Lightning accepts tools and tool_choice , so you can also use it inside an agent loop. See our tool-calling agent loop guide for the full request, tool execution, and iteration flow.
Should you use Nemotron 3.5 Lightning?
Nemotron 3.5 Lightning is worth evaluating when you need a fast, open-weight model for frequent, well-defined agent steps. Its 30B-A3B architecture, tool support, and low listed token price make it a candidate execution model for agents that would otherwise send every call to a much larger reasoning model.
Use the model name as a starting point, not the decision. Build an evaluation set from the tool calls and tasks your agent performs. Measure task success, retries, latency, and total cost across the full run. A cheaper call does not help if the agent repeats it or escalates the result often enough to erase the savings.
Start with nvidia/nemotron-3.5-lightning , or use nvidia/nemotron-3.5-lightning:free to test the model before adding paid traffic.
FAQ
What is Nemotron 3.5 Lightning?
Nemotron 3.5 Lightning is NVIDIA’s open-weight 30B mixture-of-experts language model with about 3B active parameters per token. NVIDIA positions it for high-volume agent execution, tool use, coding, instruction following, and specialized tasks.
Is Nemotron 3.5 Lightning the same as Nemotron 3 Ultra?
No. Nemotron 3.5 Lightning has 30B total parameters and 3B active parameters. Nemotron 3 Ultra has 550B total parameters and 55B active parameters. Lightning targets frequent execution steps. Ultra targets complex reasoning and orchestration.
What does 30B-A3B mean?
30B-A3B means the model has 30 billion parameters in total and activates about 3 billion parameters per token. The mixture-of-experts router selects a subset of the model’s experts for each token instead of running every expert.
The standard OpenRouter model, nvidia/nemotron-3.5-lightning , lists tools , tool_choice , response_format , and structured outputs among its supported parameters. Support varies by provider, so set require_parameters to true if your request depends on one of them. The free model lists tools and tool_choice but not response_format or structured outputs.
Is there a free Nemotron 3.5 Lightning API?
Yes. We list nvidia/nemotron-3.5-lightning:free , served by NVIDIA at no token cost. Free models have different rate limits and availability from paid models, and this endpoint carries an NVIDIA data notice. Do not send confidential information or personal data through it.
Can I run Nemotron 3.5 Lightning locally?
Yes. NVIDIA publishes BF16, NVFP4, and GGUF weights under the OpenMDW-1.1 license. The BF16 reference checkpoint targets a single 80GB H100 or A100. The
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-09-28 | 8.4 | 17 | 入选 |
| 2026-09-27 | 8.4 | 28 | 未入选 |
| 2026-09-25 | 8.77 | 46 | 未入选 |
| 2026-09-24 | 9.22 | 49 | 未入选 |
| 2026-09-23 | 9.95 | 36 | 未入选 |