对于大规模的工作负载或无需立即完成的请求,您现在可以使用新的批量 API(Batch API)。在提交批量任务时,提供商可以选择在 24 小时窗口内的任何时间完成请求,作为交换,他们通常收取正常按 token 定价的 50%(有时甚至更低)。
目前该功能支持超过 70 种模型。请参阅 Batch API 文档了解如何使用它。
在实际操作中,我们观察到您很少需要等待接近 24 小时的时间。在我们为期两周的测试期间完成的 23 万多个批量任务中,中位完成时间为 7 分钟,90% 的任务在一小时内完成。
Batch 是一种用于工作负载的异步 API,适用于您可以接受高度可变响应时间的场景,例如标注语料库、回填嵌入向量、评分评估集、总结积压工单,或在夜间对几千行数据运行相同的提示词。
批量 API 的工作原理
调用 api/v1/batches 并传入您的请求列表,同时指定要使用的端点格式。Chat completions、responses、messages 和 embeddings 均受支持。例如:
curl https://openrouter.ai/api/v1/batches \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENROUTER_API_KEY " \
-d '{
"endpoint": "/v1/chat/completions",
"model": "google/gemini-3.8-flash",
"requests": [
{ "custom_id": "ticket-0001", "body": { "messages": [{ "role": "user", "content": "Summarize this ticket in one sentence." }] } },
{ "custom_id": "ticket-0002", "body": { "messages": [{ "role": "user", "content": "Summarize this ticket in one sentence." }] } }
]
}'
提交请求后,开始轮询 GET /api/v1/batches/:id,直到状态变为 completed(已完成)、failed(失败)、expired(过期)或 cancelled(已取消)。已完成的批量任务会在同一响应中内联返回其结果。
大多数批量任务在几分钟内完成
我们分析了在为期两周的测试期间完成的批量任务,并测量了从接受到结果完成的时间。中位数为 7 分钟,第 90 百分位数为 1.0 小时,第 99 百分位数为 10.3 小时。我们还发现,提交请求的时间段比发送的请求数量对速度影响更大。
在太平洋时间上午 5 点到中午之间提交的批量任务明显比其他时间段慢。最慢的 10% 需要 2 到 4.5 小时。如果在其他任何时段提交,第 90 百分位数将降至 1.1 小时以下,而在太平洋时间晚上 6 点之后则低于 50 分钟。
单个请求的批量任务根据时段不同,完成时间为 5 到 11 分钟;包含 1,000 个或更多请求的批量任务完成时间为 12 到 21 分钟。大型批量任务确实需要更长的处理时间。在太平洋时间午夜至中午之间提交的、包含超过 100 个请求的最慢 10% 的批量任务,耗时长达 6.8 小时。
批量功能支持的内容
请求格式:您已发送给 OpenRouter 的任何文本请求体格式都将得到支持,包括 chat completions、responses、messages 和 embeddings。
路由:每个批量任务将在单个提供商上执行。默认情况下,在应用您的提供商白名单、数据策略和 BYOK(自带密钥)设置后,我们会为模型选择最便宜的批量端点。
BYOK:配置了提供商密钥后,支持该功能的提供商上的批量任务将通过您的密钥路由,您只需支付 BYOK 费用。
逐请求结果:每个结果独立返回,因此少数错误数据行永远不会导致整个作业失败。
保留期:输入和结果将保留 30 天,直到您 DELETE(删除)该批量任务为止。
日志记录:每个批量任务都会出现在您的日志的 Batches 标签页中,包含模型、提供商、状态和成本信息。
定价:折扣适用于按 token 定价,并因模型而异。Web search(网页搜索)调用按标准费率计费。
输入要求:图片和文件必须是公开 URL。音频、视频以及 OpenRouter 自带的网页搜索插件在批量功能中不可用(请参阅文档中的限制部分)。
开始使用
选择一个支持批量的模型,获取 API 密钥,并将您的第一个批量任务发布到 https://openrouter.ai/api/v1/batches 。快速入门指南提供了完整的请求和响应格式。
请在Discord的#feedback频道告诉我们你正在批量处理什么。
For large batches of work or requests that don’t need to be completed immediately, you can now use the new batch API . When submitting a batch, a provider gets to choose when during a 24 hour window they will complete the request, and in exchange they generally charge 50% (and sometimes less) of their normal per-token price.
It works today on more than 70 models . Learn how to use it in the Batch API docs .
In practice, we’ve observed you rarely wait anywhere near 24 hours. Across 230k+ batches that completed over our two week beta period, the median finished in 7 minutes and 90% finished within an hour.
Batch is an asynchronous API for workloads where you can accept highly variable response times, such as labeling a corpus, back-filling embeddings, scoring an eval set, summarizing a backlog of tickets, or running the same prompt across a few thousand rows overnight.
How the Batch API works
Call api/v1/batches with your list of requests and specify what endpoint shape to use. Chat completions, responses, messages, and embeddings are all supported. For example:
curl https://openrouter.ai/api/v1/batches \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENROUTER_API_KEY " \
-d '{
"endpoint": "/v1/chat/completions",
"model": "google/gemini-3.8-flash",
"requests": [
{ "custom_id": "ticket-0001", "body": { "messages": [{ "role": "user", "content": "Summarize this ticket in one sentence." }] } },
{ "custom_id": "ticket-0002", "body": { "messages": [{ "role": "user", "content": "Summarize this ticket in one sentence." }] } }
]
}'
After the request, begin polling GET /api/v1/batches/:id until the status is completed , failed , expired , or cancelled . Completed batches return their results inline in the same response.
Most batches finish in minutes
We looked at the batches that completed during our two week beta period and measured the time from acceptance to finished results. The median was 7 minutes, the 90th percentile was 1.0 hour, and the 99th percentile was 10.3 hours. We also found that the time of day you submit matters more than how many requests you send.
Batches submitted between 5am and noon Pacific are significantly slower than other times of day. The slowest tenth take 2 to 4.5 hours. Submit at any other hour and the 90th percentile drops under 1.1 hours, and after 6pm Pacific it’s under 50 minutes.
A single-request batch finishes in 5 to 11 minutes depending on the hour, and a batch of 1,000 or more requests finishes in 12 to 21 minutes. Large batches do take longer to process. The slowest tenth of batches with more than 100 requests submitted between midnight and noon Pacific took as long as 6.8 hours.
What’s supported with batches
Request shapes: any text request body shape you already send to OpenRouter will be supported, including chat completions, responses, messages, and embeddings.
Routing: Each batch is executed on a single provider. By default we choose the cheapest batch endpoint for the model after your provider allowlist, data policy, and BYOK settings are applied.
BYOK: with a provider key configured, batches on providers that support it route through your key and you pay only the BYOK fee.
Per-request results: every result returns independently, so a few bad rows never fail the rest of the job.
Retention: inputs and results are kept for 30 days, or until you DELETE the batch.
Logging: every batch shows up in the Batches tab of your logs with model, provider, status, and cost.
Pricing: the discount applies to per-token pricing and varies by model. Web search calls bill at standard rates.
Inputs: images and files must be public URLs. Audio, video, and OpenRouter’s own web search plugin aren’t available in batch (see the limitations section of the docs ).
Get started
Pick a batch-capable model , grab an API key , and post your first batch to https://openrouter.ai/api/v1/batches . The Quickstart has the full request and response shapes.
Tell us what you’re batching in #feedback on Discord.
首次收录 · 2026-09-23 · 9.95 分