2026年9月22日
产品
GPT‑6 的提示词缓存优化
更高的缓存命中率,以及帮助持久化智能体运行更快、成本更低的新工具。
加载中……
分享
监控缓存并诊断缓存未命中
为您的应用优化缓存
开始使用
GPT‑6 使持久化智能体能够在复杂任务上持续工作数小时,从重构代码库到生成经过充分研究的文档和演示文稿。这些智能体背后的应用程序会发起一系列相互关联的 API 请求,通常会将前几轮交互中的相同指令、工具定义和上下文延续下来。OpenAI 会对这些共享上下文进行缓存,以便在请求间复用计算,从而缩短响应时间,并为开发者提供高达 90% 的缓存输入令牌折扣。
随着 GPT‑6 系列的发布,我们推出了改进后的提示词缓存系统,默认情况下可提供更高的缓存命中率。我们现在对符合资格的共享前缀在 30 分钟窗口内被重用时给予缓存折扣。我们还引入了新工具,帮助开发者监控缓存性能、诊断未命中原因,并决定对提示词的哪一部分进行缓存。
“
OpenAI 的提示词缓存在帮助 GitHub Copilot 大规模提供快速、高效的体验方面发挥着关键作用。在过去的几个月里,与之前的基线相比,我们在数十亿次向 OpenAI 模型发出的请求中,将需要重新处理的提示词令牌比例降低了 50% 以上。其结果是推理堆栈更加高效,开发者获得首次响应的时间也更短。”
——Mario Rodriguez,首席产品官
监控缓存并诊断缓存未命中
新的提示词缓存仪表板
(在新窗口中打开)
显示您的应用程序输入中有多少内容是通过缓存提供的。跟踪随时间变化的命中率,并使用输入组成图表比较缓存和未缓存的令牌。这些视图可帮助您发现缓存命中率的下降,并评估对应用程序所做的更改如何影响缓存性能。
当您遇到意外的缓存未命中时,请使用提示词缓存诊断工具
(在新窗口中打开)
来了解发生的情况。将请求与最近的响应进行比较,以识别导致无法复用的模型、工具、设置或输入方面的变更。受影响的令牌估算数量可帮助您评估影响范围,并决定如何优化集成以最大化缓存命中率。
{
"prompt_cache_diagnostics": {
"type": "cache_miss",
"reason": "tools_changed",
"comparison_reusable_tokens": 5629,
"cache_missed_tokens": 5629
}
}
为您的应用优化缓存
选择要缓存的内容。显式缓存断点让您可以选择复用哪些提示词前缀。更新的提示词缓存指南
(在新窗口中打开)
解释了如何使用它们、缓存的前缀保持有效的时间长度,以及工具和输入的更改如何影响复用。
在不破坏缓存的情况下调整推理力度。在 GPT‑6 模型上,您现在可以在响应之间更改推理力度
(在新窗口中打开)
而不会破坏缓存。对于更复杂的任务可以提高力度,对于常规后续步骤可以降低力度,方法是附加一个 configuration_update,同时保持请求级别的推理力度不变。这使您能够调整任务所需的推理量,同时保留可复用的上下文。
在工具和指令变更时保持缓存有效。随着智能体的工具使用需求发生变化,请保持工具定义、模式(schemas)和顺序的稳定,以便之前的上下文保持可复用状态。使用 allowed_tools 仅使相关工具可被调用,或者在没有需要工具时将 tool_choice 设置为 none,而不是移除定义。使用新的开发者消息在上下文的末尾附加新指令以覆盖旧指令。请参阅我们关于管理工具更改的指南
(在新窗口中打开)
。
预热缓存以降低延迟。预热(在新窗口中打开)可提前准备已知的上下文,以便在请求到达时模型能够更快地开始响应。例如,应用程序可以在启动期间、用户提出第一个问题之前,预热共享指令、工具定义或参考资料。这将处理过程从用户的等待时间中移出。
这些可选控制功能建立在引擎的默认性能之上,帮助您根据工作负载定制缓存策略。
1 of 3
“
OpenAI 的提示词缓存诊断和仪表板帮助我们提高了几个百分点的缓存命中率,从而降低了 20% 的成本。当缓存意外中断时,我们会收到警报,并使用 Codex 代理来诊断根本原因。明确的断点还允许我们缓存稳定的上下文,同时让频繁更改的内容位于提示词的末尾。这使得在重用几乎所有共享上下文的同时为后台任务分叉对话在经济上变得可行。”
——Arian Hanifi,首席技术官
Strawberry Browser
Manus
Wordsmith
Strawberry Browser
Manus
Wordsmith
开始使用
在提示词缓存仪表板(在新窗口中打开)中监控缓存命中率。
使用诊断工具(在新窗口中打开)调查意外的未命中情况。
遵循提示词缓存指南(在新窗口中打开)以改进您的设置,或使用 Codex(在新窗口中打开)来审查代码、应用改进并衡量结果。
API
2026 年
作者
OpenAI
继续阅读
查看全部
推出 GPT-6 Sol 和 Luna
产品
2026 年 9 月 22 日
用 AI 重塑广告
产品
2026 年 9 月 16 日
如何将 AI 使用与业务价值联系起来
产品
2026 年 9 月 16 日
September 22, 2026
Product
Better prompt caching for GPT‑6
Higher cache hit rates and new tools to help persistent agents run faster and cost less.
Loading…
Share
Monitor caching and diagnose cache misses
Optimize caching for your application
Get started
GPT‑6 enables persistent agents to work for hours on complex tasks, from refactoring codebases to producing well-researched documents and presentations. The applications behind these agents make a series of API requests that build on one another, often carrying forward the same instructions, tool definitions, and context from earlier turns. OpenAI caches that shared context to reuse computation across requests, reducing response times and giving developers discounts of up to 90% on cached input tokens.
With the GPT‑6 family, we launched an improved prompt caching system that delivers higher cache hit rates by default. We now give cache discounts for eligible shared prefixes reused within a 30-minute window. We’re also introducing new tools to help developers monitor cache performance, diagnose misses, and choose how much of a prompt to cache.
“
OpenAI’s prompt caching plays a critical role in helping GitHub Copilot deliver fast, efficient experiences at scale. Over the past several months, we’ve reduced by more than 50% the share of prompt tokens requiring fresh processing across billions of requests to OpenAI models, relative to our previous baseline. The result is a more efficient inference stack and faster time to first response for developers.”
—Mario Rodriguez, Chief Product Officer
Monitor caching and diagnose cache misses
The new Prompt Caching Dashboard
(opens in a new window)
shows how much of your application’s input is served from cache. Track hit rates over time and use the input composition chart to compare cached and uncached tokens. These views help you spot drops in cache hits and evaluate how changes to your application impact caching performance.
When you see an unexpected cache miss, use the prompt caching diagnostics tool
(opens in a new window)
to understand what happened. Compare a request with a recent response to identify changes to the model, tools, settings, or input that prevented reuse. The estimated number of affected tokens helps you assess the size of the impact and decide how you can optimize your integration to maximize cache hit rates.
{
"prompt_cache_diagnostics": {
"type": "cache_miss",
"reason": "tools_changed",
"comparison_reusable_tokens": 5629,
"cache_missed_tokens": 5629
}
}
Optimize caching for your application
Choose what to cache. Explicit cache breakpoints let you choose which prompt prefixes to reuse. The refreshed prompt caching guide
(opens in a new window)
explains how to use them, how long cached prefixes remain eligible, and how changes to tools and inputs affect reuse.
Adjust reasoning effort without breaking cache. On GPT‑6 models, you can now change reasoning effort
(opens in a new window)
between responses without breaking cache. Raise effort for a harder task or lower it for a routine follow-up by appending a configuration_update while leaving request-level reasoning effort unchanged. This lets you adjust how much reasoning a task needs while preserving reusable context.
Preserve cache as tools and instructions change. As your agent’s tool use needs change, keep tool definitions, schemas, and ordering stable so earlier context stays reusable. Use allowed_tools to make only the relevant tools callable, or set tool_choice to none when no tools are needed, instead of removing definitions. Use new developer messages to append new instructions towards the end of the context to override older ones. See our guidance on managing tool changes
(opens in a new window)
.
Prewarm the cache to reduce latency. Prewarming
(opens in a new window)
prepares known context ahead of time so the model can start responding sooner when a request arrives. For example, an application can prewarm shared instructions, tool definitions, or reference material during startup, before the user asks their first question. This moves processing out of the user’s wait time.
These optional controls build on the engine’s default performance, helping you tailor caching to your workload.
1 of 3
“
OpenAI’s prompt caching diagnostics and dashboard helped us improve cache hit rates by a few percentage points, reducing costs by 20%. We now get alerts when caching breaks unexpectedly and use Codex agents to diagnose the root cause. Explicit breakpoints also let us cache stable context while keeping frequently changing content at the end of the prompt. That’s made it economically viable to fork conversations for background tasks while reusing nearly all of the shared context.”
—Arian Hanifi, Chief Technology Officer
Strawberry Browser
Manus
Wordsmith
Strawberry Browser
Manus
Wordsmith
Get started
Monitor cache hit rates in the Prompt Caching Dashboard
(opens in a new window)
.
Investigate unexpected misses with the diagnostics tool
(opens in a new window)
.
Follow the prompt caching guide
(opens in a new window)
to improve your setup, or use Codex
(opens in a new window)
to review your code, apply improvements, and measure results.
API
2026
Author
OpenAI
Keep reading
View all
Introducing GPT-6 Sol and Luna
Product
Sep 22, 2026
Reimagining advertising with AI
Product
Sep 16, 2026
How to connect AI usage to business value
Product
Sep 16, 2026
首次收录 · 2026-09-23 · 12.91 分