2026年10月2日
产品
GPT‑6 系列模型指南
在管理时间和成本的同时,从 GPT‑6 模型中获得最佳效果的实用技巧
收听文章
8:15
分享
TL;DR(摘要)
1. 在生产环境中高效运行
为生产环境做好准备的工作流
将模型与工作负载相匹配
2. 调整提示词和技能
给模型清晰的指令
定义所需的输出
3. 优化长时间运行的任务
保持复杂工作的推进
API
Codex
利用计算机使用功能完成更多工作
从测试到生产:团队如何基于 GPT-6 Astra 进行构建
GPT‑6 是我们迄今为止最先进的模型套件,提供了针对不同工作类型的多种模型选择。
无论您是将想法转化为可工作的原型,构建和测试功能,还是在代码仓库、数据库和外部 API 之间编排多步工作流,本指南都解释了如何选择 GPT‑6 模型、给出有效指令、管理长时间运行的工作以及为生产环境做好准备。
TL;DR(摘要)
在生产环境中高效运行。使用缓存
(在新窗口打开)
和压缩
(在新窗口打开)
来管理上下文和成本。衡量任务成功率和延迟,并规划监控和数据控制措施。
将模型与工作负载相匹配。通过选择适合任务的模型、推理力度
(在新窗口打开)
和速度,在能力、成本和延迟之间取得平衡。
调整提示词和技能。保持提示词、技能和仓库指令的一致性,明确模型应交付的内容、可独立执行的操作以及被视为完成的标准。
让长时间运行的工作保持在正轨上。使用引导
(在新窗口打开)
、异步工具
(在新窗口打开)
和委托
(在新窗口打开)
来处理更新和独立工作。设定清晰的边界,规定模型何时应请求输入。
在部署之前,您需要落实几项检查措施和最佳实践。
保持效率意识。削减任务不需要的上下文
(在新窗口打开)
,同时保留必要的证据。如果您的应用支持,请并行运行独立任务
(在新窗口打开)
,以免一个缓慢的步骤阻碍无关工作的进行。
通过提示词缓存
(在新窗口打开)
重用共享上下文以处理重复性工作。根据模型不同,缓存的输入令牌成本比未缓存的输入令牌低多达 95%
(在新窗口打开)
。将稳定的指令和参考资料放在变化的任务细节之前,并保持工具定义的一致性。缓存仪表板
(在新窗口打开)
和诊断指南
(在新窗口打开)
可帮助您查看重用在何处失效。在估算完整工作流的成本时,请包含缓存写入量和任何长上下文费率。
对于较长的对话,压缩
(在新窗口打开)
可在保持继续运行所需状态的同时减少上下文大小。
决定如何监控行为
(在新窗口打开)
并审查应用的数据控制措施
(在新窗口打开)
。
部署前进行测试:运行具有代表性的任务,并衡量任务成功率、延迟以及每个成功任务的成本。查看我们的 API 部署检查清单
(在新窗口打开)
。
将模型与工作负载相匹配
将模型选择和推理水平视为智能/价格的权衡。
模型:
GPT‑6 Astra
(在新窗口打开)
用于需要最高智能的最复杂推理工作。
GPT‑6.1 Sol
(在新窗口打开)
用于复杂的编码、研究和计算机使用。
GPT‑6 Luna
(在新窗口打开)
用于大规模聚焦任务以及目标明确的日常重复性工作,例如提取发票字段、分类请求或生成结构化摘要。
在评估最适合任务的模型时,请比较各模型的定价
(在新窗口打开)
。
推理水平:在 API 中,选择模型为任务投入的精力程度。
低:常规任务,例如提取事实或进行小幅编辑。
中:需要判断力的工作,例如规划功能或比较选项。
高:困难的调试、更深入的分析或仔细的审查。
极高/最高:在“高”级别不足时,在支持的情况下进行测试,并且仅在改进足以证明所增加的时间和成本合理时才保留。
在 API 中,您可以在对话中途更改推理力度
(在新窗口中打开)
而不会破坏缓存。
在 Codex 中,从该模型的默认推理级别开始,然后针对更简单的任务降低它,或针对更深入的分析提高它。
速度:
在 API 中,当响应时间至关重要时(例如在聊天应用或编码工具中),请使用快速模式
(在新窗口中打开)
。与标准处理相比,它提供更快、更一致的响应时间,但每个令牌的成本更高。
在 Codex 和 API 中,当更快的响应值得支付溢价时(例如快速编码迭代),请使用超快模式
(在新窗口中打开)
。它独立于推理力度加速令牌生成。适用于 GPT‑6 Astra
(在新窗口中打开)
。
给模型一个清晰的指令
“模型在理解细微差别和歧义方面已经变得更好,因此过于具体的指导现在可能会阻碍结果,而以前它可能会有所帮助。”
(在新窗口中打开)
——Eric Provencher,OpenAI 开发者体验团队
从一个清晰的指令开始:你想要的结果、目标受众、相关背景和约束条件,以及什么算作完成。然后回顾以下四个领域,这些内容源自《重新思考 GPT‑6 Astra 的技能和提示》
(在新窗口中打开)
,以便更深入地了解如何更新您的指令:
创建更好的技能:保持描述简短,并明确说明每个技能何时运行,仅在需要时加载支持细节,并用适合团队所用模型的指导原则取代僵化的食谱。
更新您的 AGENTS.md:解释特定文档和测试何时相关,并明确授权安全的工作流程,例如使用一次性数据和本地测试,且无生产环境访问权限。
设定决策边界:说明哪些操作可以独立进行,哪些需要批准,用清晰的边界取代笼统的“始终询问”规则。
对持久性做出规定:定义“完成”包括什么——实施更改、运行它、检查结果以及修复失败——并确定任何需要您审查的决策。
如需更多指导,请参阅推理最佳实践
(在新窗口中打开)
。
定义您需要的输出
无论您是在 Codex 中工作还是使用 API 进行构建,都要指定模型可以做出哪些决策、何时应请求输入,以及什么样的回复是有用的。
给模型足够的方向以推动工作进展,而不必对重要的决策进行猜测。告诉它可以做出哪些选择,以及何时请求您的输入
(在新窗口中打开)
,例如,它可以选择如何组织摘要,但在更改项目范围之前应与您确认。描述有用的回复是什么样的
(在新窗口中打开)
,例如通俗语言、适合您受众的技术细节,以及简短的交接说明,涵盖已更改的内容、已检查的内容以及仍需关注的部分。
保持复杂工作进展
借助 GPT‑6 系列模型,您现在可以处理跨越数小时或数天的任务。使用以下功能更好地管理长时间运行任务中的智能体。
API
在 API 中,使用引导、异步工具和并行工作来保持长时间运行任务的进展。
在运行期间更新指令:中途转向
(在新窗口中打开)
允许您在模型工作时通过 Responses WebSocket API
(在新窗口中打开)
发送更正。更新会被排队;它们不会取消正在运行的工具或撤销已完成的动作。
在工具运行时继续工作:异步工具调用(在新窗口中打开)允许模型在您的应用程序运行较慢的任务(如测试)时继续独立工作。您的应用程序在准备好结果后返回该结果。在开始依赖于该结果的工作之前,请等待该结果。
委派独立的子任务:GPT‑6.1 Sol 支持 Responses API(在新窗口中打开)中的多智能体工作流。它可以将独立的工作分配给子智能体(例如调查代码库的不同部分),并将它们的发现结果整合到最终响应中。多智能体现在处于测试阶段。
Codex
长时间运行的任务可能会揭示您在初始提示中未必能预见到的决策。使用澄清和引导功能使工作保持在正确的轨道上。
在工作进展过程中回答问题:借助 GPT‑6 Astra,Codex 可以在工作时请求澄清(在新窗口中打开)。解决影响下一步的问题,并指定在您做出决定时可以继续进行的独立工作。如果您会离开,请告诉 Codex 哪些任务可以继续,以及何时应暂停以等待您的回答。
当需求发生变化时重定向工作:使用新信息引导当前任务(在新窗口中打开),解释哪些内容需要更改,哪些内容保持不变。这有助于避免在不再满足您需求的方法上花费更多时间。
利用计算机操作完成更多工作
计算机操作(在新窗口中打开)允许 GPT‑6 Astra、GPT‑6.1 Sol 和 GPT‑6 Luna 直接与网站和桌面应用程序交互,甚至包括没有 API 的应用程序。例如,您可以要求模型调查错误、修复代码,并在浏览器中打开您的产品以检查修复是否有效。
选择最简单可靠的方式完成每一步:
当 API 或连接的工具可以直接完成任务时,请使用它们。
当模型需要读取屏幕、点击按钮或为您填写表单时,使用计算机操作。
如果您正在将计算机操作集成到自己的应用程序中,请为模型提供一个可以运行代码以控制浏览器或桌面的工具。Playwright 适用于浏览器(在新窗口中打开);PyAutoGUI 适用于桌面应用程序。
从测试到生产:团队如何使用 GPT‑6 Astra 构建产品
1/4
Harvey:更多上下文,更有用的草稿。Harvey 结合法院信息、案例法、公司文件以及律师的偏好来定制其草稿。“我们可以向模型提供更多上下文,并生成结构更好、更好的输出,”联合创始人 Gabe Pereyra 表示。
Harvey
Cognition
Hex
Invideo
ChatGPT
2026
作者
OpenAI
继续阅读
查看全部
DevDay 2026 回顾
公司
2026年9月29日
介绍 GPT-6.1 Sol
产品
2026年9月29日
介绍 dots
产品
2026年9月29日
October 2, 2026
Product
A model guide for the GPT‑6 family
Practical tips for getting the best results from GPT‑6 models while managing time and cost
Listen to article
8:15
Share
TL;DR
1. Run effectively in production
Prepare your workflow for production
Match the model to the workload
2. Adjust your prompts and skills
Give the model a clear assignment
Define the output you need
3. Optimize long-running tasks
Keep complex work moving
API
Codex
Leverage computer use to do more of the job
From testing to production: How teams are building with GPT-6 Astra
GPT‑6 is our most advanced suite of models yet, and offers you a choice of models for different kinds of work.
Whether you’re turning an idea into a working prototype, building and testing a feature, or orchestrating multi-step workflows across code repositories, databases, and external APIs, this guide explains how to choose a GPT‑6 model, give it effective instructions, manage long-running work, and prepare for production.
TL;DR
Run effectively in production. Use caching
(opens in a new window)
and compaction
(opens in a new window)
to manage context and cost. Measure task success and latency, and plan for monitoring and data controls.
Match the model to your workload. Balance capability, cost, and latency by choosing the model, reasoning effort
(opens in a new window)
, and speed that fit the task.
Adjust your prompts and skills. Keep prompts, skills, and repository instructions consistent about what the model should deliver, what it can do independently, and what counts as done.
Keep long-running work on track. Use steering
(opens in a new window)
, async tools
(opens in a new window)
, and delegation
(opens in a new window)
to handle updates and independent work. Set clear boundaries for when the model should ask for input.
Before deploying, there are several checks and best practices you’ll want to put into place.
Keep efficiency in mind. Cut context the task doesn’t need
(opens in a new window)
while keeping the evidence it does. Where your application supports it, run independent tasks together
(opens in a new window)
so one slow step doesn’t hold up unrelated work.
Reuse shared context through prompt caching
(opens in a new window)
for recurring work. Cached input tokens cost up to 95% less
(opens in a new window)
than uncached input tokens, depending on the model. Put stable instructions and reference material before changing task details, and keep tool definitions consistent. The caching dashboard
(opens in a new window)
and diagnostics guide
(opens in a new window)
help you see where that reuse breaks down. Include cache writes and any long-context rates when estimating the cost of a complete workflow.
For longer conversations, compaction
(opens in a new window)
reduces context size while preserving the state needed to continue.
Decide how you’ll monitor behavior
(opens in a new window)
and review the data controls
(opens in a new window)
for your application.
Test before deploying: Run representative tasks and measure task success, latency, and cost per successful task. Check out our API deployment checklist
(opens in a new window)
.
Match the model to the workload
Think of the model choice and reasoning level as an intelligence/ price tradeoff.
Model:
GPT‑6 Astra
(opens in a new window)
for the hardest reasoning work where maximum intelligence is needed.
GPT‑6.1 Sol
(opens in a new window)
for complex coding, research, and computer use.
GPT‑6 Luna
(opens in a new window)
for focused tasks at scale and everyday, repeated work with a clear goal, such as extracting invoice fields, classifying requests, or producing structured summaries.
When evaluating the best model for the task, compare pricing
(opens in a new window)
for each model.
Reasoning level: In the API, choose how much effort the model spends on the task.
Low: Routine tasks, such as extracting facts or making small edits.
Medium: Work requiring judgment, such as planning a feature or comparing options.
High: Difficult debugging, deeper analysis, or careful review.
Extra high / Max: Test where supported when High falls short, and keep only if the improvement justifies the added time and cost.
In the API, you can change reasoning effort mid-conversation
(opens in a new window)
without breaking cache.
In Codex, start with the default reasoning level for that model, then lower it for simpler tasks or increase it for deeper analysis.
Speed:
In the API, use Fast mode
(opens in a new window)
when response time matters, such as in chat apps or coding tools. It provides faster, more consistent response times at a higher per-token cost than Standard processing.
In Codex and the API, use Ultrafast
(opens in a new window)
when faster responses are worth the premium, such as rapid coding iterations. It speeds up token generation independently of reasoning effort. Available for GPT‑6 Astra
(opens in a new window)
.
“Models have gotten much better at understanding nuance and ambiguity, so overly specific guidance can now hinder results where it previously helped.”
(opens in a new window)
—Eric Provencher, Developer Experience at OpenAI
Start with a clear assignment: the result you want, who it’s for, the relevant context and constraints, and what counts as done. Then review these four areas, summarized from Rethinking skills and prompts for GPT‑6 Astra
(opens in a new window)
, for a deeper dive into updating your instructions:
Create better skills: Keep descriptions short and explicit about when each skill should run, load supporting details only when needed, and replace rigid recipes with guidance suited to the models your team uses.
Update your AGENTS.md: Explain when particular documents and tests are relevant, and explicitly authorize safe routine workflows, such as running local tests with disposable data and no production access.
Set decision boundaries: State which actions can proceed independently and which require approval, replacing blanket “always ask” rules with clear boundaries.
Be prescriptive about persistence: Define what “done” includes—implementing the change, running it, inspecting the result, and fixing failures—and identify any decisions that require your review.
For additional guidance, see reasoning best practices
(opens in a new window)
.
Define the output you need
Whether you’re working in Codex or building with the API, specify which decisions the model can make, when it should ask for input, and what a useful response looks like.
Give the model enough direction to keep work moving without guessing at decisions that matter. Tell it which choices it can make and when to ask for your input
(opens in a new window)
, for example, it can choose how to organize a summary, but should check with you before changing the project’s scope. Describe what a useful response looks like
(opens in a new window)
, for example plain language, technical detail suited to your audience, and a short handoff covering what changed, what was checked, and what still needs attention.
With the GPT‑6 family of models, you can now take on tasks that span hours or days. Use the following features to better manage agents on long-running tasks.
API
In the API, use steering, asynchronous tools, and parallel work to keep long-running tasks moving.
Update instructions during a run: Mid-turn steering
(opens in a new window)
lets you send a correction through the Responses WebSocket API
(opens in a new window)
while the model works. Updates are queued; they don’t cancel running tools or undo completed actions.
Keep working while a tool runs: Asynchronous tool calling
(opens in a new window)
lets the model continue independent work while your app runs a slower task, such as tests. Your app returns the result when it’s ready. Wait for that result before starting work that depends on it.
Delegate independent subtasks: GPT‑6.1 Sol supports multi-agent workflows in the Responses API
(opens in a new window)
. It can assign independent work to subagents (such as investigating different parts of a codebase) and combine their findings into a final response. Multi-agent is currently in beta.
Codex
Long-running tasks can uncover decisions you wouldn’t necessarily anticipate in the initial prompt. Use clarification and steering to keep the work on course.
Answer questions as work progresses: With GPT‑6 Astra, Codex can ask for clarification while it works
(opens in a new window)
. Resolve questions that affect the next step, and specify which independent work can continue while you decide. If you’ll be away, tell Codex which tasks can continue and when it should pause for your answer.
Redirect work when requirements change: Steer the active task
(opens in a new window)
with new information, explaining what should change and what should stay the same. This helps avoid spending more time on an approach that no longer meets your needs.
Leverage computer use to do more of the job
Computer use
(opens in a new window)
lets GPT‑6 Astra, GPT‑6.1 Sol, and GPT‑6 Luna interact directly with websites and desktop apps, even applications without an API. For example, you can ask the model to investigate a bug, fix the code, and open your product in a browser to check that the fix works.
Choose the simplest reliable way to do each step:
Use an API or connected tool when it can do the job directly.
Use computer use when the model needs to read a screen, click buttons, or fill in a form for you.
If you’re building computer use into your own app, give the model a tool that can run code to control a browser or desktop. Playwright works with browsers
(opens in a new window)
; PyAutoGUI works with desktop apps.
From testing to production: How teams are building with GPT‑6 Astra
1 of 4
Harvey: more context, more useful drafts. Harvey combines court information, case law, firm documents, and a lawyer’s preferences to tailor its drafts. “We can give more context to the model and produce better and better structured outputs,” says cofounder Gabe Pereyra.
Harvey
Cognition
Hex
Invideo
ChatGPT
2026
Author
OpenAI
Keep reading
View all
DevDay 2026 Recap
Company
Sep 29, 2026
Introducing GPT-6.1 Sol
Product
Sep 29, 2026
Introducing dots
Product
Sep 29, 2026
首次收录 · 2026-10-03 · 12.66 分