企业需要能够在其自身环境中良好运行的智能体。它们要求这些智能体完成的工作,受到其所用系统、遵循的规则以及数据状态的影响。一个模型可能具备广泛的能力,但仍可能在特定环境中表现不佳:例如处理工作流不力、误用工具组合,或未能遵守某些约束。这些正是企业需要改进的弱点。
难点在于将这些弱点转化为训练数据。单次失败能告诉我们一些信息,但训练模型需要大量新任务,在不同情境下锻炼相同的能力。这些任务还必须能够在该环境中完成, resemble 人们实际会提出的请求,并且拥有可靠的方法来检查智能体是否成功。
在 ServiceNow CoreAI,我们构建了 AutoSynthData,将这些能力差距转化为训练数据。它利用目标模型的失败和更强教师模型的成功,来决定模型下一步应学习什么,然后生成并验证能够锻炼这些能力的新任务。随着模型的改进,课程重心会转向它仍然感到困难的部分。我们通过 EnterpriseOps Gym(Malay 等人,2026)及其发布的数据集来说明这一流程。我们首先描述智能体所处的环境,以及什么样的任务对训练有用。
什么构成了有用的智能体任务?
智能体环境定义了智能体运作的“世界”:它可以观察和修改的状态、可以调用的工具和 API,以及由其动作产生的状态转换。
任务是在此环境中实例化的。我们使用以下抽象定义:
task = (系统规范, 用户提示词, 验证器)
系统规范
系统规范定义了智能体运作的约束条件,包括系统指令、环境策略,以及在适用情况下的特定于任务的初始化,例如种子数据库状态或一组知识文章。
该规范必须与环境的工具、状态和支持的操作兼容。其指令应清晰明了,避免引入仅为了制造难度而设定的任意约束。
面向智能体的任务
用户提示词指定了用户希望智能体完成的目标,以及任何用户层面的约束。生成的任务应满足三个属性。
可行性。在当前环境中,至少应存在一条轨迹能够满足用户提示词并遵守系统规范。这排除了那些依赖于不可用工具、无法访问的知识、不可能的状态转换或受策略禁止的动作的任务。
真实性。用户提示词应类似于用户在目标环境中可能合理提出的请求。可执行行为的范围通常远大于真实工作流的范围。
难度。对于训练而言,任务应暴露当前智能体的弱点。那些已经被可靠解决的任务提供的训练信号很少。因此,有用的区域是那些可行且真实,但尚未被一致解决的任务。
验证器
验证器决定生成的轨迹是否成功完成了任务。它应满足三个属性。
一致性。它应与用户提示词、系统规范以及特定于任务的環境状态保持一致。
健全性。它应拒绝未能满足任务要求或违反相关约束的轨迹。
完备性。它应接受有效的解决方案,而不是编码某一条特定的参考轨迹。
这些属性在训练过程中至关重要。宽松的验证器可能会奖励错误的行为,而过于严格的验证器则可能惩罚有效的解决方案。
概述
给定一个环境和目标模型,AutoSynthData 会生成由系统规范、用户提示词和验证器组成的训练任务。生成的任务以环境为基础,并经过筛选,以为当前模型提供有用的训练信号。
AutoSynthData 首先使用诊断任务在环境中评估目标模型,并识别出它在难以完成的任务中呈现出的模式。一个更强的教师模型有助于刻画哪些任务是可解决的,以及成功的行为表现为何种形态。AutoSynthData 将由此产生的能力差距转化为新的可执行任务,在环境中检查每个任务,并使用被接受的样本来进行后训练。对更新后的模型进行评估可以揭示剩余的能力差距,并指导下一轮的任务生成。
从模型失败到课程学习
AutoSynthData 利用目标环境中的评估运行来识别模型下一步需要学习的内容。在我们的 EnterpriseOps Gym 实验中,我们在评估任务上同时运行目标模型和更强的教师模型。我们检查这些运行结果以识别:
所测试的能力;
涉及的工具和工作流结构;
目标模型失败之处以及教师模型成功之处;
正确最终状态必须满足的属性;
在保持所测试能力不变的情况下可以变化的维度。
我们将这些发现提炼为经过脱敏处理的能力规范卡片。评估任务指导模型应该学习什么,但生成器不会接收其原始提示词、实体、轨迹或验证器细节。它接收的是这些卡片,并利用它们创建具有不同提示词、状态和解决路径的新任务。
生成与扩展任务
识别出能力差距告诉我们要教什么,但训练需要大量多样化的任务来锻炼该能力。AutoSynthData 使用规范卡片来生成这些任务。
假设目标模型在需要以下工作流的任务中表现不佳:
生成器会创建执行该工作流的新任务,并变化实体、初始环境状态、工作流组成、工具组合、措辞和难度。随后,更强的教师模型为每个任务演示一条成功的轨迹。对于监督微调(SFT),这些演示教导目标模型如何在新情境中应用该能力。
AutoSynthData 分两个阶段构建数据集:首先生成并验证核心样本,然后将其扩展为新颖的变体。
目标阶段
目标阶段从能力规范中创建核心训练样本集。工作人员并行生成独立的任务,完成一个后便领取一个新的任务。每个候选样本都要经过验证、执行、求解器评估和修复,之后才会被接受。结果是一批围绕目标模型需要学习的内容而精心筛选的示例。
乘数阶段
乘数阶段通过为已接受的目标样本创建新颖变体来扩展数据集。每个变体都有自己独立的用户请求、环境状态、实体配置、参考轨迹和验证器,并且必须通过相同的验证和执行检查。一个被乘数的样本不能作为另一个被乘数样本的种子。这将扩展过程锚定在经过验证的目标集上,并限制跨代生成的漂移。
实现细节
为了支持这两个阶段,AutoSynthData 将生成控制与特定环境的执行分离开来。共享控制器协调生成、质量控制、覆盖率和数据集构建,而适配器则处理环境执行、任务和状态管理、参考重放、确定性验证、求解器执行和任务分析。
并行目标生成与乘数扩展相结合,为达到训练规模的数据集提供了一条路径。其有效性取决于应用于每个候选样本的检查:任务必须可执行,解决方案必须有效,且验证器必须能够区分成功与失败。
高质量合成数据不仅仅需要生成
仅仅生成一个看似合理的请求不足以产生有用的训练数据。任务可能在目标环境中无法执行,其参考解决方案在执行时可能失败,或者其验证器可能会奖励错误的最终状态。AutoSynthData 在接受任务用于训练之前会检查这些属性。
AutoSynthData 在两个层面上审查质量:单个候选者必须通过验证,而批次必须提供有用的覆盖范围和多样性。
样本级验证与修复
每个候选者在进入训练数据集之前都必须通过质量控制循环。我们从求解器评估开始以衡量难度。在此使用的配置中,我们倾向于选择目标模型在三次尝试中最多解决一次、且更强求解器至少解决两次的任务。候选者还要经过正向和负向验证以及有界修复过程。
正向验证
正向关卡问的是:预期的解决方案是否解决了生成的任务?
管道在目标环境中执行参考轨迹,并将结果状态与候选者的验证器进行检查。这揭示了提示词、初始状态、解决方案和成功标准之间的不匹配。
负向验证
负向关卡问的是:相关的错误结果是否失败?
例如,它可以突变预期结果的某些部分,并确认这些状态不再通过验证。这能捕捉到那些在不要求预期行为的情况下就奖励成功的弱验证器。
批评与修复
失败的候选者在被丢弃之前会经过批评者审查。批评者检查样本及其失败原因,寻找不一致的状态、不可能的工作流、错误的任务构建、糟糕的参考轨迹、薄弱的验证器逻辑或与预期能力的不匹配。批评者的发现指导修复工作,重试次数有固定限制:
候选者
↓
失败
↓
批评/诊断
↓
针对性修复
↓
再次运行关卡检查
↓
接受或重试
修复后的任务必须再次通过相关检查。诊断指导对现有候选者的修复,而不是要求从头开始生成。
通过这些检查使样本具备训练资格,但 individually valid 的样本仍可能形成重复或不平衡的数据集。因此,AutoSynthData 还在批次层面审查生成过程。
批次级审查
一个批次可能会过度代表少数简单的任务家族,遗漏某种能力,或者反映出在低产模式上花费了过多的生成精力。
元审查检查每个批次中被接受的样本、被拒绝的样本以及生成行为。它问:
哪些任务家族过度代表,哪些能力维度缺失?
是否反复出现相同类型的示例?
是否有特定目标持续无法生成成功结果?
批评中是否出现了系统性问题?
下一批次的指导方针应如何调整?
控制器跟踪已接受数据集的覆盖范围,减少在过度代表区域的生成工作,并将更多精力引向空白区域。当一个区域反复产生糟糕的候选者时,批评和元审查会指导生成策略的改变。这些调整在可用的生成预算和数据集规模要求内,平衡有用的学习信号、任务质量、覆盖范围、多样性和低冗余性。
这些反馈循环共同提高了单个任务和它们组成的数据集的质量:样本级检查指导候选者修复,而批次级审查指导未来的生成。
推动训练前沿
随着模型的改进,有用的训练分布也会发生变化。AutoSynthData 将合成数据生成视为在目标模型能力边界附近寻找任务的过程:难度足以暴露弱点,但又足够可解以便教师提供可靠的演示。
在完成后训练(post-training)后,我们在相同的环境中评估更新后的模型。现在能够可靠解决的任務对下一轮训练的效用较低;而持续存在的失败则指出了仍需关注的薄弱环节。这些结果可以指导下一代模型的迭代方向。
我们的实验主要侧重于监督微调(SFT),但相同的机制也可以支持强化学习(RL):生成挑战当前策略的任务并提供可靠的信号,进行训练,然后随更新后的策略调整生成目标。我们计划在这一超出 SFT 范围的、难度动态校准的前沿领域进行测试。
EnterpriseOps Gym 实验
我们使用 EnterpriseOps Gym 来测试这种方法是否能在状态化企业环境中提升模型在特定任务上的表现。我们在 Gym 的 Hybrid 和 ITSM 环境中生成训练任务,在接受的样本上微调目标模型,并对生成的检查点进行评估。
Hybrid
我们在 EnterpriseOps Gym 的 Hybrid 领域对该管道进行了测试,使用 Gemma-4-26B-A4B-it 作为目标模型,Qwen3.8-27B 作为教师模型。
AutoSynthData 在约 18 小时内生成了 2,000 个合成训练样本。我们在此数据集上对 Gemma 进行了微调,并在基准测试中评估了生成的检查点。最佳检查点出现在第 5 轮(epoch 5)。
Hybrid 结果
合成 SFT 检查点将平均 Pass@1 提升了 7.2 个百分点,实现了 35% 的相对提升,并将验证器成功率从 63.01% 提高至 68.55%。它缩小了 Gemma 与参考模型之间原始 Pass@1 差距的 59%。
这些训练任务是根据能力规范全新生成的;生成器并未接收原始评估任务。这一结果证明了该方法在 EnterpriseOps Gym Hybrid(本实验所用的环境)中的改进效果。
ITSM
我们还将 AutoSynthData 应用于 EnterpriseOps Gym 的 ITSM 领域,使用 Gemma-4-26B-A4B-it 作为目标模型,DeepSeek-V4.1-Flash 作为教师模型。AutoSynthData 在 66 小时内生成了 1,994 个合成训练样本。生成时间比上述随后的 Hybrid 运行更长,主要是因为 ITSM 运行使用了更大的教师模型,且发生在提升吞吐量的管道优化之前。
在 ITSM 领域,合成 SFT 将平均 Pass@1 从 18.77% 提升至 27.18%,表明该方法也能在第二个领域中提升性能。
闭环
对训练最有用的任务取决于环境以及在其中工作的模型。AutoSynthData 利用模型的失败来选择生成内容,根据环境验证新任务,并使这些任务可用于后训练。我们的 EnterpriseOps Gym 结果展示了该方法在受控环境中的价值。随着模型的演变,同样的流程可以专注于剩余的能力差距。
Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve.
The difficulty is turning those weaknesses into training data. An individual failure tells us something, but training a model requires many new tasks that exercise the same capability in different situations. Those tasks must also be possible to complete in the environment, resemble work someone would actually request, and have a reliable way to check whether the agent succeeded.
At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data. It uses a target model’s failures and a stronger teacher’s successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities. As the model improves, the curriculum shifts toward what it still finds difficult. We illustrate the pipeline with EnterpriseOps Gym ( Malay et al., 2026 ), using the released dataset . We begin by describing the environment an agent operates in and what makes a task useful for training.
What makes a useful agentic task?
An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can invoke, and the state transitions produced by its actions.
A task is instantiated within this environment. We use the following abstraction:
task = (system specification, user prompt, verifier)
System specification
The system specification defines the constraints under which the agent operates, including system instructions, environment policies, and, when applicable, task-specific initialization such as a seeded database state or a set of knowledge articles.
The specification must be compatible with the environment’s tools, state, and supported actions. Its instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty.
Agent-facing task
The user prompt specifies what the user wants the agent to accomplish, together with any user-level constraints. A generated task should satisfy three properties.
Feasibility. There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification. This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.
Realism. The user prompt should resemble something a user would plausibly ask in the target environment. The space of executable behaviors is usually much larger than the space of realistic workflows.
Difficulty. For training, the task should expose a weakness of the current agent. Tasks that are already solved reliably provide little new training signal. The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved.
Verifier
The verifier determines whether the resulting trajectory successfully completes the task. It should satisfy three properties.
Consistency. It should agree with the user prompt, the system specification, and the task-specific environment state.
Soundness. It should reject trajectories that fail to satisfy the task or violate relevant constraints.
Completeness. It should accept valid solutions rather than encode one particular reference trajectory.
These properties matter directly during training. A lax verifier can reward incorrect behavior, while an overly restrictive verifier can penalize valid solutions.
Overview
Given an environment and a target model, AutoSynthData generates training tasks consisting of a system specification, user prompt, and verifier. The generated tasks are grounded in the environment and selected to provide useful training signal for the current model.
AutoSynthData first evaluates the target model in the environment using diagnostic tasks and identifies patterns in the tasks it struggles to complete. A stronger teacher helps characterize which of those tasks are solvable and what successful behavior looks like. AutoSynthData turns the resulting capability gaps into new executable tasks, checks each task in the environment, and uses accepted samples for post-training. Evaluating the updated model reveals which gaps remain and can guide the next round of generation.
From model failures to a curriculum
AutoSynthData uses evaluation runs in the target environment to identify what the model needs to learn next. In our EnterpriseOps Gym experiment, we run both the target model and a stronger teacher on the evaluation tasks. We examine those runs to identify:
the capability being tested;
the tools and workflow structure involved;
where the target model fails and how the teacher succeeds;
the properties that a correct final state must satisfy;
the dimensions that can vary while preserving the capability being tested.
We distill these findings into sanitized capability specification cards. The evaluation tasks guide what the model should learn, but the generator does not receive their original prompts, entities, trajectories, or verifier details. It receives the cards and uses them to create new tasks with different prompts, states, and solution paths.
Generating and scaling tasks
Identifying a capability gap tells us what to teach, but training requires many varied tasks that exercise it. AutoSynthData uses the specification card to generate those tasks.
Suppose the target model struggles with tasks that require the following workflow:
The generator creates new tasks that exercise this workflow, varying the entities, initial environment state, workflow composition, tool combinations, wording, and difficulty. The stronger teacher then demonstrates a successful trajectory for each task. For supervised fine-tuning (SFT), these demonstrations teach the target model how to apply the capability in new situations.
AutoSynthData builds the dataset in two phases: first generating and validating core samples, then expanding them into novel variants.
Target
The target phase creates the core set of training samples from the capability specifications. Workers generate independent tasks in parallel, picking up a new target when they finish. Each candidate goes through validation, execution, solver evaluation, and repair before acceptance. The result is a batch of vetted examples built around what the target model needs to learn.
Multiply
The multiply phase expands the dataset by creating novel variants of accepted target samples. Each variant has its own user request, environment state, entity configuration, reference trajectory, and verifier, and must pass the same validation and execution checks. A multiplied sample cannot seed another multiplied sample. This anchors expansion to the vetted target set and limits drift across generations.
Implementation details
To support both phases, AutoSynthData separates generation control from environment-specific execution. A shared controller coordinates generation, quality control, coverage, and dataset construction, while an adapter handles environment execution, task and state management, reference replay, deterministic verification, solver execution, and task profiling.
Together, parallel target generation and multiplication provide a path to training-scale datasets. Their usefulness depends on the checks applied to every candidate: the task must be executable, the solution must work, and the verifier must distinguish success from failure.
High-quality synthetic data needs more than generation
Generating a plausible request is not enough to produce useful training data. A task may be impossible in the target environment, its reference solution may fail when executed, or its verifier may reward the wrong final state. AutoSynthData checks these properties before accepting a task for training.
AutoSynthData reviews quality at two levels: individual candidates must pass verification, and batches must provide useful coverage and diversity.
Sample-level verification and repair
Each candidate must clear a quality-control loop before entering the training dataset. We begin with solver evaluation to measure difficulty. In the configuration used here, we favor tasks the target model solves on no more than one of three trials and the stronger solver solves on at least two of three trials. Candidates also undergo positive and negative verification and a bounded repair process.
Positive verification
The positive gate asks: Does the intended solution solve the generated task?
The pipeline executes the reference trajectory in the target environment and checks the resulting state against the candidate’s verifier. This reveals mismatches among the prompt, initial state, solution, and success criteria.
Negative verification
The negative gate asks: Do relevant incorrect outcomes fail?
For example, it can mutate parts of the expected outcome and confirm that those states no longer pass verification. This catches weak verifiers that award success without requiring the intended behavior.
Critique and repair
Failed candidates go to a critic before being discarded. The critic examines the sample and its failure, looking for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic, or a mismatch with the intended capability. The critic’s findings guide repairs, with a fixed limit on retries:
candidate
↓
failure
↓
critique / diagnosis
↓
targeted repair
↓
run the gates again
↓
accept or retry
A repaired task must pass the relevant checks again. The diagnosis guides repairs to the existing candidate rather than requiring generation to start over.
Passing these checks makes a sample eligible for training, but individually valid samples can still form a repetitive or unbalanced dataset. AutoSynthData therefore also reviews generation at the batch level.
Batch-level review
A batch may overrepresent a few easy task families, miss a capability, or reflect too much generation effort spent on a low-yield pattern.
A meta-review examines accepted samples, rejected samples, and generation behavior across each batch. It asks:
Which task families are overrepresented, and which capability dimensions are missing?
Are the same kinds of examples appearing repeatedly?
Do particular targets keep failing generation?
Are systematic problems appearing in critiques?
What guidance should change for the next batch?
The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions, and directs more work toward gaps. When a region repeatedly produces poor candidates, critiques and meta-review guide changes to the generation strategy. These adjustments balance useful learning signal, task quality, coverage, diversity, and low redundancy within the available generation budget and dataset size requirements.
Together, these feedback loops improve both individual tasks and the dataset they form: sample-level checks guide candidate repair, while batch-level review guides future generation.
Moving the training frontier
The useful training distribution changes as the model improves. AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary: difficult enough to expose weaknesses, but solvable enough for the teacher to provide reliable demonstrations.
After post-training, we evaluate the updated model in the same environment. Tasks it now solves reliably are less useful for the next training round; persistent failures point to capabilities that still need attention. Those results can guide the next generation round.
Our experiments focus on SFT, but the same mechanism could support reinforcement learning (RL): generate tasks that challenge the current policy and provide reliable learning signal, train, then move the generation target with the updated policy. We plan to test this moving, difficulty-calibrated frontier beyond SFT.
EnterpriseOps Gym experiments
We use EnterpriseOps Gym to test whether this approach improves a model on tasks in a stateful enterprise environment. We generate training tasks in the Gym’s Hybrid and ITSM environments, fine-tune the target model on accepted samples, and evaluate the resulting checkpoints.
Hybrid
We tested the pipeline on the Hybrid domain of EnterpriseOps Gym using Gemma-4-26B-A4B-it as the target model and Qwen3.8-27B as the teacher.
AutoSynthData generated 2,000 synthetic training samples in about 18 hours. We fine-tuned Gemma on this dataset and evaluated the resulting checkpoints on the benchmark. The best checkpoint was epoch 5.
Hybrid results
The synthetic SFT checkpoint improves mean Pass@1 by 7.2 percentage points, a 35% relative improvement, and raises verifier success from 63.01% to 68.55%. It closes 59% of the original Pass@1 gap between Gemma and the reference model.
The training tasks were newly generated from capability specifications; the generator did not receive the original evaluation tasks. The result demonstrates improvement in EnterpriseOps Gym Hybrid, the environment used for this experiment.
ITSM
We also applied AutoSynthData to the ITSM domain of EnterpriseOps Gym, using Gemma-4-26B-A4B-it as the target model and DeepSeek-V4.1-Flash as the teacher. AutoSynthData generated 1,994 synthetic training samples in 66 hours. Generation took longer than in the subsequent Hybrid run described above, primarily because the ITSM run used a larger teacher model and preceded pipeline optimizations that improved throughput.
On ITSM, synthetic SFT raises mean Pass@1 from 18.77% to 27.18%, showing that the approach also improves performance in a second domain.
Closing the loop
The tasks most useful for training depend on both the environment and the model working in it. AutoSynthData uses the model’s failures to choose what to generate, validates new tasks against the environment, and makes those tasks available for post-training. Our EnterpriseOps Gym results show the value of that approach in a controlled setting. As the model changes, the same process can focus on the gaps that remain.
首次收录 · 2026-10-03 · 10.1 分