Briefing

AutoSynthData: Generating Training Data for Enterprise Agent
How to turn an AI agent's failures in your own systems into validated, targeted training tasks, and keep improving as the model gets better.
Key takeaways
The problem: a capable general model can still fail in your specific environment, with your tools, rules and data.
The approach: AutoSynthData uses the target model's failures and a stronger teacher's successes to decide what to teach next, then generates and validates new tasks for those skills.
The safeguard: every task is checked for feasibility, the reference solution is replayed, and the verifier must pass correct outcomes and fail incorrect ones.
The result: on EnterpriseOps Gym, synthetic SFT raised Pass@1 by 7.2 points on Hybrid and from 18.77% to 27.18% on ITSM.
The loop: after training, the model is re-evaluated and the next round targets whatever it still gets wrong.
What is AutoSynthData?
AutoSynthData is a pipeline built at ServiceNow CoreAI that automatically creates training data for AI agents working inside a specific enterprise environment. Instead of generating generic examples, it starts from what the model actually gets wrong and builds new, verified tasks that practice exactly those skills.
Each generated task has three parts:
System specification: the rules and starting conditions, such as policies, instructions and seeded data.
User prompt: the request the agent must complete, written as a real user would write it.
Verifier: an automated check that decides whether the agent's final result is correct.
An analogy
Think of a tutor who reviews a student's graded exams, identifies the topics they keep missing, writes fresh practice problems on those topics, checks that each problem has a correct answer key, and then repeats the process after every study session.
Why enterprises need it
Enterprises need agents that work well in their own environments. The work agents are asked to do is shaped by the systems they use, the rules they follow and the state of their data. A model can be broadly capable and still struggle with a particular environment.
Typical weaknesses
A workflow it handles poorly, for example a multi-step process that must happen in a specific order.
A combination of tools it misuses, such as calling the right tools with the wrong arguments or sequence.
A constraint it fails to respect, such as a policy that forbids certain actions.
Why this is hard to fix
A single failure is informative, but training needs many new tasks that exercise the same capability in different situations. Those tasks must also meet three requirements:
Possible to complete in the environment.
Realistic, meaning they resemble work someone would actually request.
Reliably checkable, with a way to confirm whether the agent succeeded.
Writing such tasks by hand does not scale. AutoSynthData automates it.
How it is used
AutoSynthData is used as a repeatable improvement cycle for a given environment and a given model. In the experiments described here, the data is used for supervised fine-tuning (SFT).
Typical workflow
Choose the environment and target model.For example, an IT service management (ITSM) environment and the model you want to improve.
Run diagnostic tasks.Evaluate the target model, and a stronger teacher, to see where each succeeds and fails.
Generate and validate new tasks.The pipeline produces varied tasks for each capability gap and keeps only those that pass quality checks.
Fine-tune the target model.Use the accepted tasks and the teacher's successful demonstrations as SFT data.
Re-evaluate and repeat.Persistent failures guide the next round of generation.
Where this applies
Adapting a general model to a company's tools: teach an agent the workflows, tool combinations and policies of one specific platform.
Closing a measured performance gap: aim training data at the exact tasks where a smaller model trails a stronger reference model.
Continuous improvement: refresh the curriculum as the model improves so training effort keeps going to what is still hard.
Reinforcement learning (planned): the same mechanism could supply tasks that challenge the current policy. This has not been tested yet.
What makes a useful agentic task?
An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can call, and the state changes its actions produce. A task is created inside that environment:
task = (system specification, user prompt, verifier)
System specification
This defines the constraints the agent works under: system instructions, environment policies and, where relevant, task-specific setup such as a seeded database or a set of knowledge articles. It must be compatible with the environment's tools and supported actions. Its instructions should be clear and should not add arbitrary constraints just to make a task harder.
The agent-facing task (user prompt)
The prompt states what the user wants and any user-level constraints. A good generated task satisfies three properties.
Feasibility
What it means: at least one trajectory exists in the current environment that satisfies the prompt and respects the system specification.
What it rules out: tasks that need unavailable tools, inaccessible knowledge, impossible state changes or policy-prohibited actions.
Realism
What it means: the prompt resembles something a user would plausibly ask in this environment.
What it rules out: odd requests that are technically executable but never occur in real workflows.
Difficulty
What it means: the task exposes a weakness of the current agent.
What it rules out: tasks the model already solves reliably, which add little training signal.
The most valuable tasks are feasible and realistic, but not yet consistently solved.
Verifier
The verifier decides whether a completed trajectory succeeded. It should have three properties.
Consistency: it agrees with the user prompt, the system specification and the task-specific environment state.
Soundness: it rejects trajectories that fail the task or violate relevant constraints.
Completeness: it accepts any valid solution, rather than encoding one particular reference trajectory.
Why verifier quality matters
A lax verifier can reward incorrect behavior. An overly restrictive verifier can penalize valid solutions. Both damage training.
How AutoSynthData works
Given an environment and a target model, the pipeline generates training tasks that are grounded in the environment and selected to give useful signal to the current model.
Evaluate the target model.Run diagnostic tasks and identify patterns in the tasks it fails.
Characterize the gaps with a teacher.A stronger model helps determine which failed tasks are solvable and what successful behavior looks like.
Turn gaps into new tasks.Capability gaps become new executable tasks.
Check every task in the environment.Only tasks that pass validation are accepted.
Post-train.Accepted samples are used to train the target model.
Re-evaluate.The updated model reveals which gaps remain and guides the next round.
From model failures to a curriculum
In the EnterpriseOps Gym experiment, both the target model and a stronger teacher are run on evaluation tasks. Their runs are examined to identify:
the capability being tested;
the tools and workflow structure involved;
where the target model fails and how the teacher succeeds;
the properties a correct final state must satisfy;
the dimensions that can vary while still testing the same capability.
These findings are distilled into sanitized capability specification cards. The generator never sees the original evaluation prompts, entities, trajectories or verifier details. It receives only the cards and uses them to create tasks with different prompts, states and solution paths.
Why sanitizing matters
Because the generator works from abstract capability descriptions and not from the evaluation tasks themselves, training data is newly generated rather than copied from the test set.
Generating and scaling tasks
Knowing what to teach is not enough. Training needs many varied tasks for each capability. For every specification card, the generator creates tasks that exercise the same workflow while varying:
entities and initial environment state;
workflow composition and tool combinations;
wording of the request;
difficulty.
The stronger teacher then demonstrates a successful trajectory for each task. For SFT, these demonstrations show the target model how to apply the capability in new situations.
Two phases: Target and Multiply
Target: creates the core set of samples from the capability specifications. Workers generate independent tasks in parallel and pick up a new target when finished. Each candidate goes through validation, execution, solver evaluation and repair before it is accepted.
Multiply: expands the dataset by creating novel variants of accepted Target samples. Each variant has its own user request, environment state, entity configuration, reference trajectory and verifier, and must pass the same checks. A multiplied sample cannot seed another multiplied sample, which anchors expansion to the vetted set and limits drift.
Implementation design
AutoSynthData separates generation control from environment-specific execution.
Shared controller: coordinates generation, quality control, coverage and dataset construction.
Environment adapter: handles environment execution, task and state management, reference replay, deterministic verification, solver execution and task profiling.
This separation is what lets the same pipeline be applied to different environments.
Quality control: more than generation
A plausible request is not automatically useful training data. A task may be impossible in the environment, its reference solution may fail when executed, or its verifier may reward the wrong final state. AutoSynthData reviews quality at two levels: individual samples and whole batches.
Sample-level verification and repair
Each candidate passes through a quality-control loop before it can enter the dataset.
A new candidate task moves through the following stages in order:
Candidate: a newly generated task enters the pipeline.
Solver check: measures how difficult the task is for the target model and the stronger solver.
Positive gate: confirms that the intended solution actually solves the task.
Negative gate: confirms that incorrect outcomes fail verification.
Accepted: the task becomes eligible for training.
If a candidate fails a check, it goes to critique and repair and is run through the gates again. Each stage is explained below.
Solver evaluation (difficulty): in the configuration used here, the pipeline favors tasks the target model solves on no more than one of three trials, and the stronger solver solves on at least two of three trials. That keeps tasks hard for the student but demonstrably solvable.
Positive verification: asks whether the intended solution solves the task. The reference trajectory is executed in the environment and the resulting state is checked by the verifier. This exposes mismatches among the prompt, initial state, solution and success criteria.
Negative verification: asks whether relevant incorrect outcomes fail. For example, parts of the expected outcome can be mutated to confirm those states no longer pass. This catches weak verifiers that award success without requiring the intended behavior.
Critique and repair: failed candidates go to a critic before being discarded. The critic looks for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic or a mismatch with the intended capability. Findings guide targeted repairs, with a fixed limit on retries. A repaired task must pass the checks again.
Batch-level review
Individually valid samples can still form a repetitive or unbalanced dataset. A batch may overrepresent easy task families, miss a capability, or waste effort on a low-yield pattern. A meta-review examines accepted samples, rejected samples and generation behavior for each batch, asking:
Which task families are overrepresented, and which capability dimensions are missing?
Are the same kinds of examples appearing repeatedly?
Do particular targets keep failing generation?
Are systematic problems appearing in critiques?
What guidance should change for the next batch?
The controller tracks coverage, reduces generation in overrepresented regions and directs more work toward gaps. When a region keeps producing poor candidates, critiques and meta-review guide changes to the generation strategy. The goal is a balance of learning signal, task quality, coverage, diversity and low redundancy within the available budget and dataset size.
Two feedback loops
Sample-level checks guide the repair of individual candidates. Batch-level review guides future generation. Together they improve both the tasks and the dataset they form.
Moving the training frontier
The most useful training distribution changes as the model improves. AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary: difficult enough to expose weaknesses, but solvable enough for the teacher to provide reliable demonstrations.
After training: the updated model is evaluated in the same environment.
Tasks now solved reliably: become less useful for the next round.
Persistent failures: point to capabilities that still need attention and guide the next generation round.
The experiments here focus on SFT. The same mechanism could support reinforcement learning: generate tasks that challenge the current policy, train, then move the generation target with the updated policy. The team plans to test this beyond SFT.
Experiment results on EnterpriseOps Gym
EnterpriseOps Gym (Malay et al., 2026) is a stateful enterprise environment used to test whether this approach improves a model. Training tasks were generated in the Gym's Hybrid and ITSM environments, the target model was fine-tuned on accepted samples, and the resulting checkpoints were evaluated on the benchmark.
Hybrid domain
Headline result
Mean Pass@1 improved by 7.2 percentage points, a 35% relative improvement.
Experiment setup
Target model: Gemma-4-26B-A4B-it.
Teacher model: Qwen3.8-27B.
Training data: 2,000 synthetic samples, generated in about 18 hours.
Best checkpoint: epoch 5.
What improved
Pass@1: up 7.2 percentage points (35% relative).
Verifier success: rose from 63.01% to 68.55%.
Gap to the reference model: 59% of the original Pass@1 gap between Gemma and the reference model was closed.
ITSM domain
Headline result
Mean Pass@1 rose from 18.77% to 27.18%.
Experiment setup
Target model: Gemma-4-26B-A4B-it.
Teacher model: DeepSeek-V4.1-Flash.
Training data: 1,994 synthetic samples, generated in 66 hours.
What improved
Pass@1: from 18.77% before fine-tuning to 27.18% after synthetic SFT.
Takeaway: the approach also improves performance in a second domain.
How to read these results
No test leakage by design: training tasks were newly generated from capability specifications, and the generator did not receive the original evaluation tasks.
Two domains improved: gains in both Hybrid and ITSM show the method is not limited to one environment.
Why ITSM took longer: generation ran 66 hours, mainly because the ITSM run used a larger teacher and preceded pipeline optimizations that later improved throughput.
Scope and next steps
Environment-specific results: the Hybrid result demonstrates improvement in the environment used for that experiment. The ITSM run shows the approach extends to a second domain.
SFT only, so far: reinforcement learning with a moving, difficulty-calibrated frontier is planned but not yet tested.
Requires an executable environment: validation depends on being able to run tasks, replay solutions and verify outcomes in the target environment, which is what the environment adapter provides.
Iterative by nature: as the model changes, the process refocuses on the gaps that remain.
Glossary
Agent
An AI system that takes actions through tools and APIs to complete tasks.
Target model
The model being improved.
Teacher
A stronger model that demonstrates successful solutions.
Trajectory
The sequence of actions an agent takes while attempting a task.
Verifier
An automated check that decides whether a task was completed correctly.
Capability specification card
A sanitized summary of a capability gap used to guide task generation.
SFT (supervised fine-tuning)
Training a model on example inputs paired with correct demonstrations.
Pass@1
The share of tasks solved on a single attempt.
Closing the loop
The tasks most useful for training depend on both the environment and the model working in it. AutoSynthData uses the model's failures to choose what to generate, validates new tasks against the environment, and makes them available for post-training. The EnterpriseOps Gym results show the value of this approach in a controlled setting, and as the model changes, the same process can focus on the gaps that remain.