ServiceNow AI Introduces AutoSynthData for Enterprise Agents
ServiceNow CoreAI details AutoSynthData, a pipeline built to convert target model weaknesses and teacher successes into dynamic training tasks for enterprise environments.

Addressing Enterprise Environment Gaps
Enterprises frequently require agents that operate reliably within their specific environments. However, the work requested is heavily shaped by unique systems, strict rules, and the state of underlying data. While a model may possess broad capabilities, it can still struggle with a particular environment, such as a poorly handled workflow, a mismanaged combination of tools, or an unrespected constraint. According to the [Hugging Face Blog](https://huggingface.co/blog/ServiceNow-AI/autosynthdata), these precise weaknesses are the ones an enterprise needs to improve through targeted training data.
Turning individual failures into effective training data presents a distinct difficulty. Although a single failure highlights a problem, training a model requires numerous new tasks that exercise the same capability across diverse situations. Furthermore, these tasks must remain possible to complete within the environment, resemble realistic work requests, and feature reliable verification methods to check for agent success.
The AutoSynthData Pipeline Architecture
To solve this challenge, [ServiceNow-AI](https://huggingface.co/ServiceNow-AI) built AutoSynthData to convert capability gaps into actionable training data. The system utilizes a target model's failures alongside a stronger teacher's successes to determine what the model should learn next. As the target model improves, the curriculum dynamically shifts toward areas that remain difficult. The pipeline is illustrated through EnterpriseOps Gym (Malay et al., 2026), utilizing the [released dataset](https://huggingface.co/datasets/ServiceNow-AI/EnterpriseOps-Gym).
An agentic environment defines the operational world, including observable and modifiable states, available tools, APIs, and state transitions produced by actions. Within this setup, tasks are instantiated using specific system specifications and user prompts. A generated task must satisfy properties of feasibility, realism, and difficulty, ensuring that the tasks expose genuine weaknesses without relying on unavailable tools or inaccessible knowledge.
Evaluation and Specification Cards
AutoSynthData evaluates target models in the target environment using diagnostic tasks to identify patterns in areas where the model struggles. A stronger teacher assists in characterizing which tasks are solvable and what successful behavior looks like. The resulting findings are distilled into sanitized capability specification cards. These cards guide the generator in creating new tasks featuring different prompts, states, and solution paths without exposing the original evaluation prompts or trajectories.
Once a capability gap is identified, the system generates multiple varied tasks that exercise the target workflow. By altering entities, initial environment states, workflow compositions, and tool combinations, the generator creates diverse learning opportunities. A stronger teacher then demonstrates successful trajectories for supervised fine-tuning, teaching the target model how to apply the capability in novel situations.
Dataset Building Phases: Target and Multiply
AutoSynthData constructs datasets through two distinct phases: creating core samples and expanding them into novel variants. The target phase builds the core training set from capability specifications. Parallel workers generate independent tasks that undergo rigorous validation, execution, solver evaluation, and repair before acceptance, yielding a vetted batch of examples tailored to what the model needs to learn.
Following the target phase, the multiply phase expands the dataset by producing novel variants of accepted target samples. Each variant incorporates its own user request, environment state, entity configuration, reference trajectory, and verifier, and must pass the same validation and execution checks. To prevent drift across generations, a multiplied sample is restricted from seeding another multiplied sample, anchoring expansion securely to the vetted target set.
Sources
- Hugging Face BlogAutoSynthData: Generating Training Data for Enterprise Agents