Skip to content

Agent, workflow and domain-reasoning data

Licensed workflow records, commissioned demonstrations or synthetic trajectories?

Quick answer

Choosing between human demonstrations and synthetic agent trajectories is a three-way decision. Licensed workflow records show how real work unfolded, with its exceptions, waits and handoffs, but rarely capture screens or reasoning. Commissioned demonstrations record every observation and action for tasks you specify, but they are staged and priced by labor; synthetic or self-generated trajectories are cheap per trajectory and scale with compute, but inherit the generator's blind spots and terms. A practical mix uses real records to define tasks, environments and test sets, and synthesis to expand them.

By SourceX Editorial · Updated

For the comparison across all data types, see licensed vs synthetic vs scraped training data; for definitions, agent trajectory; for the category, the agent training data hub.

Three routes, defined by who performed the steps

The routes differ in who performed each step and why: an employee doing real work for a business, a paid demonstrator performing a task you wrote, or a model acting in an environment. That origin decides what the record can teach.

Licensed workflow records are logs that business systems already keep: ticket and case histories, ERP document flows, CRM activities, integration and API logs, and approval chains. They were written for operations, not training, and can usually be extracted as event logs keyed by case, activity and timestamp; IEEE 1849-2023 (XES) is an IEEE standard XML format for exchanging such logs [1]. Turning them into trajectories is a conversion job; see process mining event logs, ticket histories as trajectories and enterprise workflow datasets.

Commissioned demonstrations are recorded on purpose: people perform specified tasks while a capture tool logs what they saw and did. Mind2Web pairs more than 2,000 open-ended tasks from 137 real websites across 31 domains with crowdsourced action sequences [2]. OpenCUA's AgentNet Tool is installed on annotators' own computers to record demonstrations and the matching computer states, and the resulting dataset holds 22.6K task trajectories across Windows, macOS and Ubuntu [3]. For text-only post-training, the equivalent is expert demonstration data.

Synthetic and self-generated trajectories come from models. Synatra converts indirect knowledge, such as online tutorials, into executable web demonstrations, and its authors fine-tuned a 7B CodeLlama on 100k of them [4]. OS-Genesis lets an agent interact with GUI environments first and derives task instructions afterwards, which its authors call reverse task synthesis [5].

BAGEL bootstraps demonstrations through language-guided exploration because collecting human demonstrations for every new environment is laborious and requires knowing user instructions in advance [6]. Model rollouts filtered by a verifier, including a policy's own (self-generated) rollouts, are a fourth variant: Together AI's announcement, a vendor page, describes CoderForge-Preview as 258k test-verified coding trajectories generated with Qwen3-Coder-480B and rejection sampling [7].

Side-by-side: what each route puts in the record

Licensed records win on task realism and real outcomes, commissioned demonstrations on observation and action detail, and synthetic trajectories on volume; no single route covers all three.

DimensionLicensed workflow recordsCommissioned demonstrationsSynthetic or self-generated
Task originReal requests at their natural frequencyA task list you writeGenerator prompts, tutorials or exploration
Screens, DOM, accessibility treeUsually absent; field values and documents onlyFull, if the capture spec asks for itFull, inside the environment
Action granularityField changes, status transitions, messages, API callsClicks, keystrokes, tool callsWhatever the environment's action space allows
ReasoningWork notes and approval comments, written after the factThink-aloud or annotation, if paid forModel-written, fluent, unverified
Exceptions and rare casesPresent: blocks, reversals, rework, escalationsOnly those you scriptOnly those the environment can produce
Waits, handoffs, multi-day spansPresentCompressed into one sessionRare
Outcome signalBusiness result: paid, reopened, reversedDemonstrator's completion and reviewer checkVerifier or reward model
Environment neededOnly for replay, RL and evaluationFor capture and replayUsually before the first trajectory exists
Main cost driverExtraction, conversion, de-identification, rights reviewLabor hours, task design, QAInference, environment build, filtering
Who grants rightsThe business; customers, staff and vendors appear in the recordsDemonstrators; third parties visible on screenGenerator provider's terms; owners of seed content
Typical failureSteps taken outside the logged systemStaged behavior, task-writer biasModel reinforces its own habits; coverage skew

What published cost figures compare, and what they leave out

In the papers cited here, model-generated trajectories cost cents each while one estimate puts human annotation at roughly $9 to $425 per task, but the figures count different things. None includes licensing, environment construction, privacy work or legal review.

  • Synatra's authors report about $0.025 per synthetic demonstration, roughly 3% of the cost of a human demonstration [4].
  • Explorer reports about $0.28 per successful web trajectory, against $0.55 per trajectory reported for AgentTrek [8].
  • Go-Browse reports $754.66 of rollout cost for 27,103 trajectories [9].
  • AgentSynth estimates human annotation at about $8.8 to $110 per task for OSWorld and $34 to $425 per task for TheAgentCompany, assuming $2 to $25 per hour [10].

Some count only API spend, some only successful trajectories, and model choice and prompt caching move them substantially. Environment construction is a fixed cost outside every per-trajectory number: WebArena, for example, ships fully functional self-hosted websites in four domains [11], and a business-application replica needs comparable engineering plus seed data and state snapshots.

For licensed records, the sources found give no per-trajectory figure; the cost sits in extraction, conversion, de-identification and rights review. SourceX does not publish prices: terms depend on scope, volume, history, rights and exclusivity, and are agreed per deal in writing. Compare routes on the cost of a trajectory that fixes a failure your model actually has; see agent trajectory data cost drivers.

Where synthetic trajectories fall short, and where human data does

Synthetic trajectories come close to human data on some web and GUI benchmarks, but the evidence is benchmark-specific. The gaps cluster where agent buyers feel them: realistic task mix, recovery behavior and long, interrupted work.

Synatra's authors report that their synthetic data improved MiniWoB++ and WebArena results more than an equal number of human demonstrations collected from limited domains [4]. OS-Genesis reports that its trajectories reach over 80% of the performance of human-annotated data, which the authors still treat as the gold standard [5].

In robotics the gap is wider: SynthDemo-RL reports a 56.7% average for supervised fine-tuning on synthetic episodes against 97.7% for human demonstrations, and attributes the gap to motion statistics rather than diversity [12]. Matching a diversity metric, in other words, is not matching behavior. A 2026 survey of embodied manipulation data adds that human demonstrations better capture adaptive behaviors and recovery strategies, while their scale is limited by episode-level human labor [13].

When a policy's own rollouts are filtered by a verifier and fed back, the training set holds only tasks the model can already partly solve, and a verifier that checks only end states can pass a trajectory that reached the right state by a wrong path. Coverage follows what the task source and verifier make easy: a Hugging Face card for CoderForge-Preview notes that its sources skew toward bug fixing [14]. Reverse-synthesized tasks share the bias, since they describe what a model happened to do rather than what users ask for.

Commissioned demonstrations fail differently. Demonstrators usually start from clean states, usually know the task is feasible, and rarely meet a locked record, a missing approval or a vendor who never replies. Licensed records hold those cases at their natural rates.

The daVinci-Agency authors argue that existing synthetic and manual methods lack cross-stage supervision and miss realistic failure modes, and they build long-horizon tasks from chains of related pull requests whose submission, feedback and refinement steps supply state changes and external verification signals [15]. The trade is that records lack the screen layer and any step taken outside the logged system; exception handling records covers what to ask for.

Rights and disclosure: a different party says yes for each route

Rights come from a different party in each route: the business holding the records, the people who performed and appear in demonstrations, or the provider whose model generated the steps. Disclosure records must show each trajectory's route.

Licensed records. The supplier must own or be permitted to share records that name customers, employees and vendors, and its earlier promises bind it. In a January 2024 staff blog post, the FTC's Office of Technology warned that AI model-as-a-service companies that break commitments not to use customer data for undisclosed purposes such as training may be liable under laws the FTC enforces [16]. When you describe the workflow records you need to SourceX, it looks for US businesses that hold them; every dataset goes through rights review (does the business own or may it share the records, and are required consents in place) and is delivered under a license defining its permitted uses. Clause-level points are in license terms for agent data.

Commissioned demonstrations. You need a contributor agreement covering recordings, action logs and any reasoning text. Capture during paid work brings in workplace monitoring and consent law, and screens show third-party software and other people's data.

Synthetic trajectories. The generator's terms travel with its outputs. Anthropic's help page, for example, says as of October 2026 that its terms do not allow outputs to be used to train models that compete with its own, while listing non-competing uses that are allowed [17]. Commentators debate whether such restrictions are enforceable, since they rest on contract rather than copyright [18].

A Mayer Brown article on acquisitions advises asking which model generated synthetic data, what it was trained on, and whether an open-weight model's license restricts commercial use of outputs [19]. Seeds carry rights too: converted tutorials and replicas of commercial software each have owners, and synthetic records derived from personal data may still be personal data.

Disclosure. As of October 2026, California AB 2013 requires developers of generative AI systems made available to Californians to post training-data documentation that includes whether synthetic data was used; the first postings were due January 1, 2026 [20]. Under the EU AI Act, providers of general-purpose AI models must publish a sufficiently detailed summary of training content using the AI Office template, a duty that has applied since August 2, 2025 [21]. Both are easier to meet when every trajectory carries its route, as described in provenance records for synthetic training data.

A five-step plan for mixing the three routes

Start from real records to define what the agent must do, buy demonstrations only for layers the records cannot supply, synthesize around verified seeds, and test on held-out real cases. Each step consumes the previous step's output.

StepRouteOwnerOutput
1. Measure the real task mixLicensed recordsResearch lead, data procurementTask types, exception rates, waits and handoffs from a record sample, written into an agent data specification
2. Build or seed the environmentLicensed recordsResearch, securitySandbox state drawn from the same systems; see RL environments from business workflows
3. Commission the missing layersCommissionedProcurement, counsel, privacyScreen-and-action demonstrations for task types the records cannot show, captured to a computer-use step record format
4. Synthesize around seedsSyntheticResearch, counselVariants anchored on real or demonstrated trajectories, each tagged with its generator and terms
5. Test on held-out real casesLicensed recordsResearch, evaluationOutcome labels fixed first, then a test set no generator or training environment has touched; see real-data holdouts

Step 4 has the clearest evidence for hybrids. PC Agent-E started from 312 human-annotated computer-use trajectories and had a frontier model synthesize alternative action decisions at each step; its ablations report that this beat training on the human trajectories alone and distilling directly from the teacher model [22]. The embodied-manipulation survey describes the same pattern, with human demonstrations used as seeds for replay, augmentation and synthetic expansion [13]. For mixing outside the agent case, see combining licensed and synthetic data.

Illustrative provenance manifest for a mixed trajectory set

One manifest entry per trajectory lets researchers filter by route and lets counsel answer disclosure questions without reconstructing history.

Illustrative example: invented to show structure; it does not describe an available dataset.

[
  {"trajectory_id": "ap-0001", "route": "licensed_record",
   "source": {"system_type": "erp", "record_type": "invoice_block", "case_ref": "inv-3c91",
              "export_format": "xes", "steps_outside_system": "narrated_in_notes"},
   "license": {"agreement_ref": "lic-014", "permitted_uses": ["training", "evaluation"]},
   "deidentification": {"method": "keyed_pseudonyms", "sample_checked": true},
   "split": "test_holdout"},
  {"trajectory_id": "ap-0412", "route": "commissioned_demonstration",
   "task_ref": "task-price-variance-07", "task_derived_from": "ap-0001 task-type counts",
   "capture": {"tool": "screen_and_input_recorder", "observations": ["screenshot", "a11y_tree"]},
   "contributor_agreement": "ca-v2", "workplace_capture": false, "third_party_screens_reviewed": true,
   "split": "train"},
  {"trajectory_id": "ap-1937", "route": "synthetic",
   "seed_trajectory": "ap-0412", "method": "alternative_actions_per_step",
   "generator": {"model": "generator-model-name", "version": "2026-06", "terms_snapshot": "terms-2026-06-10.pdf"},
   "verifier": "end_state_check_v3", "split": "train"}
]

The licensed record is held out, so evaluation reflects real accounts-payable cases (procure-to-pay records describes them). The synthetic entry points to its seed, so if a demonstration is withdrawn its derivatives can be removed, and terms_snapshot keeps the generator terms in force at generation time.

Red flags in each route's offer

Each of these signs means a dataset may not deliver what its route promises.

  • A synthetic vendor cannot name the generator model, version or the terms in force at generation time.
  • "Human demonstrations" were model-generated and then human-reviewed, without a per-trajectory flag saying so.
  • Demonstrators could see the success checker, or every commissioned task was completed.
  • Record-based "trajectories" are only status transitions, with no messages, field changes or outcomes.
  • Test tasks came from the same model or environment as the training data.
  • A self-generated set reports pass rates but not which task types were never solved.
  • A record supplier cannot say which personal data fields it removed or replaced, or how.

Building agents that need real work histories?

Describe the workflows, systems, exception cases and outcome fields your agents need, and which parts you plan to commission or synthesize around them. SourceX sources operational datasets from US companies on request, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your workflow dataset with SourceX.

Sources

  1. IEEE Standards Association, "IEEE 1849-2023 - IEEE Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams" (2023). https://standards.ieee.org/ieee/1849/10907
  2. Deng, Su et al., "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
  3. Wang et al., "OpenCUA: Open Foundations for Computer-Use Agents" (2025). https://arxiv.org/abs/2508.09123
  4. Ou, Xu, Madaan et al., "Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale" (NeurIPS 2024). https://nips.cc/virtual/2024/poster/95647
  5. arXiv, "OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis" (2024). https://arxiv.org/pdf/2412.19723
  6. arXiv, "BAGEL: Bootstrapping Agents by Guiding Exploration with Language" (2024). https://arxiv.org/html/2403.08140v2
  7. Together AI (vendor page), "CoderForge-Preview". https://together.ai/blog/coderforge-preview
  8. arXiv, "Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents" (2025). https://arxiv.org/pdf/2502.11357
  9. arXiv, "Go-Browse: Training Web Agents with Structured Exploration" (2025). https://arxiv.org/pdf/2506.03533
  10. arXiv, "AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents" (2025). https://arxiv.org/pdf/2506.14205
  11. Zhou, Xu et al., "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
  12. arXiv, "SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations" (2026). https://arxiv.org/pdf/2609.21650
  13. arXiv, "Data Pyramid for Embodied Manipulation: A Survey" (2026). https://arxiv.org/pdf/2607.24744
  14. Hugging Face dataset card, "CoderForge-Preview / README.md". https://huggingface.co/datasets/kshitijthakkar/CoderForge-Preview/blob/main/README.md
  15. arXiv, "daVinci-Agency" (arXiv:2602.02619, 2026). https://arxiv.org/html/2602.02619v2
  16. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  17. Anthropic, Claude Help Center, "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  18. SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
  19. Mayer Brown, "Synthetic data as a deal asset: ownership, provenance and diligence considerations in AI acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
  20. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  21. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  22. He, Jin and Liu, "Efficient Agent Training for Computer Use" (2025). https://arxiv.org/html/2505.13909v2

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data