Agent, workflow and domain-reasoning data
What drives the cost of agent trajectory data
Quick answer
Agent trajectory data cost is driven by five things: how the trajectories are captured (extracted from existing business systems, commissioned from human operators, or synthesized by models), how deeply each step is labeled, how much screen and text de-identification is needed, what rights the license grants over the data and its environment, and whether delivery is one-time or refreshed. Published research figures span from cents per synthetic trajectory to hundreds of dollars per human-built task, so compare quotes on cost per usable trajectory, not headline price.
By SourceX Editorial · Updated
Why published per-trajectory costs vary by three orders of magnitude
Published costs vary so much because research papers measure different things: model inference spend for synthetic data versus paid human hours for demonstrations and benchmark tasks. Explorer reports about $0.28 per successful synthetic web trajectory, roughly half the $0.55 reported for AgentTrek [1]. Fara-7B's FaraGen pipeline reports roughly $1 per task with premium models [3], and the Synatra authors put a synthetic web demonstration at about 3% of the cost of a human-annotated one [5].
The human anchor sits far higher. The AgentSynth authors estimate that human-built benchmark tasks cost tens to hundreds of dollars each, depending on hourly rate and task difficulty, with tasks in the style of TheAgentCompany costing more than OSWorld-style tasks [2]. Treat all of these as research estimates as of October 2026, not market prices; none is a commercial quote, and SourceX publishes no price list.
The practical lesson is that "cost per trajectory" is meaningless until you fix the denominator. A synthetic trajectory that only succeeds on a simulated site, a commissioned human demonstration on your target app, and a reconstructed record from a ticketing system are three different products. The licensed vs commissioned vs synthetic trajectories comparison covers when each is the right input; this page covers what moves the price within each.
Capture method: extraction versus new capture
Capture method is usually the largest single cost driver, because commissioned capture pays for operator time on every task while extraction reuses records a business already holds. Extracted trajectories come from systems such as Salesforce field history, ServiceNow or Zendesk audit logs, Jira issue changelogs, ERP document flows or RPA run logs. The supplier's cost is engineering to export, join and clean them, which amortizes across many records.
Commissioned capture means recording operators doing tasks in an instrumented environment: screenshots or video, accessibility-tree snapshots, DOM state, and action events with coordinates and keystrokes. Every trajectory carries setup time, operator time, failed attempts and QA. OSWorld shows why environment setup matters: each of its 369 tasks needs a defined initial state and an execution-based checker [6], and both have to be built before a single demonstration is usable for training or evaluation.
The trade-off is detail. Historical extracts typically record the field-level outcome (status changed, amount approved, ticket reassigned) but not the pixels, scrolls and hesitations in between, so they cost less per task but carry less UI detail. See reconstructing trajectories from ticket and case histories and field-level audit trails for what extraction can and cannot recover.
Labeling depth: outcome labels to step-level rationale
Labeling cost scales with how many judgments a human makes per step, so specify the minimum depth your training method needs. A trajectory with only a task description and a final success flag is cheap to label. One with per-step action types, target element IDs, intermediate state checks, error tags and written rationale can multiply annotation time several times over.
Research shows the leverage of a small, well-labeled human core. PC Agent-E starts from 312 human-annotated computer-use trajectories and uses a frontier model to add alternative actions at each step [4]. Buyers can apply the same logic: pay for deep human labels on a seed set, then decide whether augmentation or cheaper outcome labels cover the rest.
Common labeling cost lines include:
- Task instruction writing, especially when the source record has no natural-language goal.
- Success and partial-success labels tied to a verifiable end state (see task success labels).
- Step segmentation for long sessions, including waits and handoffs between people.
- Rationale or decision notes, which require a reviewer who understands the domain policy.
- Inter-annotator agreement checks, which add a second pass on a sample.
De-identification of screens, text and linked records
De-identification adds cost in proportion to how many modalities carry personal data, and screen recordings are the most expensive case. Structured logs need field-level redaction or pseudonymization of names, emails, phone numbers and account IDs, with consistent replacement so the same customer stays linked across steps. Screenshots and video need OCR-based detection plus visual redaction, and every redaction has to stay aligned with the action coordinates that reference it.
Free text (ticket comments, chat transcripts, email bodies) needs entity detection and human review on a sample, since automated detectors miss context-dependent identifiers. If records include protected health information, treating them as de-identified under HIPAA requires either Safe Harbor removal of 18 identifier types or an Expert Determination, and HHS notes neither method eliminates all re-identification risk [7]. Expert Determination adds the fee of a qualified expert and documentation of the analysis to the budget.
Linkage is the hidden multiplier. A trajectory that joins CRM, billing and ticketing records has to be pseudonymized consistently across all three systems, so ask suppliers how they keep join keys stable after replacement.
Rights scope: training, evaluation, environments and exclusivity
Rights scope changes price because each additional permitted use is something the supplier gives up or takes risk on. A license for internal evaluation only is narrower than one that allows training production models; rights to rebuild the supplier's environment as a replayable sandbox, or to derive public benchmarks from the tasks, are separate grants again. Exclusivity, field of use and term all move price in the same direction: more scope, more cost.
Environment rights deserve their own line in the quote. A trajectory recorded inside a third-party SaaS application may depend on that vendor's terms of service for screenshots and replay, and seed data for a sandbox may contain the supplier's customer records. The license terms for agent data guide and the seed data for agent sandboxes guide cover the clauses in detail.
Risk reviewers can record these choices under the NIST AI RMF GOVERN and MAP functions, which cover accountability policies and the context and intended use of an AI system [8].
One-time delivery versus recurring refresh
Recurring refreshes turn agent data from a one-time purchase into a subscription-like cost, and that structure is often worth paying for. Business applications change their UI and workflows, so trajectories captured against last year's screens drift out of distribution. A refresh schedule (monthly, quarterly, per release) adds pipeline maintenance, repeat de-identification QA and contract administration, but avoids re-negotiating from scratch.
Ask whether the quote prices refresh deltas or full re-deliveries, and whether schema changes on the supplier side are included. The license pricing structures comparison sets out flat-fee, per-record and subscription models side by side.
A cost-driver worksheet for comparing agent data quotes
Use a single worksheet to normalize quotes, so a cheap synthetic offer and an expensive commissioned one can be compared on usable output. Fill one column per vendor and divide total cost by trajectories that pass your acceptance test, as described in normalizing vendor quotes to cost per usable record.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Driver | Question to ask the supplier | What raises cost | What lowers cost |
|---|---|---|---|
| Capture method | Extracted, commissioned or synthetic? | New operator capture on your target apps | Export of existing audit logs or RPA run logs |
| Observation detail | Screenshots, video, accessibility tree, DOM, or field changes only? | Full-frame video plus a11y tree per step | Field-level before/after values |
| Environment | Is there a reproducible initial state and checker? | Building per-task setup and execution checks | Outcome taken from the system of record |
| Labeling depth | Which labels per step versus per task? | Step rationale, error tags, double annotation | Final-state success flag only |
| De-identification | Which modalities, which method, what QA sample? | Visual redaction aligned to coordinates; Expert Determination | Structured fields with stable pseudonyms |
| Linkage | How many systems are joined per trajectory? | CRM plus billing plus ticketing with consistent keys | Single-system records |
| Rights scope | Training, evaluation-only, environment replay, derived benchmarks? | Production training plus replay and derivatives | Internal evaluation only |
| Exclusivity and term | Exclusive? For how long, in what field? | Exclusive within your field of use | Non-exclusive, fixed term |
| Refresh | One-time or recurring, full or delta? | Frequent full re-deliveries | Scheduled deltas |
| Acceptance yield | What share passes your acceptance test? | Low success rate on replay | High share of verifiable successes |
Before signing, run the acceptance test on a sample; the agent data sample evaluation guide lists checks for step completeness, replayability and label accuracy.
Budgeting across a mixed data program
Most agent programs budget for a blend rather than a single source, because each source covers a different gap. A typical structure pairs a modest set of deep human demonstrations or human baselines with larger volumes of extracted workflow records and synthetic augmentation. Human baselines matter for evaluation as well as training; human baseline data for agent evaluation covers time and quality per task.
Write the blend into your agent data specification before requesting quotes, so each supplier prices the same task list, observation format and acceptance rule. For the general economics of licensed enterprise records beyond agents, see what drives the price of licensed enterprise data, the total cost of ownership guide, and the agent trajectory glossary entry. The agent data hub links every related guide.
How SourceX handles agent workflow data requests
SourceX sources operational datasets, such as support and sales histories, engineering records and finance and legal workflows, from US companies on request; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery and a sample checked. SourceX does not publish prices; terms are agreed per deal. You can describe the workflow data your agents need as a starting point.
Budgeting agent trajectory data with SourceX
If extracted workflow records from real US businesses fit your agent program, describe the tasks, systems and allowed uses you need, and SourceX will look for companies that hold that data and manage licensing through an agreement each supplier approves. Start with the buyer request form.
Sources
- arXiv, "Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents" (2025). https://arxiv.org/pdf/2502.11357
- arXiv, "AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents" (2025). https://arxiv.org/pdf/2506.14205
- arXiv, "Fara-7B: An Efficient Agentic Model for Computer Use" (2025). https://arxiv.org/pdf/2511.19663
- arXiv, "Efficient Agent Training for Computer Use" (2025). https://arxiv.org/pdf/2505.13909
- alphaXiv, "Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale (overview)" (2024). https://arxiv.org/abs/2409.15637
- arXiv (Xie et al.), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.