Skip to content

Agent, workflow and domain-reasoning data

Evaluating an agent data sample before you buy

Quick answer

To evaluate an agent trajectory sample, ask for a randomly drawn slice plus the method used to draw it, then run four tests in order: schema and field completeness, coverage against your specification, replay of a subset in your own harness, and a small fine-tune or evaluation-delta run. Fix pass/fail thresholds before the sample arrives, and agree sample terms (use limited to evaluation, deletion if no purchase follows) before any record moves.

By SourceX Editorial · Updated

Agent data fails in ways that text or image data does not. A trajectory can look complete while its observations are stale, its actions reference UI elements that no longer exist, or its success label reflects a downstream ticket status rather than whether the task was actually done. A sample test is how you find those defects at 200 records instead of 200,000. For background on the format itself, see what every computer-use step record must contain and the agent trajectory glossary entry.

What to request: a random draw, not a showcase

The most useful sample is a random draw from the population you would actually license, delivered with the query or script that selected it. Hand-picked examples show what a supplier's best records look like; they say nothing about the median record or the failure tail. Ask for the sampling frame (which systems, date range and task types the draw came from), the seed or selection logic, and the total population count the sample was drawn from.

Request failures and partial runs, not only successes. Many agent training recipes use failed and recovered trajectories, and a sample of only clean completions hides how the supplier labels abandonment, escalation and retries. If the data comes from business records, such as ticket and case histories reconstructed as trajectories or field-level audit trails, ask for the raw source rows behind a handful of trajectories so you can check the reconstruction yourself.

Size the sample to the tests you plan to run. Schema and coverage checks work on a few hundred records; a fine-tune signal needs more, though published work shows small sets can still be informative: one computer-use study started from 312 human-annotated trajectories [4], and LIMA fine-tuned on 1,000 curated pairs [5]. The general mechanics of asking are covered in how to request a training data sample.

Agree sample terms before records move

Sample terms should be signed before delivery, because agent samples often contain real customer text, internal tool names and screenshots. At minimum, the NDA or sample agreement should state that the sample is for evaluation only, name who may access it, bar training of production models on it, and require deletion with written confirmation if no license follows. Treat any sample that arrives by email attachment as a process failure.

Ask how personal data was handled before the sample was cut. Names, emails, phone numbers and account numbers inside tool arguments and screenshots are the usual leak points. If the workflow touches patient records, the HIPAA standard is either Safe Harbor removal of 18 identifiers or Expert Determination [8], and the method should be recorded alongside the sample.

Test 1: schema, completeness and internal consistency

Run deterministic checks first, because they are cheap and their failures are unambiguous [2]. Parse every record against your agent data specification and count, per field, how often it is missing, null, malformed or out of range. A trajectory with an action but no following observation, or a timestamp that goes backward, is a structural defect regardless of how plausible the text looks.

Check that actions and observations refer to each other. Every action target (a DOM selector, accessibility node ID, API endpoint or form field) should appear in the preceding observation. Every tool call should carry arguments that match the tool's declared schema. Every task_id should resolve to exactly one task description and one outcome label.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckHow to measureExample pass threshold (set your own)
Required fields presentShare of steps with non-null action, observation, timestamp99% or higher
Action grounded in observationShare of actions whose target exists in prior observation97% or higher
Monotonic timestampsShare of trajectories with no backward time steps100%
Outcome label present and definedShare of trajectories with a label from the documented label set100%
Tool-call argument validityShare of tool calls that validate against the declared JSON schema98% or higher
Duplicate or near-duplicate trajectoriesShare of pairs above your similarity cutoffBelow 2%
Residual personal dataHits from your PII scanner per 1,000 steps, manually confirmedZero confirmed hits in sampled review

Test 2: coverage against your specification

Coverage checks answer whether the full dataset can teach the behaviors you need, not whether individual records are clean. Tabulate the sample by task type, application or system, number of steps, outcome (success, failure, escalation, abandonment) and any policy or branch condition your specification names. Compare that distribution with the target mix in your spec and with the supplier's stated population figures.

Look for the tails. Long-horizon tasks with waits and handoffs are usually underrepresented, and they are often the reason you are buying real workflow data in the first place; see measuring horizon length, waits and handoffs. If the sample's step-count histogram stops at 12 while your production tasks run to 40, a larger purchase will not fix that. Whether a random sample of a few hundred is enough to judge the full set is the subject of checking a sample against the full dataset.

Test 3: replay or render a subset in your harness

Replay is the agent-specific test: load a subset of trajectories into your own environment and confirm that each recorded action, applied to the recorded state, produces the recorded observation. Where full replay is impossible because the source system is a production CRM or ERP you cannot access, render the trajectory step by step and have a reviewer confirm that each action is plausible given what the screen or API response showed.

Replay surfaces failure modes that schema checks miss. Common ones include observations captured after, not before, the action; coordinates recorded at a different screen resolution than the screenshot; tool outputs truncated by a logging limit; and environment drift, where the UI changed between capture and delivery. OpenAI reported the same class of problem, including unreliable environment setup and underspecified task descriptions, in the original SWE-bench coding benchmark [7]. If you plan to stand up a sandbox from the purchased data, review seed data and state snapshots for agent sandboxes alongside this test.

Score replayed trajectories against a reference where one exists. Trajectory-matching metrics such as exact match, in-order match and any-order match on tool-call sequences, as implemented in Vertex AI's agent evaluation [1], let you quantify how consistently the supplier's records follow the expected path for the same task. For semantic judgments, such as whether a rationale supports the action, use an LLM judge calibrated on a human-labeled subset, and expect disagreement: step-level labeling is hard even for trained annotators [2]. Research on environment-aware agent judges [3] is a reason to give any automated judge access to state, not only the transcript.

Test 4: a small fine-tune or evaluation-delta run

The value test asks whether the sample moves a metric you care about. Hold out an evaluation set that you built before seeing the sample, fine-tune a small or mid-sized checkpoint on the sample (or mix it into an existing training run at a fixed ratio), and compare task success, step efficiency and policy violations against the same run without it. Run at least two seeds; a single-seed lift on an agent benchmark is often noise.

Divide the lift by the number of trajectories to get a rough value per trajectory, then compare it with alternatives such as commissioned demonstrations or synthetic generation, covered in licensed vs commissioned vs synthetic trajectories. Subset-based attribution methods such as datamodels [6], which predict training outcomes from which examples a subset contains, can in principle show which slices drive the effect, though they need many training runs. Treat the extrapolation to full volume as a hypothesis: returns usually diminish, and estimating a dataset's value before purchase covers how to bound that.

Keep the evaluation set clean of the supplier's data. If the sample overlaps your benchmark tasks, you are measuring memorization; check for near-duplicate task descriptions before training.

Turning results into acceptance criteria

The sample results should become written acceptance criteria for full delivery, so the same tests run against every batch. Carry the thresholds from your completeness table, the target coverage mix, the replay pass rate and the outcome-label definitions into the license schedule or statement of work. Name the sampling method the supplier will use for delivery-time checks and what happens when a batch fails, such as replacement of failing records.

Illustrative example: invented to show structure; it does not describe an available dataset.

sample_acceptance_record:
  sample_id: S-2026-10-A
  population_frame: "support desk trajectories, 3 systems, Jan-Jun"
  draw_method: "seeded uniform random by task_id"
  records_received: 400
  schema_pass_rate: 0.991
  grounded_action_rate: 0.974
  replay_subset: 50
  replay_pass_rate: 0.88
  replay_failure_modes: ["post-action screenshot", "truncated tool output"]
  coverage_gaps: ["no tasks over 25 steps", "escalation outcome under 3%"]
  eval_delta: "+2.1 pts task success, 2 seeds, held-out set v3"
  decision: "proceed with fixes to capture timing and long-horizon coverage"
  sample_deletion_confirmed: pending

Licensing questions that come up during the sample, such as whether you may replay records in a sandbox or derive benchmark tasks from them, belong in the license; see license terms for agent workflow data. For a supplier-neutral walkthrough of pilot logistics, read how to run a data pilot with a supplier. The broader cluster is at AI agent training data, and the full library at AI data for buyers.

How SourceX handles agent data requests

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and finance and legal workflows that agent teams turn into trajectories, and manages the licensing process. Nothing is held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed, personal details are removed or replaced with the method recorded, and delivery runs through private, access-controlled workflows only after an executed agreement. You can describe the agent data you need and the tests it must pass.

Source agent trajectory data with SourceX

Tell us the tasks, systems, fields and coverage your agent needs, and SourceX will look for US businesses that hold that data and assess their data and licensing permissions before anything is agreed. Diligence materials on source, rights, preparation and allowed use are prepared per dataset. Start a buyer request.

Sources

  1. Google Cloud (Vertex AI documentation), "Evaluate an agent". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/agent-engine/evaluate
  2. Langfuse, "AI agent evaluation". https://langfuse.com/resources/engineering/ai-agent-evaluation.md
  3. arXiv, "AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation" (2026). https://arxiv.org/pdf/2604.18240
  4. arXiv (He, Jin and Liu), "Efficient Agent Training for Computer Use" (2025). https://arxiv.org/pdf/2505.13909
  5. arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  6. PMLR (Ilyas, Park, Engstrom, Leclerc and Madry), "Datamodels: Understanding Predictions with Data and Data with Predictions" (2022). https://proceedings.mlr.press/v162/ilyas22a.html
  7. OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  8. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data