Agent, workflow and domain-reasoning data
Long-horizon task records: measuring horizon length, waits and handoffs
Quick answer
A task record is long-horizon when finishing it takes a skilled human hours to weeks, spans many dependent actions, includes waiting on other people or systems, passes between owners, and accumulates context the next step depends on. Specify it with five measurable fields: expert completion time, action count, elapsed calendar time with wait intervals, handoff count, and context size at each step. Require a verifiable end state, because a long record without an outcome is just a long log.
By SourceX Editorial · Updated
Why "long" needs a measurement, not an adjective
Horizon length is best measured in human time, because that is the scale on which agent autonomy research now calibrates tasks. METR's HCAST benchmark has experienced human baseliners work in the same environment as the agent, so each task carries a measured human completion time and agent success can be read against task length [1]. Published time-horizon results build on this idea: the task length, in skilled-human minutes, at which a model succeeds about half the time [8].
The practical consequence for buyers is that "long-horizon data" without a time distribution cannot be placed against the capability you are training or evaluating. A dataset whose median task takes an expert 20 minutes does not stress a model whose measured horizon is already above that. Benchmark suites of this kind are mostly software tasks with clean automatic scoring, which is the gap real business records can fill and also the gap they make harder to measure.
Recent work treats long-horizon data as a bottleneck. The daVinci-Agency authors argue that synthetic trajectories and manual annotation tend to lack cross-stage supervision and realistic feedback, and build long-horizon data by chaining interdependent pull requests through submission, review feedback and refinement [2]. That is a template for what operational records offer: chains of dependent work with real feedback between stages.
The five horizon dimensions to specify
Each dimension answers a different question about difficulty, so request them separately rather than as one "complexity" score. They are correlated but not interchangeable: a procurement exception can take 40 actions over 11 calendar days with only 90 minutes of hands-on work.
- Expert active time. Minutes of hands-on work by a competent practitioner, excluding waits. This is the dimension comparable to human-calibrated benchmarks [1]. In business records it is reconstructed from worklog entries (Jira
timespent, ServiceNow time cards, Salesforce activity durations) or measured directly in commissioned recordings. - Action count. Discrete state-changing operations: field updates, status transitions, messages sent, files edited, API calls. Count reads separately, since an agent that reads 200 records to change three is solving a different problem from one that edits 200.
- Elapsed calendar time and wait states. Wall-clock span from trigger to verified completion, with each wait typed: waiting on customer, approver, vendor, batch job, or scheduled event. Waits are part of the task because the agent must persist state and resume correctly.
- Handoffs. Changes of owner, team or system. Each handoff is a point where context is summarized, lost or re-requested, and it is where long workflows tend to break.
- Context accumulation. Tokens or documents in the working set at each step: emails, attachments, prior comments, intermediate spreadsheets. The context at step 30 is a function of everything produced in steps 1 to 29.
Waits, resumption and external state
Wait states are the feature most often stripped out of task data, and they are where long-horizon agents fail. A record that collapses "submitted PO, vendor replied four days later with revised pricing, buyer re-approved" into a single success row removes the resumption problem entirely. Ask for the event stream with timestamps, not a summarized case.
LongHorizon-Harness illustrates the design pressure: it updates task state only with environment-verified facts, so that state stays durable across long tasks instead of drifting with the model's own assumptions [3]. Training and evaluation records should support the same discipline. Each wait should end with an observable external event (an inbound email, a webhook, a status change by another user) that the agent can check, rather than a silent jump in the log.
Common failure modes to screen for in samples:
- Timestamp collapse. Bulk imports or migrations stamp hundreds of events with the same second, erasing wait structure.
- Status-only histories. Audit tables that keep state transitions but drop the message or document that caused them.
- Orphaned resumptions. A task resumes after a wait with no record of what triggered it, so the causal link is invisible.
- Clock skew across systems. CRM, ERP and email timestamps in different time zones or without offsets, which reorders events when merged.
Handoffs and cross-system timelines
Handoffs make a record long in a way action count misses, because each one forces a context transfer. Count handoffs by owner change (assignee field changes), team change (assignment group or queue), and system change (the task moves from a ticketing tool to email to an ERP). For linking these into one timeline, see the guidance on cross-system workflow records.
For each handoff, ask whether the record captures what was passed: the handoff note, the summary comment, the attachments forwarded. Customer service benchmarks show why this matters. In tau-bench, an agent must satisfy a simulated user while following domain policy documents and acting through APIs on a realistic database, and success is judged on the resulting database state [4]. Real service records add what the benchmark simplifies: escalations between tiers, partial notes, and customers who restate the problem to a new owner.
Context size and the documents produced along the way
Long tasks generate their own inputs, so a long-horizon record must preserve intermediate artifacts, not only final outputs. A month-end close task produces reconciliations, adjusting entries and reviewer comments; a contract renewal produces redlines, approval emails and revised order forms. Without them, an agent trained on the record learns to jump to conclusions it could not have reached.
Specify how intermediate artifacts are delivered: file versions linked to the step that produced them, with stable identifiers and content hashes. Report context size per step (count of documents and approximate tokens) so you can filter for tasks that exceed your model's effective context. This also lets you test retrieval and memory components separately from planning.
Outcomes and how long tasks are graded
A long-horizon record needs a verifiable terminal state, or it cannot be used for reward or evaluation. Execution-based benchmarks set the standard: OSWorld checks task completion against the resulting machine state across 369 tasks that include multi-application workflows [5], and WebArena grades end-to-end task success in self-hosted web environments, where the best GPT-4 agent reached 14.41% against 78.24% for humans in its v4 results [6]. As of October 2026, OSWorld-Verified and OSWorld 2.0 exist, so older scores are not current comparisons [5].
Business records rarely carry a clean pass flag. Define success from downstream evidence: the invoice was paid and not reversed, the ticket stayed closed for 30 days, the change was not rolled back. For labeling methods, see task success labels for agent trajectories. Also ask for partial-credit milestones, since a 15-step task that fails at step 14 is far more informative than one that fails at step 2.
Request template for long-horizon task records
A precise request describes the horizon distribution and the fields, not the supplier. The template below can go into a data specification; the broader structure is covered in writing an agent data specification.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to request | Acceptance check on a sample |
|---|---|---|
task_id, task_type | Stable ID; type from a fixed taxonomy (for example, vendor onboarding, renewal, incident) | No duplicate IDs; type coverage matches request |
expert_active_minutes | Hands-on time from worklogs or measured recording; method recorded per task | Distribution reported as p10, p50, p90; method field populated |
action_count, read_count | State-changing operations and reads counted separately | Recount on 20 sampled tasks matches within tolerance |
elapsed_hours | Trigger to verified completion | Matches first and last event timestamps |
waits[] | {start, end, wait_type, resumed_by_event_id} | Every wait ends with a linked external event |
handoffs[] | {at_event_id, from_role, to_role, system, handoff_note_ref} | Role fields use pseudonymous role labels, not names |
events[] | Ordered events with ISO 8601 timestamps including offset, actor role, system, action, payload ref | No same-second bulk clusters above threshold |
artifacts[] | Versioned intermediate documents with hash and producing event_id | Every artifact referenced by an event |
context_tokens_by_step | Approximate working-set size per step | Present for all steps |
outcome, milestones[] | Terminal state with evidence reference; partial milestones | Outcome evidence resolvable in the record |
Ask for the horizon distribution before you ask for volume. A useful target for evaluation work is a spread across expert active time bands (under 1 hour, 1 to 4 hours, 4 to 16 hours, multi-day) so you can fit a success curve against task length rather than a single pass rate [1]. Treat these as data quality measures in the sense of ISO/IEC 5259-2, which defines measurable data quality characteristics for ML data, and report them per delivery [7].
Where long-horizon records come from
The richest sources are operational systems that already record dependent work over days: project and ticket trackers, procure-to-pay and order-to-cash flows, engineering change histories, and case management. SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, finance and legal workflows, and new recordings of hands-on work, on request rather than from stock. Buyers describe the data they need, not the businesses, and every release is approved by the supplying company; a request does not guarantee a match. For which record types help most, see SourceX's overview of training data for long-horizon task agents, project records and enterprise workflow datasets and agent trajectories.
Two trade-offs to plan for. Historical records carry real waits and handoffs but rarely carry expert active time directly, so it must be estimated from worklogs. Commissioned recordings measure active time precisely but compress calendar time, so waits must be staged or reconstructed; the comparison is covered in licensed, commissioned or synthetic trajectories. If you need human timing baselines alongside the records, see human baseline data for agent evaluation.
Privacy handling matters more in long records because identifiers recur across dozens of events and attachments. Before delivery, SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers, records the method, and checks a sample; no method is perfect, so test re-identification risk on the joined timeline, not on single events. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. License scope for replay, derived tasks and benchmarks is covered in license terms for agent data, and the full cluster is at AI agent training data. If you are ready to scope a request, describe your long-horizon records to SourceX.
Source long-horizon task records for agent training and evaluation
SourceX finds US companies that hold the operational records you describe, assesses the data and its licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe the long-horizon task records you need.
Frequently asked questions
Is step count a good proxy for horizon length?
Not on its own. Action count ignores waits, handoffs and the reasoning time per step, and it varies with how finely a system logs events. Use expert active time as the primary measure for comparison with human-calibrated benchmarks [1], and report action count, elapsed time and handoffs alongside it.
Should wait time be removed before training?
Keep waits as typed events with durations even if you compress them during training. Removing them hides the resumption problem, which is a core long-horizon failure, and makes it impossible to evaluate state persistence later [3].
How do I evaluate on records that have no environment to replay?
Grade against evidence in the record: terminal state, milestones and downstream confirmations. If you need execution-based grading, pair records with sandbox environments seeded from them; see agent evaluation task suites.
Sources
- METR (arXiv:2503.17354), "HCAST: Human-Calibrated Autonomy Software Tasks" (2025). https://arxiv.org/pdf/2503.17354
- arXiv:2602.02619, "daVinci-Agency: long-horizon agent data from chained pull requests" (2026). https://arxiv.org/html/2602.02619v2
- arXiv, "LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks" (2026). https://arxiv.org/abs/2608.01964
- Sierra Research (arXiv:2406.12045), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Xie et al., XLANG Lab (arXiv:2404.07972), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- Zhou, Xu et al., Carnegie Mellon University (arXiv:2307.13854), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- METR (arXiv:2503.14499), "Measuring AI Ability to Complete Long Tasks" (2025). https://arxiv.org/abs/2503.14499
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.