Skip to content

Agent, workflow and domain-reasoning data

Task success labels: defining outcomes for agent trajectories from business records

Quick answer

A task success label for an agent trajectory is a recorded, verifiable judgment that the business goal behind a task was met: the incident stayed resolved, the invoice was paid on time, the request was approved without rework. Derive it from downstream outcome fields rather than status flags, wait out an observation window before labeling, score multi-part tasks per subgoal, and audit a stratified sample with domain experts before using the labels for reward modeling, RL or evaluation.

By SourceX Editorial · Updated

Why "closed" is not a success label

A status field records that someone stopped working on a task, not that the task succeeded. In ServiceNow, an incident's state can move to Closed after an auto-close timer even if the caller gave up; in Jira, a resolution of "Won't Do" or "Duplicate" also closes the issue; in Salesforce, Case.IsClosed is true for any status an admin marks as closed, including statuses used for spam or duplicates. Training a reward model on these flags teaches agents to close things, which is exactly the shortcut you do not want.

Public agent benchmarks show the alternative. Software-engineering trajectory studies define success as a binary label from each benchmark's own check: the patch passes the tests, or, in a stricter variant, also passes manual review as equivalent to the developer's fix [1]. OSWorld grades 369 tasks by executing checks against the final machine state [3], and τ-bench judges customer-service agents on the resulting database state under domain policies, across repeated trials [2]. Business records rarely carry a test suite, so the job is to find the fields that play the same role.

Proxy outcomes and the failure modes to check

Every usable outcome label is a proxy, and each proxy has a known way of lying. Pair the obvious field with a corroborating signal recorded later, and write down what the pair still misses. The table below is a starting map for common enterprise task families; see verifying outcome fields in operational records for field-level reliability tests.

Illustrative example: invented to show structure; it does not describe an available dataset.

Task familyNaive labelBetter success definitionCorroborating fieldsFailure mode it still misses
IT incident resolutionstate = ClosedResolved, no reopen and no linked repeat incident from the same caller and CI within 14 daysreopen_count, resolved_at, close_code, parent/child linksUsers who stop reporting and work around the issue
Customer support caseStatus "Solved"Solved, no follow-up contact on the same order within 30 days, no refund escalationContact history, CSAT response, refund or credit memo recordsLow CSAT response rates; silent churn
Accounts payable invoicePayment postedPaid on or before due date, three-way match passed, no later credit memo or vendor disputeDue date, clearing date, match exceptions, dispute logEarly payment that forfeited a discount
Purchase or access approvalApprovedApproved on first submission, no rework loop, not revoked within 90 daysApproval history, resubmission count, revocation eventsApprovers who rubber-stamp
Contract redlineSignedSigned without escalation past the playbook, no amendment within 6 monthsVersion history, escalation flag, amendment recordsConcessions buried in accepted language

Two failure modes recur across families. Label leakage happens when fields written after the trajectory ends, such as a close_code or a reviewer note, appear inside the observation the agent sees; strip them from the trajectory and keep them only in the label table. Survivorship happens when abandoned or merged tasks drop out of the export, which inflates success rates and hides the hardest cases.

Observation windows for delayed outcomes

Many business outcomes only become knowable weeks after the last action, so labels need an explicit observation window. A support ticket solved yesterday cannot yet be "solved without reopen"; an invoice approved today has no payment outcome until its terms expire. Treat tasks whose window has not elapsed as right-censored and exclude them, or label them pending, rather than counting them as successes.

Set the window from the data, not from habit. Plot the cumulative share of reopens or disputes against days since closure and pick the point where the curve flattens; record that choice as a versioned parameter. Then set a maturity cutoff for the extract, for example tasks closed at least 45 days before the snapshot date, so every labeled task had the same chance to fail. Long-horizon task records covers how waits and handoffs stretch these windows further.

Partial credit for multi-part tasks

Multi-part tasks need subgoal labels, because a single binary flag discards most of the reward signal. An onboarding ticket that provisions four systems but misses one, or an order change that updates price and quantity but not the ship date, is neither a clean success nor a clean failure. Decompose each task type into checkable subgoals, label each one from its own field, and then decide how to aggregate.

Keep both views. A strict all_subgoals_met flag mirrors how execution-based benchmarks grade final state [3] and should drive headline evaluation metrics. A weighted subgoal score is more useful as a dense reward for RL or reward-model training, where an all-or-nothing signal is sparse. Document weights in the label specification so a later team can recompute either view from the same records, and reuse the structure from an agent data specification.

A label record that grades trajectories

A label table should be separate from the trajectory, keyed to it, and carry its own provenance. The record below shows the minimum fields an evaluation lead needs to reproduce, audit and re-version a label.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "trajectory_id": "inc-2025-08-14-0193",
  "task_type": "it_incident_resolution",
  "label_spec_version": "v3.1",
  "trajectory_end": "2025-08-14T16:42:00Z",
  "observation_window_days": 14,
  "window_closed": true,
  "outcome": "success",
  "subgoals": [
    {"id": "restore_service", "met": true, "evidence_field": "resolved_at"},
    {"id": "no_reopen", "met": true, "evidence_field": "reopen_count"},
    {"id": "no_repeat_incident", "met": true, "evidence_field": "related_incidents"}
  ],
  "subgoal_score": 1.0,
  "label_source": "rule",
  "expert_review": {"sampled": true, "agrees": true, "reviewer_role": "service_desk_lead"},
  "excluded_from_observation": ["close_code", "close_notes", "reopen_count"]
}

The label_source field matters once you mix methods. Rule-derived labels, expert labels and model-judged labels have different error profiles; research on learned success detectors shows model judges can scale labeling [6], but their outputs should be tagged and audited separately rather than blended silently.

Auditing labels with domain experts

Rule-derived labels are only as good as the rule, so audit a stratified sample with people who did the work. Draw the sample across task types, outcome classes, sites and time periods, oversampling borderline cases such as tasks closed just outside the window or with one failed subgoal. Ask two reviewers to label independently from the full record, measure agreement, and adjudicate disagreements into a written decision log.

The SWE-bench Verified effort is the clearest public warning: human annotators found overly specific tests, underspecified problem statements and unreliable environments in a benchmark whose labels looked mechanical [5]. As of October 2026, OpenAI no longer reports SWE-bench Verified, citing contamination, a reminder that reused evaluation labels also age. Expect similar findings in business records, such as tickets closed as resolved because the requester left the company. Treat labeling as a managed data-quality process with documented requirements and review steps, which is the framing ISO/IEC 5259-4 gives for training-data labeling and reinforcement learning [8].

Historical outcomes are not the ceiling

A label derived from what humans did tells you whether that run succeeded, not what success was possible. Human success rates provide a reference point: WebArena reported 78.24% end-to-end success for humans against 14.41% for its best GPT-4-based agent [4], and HCAST calibrates tasks with human baseliners working in the same environment as agents [7]. Pair outcome labels with human baseline data for time, cost and quality so graders can separate "the agent failed" from "the task usually fails."

Historical outcomes also encode past policy. Approval and claim decisions reflect the approvers' habits and constraints, so a model rewarded for reproducing them inherits that bias; auditing historical decision bias in operational labels covers the checks. For evaluation sets built on the same idea, see outcome-labeled evaluation data.

What to request from a data supplier

Outcome labels depend on fields that sit outside the trajectory, so a request must name them explicitly. Ask for the event history that forms the trajectory plus the downstream outcome tables: reopen and repeat-incident links, payment clearing and dispute records, approval and revocation history, amendments. Ask for the extract's snapshot date so you can apply your maturity cutoff, and confirm that abandoned, merged and auto-closed tasks are included.

Ask how the supplier's systems define each status value, since the same label means different things across ServiceNow, Zendesk, Jira and ERP configurations. Confirm that the license covers deriving labels and training reward models, not only supervised fine-tuning. Reconstructing trajectories from ticket and case histories and RL environments from business workflows cover the trajectory side of the same request.

SourceX sources operational datasets such as support and engineering records and finance and legal workflows from US companies, on request rather than from stock, and each dataset is rights-reviewed before it is licensed. Buyers describe the records and outcome fields they need on the SourceX buyer page; a request does not guarantee a match. Related SourceX pages cover training data for reward models, RL environments and enterprise workflow task histories, and the agent data hub maps the rest of this cluster.

Getting outcome-labeled workflow records

SourceX manages the commercial process for operational datasets from US companies: it assesses the data and licensing permissions, agrees pricing and allowed uses in a license, and prepares diligence materials per dataset. Personal details are removed or replaced before delivery, and every release is approved by the supplying company. Describe the trajectories and outcome fields you need at sourcex.si/buyers.

Frequently asked questions

Should a failed human trajectory be labeled a failure for the agent?

Label the trajectory, not the agent. A failed human run is a valid negative example for reward modeling, but for evaluation you need to know whether the task was achievable, which is why human baselines and subgoal labels matter.

Can an LLM judge replace rule-derived outcome labels?

It can supplement them where no outcome field exists, such as judging whether a reply addressed the question. Tag judge labels with their source and model version, and audit them against expert labels on the same sample before mixing them with rule-derived labels.

How should personal data in outcome fields be handled?

Outcome tables often carry names, account numbers and contact details. Ask suppliers to remove or replace them before delivery while preserving the join keys you need to connect outcomes to trajectories, and to record the method used.

Sources

  1. arXiv, "Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories" (2025). https://arxiv.org/pdf/2506.18824
  2. arXiv (Sierra Research), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  3. arXiv (XLANG Lab, HKU and collaborators), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
  4. arXiv (Carnegie Mellon University), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
  5. OpenAI, "Retiring SWE-bench Verified" (2026). https://openai.com/news/retiring-swe-bench-verified/
  6. arXiv, "Vision-Language Models as Success Detectors" (2023). https://arxiv.org/pdf/2303.07280
  7. arXiv (METR), "HCAST: Human-Calibrated Autonomy Software Tasks" (2025). https://arxiv.org/pdf/2503.17354
  8. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data