Skip to content

Agent, workflow and domain-reasoning data

Human baseline data for agent evaluation: time, cost and quality per task

Quick answer

Human baseline data for agent evaluation is a set of per-task records showing how long qualified people took, what that time cost, and how good the result was, captured under conditions comparable to the agent's. Without it, a pass rate cannot be translated into hours saved or work replaced. The strongest baselines combine timed expert runs on your task suite with operational records (ticket handle time, SLA timers, QA scores) that show how the same work is done in production, aggregated so individual workers are not identifiable.

By SourceX Editorial · Updated

What a human baseline must measure

A usable baseline records three separate quantities per task: time, cost and quality, each with its own definition and unit. Collapsing them into one "human score" is the most common reason agent-versus-human comparisons fall apart under review.

  • Time. Active working time (minutes of hands-on effort) and elapsed time (open to close, including waits, handoffs and queue time) are different measures. METR's HCAST calibrates each task by how long qualified humans take to complete it, which only works if the human timing is clean [1].
  • Cost. Loaded labor cost per task (time multiplied by a fully loaded rate for the role), plus rework and escalation cost. Agent cost is inference plus tool calls plus human review, so the human side needs the same scope.
  • Quality. A graded outcome against an explicit rubric: pass or fail, a QA scorecard score, error counts, or a blinded expert preference. GDPval grounds quality in deliverables built by professionals averaging 14 years of experience across 44 occupations [2].

Baselines also need a success definition that matches the agent's grader. If agents are scored by end-state checks, as in the task suites described in agent evaluation task suites, humans must be scored by the same checks, not by a supervisor's impression.

Handle time versus elapsed time

Handle time measures effort, elapsed time measures throughput, and an agent comparison must state which one it uses. A support ticket can show 12 minutes of agent handle time in Zendesk or ServiceNow and three days of elapsed time because it waited on a customer reply, a refund approval and a shift change.

Agents do not wait for shift changes, but they do wait for tool latency, rate limits and human approvals. Comparing agent wall-clock time against human elapsed time inflates the agent's advantage; comparing it against handle time is closer to fair but ignores the coordination work that elapsed time captures. Our page on long-horizon task records covers how to separate waits and handoffs in the timeline.

Illustrative example: invented to show structure; it does not describe an available dataset.

MeasureTypical source fieldWhat it answersFailure mode
Active handle timeITSM time_worked, CRM activity duration, time-tracking entriesHow much human effort did the task consume?Under-logged when staff batch entries at end of day
Elapsed timecreated_at to resolved_atHow long did the requester wait?Includes queue, pending-customer and weekend time
SLA timerPause-aware SLA clockTime inside the service commitmentPause rules differ by ticket type and change over time
Timed expert runStopwatch or session log in a controlled environmentTime under conditions matched to the agentExperts behave differently when observed and paid per task
Task-mining durationDesktop interaction logsTime spent per application and stepMisses offline work, phone calls and thinking time

Where human baseline data comes from

There are two families of source, controlled expert runs and operational records, and serious evaluation programs use both. Controlled runs give comparable conditions; operational records give real distributions at scale.

Controlled expert runs. HCAST is the clearest published template: baseliners typically hold a degree from a top-100 global university or have more than three years of relevant professional experience, and humans and agents work in an identical environment, with humans connecting over SSH [1]. The costs are recruiting, paying and supervising experts, and a sample that is small per task.

Operational records. Enterprises already log time and quality: ticket histories with handle time, SLA breach flags and reopen counts; QA scorecards from contact-center review; code review cycle times in Git hosts; claims or invoice processing times in ERP workflow tables; and approval timestamps. These records show how long real work takes across thousands of instances, including the long tail that controlled runs miss. See reconstructing agent trajectories from ticket histories and task mining data for how these logs are structured, and SourceX's pages on QA scorecards and workflow task histories.

The trade-off is control. Operational tasks are rarely identical to your benchmark tasks, so you match on task type, complexity band and outcome rather than on the exact prompt.

Designing a defensible baseline protocol

A defensible baseline documents who the humans were, what they had access to, how they were paid and how their work was graded. Published baselines vary widely in how much of this they report, which makes "superhuman" claims hard to audit, so treat missing protocol details as a red flag.

Protocol decisions that change the result:

  1. Expertise level. Novice, median practitioner and expert baselines answer different questions. Economic-value claims need practitioners representative of who does the work today.
  2. Tool parity. Give humans the tools they normally use (IDE, search, spreadsheet, internal knowledge base) and record them. Removing search from humans but not agents, or the reverse, invalidates the comparison.
  3. Incentives. Hourly pay rewards slow work; per-task pay rewards rushing. Record the scheme and consider a quality bonus tied to the same grader the agent faces.
  4. Failure handling. Record abandoned attempts and timeouts rather than dropping them. Dropping human failures makes the human baseline look both faster and more accurate.
  5. Repeated trials. Agents are scored over multiple attempts in benchmarks such as tau-bench, which measures reliability across repeated trials on policy-bound customer-service tasks [4]. Collect more than one human per task where budget allows, so variance is visible on both sides.

Converting baselines into economic value

Economic value per task is human cost avoided minus agent cost incurred, adjusted for quality differences and review overhead. The adjustment is where most ROI claims overreach.

A worked approach: take the median human active time for a task class, multiply by a loaded hourly rate for that role, and add expected rework cost (rework rate times rework time). On the agent side, add inference and tool costs, the time a human spends reviewing the agent's output, and the cost of the agent's failures, which still go to a human. If the agent passes 70% of tasks, the remaining 30% carry the full human cost plus the time already spent reviewing the failed attempt.

Quality must enter the formula explicitly. A GDPval-style blinded comparison against expert deliverables [2] tells you whether the agent's output is accepted as-is; a QA scorecard tells you how often it would fail audit. Report the two numbers separately from time and cost, and normalize vendor or supplier pricing for baseline data to cost per usable task, as described in comparing data vendor quotes.

Illustrative baseline record schema

Each baseline record should carry task identity, conditions, time, cost, quality and provenance fields so it can be joined to agent runs on the same task.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "task_id": "inv-recon-0417",
  "task_class": "invoice_three_way_match",
  "complexity_band": "medium",
  "source": "operational_record",
  "performer_role": "AP specialist",
  "performer_experience_band": "3-5y",
  "performer_id": "pseudonymous-hash",
  "tools_available": ["ERP", "email", "spreadsheet"],
  "active_minutes": 14.5,
  "elapsed_minutes": 2880,
  "wait_reasons": ["pending_vendor_reply"],
  "loaded_rate_band_usd_per_hr": "45-60",
  "outcome": "resolved",
  "qa_score": 0.92,
  "qa_rubric_version": "ap-qa-v3",
  "rework_flag": false,
  "grader": "end_state_check+qa_review",
  "captured_at": "2025-11-03",
  "consent_and_rights_basis": "supplier-approved license"
}

Use bands rather than exact salaries and pseudonymous performer IDs. Version the rubric so quality scores from different periods are not silently mixed, and map fields to measurable characteristics such as completeness and accuracy from ISO/IEC 5259-2 when documenting quality [6].

Privacy, contamination and validity risks

Worker-level time and quality records are employee personal data, benchmark tasks leak, and lab baselines may not predict field outcomes; each needs its own control.

  • Worker privacy. Handle time and QA scores tied to a named employee are performance data. Aggregate to task class or team, pseudonymize performer IDs, strip names and emails from free text, and suppress small cells. California's CCPA defines deidentified information as data that cannot reasonably be linked to a particular consumer, with obligations on the business holding it [7]; your counsel should confirm how that applies to workforce data in your context.
  • Contamination. Public tasks and their human solutions end up in training corpora. As of October 2026, OpenAI no longer reports SWE-bench Verified scores because it judged the benchmark increasingly contaminated [5]. Keep baseline tasks and human deliverables private and time-stamped.
  • External validity. A study of 15k+ OpenHands users found substantial gaps between in-the-wild satisfaction and benchmark scores [3]. Pair controlled baselines with operational outcomes before claiming production value.

Buyer checklist for human baseline data

Ask any supplier or internal team for these items before relying on a baseline, and use the same list when you describe operational records to SourceX's buyer desk:

  • Definition of time (active, elapsed or SLA) and how pauses are handled
  • Performer qualification bands and the tools they had
  • Pay scheme and whether failures and abandons are included
  • Grader identity and rubric version, matched to the agent's grader
  • Sample size per task class and variance, not just medians
  • De-identification method, aggregation level and small-cell rules
  • Rights basis for using the records in evaluation, and whether the tasks have ever been public

For more on task outcomes, see task success labels, the agent data hub, SourceX's AI evaluation data page, and training data quality assessment.

Sourcing human baseline records for agent evaluation

SourceX sources operational datasets on request from US companies, including support and sales histories, engineering records and finance workflows, and every release is approved by the supplying company. Datasets are rights-reviewed, personal details are removed or replaced before delivery with the method recorded, and each is delivered under a license defining records, uses, term and delivery; a request does not guarantee a match. Describe the task classes and time, cost and quality fields you need at SourceX for buyers.

Sources

  1. METR (arXiv), "HCAST: Human-Calibrated Autonomy Software Tasks" (2025). https://arxiv.org/pdf/2503.17354
  2. OpenAI (arXiv), "GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks" (2025). https://arxiv.org/pdf/2510.04374
  3. arXiv, "PULSE: comparing in-the-wild user satisfaction with benchmark scores" (2025). https://arxiv.org/pdf/2510.09801v1
  4. Sierra Research (arXiv), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  5. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  6. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and ML, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  7. California Legislative Information, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data