Agent, workflow and domain-reasoning data
Human baseline data for agent evaluation: time, cost and quality per task
Quick answer
Human baseline data for agent evaluation is a set of per-task records showing how long qualified people took, what that time cost, and how good the result was, captured under conditions comparable to the agent's. Without it, a pass rate cannot be translated into hours saved or work replaced. The strongest baselines combine timed expert runs on your task suite with operational records (ticket handle time, SLA timers, QA scores) that show how the same work is done in production, aggregated so individual workers are not identifiable.
By SourceX Editorial · Updated
What a human baseline must measure
A usable baseline records three separate quantities per task: time, cost and quality, each with its own definition and unit. Collapsing them into one "human score" is the most common reason agent-versus-human comparisons fall apart under review.
- Time. Active working time (minutes of hands-on effort) and elapsed time (open to close, including waits, handoffs and queue time) are different measures. METR's HCAST calibrates each task by how long qualified humans take to complete it, which only works if the human timing is clean [1].
- Cost. Loaded labor cost per task (time multiplied by a fully loaded rate for the role), plus rework and escalation cost. Agent cost is inference plus tool calls plus human review, so the human side needs the same scope.
- Quality. A graded outcome against an explicit rubric: pass or fail, a QA scorecard score, error counts, or a blinded expert preference. GDPval grounds quality in deliverables built by professionals averaging 14 years of experience across 44 occupations [2].
Baselines also need a success definition that matches the agent's grader. If agents are scored by end-state checks, as in the task suites described in agent evaluation task suites, humans must be scored by the same checks, not by a supervisor's impression.
Handle time versus elapsed time
Handle time measures effort, elapsed time measures throughput, and an agent comparison must state which one it uses. A support ticket can show 12 minutes of agent handle time in Zendesk or ServiceNow and three days of elapsed time because it waited on a customer reply, a refund approval and a shift change.
Agents do not wait for shift changes, but they do wait for tool latency, rate limits and human approvals. Comparing agent wall-clock time against human elapsed time inflates the agent's advantage; comparing it against handle time is closer to fair but ignores the coordination work that elapsed time captures. Our page on long-horizon task records covers how to separate waits and handoffs in the timeline.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Measure | Typical source field | What it answers | Failure mode |
|---|---|---|---|
| Active handle time | ITSM time_worked, CRM activity duration, time-tracking entries | How much human effort did the task consume? | Under-logged when staff batch entries at end of day |
| Elapsed time | created_at to resolved_at | How long did the requester wait? | Includes queue, pending-customer and weekend time |
| SLA timer | Pause-aware SLA clock | Time inside the service commitment | Pause rules differ by ticket type and change over time |
| Timed expert run | Stopwatch or session log in a controlled environment | Time under conditions matched to the agent | Experts behave differently when observed and paid per task |
| Task-mining duration | Desktop interaction logs | Time spent per application and step | Misses offline work, phone calls and thinking time |
Where human baseline data comes from
There are two families of source, controlled expert runs and operational records, and serious evaluation programs use both. Controlled runs give comparable conditions; operational records give real distributions at scale.
Controlled expert runs. HCAST is the clearest published template: baseliners typically hold a degree from a top-100 global university or have more than three years of relevant professional experience, and humans and agents work in an identical environment, with humans connecting over SSH [1]. The costs are recruiting, paying and supervising experts, and a sample that is small per task.
Operational records. Enterprises already log time and quality: ticket histories with handle time, SLA breach flags and reopen counts; QA scorecards from contact-center review; code review cycle times in Git hosts; claims or invoice processing times in ERP workflow tables; and approval timestamps. These records show how long real work takes across thousands of instances, including the long tail that controlled runs miss. See reconstructing agent trajectories from ticket histories and task mining data for how these logs are structured, and SourceX's pages on QA scorecards and workflow task histories.
The trade-off is control. Operational tasks are rarely identical to your benchmark tasks, so you match on task type, complexity band and outcome rather than on the exact prompt.
Designing a defensible baseline protocol
A defensible baseline documents who the humans were, what they had access to, how they were paid and how their work was graded. Published baselines vary widely in how much of this they report, which makes "superhuman" claims hard to audit, so treat missing protocol details as a red flag.
Protocol decisions that change the result:
- Expertise level. Novice, median practitioner and expert baselines answer different questions. Economic-value claims need practitioners representative of who does the work today.
- Tool parity. Give humans the tools they normally use (IDE, search, spreadsheet, internal knowledge base) and record them. Removing search from humans but not agents, or the reverse, invalidates the comparison.
- Incentives. Hourly pay rewards slow work; per-task pay rewards rushing. Record the scheme and consider a quality bonus tied to the same grader the agent faces.
- Failure handling. Record abandoned attempts and timeouts rather than dropping them. Dropping human failures makes the human baseline look both faster and more accurate.
- Repeated trials. Agents are scored over multiple attempts in benchmarks such as tau-bench, which measures reliability across repeated trials on policy-bound customer-service tasks [4]. Collect more than one human per task where budget allows, so variance is visible on both sides.
Converting baselines into economic value
Economic value per task is human cost avoided minus agent cost incurred, adjusted for quality differences and review overhead. The adjustment is where most ROI claims overreach.
A worked approach: take the median human active time for a task class, multiply by a loaded hourly rate for that role, and add expected rework cost (rework rate times rework time). On the agent side, add inference and tool costs, the time a human spends reviewing the agent's output, and the cost of the agent's failures, which still go to a human. If the agent passes 70% of tasks, the remaining 30% carry the full human cost plus the time already spent reviewing the failed attempt.
Quality must enter the formula explicitly. A GDPval-style blinded comparison against expert deliverables [2] tells you whether the agent's output is accepted as-is; a QA scorecard tells you how often it would fail audit. Report the two numbers separately from time and cost, and normalize vendor or supplier pricing for baseline data to cost per usable task, as described in comparing data vendor quotes.
Illustrative baseline record schema
Each baseline record should carry task identity, conditions, time, cost, quality and provenance fields so it can be joined to agent runs on the same task.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"task_id": "inv-recon-0417",
"task_class": "invoice_three_way_match",
"complexity_band": "medium",
"source": "operational_record",
"performer_role": "AP specialist",
"performer_experience_band": "3-5y",
"performer_id": "pseudonymous-hash",
"tools_available": ["ERP", "email", "spreadsheet"],
"active_minutes": 14.5,
"elapsed_minutes": 2880,
"wait_reasons": ["pending_vendor_reply"],
"loaded_rate_band_usd_per_hr": "45-60",
"outcome": "resolved",
"qa_score": 0.92,
"qa_rubric_version": "ap-qa-v3",
"rework_flag": false,
"grader": "end_state_check+qa_review",
"captured_at": "2025-11-03",
"consent_and_rights_basis": "supplier-approved license"
}
Use bands rather than exact salaries and pseudonymous performer IDs. Version the rubric so quality scores from different periods are not silently mixed, and map fields to measurable characteristics such as completeness and accuracy from ISO/IEC 5259-2 when documenting quality [6].
Privacy, contamination and validity risks
Worker-level time and quality records are employee personal data, benchmark tasks leak, and lab baselines may not predict field outcomes; each needs its own control.
- Worker privacy. Handle time and QA scores tied to a named employee are performance data. Aggregate to task class or team, pseudonymize performer IDs, strip names and emails from free text, and suppress small cells. California's CCPA defines deidentified information as data that cannot reasonably be linked to a particular consumer, with obligations on the business holding it [7]; your counsel should confirm how that applies to workforce data in your context.
- Contamination. Public tasks and their human solutions end up in training corpora. As of October 2026, OpenAI no longer reports SWE-bench Verified scores because it judged the benchmark increasingly contaminated [5]. Keep baseline tasks and human deliverables private and time-stamped.
- External validity. A study of 15k+ OpenHands users found substantial gaps between in-the-wild satisfaction and benchmark scores [3]. Pair controlled baselines with operational outcomes before claiming production value.
Buyer checklist for human baseline data
Ask any supplier or internal team for these items before relying on a baseline, and use the same list when you describe operational records to SourceX's buyer desk:
- Definition of time (active, elapsed or SLA) and how pauses are handled
- Performer qualification bands and the tools they had
- Pay scheme and whether failures and abandons are included
- Grader identity and rubric version, matched to the agent's grader
- Sample size per task class and variance, not just medians
- De-identification method, aggregation level and small-cell rules
- Rights basis for using the records in evaluation, and whether the tasks have ever been public
For more on task outcomes, see task success labels, the agent data hub, SourceX's AI evaluation data page, and training data quality assessment.
Sourcing human baseline records for agent evaluation
SourceX sources operational datasets on request from US companies, including support and sales histories, engineering records and finance workflows, and every release is approved by the supplying company. Datasets are rights-reviewed, personal details are removed or replaced before delivery with the method recorded, and each is delivered under a license defining records, uses, term and delivery; a request does not guarantee a match. Describe the task classes and time, cost and quality fields you need at SourceX for buyers.
Sources
- METR (arXiv), "HCAST: Human-Calibrated Autonomy Software Tasks" (2025). https://arxiv.org/pdf/2503.17354
- OpenAI (arXiv), "GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks" (2025). https://arxiv.org/pdf/2510.04374
- arXiv, "PULSE: comparing in-the-wild user satisfaction with benchmark scores" (2025). https://arxiv.org/pdf/2510.09801v1
- Sierra Research (arXiv), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and ML, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- California Legislative Information, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.