Skip to content

Human feedback datasets from real QA reviews and corrections

A human feedback dataset is a set of real work products paired with the judgments people made about them: QA scorecards, approvals and rejections, reviewer comments and corrected final versions. SourceX sources these from companies whose reviewers grade work as part of the job, such as support QA teams, editorial desks and review-and-approve workflows, keeping rubric versions, reviewer pseudonyms and, where they exist, calibration records so the signal can be tested before it is used for reward models or evaluation.

Dataset manifest

Sourced to your spec
What it is
Work items with scores, verdicts and corrections from the people who reviewed them
Typical systems
QA platforms, help desks, CRM approvals, code review, document review tools
Typical history
Varies by partner and by when the rubric was introduced
Modality
Structured scores and verdicts plus original and corrected text
Delivery formats
Agreed per order; JSONL items with linked score and edit tables
Preparation
Reviewers pseudonymized; personal data in work items and comments de-identified
Licensing
Non-exclusive or exclusive snapshot; license can bar use for assessing individual staff
Availability
Depends on partners with a consistent, documented review practice; not guaranteed

What a delivery contains

Fields vary by source system and are fixed per order. A typical delivery includes:

FieldTypeWhat it holds
item_idstringPseudonymous ID of the reviewed work item, linkable to its source record where licensed.
item_typeenumWhat was reviewed, such as a support reply, call, document draft, code change or data entry.
work_productobjectThe version that was reviewed, de-identified, with its author role.
final_versionobjectThe version that shipped after corrections, where one exists.
edit_opsarraySpans changed between reviewed and final versions, with operation type and position.
edit_distanceobjectEdit distance between reviewed and final versions in characters and tokens, with normalization stated.
verdictenumApprove, approve with changes, reject or escalate, as recorded in the workflow.
scoresarrayRubric scores per dimension, including auto-fail and not-applicable markers.
rubricobjectRubric ID, version and effective dates, with dimension definitions and score anchors.
reviewerobjectPseudonymous reviewer ID, role and tenure band, kept stable so reviewer effects can be modeled.
commentsarrayReviewer comments, anchored to spans where the tool supports it.
calibrationobjectFor items scored by several reviewers, every score and the agreed consensus result.
disputesarrayAppeals by the author and the result, upheld or score changed.
downstreamobjectWhat happened after review, such as customer satisfaction, reopens or rework.

Example record

{
  "item_id": "qa_8a40d7",
  "item_type": "support_reply",
  "work_product": {
    "author_role": "agent_t1",
    "text": "Your booking is non-refundable, so unfortunately we can't change the dates."
  },
  "final_version": {
    "text": "Your rate is non-refundable, but it allows one date change for a fee of [AMOUNT]. Would you like me to request [DATE_1] instead?"
  },
  "edit_ops": [
    { "op": "replace", "span": [5, 12], "text": "rate" },
    { "op": "replace", "span": [32, 75],
      "text": "but it allows one date change for a fee of [AMOUNT]. Would you like me to request [DATE_1] instead?" }
  ],
  "edit_distance": { "chars": 88, "tokens": 19, "tokens_normalized": 0.83, "basis": "max_token_length" },
  "verdict": "approve_with_changes",
  "rubric": { "id": "travel_cx_qa", "version": "3.2", "effective": "2023-09-01" },
  "reviewer": { "id": "rev_17c", "role": "qa_specialist", "tenure_band": "2-5y" },
  "scores": [
    { "dim": "policy_accuracy", "score": 1, "max": 4, "auto_fail": false },
    { "dim": "resolution", "score": 2, "max": 4 },
    { "dim": "empathy", "score": 3, "max": 4 },
    { "dim": "compliance_disclosure", "score": null, "na": true }
  ],
  "comments": [
    { "on": "work_product", "span": [49, 74], "text": "Rate allows one paid date change. Offer it before refusing." }
  ],
  "calibration": { "session": "cal_2024w07", "other_scores": { "policy_accuracy": [1, 2] },
                   "consensus": { "policy_accuracy": 1 } },
  "disputes": [ { "by": "author", "outcome": "upheld" } ],
  "downstream": { "csat": 4, "reopened": false }
}

Synthetic record for illustration. Field names, structure and format are agreed per order.

What AI teams use it for

Train reward models on expert judgment

Scores against a written rubric, given by reviewers whose job is quality, provide graded labels grounded in the company's own standards rather than an annotator's general taste.

Learn from corrections, not only scores

Draft-to-final pairs show exactly what an expert changed and why, which supports supervised fine-tuning on the corrected version and preference pairs built from the edit.

Build LLM judges and auto-QA

Rubric scores with reviewer comments let you train or calibrate an automated grader and measure its agreement with human reviewers on held-out items.

Calibrate evaluation sets

Items scored by several reviewers show where humans disagree, so you can set realistic targets and leave contested items out of a pass/fail eval.

Predict rework and escalation

Rejections, disputes and downstream signals label which work products caused rework later, useful for routing drafts to human review.

Use-case guides: Private evaluation sets, Customer support agents, Enterprise and computer-use agents, Coding agents, Finance and accounting agents, Legal AI

What makes this data valuable

Versioned rubrics

Every score tied to the rubric version and anchors in force when it was given.

Calibration data

Multiple independent scores on shared items, so reviewer agreement can be measured rather than assumed.

Stable reviewer IDs

Pseudonymous but consistent, so severity and drift can be modeled per reviewer.

Paired versions

Reviewed and final text together, with edits kept as spans rather than only a score.

Disputes and appeals

Overturned scores mark where the rubric or the reviewer was wrong.

Downstream outcomes

Customer satisfaction, reopens or rework show whether a high score predicted a good result.

What operational reviewers record that annotators cannot

Most human feedback in model training is produced for the purpose: annotators compare two model outputs, under instructions written for the study, with little at stake. That is effective for general helpfulness. It is weak for domain work, where the right answer depends on a company's policies, its customers and an expert's sense of acceptable risk. A QA specialist who has scored a team's work for years applies a standard no annotation guideline captures, and the rubric, the calibration sessions and the appeals behind each score make that standard visible.

Corrections are often more useful than scores. A score says a reply was poor on policy accuracy; the edited version shows what policy-accurate looks like in that exact situation. Pairs of reviewed and corrected work are close to ideal fine-tuning targets and natural preference pairs, and they are generated as a by-product of ordinary review.

Biases built into operational review

  • Selection bias. QA programs rarely review a random sample. New staff, complaint cases and high-value accounts are reviewed more often, so scores describe the reviewed population, not all work.
  • Rubric drift. Rubrics get rewritten, dimensions are merged, and scales change from 1–4 to pass/fail. Treat a rubric change as a label-schema change, not a continuation.
  • Incentives. Where QA scores affect pay or performance reviews, authors appeal and reviewers soften borderline cases. Disputes data shows how often this happens.
  • Survivorship in corrections. Drafts that were discarded and rewritten from scratch often leave no clean pair, so the edits you receive overrepresent fixable work.

None of these rule the data out, but each changes how it should be used, so a request is qualified on the review practice rather than the number of scores: which work is reviewed and how it is sampled, which rubric versions apply over the period, whether calibration and appeal records exist, and whether reviewed and final versions link back to the original work items.

What to check before licensing

  • Ask how items were chosen for review. Random sampling, targeting new staff and reviewing complaints produce very different score distributions.
  • Get the rubric history and map each score to its version. Rescaling, merged dimensions and changed anchors make raw scores incomparable across periods.
  • Measure inter-rater agreement on the calibration items, per dimension, with a chance-corrected statistic such as Cohen's kappa or Krippendorff's alpha.
  • Check reviewer concentration. If a few reviewers produced most scores, a model learns their habits, and the stable pseudonymous IDs let you test for it.
  • Look for coaching or performance-management context in comments, and agree whether employee evaluation data stays in scope and how comments are de-identified.
  • Confirm that the final version was the one actually sent or filed, not a later edit, and that the time between review and correction is known.
  • If you will build preference pairs from corrections, check that edits are substantive and not dominated by formatting or template changes.

How licensing works through SourceX

  1. 1

    Define

    Send the domain, modality, volume, format, timeline and permitted use you need.

  2. 2

    Source

    SourceX identifies businesses that hold matching data and are open to licensing it.

  3. 3

    Qualify

    Fit, rights and quality are checked, and you review samples before committing.

  4. 4

    License

    Scope, permitted use, exclusivity, price and obligations are agreed in writing.

  5. 5

    Deliver

    Approved data is prepared, de-identified where required and transferred securely.

Questions buyers ask

How is real-work feedback different from RLHF preference data?

Real-work feedback judges human work against a company's own standard, while RLHF preference data usually comes from annotators comparing model outputs for a study. Real-work judgments have consequences attached: a rejected document went back for rework and a low-scored reply was coached. They carry domain expertise and a written rubric, but they are graded against one company's policies and were not designed for model training.

Can QA scores be converted into preference pairs?

Yes, in two common ways. A reviewed draft and its corrected final version form a natural pair in which the final is preferred, and two items in the same task category scored under the same rubric version can be paired by score. Pairs across rubric versions or very different tasks are unreliable, so keep version and task type with every pair.

What is edit distance between draft and final, and why include it?

It measures how much a reviewer changed a work product, for example the number of token insertions, deletions and substitutions, normalized by length. A small distance with an approve verdict suggests the draft was close to acceptable; a large one marks substantive correction. It is a cheap continuous signal, but formatting changes inflate it, so read it next to the edit spans.

How reliable are QA scores from real operations?

Reliability varies by team and by dimension, which is why calibration data matters. Objective dimensions, such as whether a required disclosure was given, tend to show higher agreement than judgments such as tone. Measure agreement per dimension on the sample, keep dimensions that clear your threshold, and treat the rest as noisy labels or leave them out.

Are reviewers and the people being reviewed identifiable?

Both are pseudonymized before delivery. Reviewers and authors appear as pseudonymous IDs with role and tenure band, and names in comments and work text are replaced. Performance scores are employee data, so a partner may require, and the license can state, that the data is never used to evaluate or re-identify individual staff. Free-text comments need extra de-identification review.

Which kinds of work come with QA scores or reviewer corrections?

Any work a company reviews routinely. Common sources are support and contact-center QA, editorial and translation review, claims and underwriting file reviews, accounting review and approval chains, code review, and document drafting with tracked changes. Supply rests on partners whose review practice is consistent and documented; name the work type and the rubric dimensions you need in the request.

Evaluating this data for procurement?

Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.

Request dataset diligence

Tell us what your models need

Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.

Updated 3 October 2026. Own data like this? See how companies license it to AI developers.

See if you qualify