Skip to content

Provenance, rights and permitted use

Provenance for Human-Annotated and Preference Data: Annotator Agreements and AI-Assistance Disclosure

Quick answer

Annotation data provenance means you can show, for every label, preference pair or written response, who produced it, under which guideline version, under what agreement assigning rights, and whether an AI tool helped. Ask suppliers for the annotator agreement template, the vendor-to-supplier assignment, a guideline changelog, per-label metadata (pseudonymous annotator ID, qualification, guideline version, timestamp) and the AI-use policy with its enforcement evidence. Without these, a "human" dataset can carry unknown rights and model-generated content.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why human judgments need their own provenance file

Human-annotated data needs a separate provenance file because the value sits in the judgments, not the underlying source text, and the judgments have their own authors. A preference dataset has at least three layers: the prompts, the candidate responses (often model-generated), and the human choices, rankings or rewrites. Each layer can have a different owner and a different license, and supplier documentation often addresses the prompts and responses while saying little about the human choices; the provenance hub for AI training data covers the wider chain.

The gap is well documented. The Data Provenance Initiative audited widely used instruction and alignment collections and found license and source information frequently missing or misattributed [1]. Classic RLHF pipelines combine labeler-written demonstrations with human rankings of model outputs [2], and Llama 2's reward model was trained on annotators picking the better of two responses [3]. Every one of those artifacts is human work product with a creator, a contract and a guideline behind it.

If you are licensing expert annotations and labels or data for reward models, treat the annotation layer as its own chain of title rather than assuming the prompt license covers it.

Annotator agreements and IP assignment

The rights question is answered by the contracts between annotators, the annotation vendor and the supplier, not by the dataset license alone. Under US law, the employer or commissioning party is treated as the author only for a work made for hire; otherwise ownership vests in the creator and can pass only by transfer [4]. A transfer of copyright ownership must be in a writing signed by the owner (17 U.S.C. 204(a)). Freelance and crowd annotators are usually contractors, so an express assignment is the safer document to request.

What to check in the annotator agreement:

  • Scope of the grant. It should cover labels, rankings, scores, rationales, free-text corrections, demonstrations and any written responses, not just "deliverables."
  • Assignment or license. An assignment of all rights, with a waiver or non-assertion of moral rights where local law allows, is stronger than a nonexclusive license.
  • Downstream flow. The vendor-to-supplier contract must pass those rights on, with the right to sublicense for model training.
  • Confidentiality and source content. Annotators often see customer documents; check that the agreement bars retention and reuse.
  • Jurisdiction. Offshore annotation workforces sign under local law, so ask which law governs and whether counsel reviewed it.

Simple binary preferences may carry little or no copyright on their own, but rationales, rewrites and demonstrations clearly can. Buy against the strongest case. The chain of title documents page lists the full set of upstream contracts, and the data rights attestation template gives wording for a supplier warranty on annotation rights.

Employee-generated scores and reviewer notes

Operational judgments created by a company's own staff, such as support QA scorecards, ticket reopen reasons or code review comments, follow employee-work rules rather than vendor contracts. Those records are typically works made for hire within the scope of employment [4], but privacy notices and employee policies still matter. See employee-authored records in training data and the owner page on human feedback QA scores and corrections.

Guideline version tracking

Guideline versions must be tracked per label, because a change in instructions silently changes what a label means. If version 3 of a helpfulness rubric told raters to prefer concise answers and version 4 reversed that, pooling both without a guideline_version field produces a reward signal that contradicts itself.

Ask for:

  • The full guideline document for every version used, with an effective date range.
  • A changelog listing rule changes, new examples and edge-case rulings.
  • Calibration and qualification test materials for each version.
  • Adjudication rules, and whether adjudicated labels overwrite or sit beside the original votes.

Documentation frameworks already expect this. Datasheets for Datasets asks who did the labeling, how they were compensated and what process was used [7]; Data Cards record annotation methods and decisions affecting model performance [6]; Data Statements describe annotator demographics and curation rationale [5]. Use those as the dataset-level layer and the per-label fields below as the record layer, as discussed in record-level provenance.

Per-label metadata to require

Each label should carry enough metadata to trace it to a person, a rule set and a moment in time without exposing the annotator's identity. The record below is the minimum we suggest buyers request for preference and SFT data.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "pref-000184",
  "prompt_id": "p-55021",
  "response_a_source": "model:internal-v7",
  "response_b_source": "human:demo-writer",
  "label": "B",
  "label_type": "pairwise_preference",
  "rationale_text": "B cites the refund policy section; A invents a fee.",
  "annotator_pid": "ann-7f3c",
  "annotator_qualification": "domain_expert_billing_ops; qual_test_v4_passed",
  "annotator_agreement_version": "contractor-ip-assign-2025-11",
  "guideline_version": "helpfulness-rubric-v4.2",
  "ai_assistance_policy": "prohibited",
  "ai_assistance_declared": false,
  "telemetry_flags": ["no_paste_events"],
  "created_at": "2026-03-14T16:22:09Z",
  "review_status": "adjudicated",
  "adjudicator_pid": "adj-02b1"
}

The annotator_pid should be a stable pseudonym, with the mapping to real identity held by the vendor. That lets you compute inter-annotator agreement, remove a bad annotator's work, and honor a deletion request without receiving names.

Did annotators use LLMs to label?

Assume some annotators used LLMs unless the supplier can show a policy and evidence of enforcement. Annotators are usually paid per task, so a chatbot that drafts a rationale or picks a winner in seconds is a direct productivity incentive. Instructions alone are weak controls; paste blocking, keystroke and timing telemetry and post-hoc synthetic-text detection are what turn a policy into evidence.

Contamination matters differently by use. For SFT demonstrations and rationales, model-written text defeats the purpose of buying human data. For reward models, an annotator who asks a chatbot which answer is better imports that model's preferences into your reward signal. For eval sets, it can make a model look better aligned with "humans" than it is.

Questions for the supplier:

  1. What was the written AI-use policy for each project and guideline version: prohibited, allowed for specific steps, or allowed with disclosure?
  2. Which controls were active: paste blocking, keystroke or timing telemetry, browser lockdown, typing-speed thresholds?
  3. Which detection ran after collection, and what share of items was flagged and removed?
  4. If AI assistance was allowed, is it recorded per label, and which model and version was used?
  5. Do the AI tool's terms allow its output to be used for training a competing model?

Then test it yourself on a sample, using the methods in detecting model-generated content in purchased human data and testing a supplier's provenance claims.

Preference data rights: the response layer

Preference data rights depend on where the compared responses came from, not only on who chose between them. If responses were generated by a third-party model, check that model's terms on using outputs to train other models; if they were written by humans, they need the same assignment as rationales. Record response_source per candidate, as in the example above.

Also confirm the prompts. Prompts drawn from real user logs or customer tickets bring privacy and contract questions covered in customer contracts and DPAs. For on-policy versus licensed off-policy pairs, see on-policy vs off-policy preference data.

Annotation provenance checklist

A buyer can close most annotation provenance gaps with one document request issued before price negotiation.

Illustrative example: invented to show structure; it does not describe an available dataset.

EvidenceWhat good looks likeRed flag
Annotator agreement templateAssignment covering labels, rationales, written responses"Deliverables" undefined; license only
Vendor-to-supplier contractRights flow with sublicense for trainingVendor retains IP or reuse rights
Guideline versions and changelogEvery version with date rangesOnly the latest version supplied
Per-label metadataPseudonymous ID, qualification, guideline version, timestampAggregated labels only
Qualification recordsTest results per annotator per version"Experts" without criteria; see verifying expert annotators
AI-use policy and enforcementPolicy, telemetry, detection results"Annotators are told not to"
Response and prompt sourcesPer-item source and termsMixed sources, unlabeled
Raw votes and adjudication logOriginal votes kept beside final labelOnly final labels

The costs of producing this evidence feed into pricing, covered in what drives the cost of human preference and SFT data.

How SourceX handles annotation and feedback data

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and new recordings of hands-on work, so human QA scores and corrections inside those workflows can be part of a request. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, with diligence materials on source, rights, preparation and allowed use prepared per dataset. Nothing is held in stock and a request does not guarantee a match; you can describe the judgments you need on the SourceX buyer page. For background, see what human-feedback data is.

Request annotation data with documented provenance

Describe the labels, preferences or QA judgments you need and the provenance evidence you require. SourceX looks for US businesses that hold the described data, assesses data and licensing permissions, and every release is approved by the supplying company; nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. Longpre et al., "A large-scale audit of dataset licensing and attribution in AI", Nature Machine Intelligence 6 (2024). https://www.nature.com/articles/s42256-024-00878-8
  2. arXiv (Ouyang et al., OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  3. arXiv (Touvron et al., Meta), "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
  4. Office of the Law Revision Counsel, U.S. House of Representatives, "17 USC 201: Ownership of copyright". https://uscode.house.gov/view.xhtml?req=%28title%3A17+section%3A201%28b%29+edition%3Aprelim%29
  5. University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
  6. arXiv (Pushkarna, Zaldivar, Kjartansson, Google Research), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  7. arXiv (Gebru et al.), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data