Skip to content

Provenance, rights and permitted use

Classifying Training Data by Copyright Status: Owned, Licensed, Public Domain or Unknown

Quick answer

Give every source in a training corpus one copyright-status class before it reaches a training run: owned by the licensor, licensed with training rights, open-licensed, public domain, unknown, or flagged as likely infringing. Verify the class from primary evidence rather than platform labels, which audits show are often missing or wrong [1]. Then attach a handling rule to each class. Unknown should default to quarantine, and flagged material should be excluded. Fair-use questions go to counsel, never to an engineer's judgment call.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

A license string records what someone typed; a copyright-status class records what your team verified about who holds rights and on what terms. The Data Provenance Initiative audited widely used fine-tuning datasets and found that license fields on GitHub and Hugging Face were frequently unspecified or misattributed, often more permissive than the original source allowed [1]. A license: apache-2.0 tag on a repackaged dataset tells you nothing about the underlying news articles, forum posts or textbook excerpts it contains.

A status taxonomy forces three questions per source: who authored it, who owns it now, and what document lets you train on it. That framing also covers material with no license at all, such as a supplier's internal records or a scanned 19th-century manual. Those cases do not fit in an open-license audit, which is why this page complements auditing open dataset licenses before commercial training rather than repeating it.

The regulatory backdrop makes the field auditable rather than optional. As of October 2026, EU AI Act Article 53(1)(c) requires providers of general-purpose AI models to maintain a copyright-compliance policy, including identifying and honoring text-and-data-mining reservations under Article 4(3) of the DSM Directive [7]. The GPAI Code of Practice copyright chapter turns that into a documented, maintained policy [8]. A per-source status class is the simplest evidence that such a policy is applied.

Each class is defined by the evidence that supports it, not by how the data was obtained. The table below is a working taxonomy; adapt names to your register, but keep the evidence column strict.

ClassDefinitionMinimum evidenceDefault handling
C1 Owned by licensorThe supplier is the author or holds title (employee work, assignment)Authorship facts plus employment or assignment recordsUsable within the license scope
C2 Licensed with training rightsA third party owns it and has granted training useExecuted license naming training, model types, term and territoryUsable within the grant; track expiry
C3 Open-licensedPublished under a public license (CC BY, CC BY-SA, MIT, ODC-By)License text traced to the original publisher, version capturedUsable if the terms fit the use; honor attribution and share-alike
C4 Public domainNo subsisting copyright (term expired, uncopyrightable, dedicated via CC0)Publication date and jurisdiction analysis, or a dedication instrumentUsable; recheck when distributed in other jurisdictions
Note: CC0 is a legal waiver and may not be recognized as a dedication to the public domain in all jurisdictions.
C5 UnknownOwner, terms or status cannot be establishedDocumented search showing what was triedQuarantine pending review
C6 FlaggedCredible signal of infringement (pirate mirrors, takedown notices, stripped notices)The signal itself, loggedExclude and record the removal

C1 is narrower than it looks. Under US law copyright vests in the author, and for a work made for hire the employer is treated as the author [4]. Work by independent contractors only qualifies as made for hire in limited categories under a written agreement [6]. Assigned rights need a signed writing [5], so "we paid the agency for it" is a C5 until the assignment is on file. For record-heavy corpora, see copyright in operational business records and employee-authored records in training data.

C3 and C4 can supply real scale. The Common Pile project assembled roughly 8TB of text using only public-domain and openly licensed sources, which shows those classes can anchor a pre-training mix when the evidence is collected per source [2]. The same work also shows the labor: each source needed its own license verification.

Unknown status is a finding, not a placeholder, so default it to quarantine and give it an owner and a deadline. Data in C5 sits in a separate storage prefix or table that training jobs cannot read. Release requires one of three outcomes: evidence that moves it to C1-C4, a new license that moves it to C2, or retirement. The decision options mirror provenance gap remediation: remediate, re-license, quarantine or retire.

Orphan works are the hardest C5 case. These are works where the rights holder cannot be identified or located after a diligent search. In most jurisdictions an orphan work is still protected, so failing to find the owner does not move it to C4. Treat orphan status as C5 with a documented search trail (registry searches, publisher contact attempts, dates) and let counsel decide whether any exception applies. Do not let "nobody will claim it" drive a reclassification.

Common failure modes that push sources into C5:

  • Repackaged aggregates. A dataset card lists one license, but the contents come from dozens of upstream sources with their own terms [1].
  • Mixed-authorship records. Support tickets with pasted vendor manuals or emails carrying third-party attachments. See third-party content inside licensed corpora.
  • Scanned material with no publication date. Public-domain analysis depends on dates and jurisdiction, so an undated scan stays C5.
  • Stripped metadata. EXIF, PDF XMP or HTML <meta name="copyright"> fields removed during extraction; restore them from source snapshots before deciding.

The same class can carry different risk depending on whether you pre-train, fine-tune or retrieve at inference time, and where you operate. For RAG, source text is copied into an index and into prompts at query time and can surface verbatim in outputs, so output-side reproduction matters more than for pre-training. For SFT, a small number of highly expressive works can dominate model behavior. Record the intended use alongside the class rather than collapsing them into one score.

Fair-use analysis is jurisdiction- and fact-specific, which the provenance audit literature itself flags [1]. In the US, the Copyright Office's Part 3 report, still a pre-publication version as of October 2026, concludes that many acts in building and training a generative model implicate reproduction rights and that fair use turns on the specific facts, including how the training data was obtained and the effect on licensing markets [3]. Court outcomes have been mixed and record-dependent. In Kadrey v. Meta, a June 2025 partial summary judgment for Meta on fair use turned on the plaintiffs' evidentiary showing, and the case continued [10]. In Bartz v. Anthropic, a class settlement received final approval in July 2026; a settlement is not a ruling on the merits [9].

Two practical conclusions follow for a governance lead. First, acquisition method matters, so C6 (pirate-origin material) should be excluded regardless of any fair-use theory. Second, exceptions do not travel: the UK text-and-data-analysis exception covers non-commercial research only, and UK IPO guidance says contract research for a company is unlikely to qualify [11]. A C3 or C4 call made for one market should be rechecked before a model ships to another.

Copyright status belongs as a per-source field in your training data register, with evidence pointers rather than free text. Each row should resolve to a source, not a whole dataset, because a single dataset often spans several classes. Link the row to the permitted-use record described in the AI training data register and to record-level tags where classes vary inside a source.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "source_id": "src-0412",
  "dataset_id": "ds-support-2024q3",
  "description": "Customer support tickets, US SaaS supplier, 2019-2024",
  "copyright_status": "C1_owned_by_licensor",
  "status_basis": "Tickets authored by supplier employees within scope of employment; customer-authored text held under ToS license terms, assignment validity under 17 U.S.C. 204 pending counsel review",
  "evidence_refs": ["ev-employment-policy-v3.pdf", "ev-tos-2019-2024.pdf"],
  "embedded_third_party": "C5_unknown",
  "embedded_third_party_handling": "attachments stripped; pasted vendor docs flagged by n-gram match and quarantined",
  "license_ref": "lic-2026-031",
  "permitted_uses": ["pre-training", "sft"],
  "excluded_uses": ["rag_verbatim_display"],
  "jurisdictions_checked": ["US"],
  "tdm_optout_checked": true,
  "reviewer": "governance-lead",
  "reviewed_at": "2026-09-30",
  "next_review": "2027-03-31"
}

Note the split: the main body is C1, but embedded third-party content gets its own status and handling. Rows like this also feed the public training-content summary required for GPAI providers in the EU; see completing the EU training content summary for licensed datasets.

A classification workflow for an existing corpus

Classify by working from manifest to evidence to decision, and treat any step you cannot complete as a reason for C5. The sequence below fits a corpus already in object storage.

  1. Build the source manifest. Enumerate every upstream source, not just datasets: crawl domains, vendor deliveries, internal exports, synthetic generators and their seed data.
  2. Collect evidence per source. Authorship facts, assignment or license documents, original license text with version, publication dates, and TDM opt-out checks for EU-relevant web content (DSM Article 4 opt-outs).
  3. Assign a class and a confidence. Two reviewers for C1, C2 and C4 calls on high-volume sources; disagreements go to C5.
  4. Detect embedded content. Hash or n-gram match against known third-party works and flag stock media, quoted manuals and attachments.
  5. Apply handling rules. Move C5 to quarantine, delete C6 with a removal log, and tag C2 with license expiry.
  6. Sample-test the claims. Pull a random sample per class and re-verify against evidence, as described in testing a supplier's provenance claims on a sample.
  7. Schedule re-review. Licenses expire, opt-outs change and case law moves; set a review date per row.

For a full corpus pass, pair this workflow with a training corpus provenance audit. The broader data provenance guide for AI training data covers how status fits with chain of title and consent records.

Ask suppliers to classify their own data in your taxonomy and to hand over the evidence behind each class. A supplier that cannot say who authored the records, or produce assignment terms for contractor work, is describing C5 data whatever the contract says. Useful requests include an authorship statement per source, copies or summaries of employee IP and contractor assignment terms, a list of embedded third-party content types and how they were treated, and a data rights attestation signed by someone with authority. Supporting documents are covered in chain of title for AI training data.

Licensed operational records from a company that authored them tend to land in C1 or C2, which makes them a candidate to replace C5 web material. For a comparison of sourcing routes, see licensed vs synthetic vs scraped training data and the data licensing glossary entry. If you need such data, describe it to SourceX: SourceX sources operational datasets from US companies on request, rights-reviews each one for ownership and consents, and delivers it under a license that defines records, uses, term and delivery.

SourceX sources operational datasets from US companies on request and manages the licensing process; it does not hold stock, and a request does not guarantee a match. Every release is approved by the supplying company and delivered under a license that defines records, uses, term and delivery. Describe the data you need on the buyers page.

Frequently asked questions

Does public availability make content public domain?

No. Content posted openly on the web is usually still protected by copyright; public domain (C4) requires expired term, uncopyrightable subject matter or an explicit dedication. Publicly accessible but unlicensed content belongs in C5 until evidence supports another class.

Can a CC0 or CC BY dataset contain C5 material?

Yes. The dataset-level license only covers what the publisher had the right to license. Audits show aggregated datasets often carry licenses that do not match their upstream sources [1], so classify the components, not the wrapper.

Should synthetic data get a copyright-status class?

Yes, because its status depends on the generator's terms and on the seed data. Record both, following provenance records for synthetic training data.

Sources

  1. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/pdf/2310.16787
  2. Kandpal et al. (arXiv), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  3. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  4. Legal Information Institute, Cornell Law School, "17 U.S. Code § 201 - Ownership of copyright". https://www.law.cornell.edu/uscode/text/17/201
  5. Legal Information Institute, Cornell Law School, "17 U.S. Code § 204 - Execution of transfers of copyright ownership". https://law.cornell.edu/uscode/text/17/204
  6. U.S. Copyright Office, "Circular 30: Works Made for Hire". https://www.copyright.gov/circs/circ30.pdf
  7. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  8. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  9. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  10. Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
  11. UK Intellectual Property Office, "Exceptions to copyright: Research" (2014). https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/375954/Research.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data