Skip to content

Data licensing for AI training

Commissioned datasets: IP assignment vs license in data collection contracts

Quick answer

When you commission new training data (collection, annotation, expert demonstrations or preference labels), the vendor and its contributors own whatever IP they create unless the contract moves it to you. Under US law that means a signed written assignment, or a valid work-for-hire arrangement, flowing from every contributor to the vendor and from the vendor to you [1][2][3]. Choose assignment when you need exclusivity and freedom to re-use; choose a license when cost or vendor reuse matters more.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Who owns commissioned training data by default

The creator owns it by default, so a buyer who pays for custom data without the right paperwork may hold only an implied license. Under 17 U.S.C. 201(b), the "employer or other person for whom the work was prepared" is the author of a work made for hire, but a vendor's contractors and freelance annotators are usually not your employees [1][3]. Commissioned work by non-employees counts as work for hire only in narrow statutory categories and only with a signed writing, which is why careful contracts pair a work-for-hire clause with a backup assignment [3]. Any assignment itself must be in writing and signed by the rights owner [2].

Copyright is also not the whole picture. Much commissioned data is factual or machine-generated (sensor logs, click trajectories, telemetry from a tele-operated robot), and raw data of that kind often attracts little or no copyright, so your control over it depends on contract terms, confidentiality and trade secret protection rather than ownership [4]. Expert-written SFT responses, rubric-graded essays and annotation guidelines are more likely to be protectable expression. A good contract allocates both: copyright where it exists, and contractual exclusivity and use rights over everything else.

Assignment, exclusive license or non-exclusive license

The right structure depends on whether you need to stop others from using the same data. Public-sector procurement guidance shows the default logic many buyers adopt: vest ownership of commissioned data, including IP, in the buyer at creation, and negotiate a license for use, reuse and sharing only where ownership is not available [5]. For general background on outright purchase of existing datasets, see buying data outright vs licensing it; this page covers data created to your specification.

StructureWhat you holdVendor can reuse data?Typical fitMain risk
Full IP assignment (plus work-for-hire language)Ownership of copyright in deliverables; contractual control of non-copyright dataNo, unless you license it backProprietary SFT, RLHF preference data, held-out evalsHigher price; chain-of-title gaps if contributor paperwork is weak
Exclusive licenseSole right to use in a defined field, often for a termNo (in the field); vendor keeps titleDomain demonstrations where the vendor wants to keep tooling rightsField and term definitions leak; reversion at term end
Non-exclusive licenseRight to use alongside othersYesGeneric annotation, commodity transcription, augmentation dataCompetitors train on the same data; weaker eval integrity
Assignment with license-backOwnership; vendor keeps narrow rightsOnly for listed purposes (QA, own tooling)When the vendor needs to retain labeling-tool improvementsLicense-back scope creep into "model improvement"

Evaluation data deserves special handling. A held-out benchmark loses value the moment it leaks into anyone's training corpus, so for eval sets assignment plus strict confidentiality is usually worth the premium, and a non-exclusive license is rarely adequate. For preference data of the kind used to train reward models, where annotators choose between two model outputs [8], your prompts and model outputs are themselves confidential inputs that the contract must protect.

Contributor chain of title: what the vendor must hold

Your assignment is only as good as the rights the vendor actually obtained from the people doing the work. If a freelance annotator never signed an assignment, the vendor cannot pass that annotator's copyright to you, whatever your master agreement says [2][3]. Require the vendor to represent that every contributor signed an agreement assigning (or exclusively licensing) work product to the vendor, and to produce the template and a signed sample on request.

Contributor agreements for AI data should cover more than copyright. Ask for these elements in the vendor's template:

  • Assignment of work product to the vendor, with present-tense language ("hereby assigns") and a further-assurances clause.
  • Moral rights waiver or covenant not to assert, where local law recognizes moral rights; this matters for international crowds.
  • Likeness and voice consent for recordings, covering training, synthesis and evaluation uses; Tennessee's ELVIS Act, for example, extends right-of-publicity protection to an individual's voice [6].
  • No third-party material: contributors may not paste copyrighted text, proprietary code or employer documents into responses.
  • AI-assistance disclosure: whether contributors may use LLMs to draft responses, and how use is logged. See provenance for human-annotated and preference data.
  • Confidentiality of task prompts, model outputs and guidelines.

Consent language for participants (what they agree their recordings or personal data may be used for) is a separate document from the IP assignment; see consent language for commissioned collection. When individuals assign copyright, US law gives authors a termination right for certain grants, which generally does not apply to true works made for hire; ask counsel whether this matters for long-lived expert content.

Preventing vendor reuse of your task designs and data

The commercial risk most buyers underestimate is not who owns the output but what the vendor keeps. Annotation vendors accumulate guidelines, rubrics, prompt taxonomies, gold sets and trained contributor pools; without restrictions, your task design can become the vendor's next product. Define "Buyer Materials" (prompts, model outputs, guidelines, rubrics, golden answers, environment specs) separately from "Deliverables," and make both confidential and non-reusable.

Practical clauses to request:

  • Use limitation: the vendor may use Buyer Materials and Deliverables only to perform the SOW.
  • No training: the vendor may not train or evaluate its own or third-party models on them, including QA or auto-labeling models.
  • No derived datasets: no synthetic data, paraphrases or "similar" tasks generated from your materials for other clients.
  • Retained tools carve-out: the vendor keeps general know-how and its pre-existing platform, listed in a schedule, not open-ended.
  • Deletion and certification at project end, with defined exceptions for legal holds and backups; see deletion and return clauses.

Expert-network engagements raise the same issues at higher stakes, because domain experts may also consult for your competitors. See contracting domain experts for post-training data for confidentiality and conflicts provisions.

Contract checklist for a commissioned dataset

Illustrative example: invented to show structure; it does not describe an available dataset.

#ClauseBuyer positionAcceptable fallback
1Ownership of deliverablesWork made for hire to the extent permitted, plus present assignment of all remaining rights [1][2]Exclusive, perpetual, irrevocable license in a defined AI field
2Non-copyright dataContractual exclusivity and confidentiality over all records, labels and metadata [4]Exclusivity for a fixed term with no vendor resale
3Contributor agreementsSigned assignment, moral-rights waiver, likeness/voice consent for each contributorRepresentation plus audit right and indemnity
4Buyer MaterialsConfidential; use only for SOW; no training or derived dataVendor QA use allowed, deleted at close
5Pre-existing vendor IPScheduled list; license to you for anything embedded in deliverablesUnscheduled carve-out limited to general platform tools
6License-backNone, or narrowly listed purposesNon-exclusive internal QA license, no third-party benefit
7DocumentationDatasheet covering collection process, composition and contributor pool [9]Summary data card at delivery
8Warranties and indemnityNon-infringement and consent warranties backed by IP indemnityCapped indemnity with carve-out for contributor-rights breaches
9Further assurancesVendor executes recordable assignments on requestPower of attorney limited to perfecting title

Pair this checklist with the scope, acceptance and QA terms in a statement of work for a custom data collection project, and with data warranties for AI training licenses and IP indemnities.

When licensing existing operational data beats commissioning

Commissioning is the right choice when the behavior you need does not exist in anyone's records: new demonstrations, adversarial red-team prompts, or labeled preferences over your own model's outputs. Where the target behavior already happens inside businesses (support resolutions, engineering tickets, contract redlines, finance workflows), licensing those records can produce more realistic distributions than staged collection, though you hold a license rather than ownership. See custom data collection vs licensing existing records and the AI training data licensing hub.

For licensed records, the questions shift from contributor assignments to the rights grant, model retention and output ownership; see who owns model outputs under a training data license. The US Copyright Office's training report, still a pre-publication version as of October 2026, discusses voluntary licensing of training material and is useful background for either route [7]. If you want operational records rather than commissioned work, you can describe the dataset to SourceX; SourceX sources operational datasets from US companies on request, and every release is approved by the supplying company.

Sourcing operational data for AI training

SourceX sources operational datasets from US companies on request and manages licensing, with each dataset rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Tell SourceX what data your team needs.

Frequently asked questions

Is a "work for hire" clause alone enough for annotation data?

Usually not. Work-for-hire status for non-employees is limited by statute, so contracts commonly add a present assignment of any rights that do not qualify, signed in writing [1][2][3]. Pair it with contributor-level assignments so the vendor actually holds what it assigns.

Can the vendor keep a copy for quality assurance?

It can if you allow it, but scope the right narrowly: internal QA only, no model training, no derived datasets, and deletion at a set date. Unlimited "service improvement" rights effectively convert your assignment into a shared asset.

Do we own labels applied to data we supplied?

Labels, rationales and corrections created by the vendor's workers are vendor work product unless assigned. Your underlying data remains yours, but the annotation layer needs its own ownership clause; see expert annotations and labels and the glossary entries on data annotation and human data.

Sources

  1. Office of the Law Revision Counsel, US House of Representatives, "17 USC 201: Ownership of copyright". https://uscode.house.gov/view.xhtml?req=%28title%3A17+section%3A201%28b%29+edition%3Aprelim%29
  2. Office of the Law Revision Counsel, US House of Representatives, "17 USC 204: Execution of transfers of copyright ownership". https://uscode.house.gov/view.xhtml?req=granuleid%3AUSC-prelim-title17-section204&num=0&edition=prelim
  3. Baker Donelson, "Contracting issue: ownership of IP created during performance of services". https://bakerdonelson.com/contracting-issue-ownership-of-ip-created-during-performance-of-services
  4. Clifford Chance, "Exploitation of data from a EU IP / contract law perspective: a three-step process" (2021). https://cliffordchance.com/expertise/services/intellectual-property/global-ip-updates/2021/q3/exploitation-of-data-threestep-process.html
  5. Victorian Government (DataVic), "Developing and procuring datasets". https://www.data.vic.gov.au/datavic-access-policy-guidelines/developing-and-procuring-datasets
  6. Tennessee General Assembly, "HB 2091 (Ensuring Likeness, Voice, and Image Security Act of 2024)" (2024). https://wapp.capitol.tn.gov/apps/Billinfo/default.aspx?BillNumber=HB2091&ga=113
  7. US Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (pre-publication version)" (2025). https://copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf?amp=&stream=top
  8. Touvron et al., Meta AI, "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
  9. Gebru et al., "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data