Skip to content

Provenance, rights and permitted use

Consent Language for Commissioned AI Data Collection: What Participants Must Agree To

Quick answer

A consent form for commissioned AI data collection must name machine learning training explicitly, cover future models and third-party licensees, say whether recordings will be shared, sold or published, define what withdrawal does to delivered data and already-trained models, and add statute-specific releases for voiceprints, faces, health signals and minors. The collection partner must also log the consent version, timestamp and participant ID against every record, or the language cannot be proven later.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Most consent failures in commissioned speech, video, image and text projects come from forms written for the collection vendor's own use, not for a buyer that will license, combine and train on the data. A form that says "research purposes" or "improving our services" rarely stretches to a commercial foundation model trained by a different company. The FTC has warned that using data for undisclosed purposes such as model training, contrary to what people were told, can create liability [6].

The second failure is scope drift. A buyer commissions 2,000 hours of conversational audio for an ASR model, then wants to use it for a speech-language model, a voice-cloning safety classifier and an evaluation set licensed to a partner. Each of those uses needs to be inside the original text, because re-contacting participants after collection is slow and response rates are partial.

This page covers prospective drafting for new collection. Auditing consents on records that already exist is a different task, covered in consent and notice records for AI training data. Who owns the output of the collection contract is covered in IP assignment vs license for commissioned datasets.

A training-ready form states, in plain language, who collects, who receives, what models may be built, and what the participant can and cannot undo. The table below is the drafting checklist to hand a collection partner before the statement of work is signed.

Illustrative example: invented to show structure; it does not describe an available dataset.

ClauseWhat it must sayCommon weak wording to reject
PurposeRecordings will be used to train, fine-tune, evaluate and test machine learning and AI models, including generative models"research", "product improvement"
RecipientsData may be licensed to the commissioning company and to other organizations that develop AI, under contract"our partners" with no AI mention
Future useUse covers future models and versions, not only the current projectProject-name-only scope
Distribution formWhether raw recordings, transcripts, derived features or only models leave the collector; whether any public release is plannedSilence on resale or publication
Identity handlingWhich identifiers are removed or replaced and that removal is imperfect"fully anonymous"
BiometricsSeparate written release for voiceprint, face geometry or other biometric identifiers, with retention periodBundled into general terms
CompensationPayment amount and that payment does not depend on remaining in the datasetPayment clawback on withdrawal
WithdrawalHow to withdraw, what is deleted, and what cannot be reversed (trained weights, aggregate statistics)"You may withdraw at any time" with no effect stated
Recording consentAll parties to any recorded conversation consentOnly the paid participant signs
Contact and versionCollector contact, form version ID, dateUnversioned PDFs

The form should also say the participant was not required to consent to AI training to receive unrelated services. Bundled consent, where training use is buried in a general platform agreement, is the weakest provenance position a buyer can inherit.

Naming AI training, future models and third-party licensees

The purpose clause should use the words "artificial intelligence" and "machine learning" and describe training, not just storage or analysis. Participants should be told the data may go to organizations other than the collector, because that is what licensing means in practice.

Sample purpose language a buyer might request (adapt with counsel):

"Your recordings and the information you provide may be used to build, train, fine-tune, test and evaluate artificial intelligence and machine learning systems, including systems that generate speech, text or images. [Collector] may license these materials to other organizations that develop AI systems, under written agreements that restrict how the materials are used. These uses may continue for future versions of those systems."

For EU or UK participants, consent is only one possible legal basis, and the EDPB's Opinion 28/2024 discusses legitimate interest, model anonymity and the consequences of unlawful processing during development [7]. If the collector relies on consent under GDPR, the participant can withdraw it at any time, which makes the withdrawal clause below more than a formality. Downstream, generative AI developers serving Californians must post training-data documentation under AB 2013, including whether datasets contain personal information [8], so the buyer needs consent records that support an accurate disclosure.

Biometric, health and children's data need separate releases

Voice, face and body data trigger statutes with their own consent mechanics, and a general consent paragraph does not satisfy them. Illinois BIPA treats voiceprints and scans of face geometry as biometric identifiers and requires a written release and a published retention and destruction schedule [1]. Texas prohibits capturing a biometric identifier for a commercial purpose without informing the individual and obtaining consent before capture [2].

Wearable, fitness and some behavioral signals can fall under Washington's My Health My Data Act, which defines consumer health data broadly, including biometric data [3]. Where a collection app, website or online service gathers personal information from children under 13 in the US, the amended COPPA Rule applies; as of October 2026 its April 22, 2026 compliance date has passed [4]. Children's collections therefore need verifiable parental consent, a separate child-appropriate assent script and tighter retention terms. See face data consent, releases and anonymization and the privacy hub for de-identification options.

Conversational recording adds a two-party problem. California Penal Code 632 requires the consent of all parties to a confidential communication [5], so a protocol where a paid participant records calls with friends or customers needs consent from every voice captured, not only the signer.

Defining withdrawal for delivered data and trained models

Withdrawal language must say exactly what happens at each stage, because "delete my data" means different things before delivery, after delivery and after training. A realistic clause separates three states: data still held by the collector (deleted), data already delivered to licensees (deletion request passed on under the license), and models already trained (generally not retrained, stated plainly).

Sample withdrawal language a buyer might request (adapt with counsel):

"You may withdraw by contacting [address] and quoting your participant ID. We will delete your recordings that we still hold and instruct organizations that received them to delete their copies. AI systems already trained using your recordings may not be changed, and statistics computed across many participants cannot be separated."

Buyers should mirror this in the license with a deletion-propagation obligation and a suppression list keyed to participant ID. Ask the collector how often it sends withdrawal batches and in what format, for example a CSV of participant_id, withdrawal_date, scope.

Consent language is only useful if each delivered file can be traced to the signed form version. The partner should capture consent as structured metadata at collection time, not as a folder of scanned PDFs matched by name later. Large consented collections such as Ego4D, with 931 camera wearers across 9 countries [9], show why: multi-site projects run several form versions and languages in parallel.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "aud_000418_s03",
  "participant_id": "P-7f3c21",
  "consent_form_id": "CF-ASR-US-EN",
  "consent_version": "2.1",
  "consent_language": "en-US",
  "consent_method": "e-signature",
  "consent_timestamp": "2026-08-14T15:22:07Z",
  "scopes": ["ai_training", "third_party_license", "evaluation"],
  "public_release": false,
  "biometric_release": {"type": "voiceprint", "statute_refs": ["740 ILCS 14", "Tex. Bus. & Com. Code 503.001"], "retention_until": "2029-08-14"},
  "minor": false,
  "all_party_consent": true,
  "withdrawal_status": "active",
  "collector_site": "site-04"
}

Map these to your permitted-use metadata schema and record-level provenance design, and pair them with the notice version in force at collection. For speech and language corpora, a data statement in the University of Washington V2 format records speaker demographics and curation rationale alongside the consent fields [11].

De-identification promises in the form must match the pipeline

The consent form should not promise anonymity that the processing pipeline cannot deliver. Voice is hard to anonymize without hurting the model: the VoicePrivacy 2024 Challenge evaluates speaker anonymization against both privacy attacks and ASR utility, and it remains an open problem [10]. A form that says "your voice cannot be linked to you" while delivering raw 16 kHz WAV files creates a mismatch a regulator or plaintiff can point to.

Use calibrated wording: name the identifiers that will be removed or replaced (name, email, phone, account numbers spoken aloud), state that removal is checked but imperfect, and state whether the voice or face itself remains in the data. Then confirm the speech-to-text transcripts are redacted with the same rules as the audio.

Buyer-side acceptance checks before the first delivery

Review the form before fieldwork begins, not at delivery, because a defective form cannot be fixed retroactively for participants already recorded. A short acceptance gate in the statement of work for a custom data collection keeps this enforceable.

  • Final consent forms, assent scripts and translations attached as SOW exhibits, each with a version ID.
  • Counsel sign-off on biometric, health, minor and all-party recording clauses per collection jurisdiction.
  • A pilot batch of 20 to 50 records whose consent metadata is joined to signed forms and spot-checked.
  • Withdrawal channel tested end to end, including propagation to the buyer.
  • Re-consent plan if the purpose expands beyond the signed scopes.

Track the outcome in a training data use register, and use the provenance hub and the consent management glossary entry for adjacent topics. If you would rather license existing operational records than commission new recordings, compare the options in custom collection vs licensing existing data, or describe your requirement to SourceX for buyers.

Sourcing consented data for AI training through SourceX

SourceX sources operational datasets and new recordings of hands-on work from US companies on request, and every dataset is rights-reviewed for ownership and consents before delivery under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, and nothing is contracted until the supplying company agrees. Describe the data you need on the SourceX buyers page.

Sources

  1. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  2. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  3. Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
  4. Federal Trade Commission, Federal Register, "Children's Online Privacy Protection Rule (Final Rule amendments), 90 FR 16918" (2025). https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule
  5. Justia, "California Penal Code Section 632" (2024). https://law.justia.com/codes/california/code-pen/part-1/title-15/chapter-1-5/section-632
  6. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  7. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  8. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  9. Grauman et al., arXiv, "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
  10. VoicePrivacy Challenge, arXiv, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
  11. University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data