Skip to content

Procurement, samples and ongoing supply

Statement of Work for a Custom Data Collection Project

Quick answer

A data collection statement of work (SOW) is the contract exhibit that turns a selected vendor's proposal into measurable obligations. It should fix six things: the collection protocol, who or what is recorded and in what quotas, how consent is captured and evidenced, the exact deliverable schema, batch-level acceptance tests with remedies, and a change-order process. If any of these stays in an email thread instead of the SOW, expect disputes when the first batch arrives.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Where the SOW sits between the RFP and the license

The SOW comes after vendor selection and describes commissioned work that does not yet exist, which is why it differs from both an RFP and a data license. Your AI training data RFP solicited bids against a scope; the SOW freezes the winning scope as numbered requirements. Licensing records a company already holds uses a different contract shape, covered in custom data collection vs licensing existing records.

Most teams attach the SOW to a master services agreement (MSA) that holds liability, confidentiality, data protection and IP terms. The SOW then carries the project-specific parts. Keep IP ownership language in one place: if the MSA assigns rights, the SOW should reference it rather than restate it differently.

Project context and model purpose

Start the SOW with a short, specific statement of what the data will train or evaluate, because vague context produces vague delivery just as it produces vague bids [1]. "Conversational ASR for field technicians dictating repair notes in noisy plant environments, English with Spanish code-switching" lets a vendor plan microphones, sites and scripts. "Speech data for ASR" does not.

State the intended use type (pre-training, supervised fine-tuning, evaluation, or retrieval test sets) and whether any portion must be held out. Evaluation sets need stricter isolation than training data, so name a held-out split and forbid the vendor from reusing those prompts, speakers or documents elsewhere. For model-goal translation, see turning model goals into training data requirements.

Collection protocol and participant quotas

The protocol section should be detailed enough that a second vendor could repeat the collection and get comparable data. Specify capture devices and settings (for audio: sample rate, bit depth, channel count, codec, microphone class; for images or video: resolution, frame rate, lens, lighting conditions), the task script or elicitation method, session length, and environment.

Quotas belong in a table, not prose. Define each stratum (age band, gender, accent or dialect region, device type, environment, document type, task category), the minimum and maximum share per stratum, and how the vendor reports progress against quotas each week. Cap per-participant contribution so a handful of prolific contributors cannot dominate the set; without a cap, a small group of fast, repeat contributors can end up supplying a large share of the hours.

Include eligibility and exclusion rules: minimum age, language proficiency checks, no employees of the buyer, and a rule against the same person enrolling under two identities. Require the vendor to record a stable pseudonymous participant ID so you can deduplicate and split by speaker.

The SOW should require consent that names AI training and evaluation as the purpose, and it should require the vendor to deliver evidence of that consent with each batch. Market practice already includes a final QA check on consent and metadata before delivery [2]; write that check into your acceptance tests rather than trusting it.

Biometric and health-adjacent captures raise the bar. Texas law defines voiceprints and records of face geometry as biometric identifiers and requires notice and consent before capture for a commercial purpose [4]. Washington's My Health My Data Act defines consumer health data broadly enough to include biometric data and data that can indicate health status [5]. If minors could be involved, the amended COPPA Rule, with a compliance date of 22 April 2026, adds separate parental consent requirements for certain disclosures [6]; most buyers simply exclude under-18 participants.

Specify the evidence format: consent form version ID, timestamp, participant ID, jurisdiction, language of the form, and a withdrawal flag. Require a withdrawal procedure with a defined window for the vendor to notify you and remove affected records from future batches. For wording, see consent language for commissioned AI data collection, and for workplace screen or keystroke capture, collecting computer-use demonstrations at work.

Deliverables, schema and documentation

Deliverables should be defined as files with a schema, not as "a dataset." Name the container formats (WAV or FLAC for audio, PNG or lossless video for vision, JSONL for conversational or SFT records, Parquet for manifests and metadata tables [9]), the directory layout, naming conventions and checksums (SHA-256 per file).

Require a manifest per batch that lists every record with its participant ID, consent reference, quota stratum, capture device, collection date and annotation status. Ask for dataset documentation in a machine-readable format; the Croissant-RAI vocabulary extends MLCommons Croissant to cover responsible-AI metadata such as collection process and labeling [8]. If annotation is in scope, attach the labeling guideline as a versioned exhibit and reference ISO/IEC 5259-4 as the process framework for data quality across labeling and training data [3].

Delivery mechanics belong here too: transfer method, access control, encryption in transit and at rest, and who may receive the data. Email attachments should be prohibited.

Illustrative example: invented to show structure; it does not describe an available dataset.

SOW sectionWhat to fix in writingExample requirement
1. ContextModel purpose, use type, held-out splitField-repair dictation ASR; 10% speaker-disjoint eval split never reused
2. ProtocolDevices, settings, scripts, environments16 kHz, 16-bit mono FLAC; headset and phone mics; 3 noise conditions
3. QuotasStrata, min/max share, per-person cap6 accent regions, each 12–22%; no speaker above 0.5% of hours
4. ConsentPurpose wording, evidence fields, withdrawalForm v3 names AI training; withdrawal notice within the agreed window
5. DeliverablesFormats, layout, manifest, checksums, docsPer-batch Parquet manifest, SHA-256 per file, Croissant-RAI metadata
6. AnnotationGuideline version, label set, QA samplingGuideline v1.2; verbatim transcripts with noise tags
7. AcceptanceTests, sample size, thresholds, remedies2% stratified sample; WER audit; failed batch re-collected at vendor cost
8. PilotSize, review period, go/no-go gatePilot batch of about 5% of total; scope frozen after sign-off
9. Change controlRequest form, pricing basis, approversWritten change order signed by both project leads before work starts
10. RightsOwnership or license, retention, deletionPer MSA; vendor deletes raw copies after acceptance plus certificate

Pilot batch and batch-level acceptance

Run a paid pilot before full-scale collection, and make sign-off on the pilot the gate that freezes protocol and guidelines. The pilot exposes protocol defects cheaply: clipped audio, scripts participants misread, quota strata that are hard to fill, annotation guidelines that two labelers interpret differently.

After the pilot, accept per batch rather than at the end. Each batch acceptance test should state the sample size and how it is drawn (stratified across quota cells), the checks run (schema validation, checksum match, duplicate and near-duplicate detection, consent evidence present for 100% of records, quota compliance, annotation accuracy against a gold set), the pass threshold, and the remedy. Remedies typically run from re-annotation to re-collection to a price reduction. Also set a review window and state whether silence after it counts as acceptance. More detail is in acceptance criteria for licensed training data.

Separate "rejected" from "accepted with deviations." Some batches will miss a quota cell by a small margin; deciding in advance what tolerance is acceptable avoids renegotiating each time.

Change orders and scope drift

Every change to protocol, quotas, guidelines, volume or schedule should go through a written change order before work starts, because collection changes ripple into data already delivered. Adding a new noise condition midway, for example, can make earlier batches unrepresentative and force a rebalancing that costs more than the change itself.

The change-order clause should define the request form, who can approve on each side, the pricing basis (unit rates from the SOW's rate card versus a new quote), the impact statement on timeline and on already-accepted batches, and whether guideline changes apply retroactively. Version every guideline and protocol document and require each delivered record to carry the version it was produced under.

Rights, ownership and data handling

Decide in the SOW whether you take ownership of the collected data or a license to it, and make sure participant consent and vendor subcontracts support that choice. Under 17 U.S.C. 201(b), the employer or other person for whom a work made for hire was prepared is considered the author [7], but commissioned data collections often do not fit that category cleanly, so counsel usually adds an express assignment or license; see IP assignment vs license for commissioned datasets.

Require the vendor to flow down consent, confidentiality and deletion obligations to subcontracted recruiters and annotators, list those subcontractors, and certify deletion of raw copies after acceptance. If the data will feed an AI system placed on the EU market as high-risk, the EU AI Act's Article 10 data governance duties ask providers to document how training data was collected and prepared (as of October 2026, most high-risk obligations apply from 2 December 2027 or later, after the Reg. (EU) 2026/1744 amendment); the SOW's protocol, quota and manifest records are where that evidence originates. Vet the vendor's controls with the data provider due diligence questionnaire.

When collection is the wrong tool

A custom collection SOW suits data that does not exist yet: scripted speech, staged scenes, task demonstrations. If the behavior you need is already captured in real business operations, such as support and sales histories, engineering records, or finance and legal workflows, licensing those records may give more realistic distributions than a staged collection. SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the commercial process through licensing and ongoing purchases; you can describe the data you need to SourceX. Datasets are not held in stock, and a request does not guarantee a match. Weigh the options in build, buy or synthesize training data and the wider AI training data procurement guide.

For writing the initial request to suppliers, the SourceX guide on how to write a data request for suppliers covers the description that precedes any SOW.

Scoping a data collection SOW with real operational data in mind

If part of your requirement could be met by real operational records or new recordings of hands-on work, SourceX can look for US businesses that hold the data you describe. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Describe the data your collection project needs.

Sources

  1. Ertas, "How to Scope an AI Data Preparation Project (RFP Template)". https://www.ertas.ai/blog/ai-data-preparation-rfp-template
  2. Pocstock, "The dataset licensing process from inquiry to delivery". https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
  3. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  4. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  5. Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
  6. Loeb & Loeb, "Children's Online Privacy in 2025: The Amended COPPA Rule" (2025). https://www.loeb.com/en/insights/publications/2025/05/childrens-online-privacy-in-2025-the-amended-coppa-rule
  7. U.S. Government Publishing Office (govinfo), "17 U.S.C. 201 - Ownership of copyright" (2024). https://www.govinfo.gov/content/pkg/USCODE-2024-title17/html/USCODE-2024-title17-chap2-sec201.htm
  8. Jain et al., MLCommons (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  9. The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data