Procurement, samples and ongoing supply
Custom data collection vs licensing existing records
Quick answer
Commission a custom collection when the data you need does not exist yet or must follow your protocol: assigned tasks, demographic or device quotas, and consent written for AI training. License existing records when you need behavior that happened for real, with outcomes, natural frequencies and years of history no protocol can stage. The custom data collection vs off-the-shelf dataset choice also changes the contract: a statement of work with per-unit acceptance on one side, a license grant over records that already exist on the other.
By SourceX Editorial · Updated
Three routes behind "custom vs off-the-shelf"
The question usually hides three routes, not two: hiring a custom data collection service to produce new data to your specification, buying a pre-built catalog dataset, and licensing operational records that a business created for its own work.
| Commissioned collection | Off-the-shelf catalog dataset | Licensed operational records | |
|---|---|---|---|
| Why the data exists | Produced for your project under a protocol | Produced earlier for sale, for example by a past collection project or an aggregation | Produced by a business doing its work: calls, tickets, claims, inspections, code review |
| Who set the specification | You, in a statement of work (SOW) | The vendor, before you arrived | Nobody; you select from what exists |
| Contract | SOW plus collection protocol; ownership of deliverables negotiated | Standard license, often non-exclusive | License grant over defined records |
| Outcomes and history | Only what happens during the collection window | Whatever the original project captured | Real downstream outcomes (refund issued, claim paid, change reverted) across the system's history |
| Rights evidence | Consent forms and contributor agreements you design | The vendor's license and representations | Notices, contracts and recording practices in force when the records were made |
Off-the-shelf sets can be inspected today, but one vendor that sells them notes they are generally oriented to foundational business areas rather than specific processes [1], which suits general skills more than a narrow workflow. Data origins are compared in licensed vs synthetic vs scraped training data; building in-house is covered in build, buy or synthesize training data.
What a collection protocol controls, and what it distorts
A protocol buys control over coverage and consent; it cannot buy real consequences, and it changes how people behave.
What you control. You set quotas (accents, ages, devices, lighting, rare classes), capture labels at collection time, and write consent for AI training from the first session. A 2022 representativity paper distinguishes data that covers the input space from data that mirrors a target population, and reports coverage as more robust to distribution shift [2].
What it distorts. Behavior performed for a recorder differs from behavior in the wild:
- Scripted speech. CASPER's authors note that most existing conversational speech datasets contain scripted dialogue, and built an elicitation method for natural conversation [3]. Ask any speech supplier how it elicits speech, not just how many hours it records.
- Assigned goals. Mind2Web collected more than 2,000 open-ended tasks on 137 real websites across 31 domains, with crowdsourced action sequences [4]. The interfaces are real; the goals were set for the dataset, not by a real need.
- No consequences. A staged support call ends without a real refund decision, a demonstration claim is never adjudicated, and an assigned coding task is never reverted in production.
- Channel mismatch. Contributors recording on their own phones do not reproduce contact-center telephone audio, which one speech recognition (ASR) study treats as a distinct acoustic channel needing channel-aware pretraining [5].
Operational records have the opposite profile: they mirror only the population the supplying business serves. Rare events stay rare, fields drift as systems change, and outcome codes are not always reliable labels; see verifying ground truth in operational records.
Matching the route to the data need
The route follows from two questions: does the behavior already happen somewhere, and do you need its consequences?
| Data need | Lean to commissioned collection when | Lean to licensed records when | Check before deciding |
|---|---|---|---|
| Supervised fine-tuning (SFT) demonstrations | Nobody produces the target style at work; InstructGPT's SFT stage used labeler-written demonstrations [6] | Professionals already write the artifact (claim notes, code reviews, support replies) | AI-tool use by writers; final accepted version kept |
| Agent trajectories | You need screen and action traces in a specific environment | Systems already log workflow steps with their results | Assigned vs real goals; outcomes logged |
| Speech and voice | You need balanced accents, wake words, read prompts or studio audio | Production audio is real telephone calls with overlap and hold music | Channel match; recording consent |
| Images and egocentric video | You need controlled viewpoints, per-wearer consent or staged rare classes | Inspection or claims photos already carry an inspector's or adjuster's judgment | Release scope; faces and location metadata |
| Evaluation sets | Items must be written against a rubric nobody has published | You need contamination-resistant real tasks | Your held-out items kept out of other deliveries |
| Rare events | Natural frequency is too low to collect in time | The rare event is routine in some business's operations | Base rates in your evaluation slices |
Two published datasets show one route each. Ego4D gathered 3,670 hours of egocentric video from 931 consenting camera wearers in 74 locations across 9 countries, with de-identification where needed [7]. SWE-Bench Pro's authors, noting that permissively licensed public repositories are prime candidates for pre-training corpora, kept a held-out set private and built a commercial set from codebases acquired from real startups [8]. Modality-specific comparisons cover speech data, robot recordings, agent trajectories and custom evaluation sets.
Lead time and cost follow different curves
Commissioned collection costs scale with the units produced, and its calendar is set by recruiting and the collection window; licensed records are priced on scope and rights, and their history already exists.
Commissioned collection carries fixed costs before the first usable unit: protocol design, consent and contributor-agreement review, app or device setup, recruiting and a pilot. Variable cost follows participant-hours, sessions or items, plus review and re-collection. History cannot be compressed: a year of seasonal variation takes a year to collect.
Licensed records front-load different work: finding a holder, reviewing a sample, rights review, de-identification and negotiation. Price tends to follow scope, volume, depth of history, permitted uses and exclusivity rather than effort, and delays come from supplier approval and legal review, not production.
Exclusivity differs too. A commissioned dataset can be exclusive by contract, catalog datasets are typically licensed to many buyers (a problem for evaluation), and exclusivity over licensed records is negotiated. Plan the calendar with a realistic licensing timeline and compare totals with total cost of ownership for licensed training data.
Rights work runs in opposite directions
With commissioned collection you write the consent and must prove it for every unit; with existing records the consent was fixed when the records were made, and diligence checks whether it reaches AI training.
Commissioned collection. Collecting voices or faces can bring state biometric laws into play. Illinois' Biometric Information Privacy Act (BIPA) lists voiceprints and face-geometry scans as biometric identifiers, requires written notice, a written release before collection and a published retention schedule, and sets liquidated damages of $1,000 per negligent and $5,000 per intentional or reckless violation; since a 2024 amendment, collecting the same identifier from the same person by the same method more than once counts as one violation [9].
Texas bars capturing a biometric identifier for a commercial purpose without informing the person and getting consent first [10]. In the EU, biometric data processed to uniquely identify someone is a special category under GDPR Article 9(1) and needs an Article 9(2) condition [11]. Voice work adds two more checks:
- Recruiting transparency. Voice actors interviewed for the PRAC3 study reported difficulty telling legitimate jobs from voice-data collection without consent, especially where platforms allow anonymous clients [12]. Require recruiting under a disclosed purpose.
- Voice rights. According to a 2024 Morgan Lewis analysis, Tennessee's ELVIS Act, effective 1 July 2024, treats a person's voice, including a simulation of it, as a protected property right [13].
Settle ownership of recordings, transcripts and labels in the SOW; see IP assignment vs license for commissioned datasets and consent language for commissioned collection.
Existing records. The questions move to how the records were made:
- Recordings: California Penal Code section 632 prohibits recording a confidential communication without the consent of all parties [14], so ask how calls were disclosed when they were recorded.
- Photos of people: one stock marketplace's contributor policy requires releases to include consent for secondary use, including AI and machine-learning training, where applicable [15]. A release signed years ago for advertising may not reach model training.
- Health records: HIPAA de-identification uses Safe Harbor, which removes 18 listed identifiers, or Expert Determination [16].
- Catalog datasets: an audit of more than 1,800 text datasets found license omission above 70% and license error rates above 50% on popular hosting sites [17]. Trace the license, not the listing.
This is the route SourceX works on: it sources operational datasets from US companies and manages the licensing agreement and ongoing purchases. Each dataset goes through rights review, which checks that the business may share the records and that required consents are in place, and is delivered under a license defining the records, uses, term and delivery. Personal details are removed or replaced before delivery, though no de-identification method is perfect. If the records sit in US companies' systems, describe them to SourceX.
Either route must document collection. For high-risk systems in the EU, AI Act Article 10(2) requires governance covering data collection processes and the origin of data [18]; as of October 2026, Regulation (EU) 2026/1744 has postponed the high-risk application dates, reported as 2 December 2027 for Annex III systems [19]. Commissioned collection produces that record directly; licensed records need it from the supplier, traced through chain of title documents.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
An SOW accepts units; a license accepts records
A statement of work buys production, so acceptance tests each delivered unit against the protocol; a license buys access to existing records, so acceptance tests the delivery against the agreed scope and approved sample.
| Contract element | Commissioned collection (SOW) | Licensed existing records |
|---|---|---|
| What you specify | Protocol, quotas, devices, scripts or scenarios, QA plan | Record types, fields, date range, volume, permitted uses |
| Unit of acceptance | Each recording, session or labeled item | The delivery, checked against a manifest and data dictionary |
| Quality controls | Qualification tasks, honeypots and consensus review; annotator agreement with Krippendorff's alpha, where 1 is perfect reliability and 0 is chance level [20] | Sample representativeness, field completeness, de-identification check |
| Typical defects | Off-script sessions, wrong device, unlinked consent, low agreement | Missing fields, date gaps, outcome codes whose meaning changed |
| Usual remedy | Re-collection at the supplier's cost; withheld milestone | Replacement records, re-delivery or credits |
| Your rights in the output | Assignment or license, negotiated in the SOW | Defined by the grant: uses, term, field of use |
| Documentation | Datasheet covering motivation, composition, collection process and recommended uses [21], plus consent records per unit | Datasheet, de-identification method, provenance records |
The per-unit records each route delivers also differ.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"commissioned_unit": {
"unit_id": "CC-PILOT-0417",
"protocol_version": "v1.3",
"participant_ref": "P-2291",
"consent_record_id": "CR-2291-01",
"consent_scope": ["model_training", "evaluation"],
"scenario_id": "billing-dispute-04",
"device": "contributor_phone_app",
"sample_rate_hz": 16000,
"qa": {"honeypot_pass": true, "reviewer": "R-07"},
"outcome": null
},
"licensed_record": {
"record_id": "CALL-2023-118842",
"source_system": "contact_center_platform_export",
"created_at": "2023-03-14T15:02:11Z",
"channel": "pstn",
"sample_rate_hz": 8000,
"recording_notice": "ivr_disclosure_v2",
"resolution_code": "credit_issued",
"repeat_contact_within_7d": false,
"deid_method": "names_and_account_numbers_replaced"
}
}
The commissioned unit carries consent and protocol but no outcome; the licensed record carries a real outcome and channel but depends on its 2023 recording notice. Write the SOW side with the data collection statement of work guide, set thresholds with acceptance criteria for licensed training data, and tie payment to results with milestone payments tied to acceptance.
Combining routes: license the base, commission the gaps
Many programs end up hybrid: real licensed records form the base, and collection or annotation is commissioned only for the slices the base cannot cover.
A workable sequence:
- Request a sample of licensed records and score your model on it; see how to request a training data sample.
- Slice the results by accent, device, product line or outcome type to find where real data is thin.
- Commission collection only for those slices, copying the real channel and conditions.
- Where records exist but labels do not, commission annotation on the licensed records instead of new data.
- Keep at least one evaluation set made of real records that no collection supplier has seen.
The kinds of data SourceX sources include support and sales histories, engineering records, documents, and finance and legal workflows, as well as new recordings of hands-on work. All of it is sourced on request rather than held in stock, so a request does not guarantee a match.
Red flags on each route
Each route has warning signs that show up before signature.
- Commissioned: consent records cannot be linked to individual units; the supplier will not name recruiting channels or subcontractors; the pilot was recorded by the supplier's staff rather than the recruited pool; the SOW is silent on whether contributors may use AI tools; the ownership clause is missing.
- Off-the-shelf: the license on the listing differs from the contract; the sample was hand-picked; it is sold to many buyers and you plan to evaluate on it; no collection date range is stated.
- Licensed records: the supplier cannot show the notices or terms in force when records were created; outcome fields changed meaning across years; the de-identification method is undocumented.
The full buying process is mapped in the AI training data procurement guide.
Need records that already exist rather than a new collection?
If the behavior you need is already recorded in US businesses' systems, describe it on the SourceX buyer page: the record types, the history and volume, and the uses you need licensed, following how to write a data request for suppliers. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery. Describe the records you need.
Sources
- Shaip, "Off-the-Shelf AI Training Data: choosing the right off-the-shelf AI training data provider" (vendor blog). https://www.shaip.com/blog/choosing-the-right-off-the-shelf-ai-training-data-provider
- Clemmensen and Kjærsgaard, "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
- arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
- Deng, Su et al., The Ohio State University, "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
- arXiv, "Channel-Aware Pretraining of Joint Encoder-Decoder Self-Supervised Model for Telephonic-Speech ASR" (2022). https://arxiv.org/pdf/2211.01669
- Ouyang et al., OpenAI, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Grauman et al., Ego4D consortium, "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
- arXiv, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- Texas Legislature, "Texas Business and Commerce Code Section 503.001: Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- arXiv, "PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy" (2025). https://arxiv.org/pdf/2507.16247
- Morgan Lewis, "Rise of Text-to-Speech AI Models, Part 1: Intellectual Property Issues" (2024). https://www.morganlewis.com/blogs/sourcingatmorganlewis/2024/07/rise-of-text-to-speech-ai-models-part-1-intellectual-property-issues
- California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- pocstock, "Model release" (contributor policy). https://pocstock.com/legal/model-release
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Gebru et al., "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.