Skip to content

Procurement, samples and ongoing supply

Custom data collection vs licensing existing records

Quick answer

Commission a custom collection when the data you need does not exist yet or must follow your protocol: assigned tasks, demographic or device quotas, and consent written for AI training. License existing records when you need behavior that happened for real, with outcomes, natural frequencies and years of history no protocol can stage. The custom data collection vs off-the-shelf dataset choice also changes the contract: a statement of work with per-unit acceptance on one side, a license grant over records that already exist on the other.

By SourceX Editorial · Updated

Three routes behind "custom vs off-the-shelf"

The question usually hides three routes, not two: hiring a custom data collection service to produce new data to your specification, buying a pre-built catalog dataset, and licensing operational records that a business created for its own work.

Commissioned collectionOff-the-shelf catalog datasetLicensed operational records
Why the data existsProduced for your project under a protocolProduced earlier for sale, for example by a past collection project or an aggregationProduced by a business doing its work: calls, tickets, claims, inspections, code review
Who set the specificationYou, in a statement of work (SOW)The vendor, before you arrivedNobody; you select from what exists
ContractSOW plus collection protocol; ownership of deliverables negotiatedStandard license, often non-exclusiveLicense grant over defined records
Outcomes and historyOnly what happens during the collection windowWhatever the original project capturedReal downstream outcomes (refund issued, claim paid, change reverted) across the system's history
Rights evidenceConsent forms and contributor agreements you designThe vendor's license and representationsNotices, contracts and recording practices in force when the records were made

Off-the-shelf sets can be inspected today, but one vendor that sells them notes they are generally oriented to foundational business areas rather than specific processes [1], which suits general skills more than a narrow workflow. Data origins are compared in licensed vs synthetic vs scraped training data; building in-house is covered in build, buy or synthesize training data.

What a collection protocol controls, and what it distorts

A protocol buys control over coverage and consent; it cannot buy real consequences, and it changes how people behave.

What you control. You set quotas (accents, ages, devices, lighting, rare classes), capture labels at collection time, and write consent for AI training from the first session. A 2022 representativity paper distinguishes data that covers the input space from data that mirrors a target population, and reports coverage as more robust to distribution shift [2].

What it distorts. Behavior performed for a recorder differs from behavior in the wild:

  • Scripted speech. CASPER's authors note that most existing conversational speech datasets contain scripted dialogue, and built an elicitation method for natural conversation [3]. Ask any speech supplier how it elicits speech, not just how many hours it records.
  • Assigned goals. Mind2Web collected more than 2,000 open-ended tasks on 137 real websites across 31 domains, with crowdsourced action sequences [4]. The interfaces are real; the goals were set for the dataset, not by a real need.
  • No consequences. A staged support call ends without a real refund decision, a demonstration claim is never adjudicated, and an assigned coding task is never reverted in production.
  • Channel mismatch. Contributors recording on their own phones do not reproduce contact-center telephone audio, which one speech recognition (ASR) study treats as a distinct acoustic channel needing channel-aware pretraining [5].

Operational records have the opposite profile: they mirror only the population the supplying business serves. Rare events stay rare, fields drift as systems change, and outcome codes are not always reliable labels; see verifying ground truth in operational records.

Matching the route to the data need

The route follows from two questions: does the behavior already happen somewhere, and do you need its consequences?

Data needLean to commissioned collection whenLean to licensed records whenCheck before deciding
Supervised fine-tuning (SFT) demonstrationsNobody produces the target style at work; InstructGPT's SFT stage used labeler-written demonstrations [6]Professionals already write the artifact (claim notes, code reviews, support replies)AI-tool use by writers; final accepted version kept
Agent trajectoriesYou need screen and action traces in a specific environmentSystems already log workflow steps with their resultsAssigned vs real goals; outcomes logged
Speech and voiceYou need balanced accents, wake words, read prompts or studio audioProduction audio is real telephone calls with overlap and hold musicChannel match; recording consent
Images and egocentric videoYou need controlled viewpoints, per-wearer consent or staged rare classesInspection or claims photos already carry an inspector's or adjuster's judgmentRelease scope; faces and location metadata
Evaluation setsItems must be written against a rubric nobody has publishedYou need contamination-resistant real tasksYour held-out items kept out of other deliveries
Rare eventsNatural frequency is too low to collect in timeThe rare event is routine in some business's operationsBase rates in your evaluation slices

Two published datasets show one route each. Ego4D gathered 3,670 hours of egocentric video from 931 consenting camera wearers in 74 locations across 9 countries, with de-identification where needed [7]. SWE-Bench Pro's authors, noting that permissively licensed public repositories are prime candidates for pre-training corpora, kept a held-out set private and built a commercial set from codebases acquired from real startups [8]. Modality-specific comparisons cover speech data, robot recordings, agent trajectories and custom evaluation sets.

Lead time and cost follow different curves

Commissioned collection costs scale with the units produced, and its calendar is set by recruiting and the collection window; licensed records are priced on scope and rights, and their history already exists.

Commissioned collection carries fixed costs before the first usable unit: protocol design, consent and contributor-agreement review, app or device setup, recruiting and a pilot. Variable cost follows participant-hours, sessions or items, plus review and re-collection. History cannot be compressed: a year of seasonal variation takes a year to collect.

Licensed records front-load different work: finding a holder, reviewing a sample, rights review, de-identification and negotiation. Price tends to follow scope, volume, depth of history, permitted uses and exclusivity rather than effort, and delays come from supplier approval and legal review, not production.

Exclusivity differs too. A commissioned dataset can be exclusive by contract, catalog datasets are typically licensed to many buyers (a problem for evaluation), and exclusivity over licensed records is negotiated. Plan the calendar with a realistic licensing timeline and compare totals with total cost of ownership for licensed training data.

Rights work runs in opposite directions

With commissioned collection you write the consent and must prove it for every unit; with existing records the consent was fixed when the records were made, and diligence checks whether it reaches AI training.

Commissioned collection. Collecting voices or faces can bring state biometric laws into play. Illinois' Biometric Information Privacy Act (BIPA) lists voiceprints and face-geometry scans as biometric identifiers, requires written notice, a written release before collection and a published retention schedule, and sets liquidated damages of $1,000 per negligent and $5,000 per intentional or reckless violation; since a 2024 amendment, collecting the same identifier from the same person by the same method more than once counts as one violation [9].

Texas bars capturing a biometric identifier for a commercial purpose without informing the person and getting consent first [10]. In the EU, biometric data processed to uniquely identify someone is a special category under GDPR Article 9(1) and needs an Article 9(2) condition [11]. Voice work adds two more checks:

  • Recruiting transparency. Voice actors interviewed for the PRAC3 study reported difficulty telling legitimate jobs from voice-data collection without consent, especially where platforms allow anonymous clients [12]. Require recruiting under a disclosed purpose.
  • Voice rights. According to a 2024 Morgan Lewis analysis, Tennessee's ELVIS Act, effective 1 July 2024, treats a person's voice, including a simulation of it, as a protected property right [13].

Settle ownership of recordings, transcripts and labels in the SOW; see IP assignment vs license for commissioned datasets and consent language for commissioned collection.

Existing records. The questions move to how the records were made:

  • Recordings: California Penal Code section 632 prohibits recording a confidential communication without the consent of all parties [14], so ask how calls were disclosed when they were recorded.
  • Photos of people: one stock marketplace's contributor policy requires releases to include consent for secondary use, including AI and machine-learning training, where applicable [15]. A release signed years ago for advertising may not reach model training.
  • Health records: HIPAA de-identification uses Safe Harbor, which removes 18 listed identifiers, or Expert Determination [16].
  • Catalog datasets: an audit of more than 1,800 text datasets found license omission above 70% and license error rates above 50% on popular hosting sites [17]. Trace the license, not the listing.

This is the route SourceX works on: it sources operational datasets from US companies and manages the licensing agreement and ongoing purchases. Each dataset goes through rights review, which checks that the business may share the records and that required consents are in place, and is delivered under a license defining the records, uses, term and delivery. Personal details are removed or replaced before delivery, though no de-identification method is perfect. If the records sit in US companies' systems, describe them to SourceX.

Either route must document collection. For high-risk systems in the EU, AI Act Article 10(2) requires governance covering data collection processes and the origin of data [18]; as of October 2026, Regulation (EU) 2026/1744 has postponed the high-risk application dates, reported as 2 December 2027 for Annex III systems [19]. Commissioned collection produces that record directly; licensed records need it from the supplier, traced through chain of title documents.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

An SOW accepts units; a license accepts records

A statement of work buys production, so acceptance tests each delivered unit against the protocol; a license buys access to existing records, so acceptance tests the delivery against the agreed scope and approved sample.

Contract elementCommissioned collection (SOW)Licensed existing records
What you specifyProtocol, quotas, devices, scripts or scenarios, QA planRecord types, fields, date range, volume, permitted uses
Unit of acceptanceEach recording, session or labeled itemThe delivery, checked against a manifest and data dictionary
Quality controlsQualification tasks, honeypots and consensus review; annotator agreement with Krippendorff's alpha, where 1 is perfect reliability and 0 is chance level [20]Sample representativeness, field completeness, de-identification check
Typical defectsOff-script sessions, wrong device, unlinked consent, low agreementMissing fields, date gaps, outcome codes whose meaning changed
Usual remedyRe-collection at the supplier's cost; withheld milestoneReplacement records, re-delivery or credits
Your rights in the outputAssignment or license, negotiated in the SOWDefined by the grant: uses, term, field of use
DocumentationDatasheet covering motivation, composition, collection process and recommended uses [21], plus consent records per unitDatasheet, de-identification method, provenance records

The per-unit records each route delivers also differ.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "commissioned_unit": {
    "unit_id": "CC-PILOT-0417",
    "protocol_version": "v1.3",
    "participant_ref": "P-2291",
    "consent_record_id": "CR-2291-01",
    "consent_scope": ["model_training", "evaluation"],
    "scenario_id": "billing-dispute-04",
    "device": "contributor_phone_app",
    "sample_rate_hz": 16000,
    "qa": {"honeypot_pass": true, "reviewer": "R-07"},
    "outcome": null
  },
  "licensed_record": {
    "record_id": "CALL-2023-118842",
    "source_system": "contact_center_platform_export",
    "created_at": "2023-03-14T15:02:11Z",
    "channel": "pstn",
    "sample_rate_hz": 8000,
    "recording_notice": "ivr_disclosure_v2",
    "resolution_code": "credit_issued",
    "repeat_contact_within_7d": false,
    "deid_method": "names_and_account_numbers_replaced"
  }
}

The commissioned unit carries consent and protocol but no outcome; the licensed record carries a real outcome and channel but depends on its 2023 recording notice. Write the SOW side with the data collection statement of work guide, set thresholds with acceptance criteria for licensed training data, and tie payment to results with milestone payments tied to acceptance.

Combining routes: license the base, commission the gaps

Many programs end up hybrid: real licensed records form the base, and collection or annotation is commissioned only for the slices the base cannot cover.

A workable sequence:

  1. Request a sample of licensed records and score your model on it; see how to request a training data sample.
  2. Slice the results by accent, device, product line or outcome type to find where real data is thin.
  3. Commission collection only for those slices, copying the real channel and conditions.
  4. Where records exist but labels do not, commission annotation on the licensed records instead of new data.
  5. Keep at least one evaluation set made of real records that no collection supplier has seen.

The kinds of data SourceX sources include support and sales histories, engineering records, documents, and finance and legal workflows, as well as new recordings of hands-on work. All of it is sourced on request rather than held in stock, so a request does not guarantee a match.

Red flags on each route

Each route has warning signs that show up before signature.

  • Commissioned: consent records cannot be linked to individual units; the supplier will not name recruiting channels or subcontractors; the pilot was recorded by the supplier's staff rather than the recruited pool; the SOW is silent on whether contributors may use AI tools; the ownership clause is missing.
  • Off-the-shelf: the license on the listing differs from the contract; the sample was hand-picked; it is sold to many buyers and you plan to evaluate on it; no collection date range is stated.
  • Licensed records: the supplier cannot show the notices or terms in force when records were created; outcome fields changed meaning across years; the de-identification method is undocumented.

The full buying process is mapped in the AI training data procurement guide.

Need records that already exist rather than a new collection?

If the behavior you need is already recorded in US businesses' systems, describe it on the SourceX buyer page: the record types, the history and volume, and the uses you need licensed, following how to write a data request for suppliers. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery. Describe the records you need.

Sources

  1. Shaip, "Off-the-Shelf AI Training Data: choosing the right off-the-shelf AI training data provider" (vendor blog). https://www.shaip.com/blog/choosing-the-right-off-the-shelf-ai-training-data-provider
  2. Clemmensen and Kjærsgaard, "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
  3. arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
  4. Deng, Su et al., The Ohio State University, "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
  5. arXiv, "Channel-Aware Pretraining of Joint Encoder-Decoder Self-Supervised Model for Telephonic-Speech ASR" (2022). https://arxiv.org/pdf/2211.01669
  6. Ouyang et al., OpenAI, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  7. Grauman et al., Ego4D consortium, "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
  8. arXiv, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  9. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  10. Texas Legislature, "Texas Business and Commerce Code Section 503.001: Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  11. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  12. arXiv, "PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy" (2025). https://arxiv.org/pdf/2507.16247
  13. Morgan Lewis, "Rise of Text-to-Speech AI Models, Part 1: Intellectual Property Issues" (2024). https://www.morganlewis.com/blogs/sourcingatmorganlewis/2024/07/rise-of-text-to-speech-ai-models-part-1-intellectual-property-issues
  14. California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  15. pocstock, "Model release" (contributor policy). https://pocstock.com/legal/model-release
  16. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  17. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  18. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  19. European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  20. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  21. Gebru et al., "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data