Procurement, samples and ongoing supply
Data vendor evaluation scorecard: knockout gates, weighted criteria and anchored scores
Quick answer
A data vendor evaluation scorecard turns competing supplier bids into a defensible award in three passes. Knockout gates first remove bids without a documented rights basis, a lawful source, the required de-identification standard, minimum security controls or a random sample. Reviewers then score six criteria families on anchored 0-4 evidence scales. Finally, weights fixed before proposals open, chosen for your training stage, produce a total that you test for sensitivity and record with its evidence.
By SourceX Editorial · Updated
The scorecard chooses among bidders answering the same AI training data RFP. Use questions to ask a training data vendor to shortlist, a supplier performance scorecard after award, and a model bake-off to choose a model provider.
Knockout gates: failures no score can offset
Apply pass/fail gates before anyone scores, because a weighted sum lets low price or large volume offset data you cannot lawfully train on. The cost can reach the model: the FTC's 2021 order against Everalbum required deletion of the models and algorithms built from users' photos and videos [2].
| Gate | Passes when | Evidence to see | Owner |
|---|---|---|---|
| G1 Rights basis | Bidder owns the records or holds upstream rights for your uses | Chain-of-title memo; resellers' upstream resale terms | Counsel |
| G2 Lawful source | No pirated copies; personal data collected under notices covering disclosure for AI training | Source list; notices in force at collection | Counsel, privacy |
| G3 De-identification | Regulated data meets your standard | Method record, such as HIPAA Safe Harbor or Expert Determination | Privacy |
| G4 Permitted uses | License names your stage: training, evaluation, retrieval, derived models | License markup | Counsel |
| G5 Security minimum | Your storage and transfer controls are met | Questionnaire, report scope | Security |
| G6 Random sample | Random draw from the offered segment, under evaluation terms | Sampling method, manifest | ML or data lead |
| G7 No conflicting grant | No exclusive license blocking your field of use | Written representation | Counsel |
The Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [3], so a dataset card's license field cannot pass G1. For G2, the class settlement in Bartz v. Anthropic received final approval in July 2026, according to Authors Alliance [4]; it is a settlement, not a merits ruling. FTC staff warned in February 2024 that adopting AI training through a surreptitious, retroactive change to terms or a privacy policy may be unfair or deceptive [5], hence the notice in force at collection. For G3, HIPAA Safe Harbor removes 18 listed identifiers; Expert Determination needs a qualified expert to find identification risk very small [6].
The FISD Alternative Data Council's industry questionnaire shows the evidence format: a data dictionary, a small sample, consent terms agreed with individuals and, for data bought from others, the terms that allow resale [7]. Collect them through a data provider due diligence questionnaire and a training data sample request.
Six criteria families and the evidence that earns points
Score what the bidder can show for this dataset, not what it says about itself; weighted criteria reward depth above the gates.
| Family | What you score | Evidence that earns points | Earns nothing |
|---|---|---|---|
| Data fit | Fields against your data dictionary, population, time window, usable volume, languages, label taxonomy | Fill rates and distributions measured on the random draw | Category names, record counts |
| Rights and provenance | Title depth, upstream licenses, notice coverage, permitted uses, disclosure support | Dataset-specific documents, then your trace of sample records | "Rights-cleared" |
| Quality evidence | Schema conformance, duplicates, label accuracy, annotator agreement, contamination | Measures with named methods, reproduced in your pilot | Accuracy without a method |
| Privacy and security | De-identification method, residual risk, access control, subprocessors, deletion | Method record, re-identification tests, security report scope | Out-of-scope certificates |
| Delivery and continuity | Formats (Parquet, JSON Lines), schema versioning, transfer, refresh cadence, capacity | Milestone plan; past schema change notices | "Flexible delivery" |
| Commercial terms | Cost per usable record, license structure, warranties, exit | Price form converted with pilot pass rates | Quoted totals |
Look beyond price per label to data lineage, security certifications, workforce specialization and operational maturity.
Quality claims score only when measurable. ISO/IEC 5259-2 sets out data quality measures [8]; require each claim to name the property and the method used to quantify it. A Hugging Face forum discussion draws the same line: "the exact-duplicate rate is below 1%" can be checked, "good for fine-tuning" cannot [9]. Treat nothing a vendor says as evidence until you test it on your own data [1], with paid pilot terms where a free sample is too small.
Score whether a supplier can feed your training-data disclosures
A supplier that cannot document sources, collection periods and processing steps leaves gaps in disclosures developers now publish, so score disclosure support within rights and provenance. As of October 2026:
- EU general-purpose AI models. AI Act Article 53(1)(d) requires providers to publish a sufficiently detailed training-content summary on the AI Office template [13].
- California. AB 2013 requires developers of generative AI systems offered to Californians to post a high-level dataset summary covering items such as sources or owners, copyrighted, licensed or personal information, and collection periods; first postings were due 1 January 2026 [14].
- EU high-risk systems. Article 10(2) requires governance of collection processes and data origin, preparation such as annotation, labelling and cleaning, and examination of possible biases [15]. Regulation (EU) 2026/1744, published 24 July 2026, reportedly moves Annex III high-risk obligations to 2 December 2027 and also amends Article 10 [16].
A datasheet covering motivation, composition, collection process and recommended uses [17] earns a 2. Machine-readable metadata checked against the sample earns a 3; NeurIPS 2026 requires Croissant-based responsible-AI fields for its Evaluations and Datasets Track [18]. A warranty on the facts you will publish earns a 4.
Anchored 0-4 scales that make reviewers agree
Anchor every score to a level of evidence so that reviewers reading the same file land within a point of each other. One ladder serves all six families:
| Score | Evidence level | What it looks like |
|---|---|---|
| 0 | Missing or contradicted | No answer, or the sample contradicts it |
| 1 | Asserted | "Verified" or "high quality", nothing dataset-specific |
| 2 | Documented | Data dictionary, datasheet, notice text, upstream license list or method record for this dataset |
| 3 | Tested by you | Confirmed on a random draw, within your threshold |
| 4 | Tested and committed | As 3, and accepted as a warranty, acceptance criterion or remedy |
Quality and rights evidence get their own anchors:
| Score | Quality evidence | Rights and provenance |
|---|---|---|
| 1 | "98% accurate", no method | License field on a dataset card |
| 2 | QA report naming measures: gold-set label accuracy, Krippendorff's alpha, near-duplicate rate | Chain-of-title memo, upstream agreements, dated notice text |
| 3 | Your pilot reproduces the measures within tolerance | Your provenance test on sample records traces them to source systems |
| 4 | Measures become acceptance criteria with remedies | Title and collection warranties, full disclosure schedule |
Krippendorff's alpha measures agreement among annotators labeling the same items: 1 is perfect reliability, 0 is chance [11]. A hand-picked sample caps quality at 2; check whether the sample is representative.
Weight profiles for pre-training, fine-tuning, evaluation and RAG
Set weights by what the purchase must do and freeze them before proposals open; these starting profiles sum to 100.
| Criterion family | Pre-training | Fine-tuning (SFT) | Evaluation | RAG |
|---|---|---|---|---|
| Data fit | 25 | 20 | 25 | 20 |
| Rights and provenance | 20 | 15 | 15 | 20 |
| Quality evidence | 15 | 30 | 30 | 15 |
| Privacy and security | 10 | 10 | 10 | 15 |
| Delivery and continuity | 10 | 10 | 5 | 20 |
| Commercial terms | 20 | 15 | 15 | 10 |
- Pre-training buys volume, so duplication and usable-record price matter. Lee et al. found one sentence repeated over 60,000 times in C4, and deduplicated training cut memorized output about tenfold [12].
- Fine-tuning is quality-led: LIMA fine-tuned a 65B-parameter model on 1,000 curated prompt-response pairs [19], so example quality counts for more than volume.
- Evaluation data must be correct and unseen. Northcutt et al. estimated average label error of at least 3.3% across test sets of 10 widely used datasets and showed that such errors can change model rankings [10]. Score contamination checks under quality.
- RAG content changes, so refresh cadence, deletion that reaches your index and display rights carry weight; see RAG content licensing.
Worked example: two bids that survive the gates
Score price per usable record, and test any close total for sensitivity.
Illustrative example: invented to show structure; it does not describe an available dataset.
A buyer scores bids for resolved support-ticket threads for supervised fine-tuning, using the SFT profile. Bidder C fails G2 (its transcripts' collection notice excluded third-party disclosure) and is not scored. A and B each offer 100,000 threads at price indexes of 100 and 75. In the pilot, 92% of A's threads and 64% of B's pass acceptance checks, so A costs 1.087 per 1,000 usable threads and B 1.172.
Price score = 4 x (lowest cost per usable record / bidder's cost per usable record).
| Family (weight) | A | B |
|---|---|---|
| Data fit (20) | 3 | 4 |
| Rights and provenance (15) | 4 | 3 |
| Quality evidence (30) | 3 | 2 |
| Privacy and security (10) | 3 | 3 |
| Delivery and continuity (10) | 2 | 4 |
| Commercial, on quoted price (15) | 3.0 | 4.0 |
| Commercial, per usable record (15) | 4.0 | 3.71 |
| Total on quoted price (of 100) | 76.3 | 78.8 |
| Total per usable record (of 100) | 80.0 | 77.7 |
Total = sum of weight x score, divided by 4. On quoted prices B wins; per usable record A wins by 2.3 points. Sensitivity check: moving 10 points from quality evidence to data fit gives A 80.0 and B 82.7, so the award rests on the SFT profile's quality emphasis, which the panel must record. Comparing vendor quotes per usable record covers price normalization.
Running the panel and breaking ties
Reviewers judge only the gates and families they are qualified for and score independently, citing a reason and evidence document ID per score, before the panel calibrates.
| Reviewer | Gates | Families scored |
|---|---|---|
| ML or data lead | G6 | Data fit, quality evidence |
| Counsel | G1, G2, G4, G7 | Rights and provenance, license terms within commercial |
| Privacy | G3 | Privacy half of privacy and security |
| Security | G5 | Security half; see the security review of a data supplier |
| Procurement | None | Cost per usable record, delivery and continuity |
- Freeze specification, gates, weights, anchors and tie-break rule; publish gates and weights in the RFP.
- Record gate failures with evidence; failed bids are not scored.
- Calibrate spreads of 2 or more points; changes need a written reason.
- Open prices after technical scores lock, where policy allows; convert with pilot pass rates.
- Run the sensitivity check, then the tie-break.
Set the tie margin in advance, for example 3 points of 100. Within it, prefer the higher rights and provenance score, then the higher quality evidence score; if still level, run a second pilot or split the award under a multi-supplier strategy. A's 2.3-point lead in the worked example falls inside it, and A also wins the first tie-break.
Score datasets sourced through SourceX against the same gates. SourceX looks for US businesses holding the data you describe; rights review checks that each business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. Score them at 2 until your sample tests earn a 3.
Personal details are removed or replaced before delivery, but no de-identification method is perfect, so privacy still scores residual risk. You can describe the records you are scoring bids for; a request does not guarantee a matching dataset.
The selection record to keep
Keep one record per award that an auditor or successor can re-run without the panel. NIST AI 600-1 lists data privacy, intellectual property, and value chain and component integration among the 12 risks it identifies for generative AI [20]; a kept scorecard shows how third-party data was assessed against them.
Illustrative example: invented to show structure; it does not describe an available dataset.
selection_id: DS-2026-031
specification: "support-thread SFT spec v1.4"
weights_frozen: 2026-08-03 # before proposals opened on 2026-08-17
weight_profile: sft # fit 20, rights 15, quality 30, privacy/security 10, delivery 10, commercial 15
gates:
bidder_A: pass
bidder_B: pass
bidder_C: {result: fail, gate: G2, evidence: "EV-C-07 privacy notice excludes third-party disclosure"}
scores: # consensus after calibration; reviewer sheets attached
bidder_A: {total: 80.0, price_basis: cost_per_usable_record}
bidder_B: {total: 77.7, price_basis: cost_per_usable_record}
sensitivity: "B leads if data-fit weight is 30 and quality weight is 20"
decision: "Award A: pilot label accuracy and rights traced on sample outweigh B's coverage"
carried_into_contract:
- "pilot thresholds become acceptance criteria"
- "title and collection warranty with disclosure schedule"
open_risks:
- "A delivery capacity scored 2: monthly milestones with remedies"
Scoring errors that pick the wrong supplier
These errors let claims or quoted prices outweigh tested evidence:
- Pricing quoted totals. The worked example flips on this alone.
- Moving weights after proposals open. The scorecard becomes a justification.
- Averaging away a 0. A 0 from privacy or counsel triggers review, not a mean.
- Double counting. One datasheet scored under rights and quality inflates documentation-heavy bids.
SourceX's guides to evaluating data supplier quality and running a data pilot cover quality signals and pilots; its data quality scorecard shows how data owners rate their own records' readiness. The procurement hub places selection among the other purchase stages.
Evaluating data suppliers for a purchase?
Describe the records, fields, volume, time window and uses your scorecard is built around. SourceX looks for US companies that hold that data, checks the data and each supplier's licensing permissions, and manages the license in which pricing and allowed uses are agreed; nothing is contracted until a supplier agrees. See how SourceX works with data buyers.
Sources
- Codebridge, "AI vendor evaluation checklist for accounting firm COOs". https://www.codebridge.tech/articles/ai-vendor-evaluation-checklist-for-accounting-firm-coos
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version 2024). https://arxiv.org/abs/2310.16787
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ), with generative AI questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Hugging Face Forums, "Is selling datasets way harder than building them? Or is it just me?". https://discuss.huggingface.co/t/is-selling-datasets-way-harder-than-building-them-or-is-it-just-me/177440
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Klaus Krippendorff, "Computing Krippendorff's Alpha-Reliability". https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- Gebru et al., "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
- NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.