Skip to content

Procurement, samples and ongoing supply

Data vendor evaluation scorecard: knockout gates, weighted criteria and anchored scores

Quick answer

A data vendor evaluation scorecard turns competing supplier bids into a defensible award in three passes. Knockout gates first remove bids without a documented rights basis, a lawful source, the required de-identification standard, minimum security controls or a random sample. Reviewers then score six criteria families on anchored 0-4 evidence scales. Finally, weights fixed before proposals open, chosen for your training stage, produce a total that you test for sensitivity and record with its evidence.

By SourceX Editorial · Updated

The scorecard chooses among bidders answering the same AI training data RFP. Use questions to ask a training data vendor to shortlist, a supplier performance scorecard after award, and a model bake-off to choose a model provider.

Knockout gates: failures no score can offset

Apply pass/fail gates before anyone scores, because a weighted sum lets low price or large volume offset data you cannot lawfully train on. The cost can reach the model: the FTC's 2021 order against Everalbum required deletion of the models and algorithms built from users' photos and videos [2].

GatePasses whenEvidence to seeOwner
G1 Rights basisBidder owns the records or holds upstream rights for your usesChain-of-title memo; resellers' upstream resale termsCounsel
G2 Lawful sourceNo pirated copies; personal data collected under notices covering disclosure for AI trainingSource list; notices in force at collectionCounsel, privacy
G3 De-identificationRegulated data meets your standardMethod record, such as HIPAA Safe Harbor or Expert DeterminationPrivacy
G4 Permitted usesLicense names your stage: training, evaluation, retrieval, derived modelsLicense markupCounsel
G5 Security minimumYour storage and transfer controls are metQuestionnaire, report scopeSecurity
G6 Random sampleRandom draw from the offered segment, under evaluation termsSampling method, manifestML or data lead
G7 No conflicting grantNo exclusive license blocking your field of useWritten representationCounsel

The Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [3], so a dataset card's license field cannot pass G1. For G2, the class settlement in Bartz v. Anthropic received final approval in July 2026, according to Authors Alliance [4]; it is a settlement, not a merits ruling. FTC staff warned in February 2024 that adopting AI training through a surreptitious, retroactive change to terms or a privacy policy may be unfair or deceptive [5], hence the notice in force at collection. For G3, HIPAA Safe Harbor removes 18 listed identifiers; Expert Determination needs a qualified expert to find identification risk very small [6].

The FISD Alternative Data Council's industry questionnaire shows the evidence format: a data dictionary, a small sample, consent terms agreed with individuals and, for data bought from others, the terms that allow resale [7]. Collect them through a data provider due diligence questionnaire and a training data sample request.

Six criteria families and the evidence that earns points

Score what the bidder can show for this dataset, not what it says about itself; weighted criteria reward depth above the gates.

FamilyWhat you scoreEvidence that earns pointsEarns nothing
Data fitFields against your data dictionary, population, time window, usable volume, languages, label taxonomyFill rates and distributions measured on the random drawCategory names, record counts
Rights and provenanceTitle depth, upstream licenses, notice coverage, permitted uses, disclosure supportDataset-specific documents, then your trace of sample records"Rights-cleared"
Quality evidenceSchema conformance, duplicates, label accuracy, annotator agreement, contaminationMeasures with named methods, reproduced in your pilotAccuracy without a method
Privacy and securityDe-identification method, residual risk, access control, subprocessors, deletionMethod record, re-identification tests, security report scopeOut-of-scope certificates
Delivery and continuityFormats (Parquet, JSON Lines), schema versioning, transfer, refresh cadence, capacityMilestone plan; past schema change notices"Flexible delivery"
Commercial termsCost per usable record, license structure, warranties, exitPrice form converted with pilot pass ratesQuoted totals

Look beyond price per label to data lineage, security certifications, workforce specialization and operational maturity.

Quality claims score only when measurable. ISO/IEC 5259-2 sets out data quality measures [8]; require each claim to name the property and the method used to quantify it. A Hugging Face forum discussion draws the same line: "the exact-duplicate rate is below 1%" can be checked, "good for fine-tuning" cannot [9]. Treat nothing a vendor says as evidence until you test it on your own data [1], with paid pilot terms where a free sample is too small.

Score whether a supplier can feed your training-data disclosures

A supplier that cannot document sources, collection periods and processing steps leaves gaps in disclosures developers now publish, so score disclosure support within rights and provenance. As of October 2026:

  • EU general-purpose AI models. AI Act Article 53(1)(d) requires providers to publish a sufficiently detailed training-content summary on the AI Office template [13].
  • California. AB 2013 requires developers of generative AI systems offered to Californians to post a high-level dataset summary covering items such as sources or owners, copyrighted, licensed or personal information, and collection periods; first postings were due 1 January 2026 [14].
  • EU high-risk systems. Article 10(2) requires governance of collection processes and data origin, preparation such as annotation, labelling and cleaning, and examination of possible biases [15]. Regulation (EU) 2026/1744, published 24 July 2026, reportedly moves Annex III high-risk obligations to 2 December 2027 and also amends Article 10 [16].

A datasheet covering motivation, composition, collection process and recommended uses [17] earns a 2. Machine-readable metadata checked against the sample earns a 3; NeurIPS 2026 requires Croissant-based responsible-AI fields for its Evaluations and Datasets Track [18]. A warranty on the facts you will publish earns a 4.

Anchored 0-4 scales that make reviewers agree

Anchor every score to a level of evidence so that reviewers reading the same file land within a point of each other. One ladder serves all six families:

ScoreEvidence levelWhat it looks like
0Missing or contradictedNo answer, or the sample contradicts it
1Asserted"Verified" or "high quality", nothing dataset-specific
2DocumentedData dictionary, datasheet, notice text, upstream license list or method record for this dataset
3Tested by youConfirmed on a random draw, within your threshold
4Tested and committedAs 3, and accepted as a warranty, acceptance criterion or remedy

Quality and rights evidence get their own anchors:

ScoreQuality evidenceRights and provenance
1"98% accurate", no methodLicense field on a dataset card
2QA report naming measures: gold-set label accuracy, Krippendorff's alpha, near-duplicate rateChain-of-title memo, upstream agreements, dated notice text
3Your pilot reproduces the measures within toleranceYour provenance test on sample records traces them to source systems
4Measures become acceptance criteria with remediesTitle and collection warranties, full disclosure schedule

Krippendorff's alpha measures agreement among annotators labeling the same items: 1 is perfect reliability, 0 is chance [11]. A hand-picked sample caps quality at 2; check whether the sample is representative.

Weight profiles for pre-training, fine-tuning, evaluation and RAG

Set weights by what the purchase must do and freeze them before proposals open; these starting profiles sum to 100.

Criterion familyPre-trainingFine-tuning (SFT)EvaluationRAG
Data fit25202520
Rights and provenance20151520
Quality evidence15303015
Privacy and security10101015
Delivery and continuity1010520
Commercial terms20151510
  • Pre-training buys volume, so duplication and usable-record price matter. Lee et al. found one sentence repeated over 60,000 times in C4, and deduplicated training cut memorized output about tenfold [12].
  • Fine-tuning is quality-led: LIMA fine-tuned a 65B-parameter model on 1,000 curated prompt-response pairs [19], so example quality counts for more than volume.
  • Evaluation data must be correct and unseen. Northcutt et al. estimated average label error of at least 3.3% across test sets of 10 widely used datasets and showed that such errors can change model rankings [10]. Score contamination checks under quality.
  • RAG content changes, so refresh cadence, deletion that reaches your index and display rights carry weight; see RAG content licensing.

Worked example: two bids that survive the gates

Score price per usable record, and test any close total for sensitivity.

Illustrative example: invented to show structure; it does not describe an available dataset.

A buyer scores bids for resolved support-ticket threads for supervised fine-tuning, using the SFT profile. Bidder C fails G2 (its transcripts' collection notice excluded third-party disclosure) and is not scored. A and B each offer 100,000 threads at price indexes of 100 and 75. In the pilot, 92% of A's threads and 64% of B's pass acceptance checks, so A costs 1.087 per 1,000 usable threads and B 1.172.

Price score = 4 x (lowest cost per usable record / bidder's cost per usable record).

Family (weight)AB
Data fit (20)34
Rights and provenance (15)43
Quality evidence (30)32
Privacy and security (10)33
Delivery and continuity (10)24
Commercial, on quoted price (15)3.04.0
Commercial, per usable record (15)4.03.71
Total on quoted price (of 100)76.378.8
Total per usable record (of 100)80.077.7

Total = sum of weight x score, divided by 4. On quoted prices B wins; per usable record A wins by 2.3 points. Sensitivity check: moving 10 points from quality evidence to data fit gives A 80.0 and B 82.7, so the award rests on the SFT profile's quality emphasis, which the panel must record. Comparing vendor quotes per usable record covers price normalization.

Running the panel and breaking ties

Reviewers judge only the gates and families they are qualified for and score independently, citing a reason and evidence document ID per score, before the panel calibrates.

ReviewerGatesFamilies scored
ML or data leadG6Data fit, quality evidence
CounselG1, G2, G4, G7Rights and provenance, license terms within commercial
PrivacyG3Privacy half of privacy and security
SecurityG5Security half; see the security review of a data supplier
ProcurementNoneCost per usable record, delivery and continuity
  1. Freeze specification, gates, weights, anchors and tie-break rule; publish gates and weights in the RFP.
  2. Record gate failures with evidence; failed bids are not scored.
  3. Calibrate spreads of 2 or more points; changes need a written reason.
  4. Open prices after technical scores lock, where policy allows; convert with pilot pass rates.
  5. Run the sensitivity check, then the tie-break.

Set the tie margin in advance, for example 3 points of 100. Within it, prefer the higher rights and provenance score, then the higher quality evidence score; if still level, run a second pilot or split the award under a multi-supplier strategy. A's 2.3-point lead in the worked example falls inside it, and A also wins the first tie-break.

Score datasets sourced through SourceX against the same gates. SourceX looks for US businesses holding the data you describe; rights review checks that each business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. Score them at 2 until your sample tests earn a 3.

Personal details are removed or replaced before delivery, but no de-identification method is perfect, so privacy still scores residual risk. You can describe the records you are scoring bids for; a request does not guarantee a matching dataset.

The selection record to keep

Keep one record per award that an auditor or successor can re-run without the panel. NIST AI 600-1 lists data privacy, intellectual property, and value chain and component integration among the 12 risks it identifies for generative AI [20]; a kept scorecard shows how third-party data was assessed against them.

Illustrative example: invented to show structure; it does not describe an available dataset.

selection_id: DS-2026-031
specification: "support-thread SFT spec v1.4"
weights_frozen: 2026-08-03          # before proposals opened on 2026-08-17
weight_profile: sft                 # fit 20, rights 15, quality 30, privacy/security 10, delivery 10, commercial 15
gates:
  bidder_A: pass
  bidder_B: pass
  bidder_C: {result: fail, gate: G2, evidence: "EV-C-07 privacy notice excludes third-party disclosure"}
scores:                             # consensus after calibration; reviewer sheets attached
  bidder_A: {total: 80.0, price_basis: cost_per_usable_record}
  bidder_B: {total: 77.7, price_basis: cost_per_usable_record}
sensitivity: "B leads if data-fit weight is 30 and quality weight is 20"
decision: "Award A: pilot label accuracy and rights traced on sample outweigh B's coverage"
carried_into_contract:
  - "pilot thresholds become acceptance criteria"
  - "title and collection warranty with disclosure schedule"
open_risks:
  - "A delivery capacity scored 2: monthly milestones with remedies"

Scoring errors that pick the wrong supplier

These errors let claims or quoted prices outweigh tested evidence:

  • Pricing quoted totals. The worked example flips on this alone.
  • Moving weights after proposals open. The scorecard becomes a justification.
  • Averaging away a 0. A 0 from privacy or counsel triggers review, not a mean.
  • Double counting. One datasheet scored under rights and quality inflates documentation-heavy bids.

SourceX's guides to evaluating data supplier quality and running a data pilot cover quality signals and pilots; its data quality scorecard shows how data owners rate their own records' readiness. The procurement hub places selection among the other purchase stages.

Evaluating data suppliers for a purchase?

Describe the records, fields, volume, time window and uses your scorecard is built around. SourceX looks for US companies that hold that data, checks the data and each supplier's licensing permissions, and manages the license in which pricing and allowed uses are agreed; nothing is contracted until a supplier agrees. See how SourceX works with data buyers.

Sources

  1. Codebridge, "AI vendor evaluation checklist for accounting firm COOs". https://www.codebridge.tech/articles/ai-vendor-evaluation-checklist-for-accounting-firm-coos
  2. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  3. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version 2024). https://arxiv.org/abs/2310.16787
  4. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  5. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  6. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  7. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ), with generative AI questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  8. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  9. Hugging Face Forums, "Is selling datasets way harder than building them? Or is it just me?". https://discuss.huggingface.co/t/is-selling-datasets-way-harder-than-building-them-or-is-it-just-me/177440
  10. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  11. Klaus Krippendorff, "Computing Krippendorff's Alpha-Reliability". https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  12. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  13. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  14. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  15. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  16. European Parliament and Council of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  17. Gebru et al., "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
  18. NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  19. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  20. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data