Procurement, samples and ongoing supply
Is the Vendor's Sample Representative? Checking a Sample Against the Full Dataset
Quick answer
A vendor sample is representative only if you can show that its distributions match the full dataset on the dimensions your model cares about. Treat any sample as unrepresentative until proven otherwise. Ask how it was drawn, request a stratified draw across time, source, outcome and edge cases, and get population statistics computed over every record. Then compare the two with simple tests, and write a sample-to-delivery tolerance into acceptance criteria so the full delivery has to match what you evaluated.
By SourceX Editorial · Updated
Why vendor samples skew
Most vendor samples are not random draws, and the skew is usually structural rather than malicious. Data marketplace documentation, such as HERE's provider guide, describes letting providers limit trial access to a geographic area or a historical time window [1]. That is a reasonable way to protect the asset, but it means the sample tells you about one slice. A sample of 2024 support tickets from one product line says little about 2019 tickets, a different CRM, or a team that tagged resolutions differently.
Four selection patterns cover most of the problems buyers meet:
- Recency bias. The supplier exports the last 30 or 90 days because that data is easiest to query. Older records often have different schemas, sparser fields or retired category codes.
- Showcase bias. Someone picks "good" examples, such as tickets with clean resolutions, contracts with full metadata or recordings with no background noise. Your SFT or eval set then inherits a difficulty level the full corpus does not have.
- Convenience bias. The sample comes from the one system, region or team that was easiest to export. A Zendesk instance sampled while the rest of the history sits in Salesforce Service Cloud is a common case.
- Survivorship bias. Deleted, merged, escalated or failed records are missing because the export query filtered on status. For evaluation data this removes the hard cases you most need.
None of these show up from reading 50 rows. They show up when you compare the sample with statistics computed over the whole population.
Ask how the sample was drawn before you analyze it
The cheapest representativeness check is a written description of the sampling method, because a vague answer is itself a finding. Ask for the query or script that produced the sample, the date range and filters applied, the total population size the sample was drawn from, and whether any rows were removed by hand afterward. Datasheets for Datasets lists sampling questions among its composition prompts, including whether the dataset is a sample from a larger set and how representativeness was validated [4]. Data Cards push the same idea toward decisions that affect model performance, such as filtering and annotation choices [5].
If the answer is "a random 1%," ask for the seed and the random function. If it is "a curated selection," you already know the sample is not representative, and you can move straight to requesting a stratified redraw. For how to frame the initial request, see how to request a training data sample from a supplier; this page covers what to do once a sample is in hand.
Request a stratified sample and population statistics together
A stratified sample plus vendor-computed population statistics gives you two things to compare instead of one thing to trust. Stratify on the variables that drive model behavior for your use case, not on whatever is easy to group by. For SFT on support conversations, those are usually year, channel, product, resolution outcome, language and conversation length. For an eval set, add the edge cases you care about: escalations, refunds, policy exceptions, multi-turn threads and records with missing fields.
Ask for proportional allocation in most strata so frequencies can be compared directly, plus a deliberate oversample of rare strata so you can judge their quality. Label the oversampled strata in a column so they do not distort your comparison.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Request item | What to ask for | Why it matters |
|---|---|---|
| Sampling method | Query or script, seed, filters, date range, manual removals | Exposes curation and status filters |
| Strata | Year, source system, team or region, outcome label, length bucket | Covers the main axes of selection bias |
| Allocation | Proportional per stratum, plus flagged oversample of rare strata | Lets you compare rates and inspect tails |
| Population row counts | Total rows and rows per stratum over the full dataset | Baseline for category proportions |
| Field fill rates | Percent non-null per field, full dataset and sample | Detects "clean sample, sparse corpus" |
| Category counts | Top-N values for each categorical field, with an "other" bucket | Detects missing or retired codes |
| Date histogram | Record counts by month across the full range | Detects recency windows and gaps |
| Length distribution | Percentiles (p5, p50, p95, p99) of tokens, turns or duration | Detects showcase selection of short, tidy records |
| Duplicate rate | Exact and near-duplicate rates on a stated key | Detects inflated volume |
| Label distribution | Counts per label and annotator agreement where labels exist | Detects class balance that differs from delivery |
Population statistics are only useful if you can reproduce them, so ask for the SQL or notebook that generated them and the snapshot date. ISO/IEC 5259-2 defines data quality measures such as completeness and accuracy that you can reference when naming these metrics, which keeps vendor and buyer measuring the same thing [3].
Compare sample statistics to the full dataset
The comparison is a short set of distribution tests, one per field type, read against a tolerance you set in advance. Fix the thresholds before you see the sample so the results cannot be argued into a pass.
- Continuous fields (token counts, durations, amounts, timestamps): run a two-sample Kolmogorov-Smirnov test when you have record-level population data. The test asks whether two samples come from the same distribution without assuming which distribution that is. When you only have population percentiles, compare p5, p50, p95 and p99 directly and flag any gap larger than your tolerance.
- Categorical fields (product, channel, outcome, language): compare proportions per category. A chi-square goodness-of-fit test against the population proportions works when expected counts are not tiny. In practice, a table of absolute percentage-point differences is easier for a vendor to act on.
- Fill rates: compare non-null percentages field by field. A field that is 98% filled in the sample and 61% filled in the population is the clearest single sign of a curated draw.
- Time coverage: overlay the sample's monthly counts on the population date histogram. Gaps, a single-quarter spike or a missing year show up immediately.
- Text and media content: embed sample and population-subset records with the same encoder and compare cluster membership. A sample that sits in a few clusters the population barely uses is showcase data.
Two caveats matter. Large samples make tiny, harmless differences statistically significant, so read effect sizes such as the KS statistic or percentage-point gaps, not only p-values. Small samples hide real differences, so for rare strata judge quality by reading records rather than testing proportions. For overlap with data you already hold, use a cross-dataset overlap check alongside a coverage gap analysis against your deployment distribution.
Worked example: a support-ticket sample that passed reading but failed statistics
A sample can read well and still fail every distribution test. Consider a hypothetical vendor that offers five years of B2B support tickets and sends 2,000 records for evaluation.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | Sample | Population statistics | Tolerance | Result |
|---|---|---|---|---|
| Records from latest 12 months | 71% | 24% | ±5 pp | Fail |
resolution_code fill rate | 97% | 64% | ±5 pp | Fail |
| Escalated tickets | 3% | 11% | ±3 pp | Fail |
| Median turns per ticket | 6 | 9 | ±1 | Fail |
product_line categories present | 4 of 7 | 7 of 7 | All present | Fail |
Exact duplicate rate on ticket_id | 0% | 0.4% | ≤1% | Pass |
The reading review saw clean, well-tagged, recent tickets. The statistics show a recent, single-system export filtered to resolved tickets. The right response is not to reject the vendor but to ask for a stratified redraw across all five years, all seven product lines and every resolution status, with the escalated stratum oversampled and flagged.
Write sample-to-delivery tolerance into acceptance criteria
A representative sample protects you only if the delivered dataset must match it. Convert the comparison table into acceptance criteria: for each metric, state the population value the vendor reported, the tolerance and what happens if delivery falls outside it, such as redelivery or exclusion of the failing stratum. Problems caught at sample stage are much cheaper to fix than problems found after payment, which is why acceptance tests should be defined before purchase [2].
Keep the criteria measurable. "Representative of the full dataset" is not testable. "Monthly record counts within ±10% of the reported histogram, resolution_code fill rate at least 60%, all seven product lines present" is. If samples are released under a separate evaluation agreement, align the tolerance language there too; see evaluation licenses and NDAs for dataset samples. When the supplier cannot release raw records at all, the same statistics can often be computed inside a secure viewing environment, covered in sample access options for sensitive data.
How this fits vendor evaluation
Representativeness is one scored line in a broader vendor review, not the whole decision. Fold the results into a data vendor evaluation scorecard and the evidence to request from data vendors before you sign. The wider sequence from requirements to renewal is in the AI training data procurement hub. For SourceX's general guidance on sample selection, see selecting a representative sample and evaluating data supplier quality.
When SourceX sources a dataset on request, diligence materials covering source, rights, preparation and allowed use are prepared for each dataset, and every release is approved by the supplying company. Buyers can describe the data they need on the buyers page, and the sampling questions above make a useful part of that description.
Testing a sample before you license a dataset
SourceX sources operational datasets from US companies on request and manages the commercial process from Find and Assess through Agree, Transact and Manage. Nothing is contracted until a supplier agrees, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe the dataset you need to test on the SourceX buyers page.
Frequently asked questions
How large does a sample need to be to test representativeness?
Size depends on the rarest stratum you need to judge, not on the total. A few thousand records can test common category proportions, but a stratum that is 1% of the population needs a deliberate oversample before you can say anything about it.
Can I test representativeness without record-level access to the full dataset?
Yes, with aggregate statistics. Population percentiles, fill rates, category counts and date histograms let you compare the sample without seeing other records, provided the vendor shares the query that produced them and the snapshot date.
What if the vendor refuses to share population statistics?
Treat that as a material gap. Without a population baseline you cannot tell a stratified sample from a curated one, so record the refusal in your scorecard and limit the deal's acceptance criteria to what you can verify on delivery.
Sources
- HERE Technologies, "Showcase your data (Marketplace provider user guide)". https://developers.here.com/documentation/marketplace-provider/user_guide/topics/showcase-data.html
- Contra, "Robot dataset acceptance audit: a verdict before you pay". https://contra.com/s/ee6vvEQU-robot-dataset-acceptance-audit-a-verdict-before-you-pay
- International Organization for Standardization (ISO), "ISO/IEC 5259-2:2024 Artificial intelligence: Data quality for analytics and machine learning (ML), Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Gebru et al. (arXiv:1803.09010), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
- Pushkarna, Zaldivar, Kjartansson, Google Research (arXiv:2204.01075), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.