Data quality, coverage and contamination
Training Data Quality Assessment: How AI Teams Judge Dataset Quality, Coverage and Contamination
Quick answer
Training data quality assessment is the set of measurements a buyer runs before a licensed dataset enters training or evaluation. It covers five families of checks: intrinsic record quality, label quality, coverage of the deployment distribution, duplication and benchmark contamination, and residual personal data or secrets. ISO/IEC 5259 supplies the shared terminology and measures [1] [2]. In practice, each family is measured on a random sample of the actual delivery, against thresholds written down before the sample arrives.
By SourceX Editorial · Updated
The five check families and what each one catches
A licensed dataset can fail in five independent ways, and passing one family says nothing about the others. Clean schemas can carry wrong labels, and accurate labels can sit in records that duplicate a public benchmark.
| Check family | Question it answers | What to measure | Failure it prevents | Deep dive |
|---|---|---|---|---|
| Intrinsic record quality | Are records correct, complete and internally consistent? | Schema conformance, missingness by field, out-of-range values, referential integrity, timestamp order | Training on truncated threads, null outcomes or mixed units | Accuracy, completeness and timeliness metrics |
| Label quality | Do labels and outcome fields mean what the supplier says? | Error rate on an adjudicated audit sample, inter-annotator agreement, per-class confusion | Learning annotator noise or auto-set status fields | Annotation quality audit |
| Coverage and representativeness | Does the data span the inputs the model will see? | Records per deployment slice, long-tail frequency, date gaps, subgroup counts | Good aggregate scores that hide failing slices | Coverage gap analysis |
| Duplication and contamination | Is the data new, and is it free of your test sets? | Exact and near-duplicate rates, overlap with owned data, n-gram and semantic benchmark overlap | Memorization, paying twice, inflated eval scores | Near-duplicate detection with MinHash and LSH |
| Safety and privacy residue | What personal data, secrets or toxic text survived preparation? | Personally identifiable information (PII) hits by type, credential and API-key matches, toxicity scores, language-ID mismatches | Memorized personal data and leaked credentials | PII scanning before fine-tuning |
The grouping is a working structure for buyers, not a standard; ISO/IEC 5259 frames the same ground as quality characteristics.
ISO/IEC 5259, Article 10 and other dataset quality standards
ISO/IEC 5259 is the reference standard series for data quality in analytics and machine learning, and its terms are the safest basis for quality thresholds in a data request or license annex. Part 1 (2024) gives the overview and terminology, Parts 2 to 4 (2024) cover measures, management requirements and a process framework [1], Part 5 (2025) adds a governance framework [22], and Technical Report 5259-6 on visualization followed in 2026 [23]. Part 2 defines a data quality model, measurable characteristics and guidance on reporting, and builds on ISO/IEC 25012 and ISO 8000 [2].
A published summary of the series lists characteristics such as accuracy, completeness, consistency, timeliness and representativeness [3]. The standards are paid, so quote the clause your team adopts, not a summary.
For high-risk AI systems in the EU, Article 10(3) of the AI Act requires training, validation and testing data sets to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose; Article 10(2) adds examination of possible biases and identification of data gaps [4]. As of October 2026, the AI Act has been amended by Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026. Secondary sources report that the amendment moved high-risk application dates to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [5]. The amending regulation also amends Article 10 [5]. The AI training data compliance hub covers who these obligations reach.
Documentation expectations are rising. Datasheets for Datasets proposed that every dataset document its motivation, composition, collection process and recommended uses [6]. NeurIPS 2026 requires Croissant-RAI-based responsible-AI metadata, including limitations, potential biases and intended use, for its Evaluations and Datasets Track [7]. Ask suppliers for measured values behind a datasheet for datasets, not prose; what a dataset quality report should contain lists the fields.
Training data quality vs quantity: what the evidence supports
Published results show that a smaller, cleaner dataset often matches or beats a larger noisy one, though the effect depends on the training stage and how quality was defined. For buyers, record counts are a weak proxy for value; the useful number is how many records survive your own filters.
- Instruction tuning. LIMA fine-tuned a 65B-parameter LLaMA model on 1,000 curated prompt-response pairs, and its authors concluded that most knowledge comes from pretraining while a small high-quality set teaches output format [8]. AlpaGasus had ChatGPT grade Alpaca's 52,000 examples, trained on about 9,000 high scorers, and outperformed the original Alpaca in GPT-4 evaluations [9].
- Tool use. Models trained on smaller sets of validated synthetic tool-use data outperformed models trained on larger unvalidated sets [10].
- Deduplication. Lee et al. found a single 61-word sentence repeated over 60,000 times in C4. Models trained on deduplicated data emitted memorized text about ten times less often and reached the same or better accuracy in fewer training steps [11].
- Filtering side effects. Across 28 pretrained 1.5B-parameter models, quality filtering raised downstream performance and toxic generation even though it removed more than 10% of the data, while toxicity filtering reduced toxic generation at some cost to generalization. The effects were not predictable from domain characteristics [12].
LIMA's authors note that curation is labor-intensive and that LIMA is less robust than product-grade systems [8]. "Less is more" is not a reason to buy less. It is a reason to ask for the expected yield after deduplication, filtering and label audit, and to run an ablation first; see estimating what a dataset adds before you buy it.
Data quality for LLM training by stage: which checks come first
The same dataset needs different checks depending on whether it feeds pre-training, supervised fine-tuning, preference tuning, evaluation or retrieval. How procurement differs by training stage covers the commercial side.
| Use | Checks to run first | Threshold to state in the request |
|---|---|---|
| Pre-training or continued pre-training | Near-duplicate rate, overlap with public web crawls, language ID, toxicity, PII residue | Maximum near-duplicate share at a named similarity method and cutoff |
| Supervised fine-tuning (SFT) | Response correctness, format conformance, instruction diversity, share of templated replies | Maximum error rate on an expert-audited sample |
| Preference tuning (RLHF or direct preference optimization) | Rater agreement on overlapping pairs, tie handling, rater calibration, position bias | Minimum agreement statistic and the share of pairs double-rated |
| Evaluation sets | Gold-label accuracy, contamination against training corpora, leak-free splits | No verbatim benchmark overlap; adjudicated gold labels |
| Retrieval-augmented generation (RAG) corpora | Duplicate and superseded document versions, conflicting documents, freshness | Version and effective-date fields on every document |
Deep dives: preference data noise and agreement, instruction-tuning data filtering, pretraining-scale text filtering and knowledge-corpus quality for RAG. Eval set design belongs to the LLM evaluation datasets hub. If a model scores your records, validate it first with the LLM-judge guide.
How to evaluate a dataset before buying it: a nine-step checklist
Evaluate a dataset before buying it by testing a random sample drawn from the exact segment you will license, with the same pipeline and thresholds you will apply at delivery. One 2026 data-centric paper describes an assessment stack that combines metric checks (format, duplication, diversity, fluency, factual accuracy), model-based scoring and human review of samples [13].
- Fix thresholds first. Write a pass line for each check family before the sample arrives, so the sample cannot set its own bar.
- Get a random sample, not a showcase. Ask how it was drawn, from which source systems and over which date range; checking a sample against the full dataset explains the tests.
- Validate structure. Check schema, types, missingness by field, value ranges, foreign keys and timestamp order.
- Deduplicate inward and outward. Measure duplicates within the sample, against data you already own and against public pretraining corpora.
- Scan for contamination against every public benchmark and private eval set you report on.
- Audit labels on an adjudicated subsample sized for the error rate you need to detect; see sample sizes for estimating error rates.
- Scan for residue: PII by type, credentials and secrets in tickets and logs, and model-generated text in purchased human data.
- Map coverage against your deployment taxonomy and record empty or thin slices.
- Write a quality report the supplier can see, so any dispute is about numbers.
Illustrative example: invented to show structure; it does not describe an available dataset.
quality_report:
dataset: support_ticket_threads_sample # supplier sample, not the full delivery
sample: {method: stratified_random, strata: [product_line, quarter], records: 2000}
intrinsic:
schema_conformance: 0.997
missing_resolution_code: 0.043 # threshold <= 0.05
truncated_threads: 0.012 # threshold <= 0.02
labels:
audited_records: 400
adjudicated_error_rate: 0.031 # threshold <= 0.05
krippendorff_alpha_resolution_code: 0.78
coverage:
taxonomy_slices: 42
slices_under_50_records: 6
duplication_contamination:
near_duplicate_rate_minhash: 0.064 # method and cutoff named in the request
overlap_with_owned_data: 0.009
benchmark_13gram_hits: 0
residue:
pii_hits: {email: 3, phone: 1, account_number: 0}
credential_pattern_hits: 2 # threshold = 0
verdict: fail
open_issues:
- "Auto-closed tickets carry resolution_code=resolved with no agent reply"
- "Credential hits must be scrubbed and the sample re-sent"
Measurement is separate from contract acceptance. Thresholds become enforceable only when they appear in acceptance criteria for licensed training data with an inspection plan. ANSI/ASQ Z1.4 attribute sampling plans, with switching between normal, tightened and reduced inspection [14], were built for manufactured lots; acceptance sampling for dataset deliveries adapts them to records and labels. When SourceX sources a dataset, it checks the data and the supplier's licensing permissions and prepares diligence materials on source, rights, preparation and allowed use for the buyer's review; your quality report sits beside them, and you can state the thresholds a dataset must meet in your data request.
Contamination and overlap: where standard checks fall short
N-gram overlap is the predominant contamination check and catches verbatim leakage cheaply, but it misses paraphrased, translated and synthetic copies of test items. A survey of benchmark contamination reports thresholds that differ across labs, from 13-gram matches for GPT-3 to other approaches in later work [15]. LMSYS researchers showed the gap: a 13B model trained on rephrased MMLU test items scored 85.9 on MMLU while n-gram overlap failed to flag it [16].
Contamination can also retire a benchmark. In a February 2026 post, OpenAI said SWE-bench Verified had become increasingly contaminated, stopped reporting it and recommended SWE-bench Pro in the interim [17].
Lee et al. also found train-test overlap affecting more than 4% of validation examples in standard benchmarks [11]. For licensed data, test overlap with data you already own, which means paying twice, and with public web crawls, which erodes the novelty you are paying for. Decontaminating a training set against benchmarks covers thresholds and reports; the guide to contamination checks for licensed evaluation data covers eval-set hygiene, and benchmark contamination has the definition.
Label and outcome-field quality in operational records
Labels in licensed data are rarely as clean as their documentation implies, so audit them on a sample instead of accepting a supplier-reported accuracy figure. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets, including at least 6% of the ImageNet validation set, and showed that such errors can reorder model rankings [18].
Agreement statistics measure consistency, not correctness. In Krippendorff's alpha, 1 indicates perfect reliability and 0 indicates agreement no better than chance [19], but high agreement can still mean annotators share a systematic error. That is why audits use golden datasets and adjudication alongside inter-annotator agreement metrics.
Operational business records add failure modes that research datasets rarely show. Support tickets carry macros and canned replies that inflate apparent diversity. A "resolved" status may be set by an auto-close rule rather than by an outcome. Email threads arrive truncated at export limits, and CRM migrations change what a field means partway through the history.
Each needs its own check: templates and boilerplate in business records, completeness of case and ticket records and outcome fields as ground truth.
Mistakes that let a weak dataset through
Weak datasets get through when checks run on the wrong sample, at the wrong level of aggregation, or only once.
- Trusting the supplier's accuracy number. Re-measure on your own adjudicated sample.
- Treating "representative" as one property. A dataset can mirror a population or cover the input space. Coverage is more robust to distribution shift and narrows accuracy gaps between subgroups, while population mirroring serves population-level inference [20]. Say which you need; assessing representativeness shows how to test each.
- Reporting only aggregates. A strong overall error rate can hide one product line or customer segment where most labels are wrong. Report per slice, and check long-tail and edge-case coverage directly.
- Deduplicating only exact matches. Templated records differing only by a ticket number or name pass exact hashing.
- Assuming de-identified means clean. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including contact details [21]. SourceX removes or replaces names, emails, phone numbers and account numbers before delivery and records the method used, but no de-identification method is perfect, so run your own scan. The de-identified data hub covers methods.
- Measuring once. Re-run the checks on every delivery and refresh; supplier systems and export jobs change between them.
To compare suppliers rather than measure a delivery, use the guide to evaluating data supplier quality. The AI data hub maps the other buyer topics, including data provenance.
Need licensed data that has to pass these checks?
Describe the data you need and the quality thresholds it must meet. SourceX looks for US businesses that hold it, checks the data and the supplier's licensing permissions, and manages a license that defines which records are included and what they can be used for; a request does not guarantee a matching dataset. Describe your dataset and quality requirements.
Guides in this section
- Acceptance Sampling for Dataset Deliveries: AQL PlansAdapt ANSI/ASQ Z1.4 acceptance sampling to dataset deliveries: define record and label defects, pick n and c, read the OC curve and apply switching rules.
- Annotation Quality Audit: Checking Labels Before AcceptanceHow buyers audit annotation quality in a delivered labeled dataset: frozen snapshot, stratified sample, blind relabel, error taxonomy, intervals.
- Dataset Bias Audit: Representation, Labels, IntersectionsHow to audit a training dataset for bias: representation tables against a deployment population, intersectional coverage, label bias and curation logs.
- Dataset Overlap Check: Measure Duplication Before You BuyHow to measure overlap between a candidate dataset and data you already own: exact hashes, MinHash, embeddings, private set intersection and thresholds.
- Decontaminating Training Data Against Public BenchmarksHow to decontaminate a licensed training set against public benchmarks: n-gram thresholds, normalization, removal rules and a decontamination report.
- Edge Case Training Data: Measuring and Sourcing the TailHow to measure long-tail coverage with worst-slice metrics, surface rare cases in your own data, and specify edge-case training data from real operations.
- Instruction-Tuning Data Quality: LIMA and AlpaGasus LessonsWhat LIMA and AlpaGasus show about SFT data quality vs quantity, how to score instruction-response pairs, and how to filter and buy instruction data.
- Inter-Annotator Agreement: Kappa vs Krippendorff's AlphaChoose between Cohen's kappa, Fleiss' kappa and Krippendorff's alpha, read supplier agreement numbers correctly, and set IAA thresholds for training data.
- Is Licensed Data Already in Common Crawl? Novelty TestsHow to test whether a dataset sold as proprietary is already in Common Crawl or open pretraining corpora: URL lookups, n-gram search and membership tests.
- Near-Duplicate Detection with MinHash and LSH for Text DataHow to find exact and near-duplicates in a licensed text dataset with hashing, MinHash and LSH, tune the Jaccard threshold and report a duplicate rate.
- Preference Data Quality: Noise, Agreement and AmbiguityHow to measure noise, rater agreement and ambiguity in pairwise preference data, and when to clean, reweight or reject pairs before reward modeling or DPO.
- Synthetic Data Quality Evaluation for AI Training SetsHow to evaluate synthetic training data: fidelity, diversity, utility and privacy checks for LLM-generated text and tables, with an acceptance scorecard.
- Training Data Coverage Analysis: Slice Matrix and GapsMap a candidate dataset against production traffic: choose coverage axes, build a slice matrix, flag thin cells and turn gaps into a supplier request.
- Training Data Quality Metrics: Formulas and ReportingComputable formulas for training data quality: completeness, validity, consistency, duplicates, label accuracy and timeliness, with confidence intervals.
- Verifying Outcome Fields as Ground-Truth LabelsHow to test whether resolution codes, claim decisions and approvals in business records are reliable labels: failure modes, leakage checks and re-review.
- Annotation Adjudication: Majority Vote vs Soft LabelsHow to resolve annotator disagreement: majority vote, expert adjudication, Dawid-Skene aggregation or soft labels, and what to require in a label delivery.
- Boilerplate Removal in Ticket, Email and Chat Training DataDetect signatures, disclaimers, quoted history and macro replies in business records, then strip, mask or label them before SFT, agent or RAG training.
- Case and Ticket Record Completeness Checks for AI DataHow to find incomplete records in ticket and case training data: define a complete case, measure truncation and missing outcomes, set thresholds.
- Dataset Diversity Metrics: Lexical, Semantic and TaskHow to measure dataset diversity: distinct-n, compression ratio, Self-BLEU, embedding dispersion, Vendi Score and task coverage, plus a buyer scorecard.
- Dataset Quality Report: Required Metrics and MethodsWhat a dataset quality report should contain: scope, stratum counts, field-level completeness, duplicate rate, label audits, PII residue and known defects.
- Dataset Representativeness: Population vs Coverage TestsHow to define and test dataset representativeness: population fidelity vs input-space coverage, ISO/IEC 5259 measures, and a statement template.
- Detecting Model-Generated Content in Human Training DataHow to check whether purchased human-written responses, preference labels or annotations were partly produced by LLMs, using process evidence and sampling.
- Drift Between Licensed Training Data and Production TrafficMeasure drift between licensed historical data and production inputs: PSI, KS, embedding and classifier tests, shift types, and what to do with the result.
- Expert Annotator Qualification: How Buyers Verify ItHow AI data buyers verify expert annotators: credential evidence, ground-truth qualification tests, per-labeler gold accuracy over time, pseudonymous IDs.
- Factual Accuracy Checks for Licensed Knowledge TextHow to estimate the share of wrong or outdated facts in licensed KB articles, documentation and Q&A data before you use it for SFT, RAG or evaluation.
- Find Label Errors in a Dataset with Confident LearningHow to rank likely mislabels in a licensed dataset with confident learning and cleanlab, using out-of-sample probabilities, review queues and remediation.
- Gold Questions and Honeypots in Annotation QAHow qualification tests, honeypot gold questions and consensus labeling work, what thresholds look like, and what to ask a labeling supplier about each.
- Historical Decision Bias in Lending, Hiring and Claims DataHow to audit past underwriting, hiring and claims decisions for label bias before training: conditional rate tests, selective labels and relabeling.
- How Much Label Noise Is Acceptable? Tolerances by UseSet label-error tolerances by use: looser for pretraining, tighter for SFT, tightest for reward and eval data, with systematic errors capped separately.
- Language Identification QA for Multilingual DatasetsHow to verify language labels and language mix in multilingual training data: LID tools, per-language thresholds, failure modes and native-reviewer audits.
- LLM Judges for Training Data Scoring: Validation and BiasHow to use an LLM judge to score or filter training examples: rubric design, thresholds, bias checks and validation against a human-labeled sample.
- PII Scan of a Training Corpus Before Fine-TuningHow to add a full-corpus PII gate before fine-tuning: scan inputs and targets, benchmark detectors on your own data, and decide on each finding.
- Pretraining Text Quality Filtering: Rules and ClassifiersHow heuristic rules, perplexity filters and quality classifiers like QuRating reshape pretraining text, and how to validate them on your tasks.
- RAG Corpus Quality: Duplicates, Stale Versions, ConflictsHow to audit and clean a RAG knowledge base: detect duplicate chunks, retire superseded versions, surface conflicting documents and measure the effect.
- Rater Calibration for Rubric and Human-Feedback DataCheck that rubric graders and QA reviewers were calibrated and stayed consistent before you use their scores as labels, rewards or eval ground truth.
- Real-Data Holdouts for Validating Synthetic Training DataHow to size, license and protect a real-data holdout that tests whether models trained on synthetic data actually work on real operational records.
- Sample Size to Estimate a Dataset's Error RateHow many records to audit: binomial confidence intervals, a sample-size table, the rule of three for zero errors, and stratified audits for AI datasets.
- Secrets in Training Data: Scanning Tickets, Chats and LogsHow to find API keys, passwords and tokens in support tickets, chat transcripts and system logs, and replace them with typed placeholders before training.
- Semantic Deduplication with Embeddings: Beyond MinHashHow embedding-based semantic deduplication catches paraphrased duplicates MinHash misses, how to calibrate cosine thresholds, and how to control its cost.
- Structured Dataset Validation Checks: Schema to Foreign KeysAn arrival validation suite for CSV and Parquet deliveries: manifest reconciliation, schema, nulls, ranges, code lists, key uniqueness and foreign keys.
- Synthetic Data Contamination: Finding Benchmark LeakageHow teacher-model outputs reproduce benchmark items as paraphrases, why n-gram checks miss them, and how to decontaminate synthetic SFT data.
- Temporal Coverage in Training Data: Gaps and SeasonalityCheck a multi-year dataset's real date range: monthly record histograms, missing months, migration seams, timestamp semantics and seasonal cycles.
- Toxicity Filtering Training Data: Trade-offs to MeasureHow toxicity filtering of licensed text trades safety against generalization and toxicity detection, and why buyers should ask for scores, not deletions.
- Training Data Freshness: Measuring Staleness and RecencyHow to measure training data staleness with record-age and lag metrics, why data age hurts newer test sets, and how to set recency rules for licensed data.
Sources
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-1:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 1: Overview, terminology, and examples" (2024). https://www.iso.org/standard/81088.html
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Management Solutions, "ISO/IEC 5259 Artificial intelligence: Data quality for analytics and machine learning (ML)". https://www.managementsolutions.com/en/node/4456
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union (Official Journal, via EUR-Lex), "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- Gebru et al. (arXiv; Communications of the ACM), "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
- NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- Zhou et al. (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Chen et al. (arXiv:2307.08701), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
- arXiv:2409.16341 (EMNLP 2024), "Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs" (2024). https://arxiv.org/abs/2409.16341v1
- Lee et al. (arXiv:2107.06499; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Longpre et al. (arXiv:2305.13169; NAACL 2024), "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
- arXiv:2603.14712, "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712
- ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables" (2003, reaffirmed 2018). https://asq.org/quality-press/display-item?item=T1164
- arXiv:2406.04244, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
- LMSYS Org, "Catch me if you can! How to beat GPT-4 with a 13B model (LLM Decontaminator, 2023-11-14)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Klaus Krippendorff, Annenberg School for Communication, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- arXiv:2203.04706, "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
- Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-5:2025 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 5: Data quality governance framework" (2025). https://www.iso.org/standard/5259-5
- ISO/IEC JTC 1/SC 42, "ISO/IEC TR 5259-6:2026 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 6: Visualization framework for data quality" (2026). https://www.iso.org/standard/5259-6
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.