Evaluation and benchmarking datasets
LLM evaluation datasets: a buyer's map of benchmarks, private eval sets and licensed test data
Quick answer
LLM evaluation datasets come from five sources: public benchmarks, privately commissioned expert sets, licensed business records with real outcomes, synthetic or model-generated items, and golden sets drawn from your own production traffic. Public benchmarks are cheap and comparable across models, but they are often contaminated and rarely match a deployment. Private and licensed sets cost more and take longer to assemble, yet they measure the task you will ship. Use benchmarks to screen models and keep held-out private data for release decisions.
By SourceX Editorial · Updated
Definitions live in what AI evaluation data is, the eval set entry and training data vs evaluation data.
Five sources of LLM evaluation data and what each can prove
Each source proves something different: benchmarks compare models on a shared task; private and licensed sets test yours. One survey groups 112 evaluation datasets into 20 domains, from reasoning and code to law, medicine and tool use [1]. Breadth is not fit: a 2025 practical guide notes that public benchmarks often lack use-case specificity, may be contaminated and need auditing for quality and licensing, and recommends human annotation for specialized knowledge [2].
| Source | Where items come from | Measures well | Main failure mode | First licensing question | Go deeper |
|---|---|---|---|---|---|
| Public benchmarks | SWE-bench, LegalBench, τ-bench and similar published sets | Cross-model comparison on a fixed task | Contamination, saturation, label errors, task mismatch | Is commercial use allowed, including by upstream sources? | Private sets vs public benchmarks |
| Commissioned private sets | Expert-written items, references and rubrics | Capabilities no benchmark covers | Templated items; high cost per item | Who owns items and rubrics; can the vendor reuse them? | Commissioning a custom eval set |
| Licensed business records | Cases, claims, code changes and redlines with recorded outcomes | Realistic long-tail inputs; labels set by real decisions | Noisy outcomes, policy drift, test-breaking redaction | Is the grant evaluation-only, and may results be published? | Outcome-labeled evaluation data |
| Synthetic or generated sets | Model-written items from documents or seed examples | Many variants of rare formats | Shares the generator's blind spots | Do the generator's terms restrict output use? | When synthetic test sets mislead |
| Production golden sets | Your logs, escalations and user corrections | Regressions on traffic you already serve | Only traffic you already see; logs need privacy review | Do your user terms permit this reuse? | Golden sets from business records |
Why public benchmark scores rarely settle a deployment decision
A public benchmark score reports how a model did on items that may already be in its training data, graded by labels that may be wrong, on a task that may not resemble yours.
Contamination. The LiveBench authors call test-set contamination a well-documented obstacle that can quickly make benchmarks obsolete, citing Codeforces results that drop after models' training cutoffs [3]. A 2024 analysis notes that a private holdout could verify published scores, but most benchmarks have none [4]. In a February 2026 post, OpenAI said SWE-bench Verified had become increasingly contaminated, stopped reporting it and recommended SWE-bench Pro in the interim [5]. SWE-Bench Pro splits tasks by access: a public set from copyleft repositories, a private held-out set for future overfitting checks, and a commercial set from startup codebases that stay private [6].
Label errors. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and showed that such errors can change model rankings [7].
License gaps. The Data Provenance Initiative audited more than 1,800 text datasets and, in its arXiv version, reported license omission above 70% and license error rates above 50% on popular hosting sites [8]. Eval sets inherit their sources' terms: one red-teaming dataset on Hugging Face is released under CC BY-NC 4.0, and all its prompts come from datasets with non-commercial licenses [9]. Before scores go into a model card or sales deck, check the eval dataset's license for commercial use.
Benchmarks remain cheap screening; dynamic ones such as LiveBench limit contamination with fresh questions scored against objective ground truth [3]. See contamination checks for licensed evaluation data for detection, contamination-resistant evaluation design for prevention, and when public scores are enough and when you need held-out data for the decision.
Match the evaluation job to a data source
The right source depends on the decision the eval drives, which also dictates what each item must carry.
| Evaluation job | Source that usually fits | Each item must carry | Start with |
|---|---|---|---|
| Model or vendor selection | Private set on your task; benchmarks for screening | Real inputs, reference outputs or rubric, slice tags | Model bake-off with a private eval set |
| Release gating and regression | Frozen, versioned golden set plus long-tail cases | Stable item IDs, expected behavior, failure-mode tags | Stratified sets for rare and high-risk cases |
| Tool-using and workflow agents | Task suites with environment state and policies | Starting state, allowed tools, policy text, goal state | Agent evaluation task suites |
| RAG and enterprise search | Question-evidence-answer triples on a corpus snapshot | Corpus version, gold passages, answerability flag | Question-answer-citation triples for RAG |
| LLM-as-a-judge validation | Items graded by several domain reviewers | Rubric version, individual grades, adjudicated label | Judge calibration sets |
| Domain capability (legal, finance, health) | Licensed records with expert or outcome labels | Source document, expert label, reviewer qualification | Contract review evaluation sets |
| Coding agents | Held-out repositories with real issues and tests | Pre-change repository state, issue text, fail-to-pass tests | Held-out coding agent sets from private repositories |
Published benchmarks show the item structure to copy. τ-bench pairs databases, APIs and domain policy documents with annotated user scenarios, grades the database state and agent-to-user messages at the end of a conversation, and reports pass^k, the probability of succeeding on all k repeated trials of a task [10].
For RAG, the WixQA authors note that end-to-end evaluation needs the knowledge-base snapshot the answers came from, not just question-answer pairs [11]. The Ragas v0.1 documentation gives test records question, contexts, answer and ground_truth fields [12]; check the field names in the version you use.
LegalBench's 162 tasks cover six types of legal reasoning and were built with legal professionals [13]. FinanceBench asks 10,231 questions about public companies with answers and evidence strings; on a 150-case sample, its authors found GPT-4-Turbo with retrieval answered incorrectly or refused 81% [14]. Only a FinanceBench subset is open, and Patronus AI's documentation says the full benchmark is licensed from the company [15].
Where licensed business records fit in an eval program
Licensed business records supply what benchmarks and synthetic sets lack: inputs that were never published and labels set by a real decision, such as a claim adjudication, a ticket's final disposition or the tests merged with a code fix. Their timestamps also let you keep only items created after a model's training cutoff; see post-cutoff temporal holdouts.
Outcome fields need checking: an auto-close rule can mark a ticket resolved with no agent reply, and a refund approved under 2022 policy may be wrong under 2025 rules, so each item needs the policy version in force. De-identification can break the test, for example when every person becomes one placeholder and speakers blur together; see de-identifying evaluation data without breaking tests.
SourceX sources operational datasets from US companies, such as support and sales histories, engineering records, documents, and finance and legal workflows, and manages the license. Datasets are sourced on request, not held in stock, so a request does not guarantee a match; every release is approved by the supplying company. Names, emails, phone numbers and account numbers are removed or replaced before delivery, with the method recorded, though no de-identification method is perfect. To source such records for a held-out set, describe the evaluation data you need; evaluation sets built from real business work lists record types.
Illustrative example: invented to show structure; it does not describe an available dataset.
eval_item:
item_id: sup-refund-000417
set_version: v3.2 # frozen; any edit creates a new version
source:
record_type: support_case
systems: [helpdesk_ticket, order_lookup, refund_ledger]
created_at: 2025-11-14 # after the cutoff of the models under test
policy_version: refunds-2025-09
input:
customer_message: "Second damaged order this month. I want a refund, not a replacement."
context_refs: [order_snapshot_8812, refund_policy_2025_09]
expected:
action: issue_refund
must_not: [offer_replacement_first]
outcome_label: refund_approved # final disposition
label_source: supervisor_qa_review # not the auto-close status
rubric_id: support-refund-rubric-v2
slices: {intent: refund, channel: email, risk: repeat_damage}
privacy: {deid_method: consistent_pseudonyms, dates_preserved: true}
provenance: {published_anywhere: false, license_scope: evaluation_only}
Commissioned sets, synthetic items and judge calibration
Commission experts when the capability needs judgment no record captures, such as drafting graded against a rubric, and generate items when you need many variants of a fixed format. Either way, validate the grader before trusting scores.
Commissioned sets. A 2025 paper examines the transparency and conflict-of-interest risks when private data curators run evaluations [16]. Ask whether the vendor's writers or items also feed training data it sells; see third-party evaluation vendor independence.
Synthetic items. Libraries such as Ragas generate test sets from your documents [12], but a model-written question with a model-written answer is a second model output, not ground truth; keep generated items human-checkable.
Judges. LiveBench avoids LLM judges because its authors say they can introduce biases and break down on hard questions [3]. Before an LLM grades your eval, compare its agreement with human graders to agreement among the humans, using a statistic like Krippendorff's alpha, where 1 is perfect reliability and 0 is chance-level agreement [17]. Human feedback and QA score datasets can supply the human reference; rubric design with domain experts covers the rubric.
Size, refresh and leakage control
A private set earns its cost only if it detects the differences you care about and stays unseen by the models it tests.
Size. A 2025 ICML position paper argues that CLT-based confidence intervals are too narrow on evals with fewer than a few hundred items [18]. One worked example finds that detecting a move from 82% to 85% accuracy at 80% power and 5% significance needs about 2,400 items per model variant, while a 100-item set cannot reliably detect gaps smaller than about 10 to 12 points [19]. Paired comparisons need fewer items. Fix the smallest gain that would change your decision, then size the set and each slice.
Refresh and leakage. Sets saturate as models improve and leak as items are reused, so budget a refresh cadence and retire leaked items. A hosted model API receives every test item you send, so check its retention and training terms, keep the held-out split access-controlled and log which version each run used. See keeping a private eval set private and, when data cannot leave the supplier, supplier-hosted and enclave evaluation.
License terms and rules that apply to test data
Evaluation licenses must protect the test from exposure, which training licenses rarely address:
- Scope: evaluation-only or also training, and whether benchmarking third-party models is permitted; see evaluation-only data license terms and licensing data for evaluation only.
- External model APIs: whether items may be sent to outside providers for scoring.
- Publication: scores, sample items or aggregates only (publishing results on licensed eval data).
- Supplier reuse: whether the same items are licensed to others, including developers of the models you test.
- Retention and derivatives: how long frozen versions may be kept, and who owns rubrics, judge prompts and adjudicated labels you create.
For high-risk AI systems in the EU, Article 10 of the AI Act applies its data-governance and quality criteria to testing data sets as well as training and validation sets [20]. As of October 2026, the Act has been amended by Regulation (EU) 2026/1744, which altered the Article 10 data quality criteria [21]. See regulatory requirements for AI validation and test data and the AI training data licensing hub.
Checklist: specifying a private evaluation dataset
A specification should let a supplier price the work and let you reject a delivery that misses it.
- Decision: what the eval decides and the smallest score difference that matters.
- Unit and format: turn, conversation, multi-step task or document set, as JSONL with stable item IDs plus any corpus or environment snapshot.
- Ground truth: label source (outcome field, expert reference or adjudicated grade) and the adjudication rule.
- Slices: minimum item counts per slice, including rare and high-risk cases.
- Freshness: never published and, where it matters, created after a stated date.
- Privacy: de-identification that preserves the property under test.
- Documentation: a datasheet or machine-readable metadata; NeurIPS 2026 requires Responsible AI (RAI) metadata included in the dataset's Croissant file for its Evaluations and Datasets Track [22].
- License: the terms listed above.
- Acceptance: a gold-label audit on a random sample before sign-off.
Writing an evaluation dataset specification and gold-label audits and adjudication expand these steps.
Start here: evaluation guides by question
Each guide below answers one sourcing or design question in depth.
- Choosing a source: private sets vs public benchmarks, what a custom evaluation dataset costs, sourcing domain-expert raters.
- Agents, RAG and long context: computer-use agent tasks with verifiable end states, faithfulness labels for RAG answers, permission-aware RAG evaluation, long-context evaluation on real documents, prompt injection sets for agents.
- Domain sets: customer support policy-following, medical coding, insurance claims, accounting agents, voice agents, text-to-SQL on enterprise schemas, document extraction.
Mistakes that make an eval set measure the wrong thing
Three mistakes produce clean scores for the wrong capability.
- Rephrasing public items and calling the set private. Paraphrases keep the overlap and hide it from string matching: LMSYS researchers showed a 13B model trained on rephrased MMLU test items scored 85.9 on MMLU while n-gram overlap checks missed it [23].
- Reporting one aggregate. A strong overall score can hide a failing slice; report per slice.
- Buying training and test data from one pool without split controls. Near-duplicates of test items in a training purchase inflate every later score.
Need real business records for evaluation?
Describe the evaluation data you need: the task, the outcome or expert labels, the slices and the license scope. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe your evaluation data needs.
Guides in this section
- AI Agent Evaluation Benchmark Design: Tasks, State, pass^kHow an AI agent evaluation benchmark is built: seeded state, tools, policy, simulated user, goal-state grading, pass^k trials and the records behind each.
- Contamination-Resistant Evaluation Set Design for LLMsHow to design LLM eval sets models are unlikely to have trained on: never-published sources, post-cutoff items, private tiers, refresh and governance.
- Contract Review AI Evaluation Datasets from Real RedlinesSource a private contract review AI evaluation set: clause gold labels, playbook risk flags and redline tasks from real negotiations, beyond CUAD.
- Custom LLM Eval Dataset: Scope, SOW and Held-Out ControlsHow to commission a custom LLM evaluation dataset: scope, sourcing options, statement of work, gold-label method, pilot batch and held-out controls.
- Customer Support AI Agent Evaluation: Policies and pass^kBuild customer support AI agent evals from real policies and resolved cases: task anatomy, end-state grading, pass^k reliability and the data to request.
- Evaluation-Only Data License Terms Buyers Should NegotiateWhat an evaluation-only data license must define: permitted eval activities, no-training scope, API recipients, disclosure, deletion, audit and pricing.
- Golden Dataset for LLM Evaluation Built from Real RecordsHow to build a golden dataset for LLM evaluation from real business records: choosing cases, expert-validated references, de-identification and versioning.
- LLM Eval Set Size: How Many Examples for Statistical PowerSize an LLM eval set from the smallest gap you must detect: power math, paired tests, small-n intervals, clustering, slice minimums and a worked example.
- LLM Evaluation Rubric Design with Domain ExpertsHow domain experts turn failure modes into an LLM evaluation rubric: pass/fail criteria, required claims, anchored scales, agreement pilots, versioning.
- LLM Regression Testing Datasets for Production CI GatesBuild an LLM regression testing dataset: turn production failures into test cases, choose graders and set per-slice CI gates for model and prompt changes.
- LLM-as-a-Judge Calibration Datasets: Sizing and SourcingHow to size, source and score the human-labeled set that validates an LLM judge: item counts, raters, kappa and alpha, splits, bias slices and licenses.
- Private Eval Sets vs Public Benchmarks: When You Need EachWhen public LLM benchmark scores are enough, when contamination and task mismatch call for a private held-out eval set, and how to source and protect one.
- RAG Evaluation Datasets: Question-Answer-Citation TriplesHow to build a RAG evaluation set of question, answer and citation triples: record fields, evidence spans, source-first construction and acceptance checks.
- Real-World Outcome Labels as AI Evaluation Ground TruthUse real business outcomes as eval ground truth: which labels to trust, fields to request, selective labels, policy drift, outcome lag and agent grading.
- Accounting Agent Evaluation: GL Coding and ReconciliationHow to build accounting AI agent evaluation sets from reviewed GL coding, bank reconciliation and close records, graded on the ledger end state.
- Computer-Use Agent Eval Tasks With Verifiable End StatesHow to turn real business software workflows into computer-use agent eval tasks with state-based checks, enough tasks per app and a UI drift plan.
- Custom LLM Evaluation Dataset Cost: Drivers and BudgetWhat drives the cost of a custom LLM evaluation dataset: item count from power analysis, expert grading, rater redundancy, context capture and refresh.
- De-identify Evaluation Data Without Breaking the TestHow to de-identify LLM evaluation items with consistent surrogates so coreference, formats, retrieval cues and answer keys still test what they should.
- Document Extraction Eval Sets: Field-Level Ground TruthHow to source document extraction evaluation sets with field-level ground truth from posted invoices and forms: metrics, normalization, slices and privacy.
- Domain-Expert Raters for LLM Evaluation: Sourcing GuideHow to find, qualify, size and manage credentialed domain-expert raters for LLM evaluation, and when expert-reviewed business records are a better fit.
- Eval Data Vendor Independence: Conflicts and VerificationHow to vet an outside evaluation data vendor for conflicts of interest, training-data overlap, label quality and reproducible scores before trusting them.
- Eval Dataset Acceptance Criteria: Gold-Label AuditsAcceptance criteria for a delivered eval set: audit sample sizes, gold-label error tolerances, expert re-labeling, adjudication, remedies before payment.
- Eval Set Refresh Cadence: Saturation, Drift and LeakageWhen to refresh an LLM eval set: saturation, drift and leakage triggers, rolling windows, frozen anchor subsets, and supply terms for recurring eval data.
- Evaluate a Model Without Taking the Test DataHow to evaluate a model on test data that never leaves the owner: supplier-hosted runs, attested enclaves, what results may leave, and how to verify them.
- Evaluation Dataset Specification: Fields, Slices, AcceptanceHow to write an evaluation dataset specification a supplier can price and build: record schema, context capture, gold formats, slices, size and acceptance.
- Financial LLM Evaluation Datasets: Analyst QA With EvidenceHow to source a financial LLM evaluation dataset: analyst questions, gold answers, evidence strings, private documents and retrieval-aware scoring.
- Function-Calling Evaluation Datasets: Items and ScoringHow to specify a function-calling eval set: tool catalogs, requests, expected calls and arguments, forbidden calls, multi-step state, and argument scoring.
- How to Keep a Private Eval Set Private: Controls That WorkAccess tiers, canary strings, API retention settings, vendor terms and a leak-response plan to keep a licensed or in-house LLM eval set from leaking.
- Insurance Claims AI Evaluation Datasets from Closed ClaimsBuild an insurance claims AI evaluation dataset from closed claim files: final adjudication labels, reopenings, severity slices and state-based grading.
- LLM Vendor Bake-Off: Pick a Model With a Private Eval SetHow to run an LLM vendor bake-off on your own private eval set: paired design, use-case slices, error bars, cost, latency, refusals and item protection.
- Long-Context Evaluation Datasets Built on Real DocumentsWhy needle-in-a-haystack scores overstate long-context ability, and how to build eval sets from real documents with distractors and controlled evidence.
- Medical Coding AI Evaluation Data with Audited CodesHow to source de-identified encounters with post-audit ICD-10 and CPT codes to test coding AI: ground truth, slices, scoring and de-identification.
- Permission-Aware RAG Evaluation: Testing ACL LeakageHow to build RAG eval items with user identities, group memberships and document ACLs to catch forbidden retrieval and answer leakage before launch.
- Post-Cutoff Evaluation Data: Building Temporal HoldoutsHow to build LLM test sets from records created after a model's training cutoff: which timestamps to trust, matching cutoffs, and controlling for drift.
- Prompt Injection Evaluation Sets for Tool-Using AgentsHow to build indirect prompt injection eval sets for agents: realistic carrier emails and documents, attacker goals, safe-behavior labels and state checks.
- Public Benchmark Licenses for Commercial EvaluationCheck whether a public benchmark's license and upstream sources allow commercial evaluation, with an audit checklist and keep-or-replace table.
- Publishing Benchmark Results on Licensed Eval DataWhich scores, slices and example items you can publish from a licensed private eval set, and the disclosure clauses to negotiate before your paper ships.
- RAG Evaluation on Outdated and Conflicting DocumentsHow to build RAG eval items over superseded and conflicting document versions: version metadata, distractors, forbidden claims and split retrieval scoring.
- RAG Faithfulness Evaluation Sets: Grounded vs UnsupportedHow to source or build RAG faithfulness evaluation sets: claim-level labels, balanced unsupported answers, and validating groundedness metrics and judges.
- Real-Work Task Benchmarks Built on Professional DeliverablesHow to source and grade real professional tasks: briefs, input files, accepted deliverables and reviewer feedback for benchmarking models on real work.
- Red-Team Eval Datasets With Commercial-Use RightsCheck commercial rights in red-team and adversarial prompt datasets, audit upstream licenses, and decide when to commission a custom safety eval set.
- Side-by-Side Human Preference Evaluation for LLMsHow to source blinded pairwise human judgments to compare LLM versions: prompt sets, rater panels, order randomization, ties, win rates and sample size.
- Stratified Eval Sets: Allocating Examples to Rare CasesHow to split an LLM eval budget across slices: per-slice floors, oversampling rare high-risk cases, weighted vs unweighted scores and release gates.
- Summarization Evaluation Sets from Real Source-Summary PairsHow to source summarization evaluation data: transcripts, calls and case files paired with summaries people wrote for real work, scored by claim.
- Synthetic Evaluation Data for LLMs: Where It MisleadsWhen LLM-generated eval items are safe to use, where synthetic test sets mislead on RAG and agents, and how to verify generated items against real data.
- Unanswerable Questions in RAG Evaluation: Abstention TestsHow to build and score unanswerable questions for RAG evaluation: answerability labels, sourcing from real logs, mix ratios and abstention metrics.
- User Simulator Scenarios for Agent EvaluationHow to ground LLM user simulators in real conversations: scenario schema, persona and hidden-information design, simulator bias, and validation checks.
- Voice Agent Evaluation Datasets from Real Call ScenariosHow to build a voice agent evaluation dataset from real calls: scenario records, outcome labels, turn-taking and latency tests, and consent checks.
Sources
- arXiv:2402.18041, "Datasets for Large Language Models: A Comprehensive Survey" (2024). https://arxiv.org/pdf/2402.18041
- arXiv:2506.13023, "A Practical Guide for Evaluating LLMs and LLM-Reliant Systems" (2025). https://arxiv.org/html/2506.13023v1
- White et al. (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- arXiv:2410.09247, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv:2509.16941, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Galtea AI (Hugging Face), "Galtea red teaming clustered data (non-commercial subset): dataset card". https://huggingface.co/datasets/Galtea-AI/galtea-red-teaming-clustered-data/blob/main/README.md
- Yao et al., Sierra (arXiv:2406.12045), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- arXiv:2505.08643, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
- Ragas documentation (v0.1.21), "Prepare your test dataset". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
- Guha et al. (arXiv:2308.11462; NeurIPS 2023), "LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models" (2023). https://arxiv.org/abs/2308.11462v1
- Islam et al. (arXiv:2311.11944), "FinanceBench: A New Benchmark for Financial Question Answering" (2023). https://arxiv.org/abs/2311.11944v1
- Patronus AI (vendor documentation, market practice), "FinanceBench". https://docs.patronus.ai/docs/financebench-1
- arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
- Klaus Krippendorff, Annenberg School for Communication, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- ICML 2025 (poster), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- tianpan.co (practitioner blog), "Statistical power for LLM evals" (2026). https://tianpan.co/blog/2026/04/15/statistical-power-llm-evals
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union (Official Journal, via EUR-Lex), "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- LMSYS Org, "LLM Decontaminator blog post (2023-11-14)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.