Fine-tuning and post-training data
RLVR datasets: prompts, reference answers and verifiers
Quick answer
An RLVR dataset is a set of prompts whose answers a program can check: each record pairs a task with a reference answer or test suite and names the verifier that turns a model's output into a reward, usually 1 or 0. Because that check replaces a learned reward model, verifier mistakes become training-signal mistakes. Before buying, judge three things: whether the verifier is right in both directions, whether your policy solves the prompts sometimes but not always, and whether the tasks trace back to public benchmarks.
By SourceX Editorial · Updated
For adjacent data types, start at the fine-tuning and post-training data hub; the glossary covers reinforcement learning basics.
What verifiable rewards ask of the data
In reinforcement learning with verifiable rewards (RLVR), a deterministic check scores each sampled completion, so the dataset ships ground truth and a way to check it, not human preference labels. AI2's Tulu 3 report named the method: it replaces the reward model with a verification function that pays a fixed reward when a completion is verifiably correct and zero otherwise, and trains with PPO [1]. Tulu 3 applied it to math problems and to instructions with checkable constraints [1].
Three consequences follow for a buyer:
- The reference answer is the label. One wrong answer key rewards wrong outputs on every rollout of that prompt, at every epoch.
- Reasoning text is optional. The policy writes its own reasoning, so a record needs only what the verifier compares. Worked solutions help audit an answer key, while reasoning trace datasets serve supervised training instead.
- Pass rate matters more than raw volume. With group-based methods such as GRPO, a prompt where every sampled output earns the same reward contributes no policy-gradient signal. The DAPO authors note that when all outputs for a prompt are correct the group advantage is zero, so they over-sample and filter out prompts with accuracy of 1 or 0 [2].
Single-turn tasks are the core of RLVR data. Multi-step tasks with tools and changing state belong to RL environments built from business workflows; SourceX also describes training data for RL environments.
Anatomy of an RLVR record
A usable record holds the prompt, a canonical reference answer with its accepted equivalents, a named and versioned verifier, a measured difficulty and the record's lineage. Platform formats ask mainly for the prompt and the reference answer; the rest lets you audit, filter and defend the data.
Platform formats set the minimum; the details below reflect vendor documentation as of October 2026. Microsoft's documentation for OpenAI reinforcement fine-tuning (RFT) on Azure requires JSONL training and validation files with a chat-format messages array whose final message has the user role, and allows extra fields that the grader reads [3]. Amazon Bedrock's guide for open-weight models expects a reference_answer field holding the expected output or evaluation criteria and states a range of 100 to 20,000 examples [5]. Schemas differ by platform, so ask for a neutral delivery format you can map rather than one vendor's layout.
OpenAI's graders guide says graders read those fields through templates whose only namespaces are item, the dataset row, and sample, the model output; as of October 2026 it also says graders are being deprecated as part of the evals and fine-tuning workflows they support [4]. Keep verifier logic in portable code you control, so the dataset does not depend on one grader API.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"task_id": "ap-recon-004112",
"messages": [
{"role": "system", "content": "Give the final amount only, inside <answer></answer>."},
{"role": "user", "content": "PO [PO_ID]: 400 units at $62.00. Receiving log: 386 units received, 6 rejected for damage. Invoice [INVOICE_ID] bills 400 units plus $85.00 freight. Terms T-7: pay accepted units only; freight is payable in full. What amount should be approved?"}
],
"reference_answer": {"value": "23645.00", "type": "decimal", "unit": "USD"},
"accepted_forms": ["23645", "23,645.00", "$23,645.00"],
"verifier": {"id": "numeric_extract", "version": "3.1", "extract": "last <answer> tag", "abs_tolerance": 0.005, "timeout_s": 2},
"difficulty": {"model": "buyer-ckpt-2026-09", "k": 16, "temperature": 1.0, "pass_count": 5},
"lineage": {"origin": "ap_reconciliation_records", "answer_source": "recomputed from PO, receipt and terms; reviewer-confirmed", "upstream_public_sources": []},
"contamination": {"checked_against": ["GSM8K", "MATH"], "methods": ["13-gram overlap", "embedding similarity"], "max_similarity": 0.41},
"privacy": {"method": "typed_placeholders", "numeric_values_preserved": true},
"split": "train",
"license_scope": ["rl_training", "internal_evaluation"]
}
Check a sample for the accepted forms, the verifier version and the model and sample count behind the difficulty figure; without them you cannot tell a hard prompt from a broken one.
Verifier types and how each one fails
Every verifier errs in two directions: false positives pay for wrong answers, which the policy learns to exploit, and false negatives withhold reward from correct answers, which wastes data and pushes the policy toward the checker's preferred format. Ask which verifier each task uses and how its error rates were measured.
| Verifier | Fits | False positive (rewards a wrong answer) | False negative (rejects a right answer) |
|---|---|---|---|
String match; OpenAI's string check offers eq, ne, like and ilike [4] | Multiple-choice letters, labels, codes, IDs | Small label spaces reward guessing; a containment check (like) passes an output that lists every option | Casing, punctuation or extra words around the answer |
| Numeric or symbolic equivalence | Math, finance and engineering quantities | Parser takes the wrong number; tolerance too loose | Equivalent forms such as 1/2, 0.5 and 50%, or unnormalized units [6] |
| Unit tests or execution | Code generation and repair | Weak tests pass hard-coded or partial fixes | Tests tied to one implementation; unreliable environments [7] |
| Execution match on a database | Text-to-SQL | A small fixture database returns the right rows for a wrong query | Row order, column aliases, ties |
| Schema or constraint check | JSON output, length and format instructions | Valid structure with wrong content | Markdown fences or whitespace around valid output |
| Text similarity: fuzzy match, BLEU, ROUGE, cosine [4] | Short extractive answers | High overlap with a factually wrong answer | Correct paraphrases score low |
| Model-based grader | Free-form answers | Judge fooled by response patterns, then exploited in training [6] | Inconsistent scores for unusual but correct answers |
A study of math verifiers found that rule-based checkers often miss equivalent answers written in different formats, a gap that hurts RL training more as the policy gets stronger, while model-based verifiers can be fooled by certain response patterns that the policy then exploits to inflate its reward [6]. A second paper notes that many RLVR systems collapse rewards to binary 0 or 1 to reduce verifier hacking, a choice that brings both false negatives and false positives, and it treats those errors as reward noise to correct [8].
In code, SWE-bench checks a model's repository edit by running tests [9], and OpenAI's review that produced SWE-bench Verified cited overly specific unit tests, underspecified problem statements and unreliable environment setup [7], the same three defects to look for in purchased code-RL tasks. Issue-to-fix pairs from private repositories and text-to-SQL training data from real enterprise schemas cover validation for those task families.
Answer normalization belongs in dataset construction, not only in the verifier: the DAPO authors converted their math answers to integers so they could be parsed reliably [2]. When an answer cannot be reduced to a checkable form, rubrics are the usual next step, and one paper frames RLVR as the special case of rubric-guided RL with a single essential criterion [10]. See rubrics as rewards for non-verifiable domains and critique and revision data for generative verifiers.
Matching difficulty to your policy's pass rate
Difficulty in RLVR is relative to the model being trained: a task is useful while your policy solves it sometimes but not always. A supplier's "hard" label means little unless it names the model, the number of samples (k) and the sampling temperature behind it, so re-measure on your own checkpoint before acceptance and again as training moves prompts into the always-solved band.
| Pass count on your checkpoint | Likely meaning | What to do |
|---|---|---|
| 0 of k | Too hard, or a wrong answer key or broken verifier | Audit the answer and verifier on a sample first; hold verified-hard tasks for later stages |
| 1 to k-1 | Learnable now | Core training pool; track how the band shifts per checkpoint |
| k of k | Already solved, or a verifier that accepts almost anything | Drop from RL batches; spot-check for false positives; reuse as regression checks |
The zero-pass band deserves the closest look because it is where answer-key errors hide: a task no model solves is also a task no model output has confirmed. Ask suppliers for pass-count distributions rather than difficulty labels. Expert worked solutions are one way to confirm answer keys on the hardest tasks.
Lineage and contamination: where RLVR tasks come from
Most public RLVR data is recycled. A 2026 lineage study reports that most RLVR datasets are variants of a small set of upstream sources, many with contamination risks, and its ATLAS framework attributes over 99.7% of 1.45 million instances to 20 atomic sources [11]. Two "different" open sets can therefore hold the same problems, and a test split can overlap the benchmarks you report.
Dataset cards show the pattern. One MIT-licensed RLVR set on Hugging Face describes its training split as preprocessed DeepScaleR-Preview data and its test split as derived from AIME 2024 [12]. A team that already trains on DeepScaleR or reports AIME 2024 results needs to know that before merging.
Contamination erodes what scores mean. In a February 2026 post, OpenAI said it had stopped reporting SWE-bench Verified because contamination meant score gains increasingly reflected training-time exposure [13]. The SWE-Bench Pro authors argue that permissively licensed public repositories are likely to sit in pre-training corpora, and built their benchmark from copyleft repositories plus commercial codebases acquired from startups [14].
Standard checks miss rewritten items. N-gram overlap is the predominant decontamination method, with thresholds that vary by lab; GPT-3's authors used a 13-gram overlap [15]. LMSYS researchers showed that paraphrased or translated test items slip past such checks: a 13B model trained on rephrased MMLU items reached 85.9 without n-gram overlap flagging it [16]. A reworded or number-perturbed benchmark problem is still that benchmark problem.
Freshness is the other defense. LiveBench uses frequently updated questions from recent sources scored against objective ground truth, and its authors cite evidence that model performance on Codeforces drops after the training cutoff [17]. Tasks created after your models' cutoffs from material never published online carry the least exposure risk. Methods are covered in decontaminating a licensed training set against public benchmarks and contamination-resistant evaluation design; the glossary defines benchmark contamination.
Verifiable tasks from business records
Operational records can supply checkable tasks beyond competition math and public code, because business systems store the outcome a task should reproduce: an approved payment amount, a reconciled balance, a routing code, a query's result set or a field a reviewer confirmed. The verifier is then an exact or numeric match against that recorded outcome.
Three conditions decide whether a record becomes a valid task:
- The answer must follow from the prompt. If the approved amount depended on a phone call or a policy missing from the prompt, the task rewards guessing. Put the governing terms or policy version in the prompt.
- The outcome must be correct, not merely recorded. Systems capture overrides, reversals and errors; verifying outcome labels in operational records covers how to test them.
- De-identification must keep the checkable values. Replace names and account numbers with consistent placeholders, but keep the quantities, dates and codes the answer depends on.
SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, the systems where such outcomes are recorded. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match; buyers describe the data, and SourceX looks for US businesses that hold it. To scope outcome-bearing records for verifiable tasks, describe the task types and answer fields you need.
Licensing questions specific to RLVR data
An RLVR dataset bundles separately owned layers: prompts, reference answers, test suites, verifier code and difficulty metadata generated with someone's model. Clear each layer, because a permissive license on the card can sit on top of restricted upstream content.
- Upstream licenses. The Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [18], so trace each record to its upstream source rather than trusting the card. See open datasets that allow commercial fine-tuning.
- Model-generated parts. Reference solutions and paraphrased prompts generated with a third-party model may carry that provider's output terms; see due diligence for synthetic fine-tuning data.
- Verifier and test code. Code usually carries a software license separate from the data license; confirm you may modify and run it in your training infrastructure.
- Derived tasks and use scope. Confirm the license lets you create variants such as rephrasings or new numeric instances, and covers RL training, evaluation and release of the trained weights.
For datasets SourceX sources, rights review checks that the business owns or may share the records and that required consents are in place, and the license defines which records are included, what they can be used for, how long the license runs and how delivery happens.
RLVR dataset acceptance checklist
A purchasable RLVR set has a verifier that is accurate in both directions and deterministic, pass rates that fit your checkpoint, and clean lineage. Run these checks on a sample before signing, with your own checkpoint and training stack.
- Schema. Every record has a prompt, a reference answer and a verifier ID and version, and the fields map to your trainer's format, including a separate validation split where your platform requires one [3].
- Self-check. The verifier passes every reference answer and every listed accepted form.
- Negative probes. The verifier fails perturbed wrong answers, empty outputs, answers that list every option, and code that hard-codes expected outputs.
- Human audit. People adjudicate a few hundred policy outputs; compute false-positive and false-negative rates per task type.
- Determinism. Re-running the verifier yields identical rewards; code runs in a pinned container with time and memory limits and no network access.
- Pass-rate distribution. Measured on your checkpoint at a stated k and temperature, with zero-pass tasks audited.
- Duplicates. Exact and near-duplicate prompts removed within the set and against your existing RL pool.
- Contamination. Checked against named benchmarks with both n-gram and paraphrase-aware methods, with method and threshold recorded.
- Lineage and rights. Upstream source and license per record, and generating models named for any model-written part.
Source verifiable tasks beyond public math and code
Tell SourceX the task types you want to train on, the outcome fields that should serve as reference answers, the volume and the uses you need licensed. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your verifiable task data.
Sources
- Lambert et al., Allen Institute for AI (arXiv:2411.15124), "Tulu 3: Pushing Frontiers in Open Language Model Post-Training" (2024). https://arxiv.org/pdf/2411.15124
- Yu et al., ByteDance Seed and collaborators (arXiv:2503.14476), "DAPO: An Open-Source LLM Reinforcement Learning System at Scale" (2025). https://arxiv.org/pdf/2503.14476
- Microsoft Learn, Foundry documentation for Azure OpenAI models, "Reinforcement fine-tuning (how-to guide)". https://learn.microsoft.com/en-my/Azure/foundry/openai/how-to/reinforcement-fine-tuning
- OpenAI API documentation, "Graders". https://developers.openai.com/api/docs/guides/graders.md
- Amazon Web Services, Amazon Bedrock User Guide, "Prepare data for open-weight models (reinforcement fine-tuning)". https://docs.aws.amazon.com/bedrock/latest/userguide/rft-prepare-data-open-weight.html
- arXiv preprint 2505.22203 (v2), "From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning" (v1 titled "Pitfalls of Rule- and Model-based Verifiers") (2025). https://arxiv.org/html/2505.22203v2
- OpenAI (with the SWE-bench authors), "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- arXiv preprint 2510.00915 (v3), "Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers" (2025). https://arxiv.org/html/2510.00915v3
- Jimenez et al., Princeton and UChicago (arXiv:2310.06770; ICLR 2024), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- Gunjal et al. (arXiv:2507.17746), "Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains" (2025). https://arxiv.org/pdf/2507.17746
- arXiv preprint 2605.26971, "RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data" (2026). https://arxiv.org/pdf/2605.26971
- Miaow-Lab on Hugging Face, "RLVR-Linearity-Dataset (dataset card README)". https://huggingface.co/datasets/Miaow-Lab/RLVR-Linearity-Dataset/blob/main/README.md
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Scale AI researchers (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Xu et al. (arXiv:2406.04244), "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
- LMSYS Org (blog post introducing the LLM Decontaminator, 14 November 2023), "Catch me if you can! How to beat GPT-4 with a 13B model" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
- White et al. (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.