Fine-tuning and post-training data
On-policy vs off-policy preference data: what licensed pairs can and cannot do
Quick answer
On-policy preference data is chosen/rejected pairs sampled from the model you are training, or a recent checkpoint, then labeled. Pre-existing licensed pairs are off-policy for your model: another model or person wrote both responses. As of October 2026, research shows off-policy pairs still help, especially on open-ended tasks and when mixed with on-policy samples, but their value falls as your policy diverges. The durable purchases are prompts, rubrics, expert judges and verifiers that can label fresh samples.
By SourceX Editorial · Updated
What makes preference data "on-policy"
Preference data is on-policy when the responses being compared were generated by the current policy, and off-policy when they came from any other source [2]. The label itself (a human or AI judgment of which response is better) can be fresh or old; what matters is who produced the candidate text. A survey of preference learning separates two axes: whether data is collected on- or off-policy, and whether feedback is gathered online during training or fixed in advance [2].
Standard DPO is offline: it fits the policy to a fixed set of (prompt, chosen, rejected) triples with no sampling from the model during fine-tuning [1]. That is why DPO is cheap and stable, and also why it inherits whatever distribution the pairs came from. The classic RLHF pipeline in Llama 2 instead had annotators pick the better of two outputs from the company's own models, then trained a reward model on those comparisons [6].
For background on the format, see the preference data glossary entry and the guide to preference datasets for DPO.
Why fixed preference pairs lose value as the policy moves
Fixed pairs lose value because the gradient signal concentrates on responses your model would rarely produce. If both the chosen and rejected answers sit far from what your policy actually samples, DPO pushes probability mass between two regions the model does not visit, and the improvement on its real outputs can be small. This is the core "static preference dataset limitation": the pairs were informative for the model that generated them.
Three failure modes show up in practice:
- Style mismatch. Pairs written by a stronger or differently tuned model teach formatting, length and tone conventions that do not match your model's register, so gains on reward-model scores may not transfer to blind evaluation.
- Easy negatives. Rejected responses from a weak generator are trivially distinguishable from your model's outputs, so the pair carries little information about the errors your model actually makes.
- Stale coverage. After one or two rounds of tuning, your model stops making the mistakes the old pairs penalize, and further epochs on them mostly add overfitting risk.
What the research says about on-policy vs off-policy DPO data
The evidence says on-policy data is not uniformly better; its advantage depends on task type, base model and training stage. SimpleMix (ICML 2025) found on-policy data helps most on objective tasks such as math and code, while off-policy data does better on open-ended tasks like creative writing, and simply mixing the two improved AlpacaEval 2.0 results by an average of 6.03% over on-policy-only and off-policy-only DPO [3].
An ICLR 2026 paper reports that the benefit of on-policy data varies sharply by model: roughly 3x for Llama-3 but about 0.4x for Zephyr in its setup [4]. The same work suggests early alignment benefits most from diverse data and later stages from high-quality data, which maps neatly to "licensed diversity first, on-policy refinement later" [4]. A systematic study of DPO data also examines which properties of the pairs drive results, including response quality, so treat the on-policy question as one variable among several rather than the whole story [5].
Read these as findings in specific setups (specific base models, judges and benchmarks), not as rules. Run your own ablation before committing a budget.
When licensed off-policy pairs are still the right purchase
Licensed pairs remain valuable when the job is reward-model training, early-stage alignment, or covering open-ended domains your model cannot yet generate well. Reward models need diverse comparisons across many response styles, and a permissively licensed human-annotated set such as HelpSteer2 was built specifically for that purpose [8]. Off-policy data can also warm-start online learning: one line of work models an offline expert dataset of unknown quality and uses it to accelerate later online preference learning [9].
Licensed pairs drawn from real operations are strongest where the "chosen" side encodes judgment your model cannot produce yet. Examples include a support reply that a QA team approved over the agent's first draft, a contract clause a reviewing attorney accepted over the redline, or an engineering design review where the approved version replaced a rejected one. Those pairs carry domain expertise as much as preference, which is why domain-expert preference data is evaluated differently from generic chat pairs.
The durable purchase: prompts, rubrics, judges and verifiers
The inputs that keep their value across training rounds are the ones that let you label fresh samples from your own model. When you run iterative or online DPO, you sample new responses from the current checkpoint, score them, and build new pairs each round [2]. Every round needs the same three things: representative prompts, a way to judge responses, and a way to check that judgment.
- Prompt sets. Real user requests with metadata (intent, domain, difficulty, language) are reusable indefinitely, because you resample responses against them each round. See real-world prompt sets for post-training.
- Rubrics and policies. Written grading criteria, ideally derived from the reviewing organization's own QA forms or style guides, let a human or AI judge label new samples consistently.
- Expert judges. Annotator time from qualified reviewers labels your model's real outputs; the trade-off against AI feedback is covered in AI feedback vs human preference data.
- Verifiers. For math, code and structured extraction, programmatic checks can replace learned labels; Tulu 3 describes reinforcement learning with verifiable rewards built on this idea [7]. Verifiers need clean reference answers and test cases, which are themselves a data purchase.
Decision table: what to buy for each training stage
Match the purchase to the method and stage rather than defaulting to the largest pair count on offer.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Your situation | Method | Highest-value purchase | Role of licensed fixed pairs |
|---|---|---|---|
| Training a reward model from scratch | Pairwise RM, then PPO or rejection sampling | Diverse human comparisons across styles and domains | Primary input |
| First DPO pass on a new SFT checkpoint | Offline DPO | Off-policy pairs plus a held-out prompt set | Primary input; diversity matters most |
| Second and later rounds | Iterative or online DPO | Prompts, rubric, judge capacity | Mix-in or regularizer only |
| Math, code, extraction | RLVR or DPO with verified pairs | Problems with reference answers and tests | Low; verifier labels dominate |
| Regulated or expert domain | DPO with expert-labeled on-policy samples | Expert reviewer hours and domain rubric | Seed set from approved-vs-rejected records |
A request spec for policy-matched preference inputs
A good request describes the inputs you will reuse every round, not just a target pair count. Use a spec like the one below when you talk to any supplier, and pair it with the checks in how to evaluate a fine-tuning dataset before you buy it.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: preference-tuning inputs, support domain
policy_model: internal 8B SFT checkpoint (responses sampled by buyer)
components:
prompts:
source: real customer support tickets, first customer message only
fields: [ticket_id_hash, product_area, intent, language, created_month]
pii: names, emails, phones, account numbers removed or replaced
rubric:
source: supplier QA scorecard converted to 1-5 criteria
criteria: [policy_accuracy, resolution, tone, escalation_correctness]
seed_pairs:
chosen: QA-approved final reply
rejected: first draft or reply flagged by QA
fields: [prompt_ref, chosen, rejected, qa_reason_code, reviewer_role]
judge_capacity: optional expert review of buyer-sampled responses
license_needs: training on prompts and pairs; reuse of rubric to label model outputs
The license matters as much as the data: confirm that prompts may be reused to generate and label new samples, and that the rubric may be applied by an AI judge if you plan to use one. The fine-tuning-only data license guide explains the scope questions, and provenance for human-annotated preference data covers annotator agreements and AI-assistance disclosure.
How SourceX fits into preference-data sourcing
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing process. Those records often contain the raw material for prompts, rubrics and approved-versus-rejected seed pairs. Datasets are not held in stock, and a request does not guarantee a match.
Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery, and every release is approved by the supplying company. SourceX does not train models. Teams can describe the preference-tuning inputs they need on the buyers page.
For the wider landscape of SFT, preference and RL data, start at the fine-tuning and post-training data hub or compare managed collection with licensed preference data.
Sourcing inputs for on-policy preference tuning
If your next rounds of preference tuning need real prompts, domain rubrics or expert-approved seed pairs, describe the data rather than the companies you think hold it. SourceX looks for US businesses that hold the described data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Start at https://sourcex.si/buyers.
Sources
- arXiv (Rafailov et al., Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- arXiv, "Towards a Unified View of Preference Learning for Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2409.02795
- Proceedings of Machine Learning Research (ICML 2025), "SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning" (2025). https://proceedings.mlr.press/v267/li25au.html
- arXiv (ICLR 2026 poster), "Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?" (2025). https://arxiv.org/html/2508.10530v2
- arXiv, "What Matters in Data for DPO?" (2025). https://arxiv.org/pdf/2508.18312
- arXiv (Meta), "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
- arXiv (Allen Institute for AI), "Tulu 3: Pushing Frontiers in Open Language Model Post-Training" (2024). https://arxiv.org/pdf/2411.15124
- arXiv (NVIDIA), "HelpSteer2: Open-source dataset for training top-performing reward models" (2024). https://arxiv.org/pdf/2406.08673
- arXiv, "Online Bandit Learning with Offline Preference Data for Improved RLHF" (2024). https://arxiv.org/pdf/2406.09574
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.