Fine-tuning and post-training data
DPO datasets: preference pair structure, sources and buying options
Quick answer
A DPO dataset is a set of preference records, each holding a prompt, a chosen response and a rejected response to that same prompt, usually delivered as JSONL. Direct preference optimization (DPO) trains on those pairs directly, without a separate reward model, so the pairs are the whole training signal. Buyers get them four ways: open datasets after a license check, commissioned comparisons of their own model's outputs, AI-judged labels, or pairs derived from business records in which a reviewer approved one version over another.
By SourceX Editorial · Updated
For the general term, see preference data; for other post-training data, the fine-tuning and post-training data hub.
How DPO uses each chosen and rejected pair
DPO raises the probability of each chosen response relative to its rejected partner, measured against a frozen reference model, so every pair acts directly on the model with no reward model in between to absorb mistakes. Rafailov et al. introduced the method as a way to fit a policy to preference data without training a reward model or sampling from the model during fine-tuning, and reported results matching or exceeding PPO-based RLHF on sentiment control, summarization and single-turn dialogue [1].
Three data requirements follow:
- Both sides answer one prompt. The loss compares the two responses' log-probabilities given the same context, so a pair whose prompt or history differs between sides teaches nothing reliable.
- The reference model shapes the signal. DPO scores each response against a reference policy, normally the SFT checkpoint that produced the samples. For public datasets whose generating model was unavailable, the DPO authors first fine-tuned on the preferred completions to reduce the distribution shift [1]. Pairs sampled from an unrelated model carry that shift into your run; on-policy vs off-policy preference data covers when it matters.
- A wrong label pushes the wrong way. A 2026 ICML paper notes that standard methods such as DPO treat pairs with high annotator disagreement the same as unanimous ones, and that inconsistent labels can severely degrade aligned models [2].
Related objectives such as IPO, ORPO and SimPO consume the same prompt, chosen and rejected records, and some of them drop the reference model. KTO-style methods instead take single responses labeled desirable or undesirable, which suits approvals and rejections with no matched alternative; see binary feedback data for KTO.
Anatomy of a DPO record: three required fields and the metadata worth paying for
A trainable DPO record needs only prompt, chosen and rejected, but a record worth buying also says where each response came from, who preferred it, how strongly, and under which rubric. Without that metadata you cannot filter noisy pairs, rebalance sources or answer a provenance question later.
The DPO paper defines its training data as triples of a prompt, a preferred completion and a dispreferred completion [1], and public preference sets store the same three fields, usually named prompt, chosen and rejected. In the plain style the prompt is a preformatted string; in the conversational style prompt is a list of messages and chosen and rejected each hold the final assistant turn. Prefer conversational records, because you apply your own chat template at training time instead of inheriting someone else's special tokens.
Files are usually JSON Lines, which requires UTF-8 without a byte order mark and one valid JSON value per line [3]. Check truncation too: a trainer that cuts records at a maximum length can drop the sentences that made a long chosen response better.
Rankings need one more field. InstructGPT's labelers ranked between four and nine responses per prompt, which yields every pairwise comparison within each ranking, and the authors trained on all comparisons from one prompt together because comparisons from the same ranking are correlated and shuffling them caused overfitting [4]. Ask for a group ID so that pairs from one prompt never straddle your train and evaluation splits; pairwise, ranking or rating formats compares the collection designs.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"id": "pair-000418",
"group_id": "ap-dispute-7731",
"prompt": [
{"role": "system", "content": "You draft accounts-payable replies. Apply dispute policy AP-DSP v4."},
{"role": "user", "content": "[SUPPLIER_1] asks why invoice [INVOICE_ID] was paid short by $1,240.00. PO quantity 400, received 380, unit price $62.00."}
],
"chosen": [
{"role": "assistant", "content": "We paid for the 380 units received against [PO_ID]. The 20-unit shortfall, $1,240.00 at $62.00 each, is on hold until we receive a signed delivery receipt for the missing units or your credit note."}
],
"rejected": [
{"role": "assistant", "content": "The invoice was short-paid because of a discrepancy. Please contact us with any questions."}
],
"preference": {"source": "supervisor_edit", "strength": 2, "scale": "-3 to 3", "tie": false, "rationale_code": "missing_quantity_math"},
"response_origin": {"chosen": "human_final_sent", "rejected": "human_first_draft"},
"labeler": {"id": "rev-114", "role": "ap_supervisor", "rubric_version": "ap-reply-rubric-2", "double_rated": false},
"policy_version": "AP-DSP v4",
"privacy": {"method": "typed_placeholders", "replaced": ["supplier_name", "invoice_id", "po_id"], "same_mapping_both_sides": true},
"split": "train",
"license_scope": "preference_tuning"
}
In a JSONL file the record sits on one line. Insist on response_origin, strength, tie and the rubric version; without them you cannot drop ambiguous pairs or separate human-written from model-written responses.
Four ways to source preference pairs, and where each breaks
Preference pairs come from open datasets, commissioned annotation of your own model's outputs, AI-generated labels, or business records in which a person chose one version over another. Each route fails differently, so the first check differs too.
| Route | Responses come from | Preference decided by | Typical defect | First check |
|---|---|---|---|---|
| Open preference datasets | Mostly other models | Crowd raters, AI judges or both | Off-policy responses, misstated licenses, mislabeled pairs | Original license text; generating models |
| Commissioned comparisons | Your SFT checkpoint, several samples per prompt | Trained raters or experts on your rubric | Cost per pair, rater disagreement, guideline drift | Pilot batch with a double-rated subset |
| AI-feedback labels | Your model or a pool of models | An LLM judge with a rubric | Position and length bias; provider output terms | Judge agreement with a human-labeled subset |
| Pairs from business records | Employees: draft and approved final, or competing versions | Supervisors, compliance reviewers, customers | Rejections driven by negotiation, timing or price; personal data | Whether both sides answer the same input |
Open datasets. Public sets differ in who wrote the responses and who judged them. The DPO paper's dialogue experiments used Anthropic's Helpful and Harmless (HH) comparisons [1], while releases such as AI2's Self-Directed Synthetic Dialogues pair model-generated conversations with a preference dataset [5]. Some open sets store a rating per response, with pairs built afterwards from the highest- and lowest-rated responses, so ask which rule was used, and compare candidate sets under your own fixed training recipe.
License metadata across open sets is weak. The Data Provenance Initiative reports license omission above 70% and license error rates above 50% on popular dataset hosting sites [6], so a dataset card's YAML license field [7] is a starting point, not evidence. AI-generated sets can also carry restrictions from the terms of whichever provider produced the completions or ratings: as of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to be used to train competing models such as general-purpose chatbots [8]. See open instruction and preference datasets that allow commercial fine-tuning.
Commissioned comparisons. Having raters choose between responses sampled from your own SFT checkpoint gives on-policy pairs aimed at your model's actual failures. Buying RLHF comparison data compares managed collection with licensing; domain-expert preference data covers specialist raters.
AI feedback. An LLM judge labels pairs cheaply but brings its own biases and its provider's terms. Confirm the judge saw both response orders and was checked against human labels; AI feedback vs human preference data covers where judge labels suffice.
Turning approvals, edits and rejections into valid DPO pairs
Business records yield valid DPO pairs only when two complete responses answer the same input and a person with authority chose between them on quality. Draft-to-final edits are the cleanest source; a bare approval flag is usually not a pair at all.
| Business record | Chosen | Rejected | Valid when | Breaks when |
|---|---|---|---|---|
| Supervisor-edited customer or supplier replies | Approved final as sent | Original draft | The edit changed substance | Edits only fixed typos, signatures or placeholders |
| Compliance review of outbound communications | The approved version | The version returned with a finding | The finding cites a rule and policy version | It was returned for a missing attachment |
| QA-scored responses to the same issue type | Higher-scored response | Lower-scored response | Same issue, policy and rubric version | The responses address different facts |
| Contract redlines | The clause the company accepted | The counterparty's proposed clause | Teaching your side's drafting standard | Read as quality: the rejection reflects negotiating position |
| Competing proposals or quotes | The selected proposal | An unselected proposal | Selection criteria were recorded | Price or relationship decided it, not the text |
A lone approval or rejection fits a binary method instead. For rewriting models, see draft-to-final document pairs; similar record families appear in human feedback datasets of QA scores and corrections and training data for reward models.
De-identify both sides with the same mapping, so [SUPPLIER_1] names the same entity in chosen and rejected and the only difference left is content. Keep the policy version with each pair, because a final approved under last year's policy can be the wrong answer today.
SourceX sources operational datasets from US companies, including support and sales histories, documents, and finance and legal workflows, the systems where edits and approvals are typically recorded. Buyers describe the records they need, not the companies; SourceX looks for US businesses that hold them, and every release is approved by the supplying company. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match. To scope approval-derived pairs, describe the edit or approval signal you need.
Checks to run on a sample before you buy
Ask for a sample with full metadata and run these checks before signing: a dataset card describes intent, while a sample shows whether the pairs are usable. Fix your pilot evaluation prompts before the sample arrives.
- Schema and prompt identity. Every record has a non-empty prompt, chosen and rejected; conversational records share identical history; no record has chosen equal to rejected.
- Trivial differences. Flag pairs whose sides differ only in whitespace, signatures, placeholders or formatting.
- Length skew. Compute the share of pairs in which the chosen response is longer. If nearly all are, length becomes a shortcut the model can learn in place of quality.
- Duplicates and contamination. Find exact and near-duplicate prompts and any overlap with your evaluation prompts.
- Agreement and ties. On a double-rated subset, compute Krippendorff's alpha, where 1 means perfect agreement and 0 means agreement no better than chance [9]. Ask whether raters could mark a tie and whether tied pairs were dropped or forced into an order; inter-annotator agreement metrics explains the choice of statistic.
- Label noise. Score pairs with reward models and read those scored strongly against their label. The Secrets of RLHF Part II study reports that, under this kind of ensemble scoring, the mean preference difference was negative for about 25% of HH-RLHF pairs, suggesting those labels may be wrong [10]. Preference data quality and noise covers cleaning methods.
- Chosen-response quality. Read 50 random chosen responses as if they were going to customers. DPO only teaches the model to prefer chosen over rejected, so a mediocre chosen answer sets a low ceiling.
- Source mix. Tally response-generating models and label sources; a set dominated by one generator or judge passes on its habits.
- Pilot run. Train on the sample or a slice and compare against your SFT baseline on the held-out prompts; evaluating a fine-tuning dataset before you buy it covers pilot design.
Rights and provenance questions specific to preference data
Preference data stacks three layers of rights: the prompts, the responses (often model outputs or employee writing), and the human judgments and rationales. A license that grants "the dataset" can leave one layer unaddressed.
- Model-written responses. Record the generating model and version for every response, and check that provider's output terms as they stood on the generation date [8]. See due diligence for purchased synthetic data.
- Rater work product. Rationales, rewritten chosen responses and rubrics can be copyrightable work product. Contracts should assign or license them and state whether raters could use AI assistance; provenance for human-annotated and preference data covers annotator agreements.
- Personal data on both sides. In business-derived pairs, the rejected draft can hold exactly what a supervisor removed, such as a wrong account number or an over-disclosure, so it needs the same de-identification as the final.
- Use scope. Confirm the license covers preference tuning, release of tuned weights and any hosted fine-tuning service; see fine-tuning-only data licenses.
- Documentation. Ask for a datasheet-style card plus machine-readable metadata. Croissant-RAI, which extends the Croissant format, lists data labeling and participatory data among its design use cases [11].
For datasets SourceX sources, rights review checks that the business owns or may share the records, and names, emails, phone numbers and account numbers are removed or replaced before delivery, with the method recorded; no de-identification method is perfect.
A request template for DPO preference data
A usable request states the behavior to change, where responses and preferences come from, the label rules, the format and the licensed uses. Copy these rows into an RFP or data request.
| Field | What to state | Why it matters for DPO |
|---|---|---|
| Target behavior | The failure to fix, with evaluation prompts held back | Pairs must disagree on it |
| Prompt source | Real user or business inputs versus synthetic, and the task mix; see real-world prompt sets | Sets what the model learns to prefer |
| Response origin | Sampled from your checkpoint, written by people, or generated by named models and versions | Reference-model fit and provider terms |
| Preference source | Raters, domain experts, an AI judge, or recorded business decisions | Sets the noise level and rights chain |
| Rubric and ties | Rubric version, tie option, strength scale, rationales | Lets you drop ambiguous pairs |
| Agreement | Share double-rated; minimum agreement set in the pilot | Makes quality an acceptance test |
| Volume | Unique prompts and pairs per prompt, for a pilot and then full delivery; see cost drivers for human preference data | Pairs from one ranking are correlated [4] |
| Format | JSONL, conversational messages, field names, group ID | Avoids template lock-in and split leakage |
| Privacy | De-identification method, same mapping on both sides | The difference must be content |
| Licensed uses | Preference tuning, weight release, hosted fine-tuning, term | Grant matches plan |
Describe the preference data your DPO run needs
Tell SourceX the behavior you want preference tuning to change, the records that capture the decision (edits, approvals or rejections with reasons), the volume and the uses you need licensed. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Submit your preference data requirements.
Sources
- Rafailov, Sharma, Mitchell, Ermon, Manning, Finn, Stanford (arXiv:2305.18290; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- ICML 2026 (poster page), "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
- jsonlines.org (community format specification), "JSON Lines". https://jsonlines.org/
- Ouyang et al., OpenAI (arXiv:2203.02155), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Lambert et al., Allen Institute for AI (arXiv:2407.18421), "Self-Directed Synthetic Dialogues and Revisions Technical Report" (2024). https://arxiv.org/pdf/2407.18421
- Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- Anthropic Help Center (provider terms guidance), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- Klaus Krippendorff, University of Pennsylvania (technical report), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Wu et al. (arXiv:2401.06080), "Secrets of RLHF Part II: Reward Modeling" (2024). https://arxiv.org/pdf/2401.06080
- Jain et al., MLCommons Croissant RAI task force (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.