Skip to content

Fine-tuning and post-training data

DPO datasets: preference pair structure, sources and buying options

Quick answer

A DPO dataset is a set of preference records, each holding a prompt, a chosen response and a rejected response to that same prompt, usually delivered as JSONL. Direct preference optimization (DPO) trains on those pairs directly, without a separate reward model, so the pairs are the whole training signal. Buyers get them four ways: open datasets after a license check, commissioned comparisons of their own model's outputs, AI-judged labels, or pairs derived from business records in which a reviewer approved one version over another.

By SourceX Editorial · Updated

For the general term, see preference data; for other post-training data, the fine-tuning and post-training data hub.

How DPO uses each chosen and rejected pair

DPO raises the probability of each chosen response relative to its rejected partner, measured against a frozen reference model, so every pair acts directly on the model with no reward model in between to absorb mistakes. Rafailov et al. introduced the method as a way to fit a policy to preference data without training a reward model or sampling from the model during fine-tuning, and reported results matching or exceeding PPO-based RLHF on sentiment control, summarization and single-turn dialogue [1].

Three data requirements follow:

  • Both sides answer one prompt. The loss compares the two responses' log-probabilities given the same context, so a pair whose prompt or history differs between sides teaches nothing reliable.
  • The reference model shapes the signal. DPO scores each response against a reference policy, normally the SFT checkpoint that produced the samples. For public datasets whose generating model was unavailable, the DPO authors first fine-tuned on the preferred completions to reduce the distribution shift [1]. Pairs sampled from an unrelated model carry that shift into your run; on-policy vs off-policy preference data covers when it matters.
  • A wrong label pushes the wrong way. A 2026 ICML paper notes that standard methods such as DPO treat pairs with high annotator disagreement the same as unanimous ones, and that inconsistent labels can severely degrade aligned models [2].

Related objectives such as IPO, ORPO and SimPO consume the same prompt, chosen and rejected records, and some of them drop the reference model. KTO-style methods instead take single responses labeled desirable or undesirable, which suits approvals and rejections with no matched alternative; see binary feedback data for KTO.

Anatomy of a DPO record: three required fields and the metadata worth paying for

A trainable DPO record needs only prompt, chosen and rejected, but a record worth buying also says where each response came from, who preferred it, how strongly, and under which rubric. Without that metadata you cannot filter noisy pairs, rebalance sources or answer a provenance question later.

The DPO paper defines its training data as triples of a prompt, a preferred completion and a dispreferred completion [1], and public preference sets store the same three fields, usually named prompt, chosen and rejected. In the plain style the prompt is a preformatted string; in the conversational style prompt is a list of messages and chosen and rejected each hold the final assistant turn. Prefer conversational records, because you apply your own chat template at training time instead of inheriting someone else's special tokens.

Files are usually JSON Lines, which requires UTF-8 without a byte order mark and one valid JSON value per line [3]. Check truncation too: a trainer that cuts records at a maximum length can drop the sentences that made a long chosen response better.

Rankings need one more field. InstructGPT's labelers ranked between four and nine responses per prompt, which yields every pairwise comparison within each ranking, and the authors trained on all comparisons from one prompt together because comparisons from the same ranking are correlated and shuffling them caused overfitting [4]. Ask for a group ID so that pairs from one prompt never straddle your train and evaluation splits; pairwise, ranking or rating formats compares the collection designs.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "id": "pair-000418",
  "group_id": "ap-dispute-7731",
  "prompt": [
    {"role": "system", "content": "You draft accounts-payable replies. Apply dispute policy AP-DSP v4."},
    {"role": "user", "content": "[SUPPLIER_1] asks why invoice [INVOICE_ID] was paid short by $1,240.00. PO quantity 400, received 380, unit price $62.00."}
  ],
  "chosen": [
    {"role": "assistant", "content": "We paid for the 380 units received against [PO_ID]. The 20-unit shortfall, $1,240.00 at $62.00 each, is on hold until we receive a signed delivery receipt for the missing units or your credit note."}
  ],
  "rejected": [
    {"role": "assistant", "content": "The invoice was short-paid because of a discrepancy. Please contact us with any questions."}
  ],
  "preference": {"source": "supervisor_edit", "strength": 2, "scale": "-3 to 3", "tie": false, "rationale_code": "missing_quantity_math"},
  "response_origin": {"chosen": "human_final_sent", "rejected": "human_first_draft"},
  "labeler": {"id": "rev-114", "role": "ap_supervisor", "rubric_version": "ap-reply-rubric-2", "double_rated": false},
  "policy_version": "AP-DSP v4",
  "privacy": {"method": "typed_placeholders", "replaced": ["supplier_name", "invoice_id", "po_id"], "same_mapping_both_sides": true},
  "split": "train",
  "license_scope": "preference_tuning"
}

In a JSONL file the record sits on one line. Insist on response_origin, strength, tie and the rubric version; without them you cannot drop ambiguous pairs or separate human-written from model-written responses.

Four ways to source preference pairs, and where each breaks

Preference pairs come from open datasets, commissioned annotation of your own model's outputs, AI-generated labels, or business records in which a person chose one version over another. Each route fails differently, so the first check differs too.

RouteResponses come fromPreference decided byTypical defectFirst check
Open preference datasetsMostly other modelsCrowd raters, AI judges or bothOff-policy responses, misstated licenses, mislabeled pairsOriginal license text; generating models
Commissioned comparisonsYour SFT checkpoint, several samples per promptTrained raters or experts on your rubricCost per pair, rater disagreement, guideline driftPilot batch with a double-rated subset
AI-feedback labelsYour model or a pool of modelsAn LLM judge with a rubricPosition and length bias; provider output termsJudge agreement with a human-labeled subset
Pairs from business recordsEmployees: draft and approved final, or competing versionsSupervisors, compliance reviewers, customersRejections driven by negotiation, timing or price; personal dataWhether both sides answer the same input

Open datasets. Public sets differ in who wrote the responses and who judged them. The DPO paper's dialogue experiments used Anthropic's Helpful and Harmless (HH) comparisons [1], while releases such as AI2's Self-Directed Synthetic Dialogues pair model-generated conversations with a preference dataset [5]. Some open sets store a rating per response, with pairs built afterwards from the highest- and lowest-rated responses, so ask which rule was used, and compare candidate sets under your own fixed training recipe.

License metadata across open sets is weak. The Data Provenance Initiative reports license omission above 70% and license error rates above 50% on popular dataset hosting sites [6], so a dataset card's YAML license field [7] is a starting point, not evidence. AI-generated sets can also carry restrictions from the terms of whichever provider produced the completions or ratings: as of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to be used to train competing models such as general-purpose chatbots [8]. See open instruction and preference datasets that allow commercial fine-tuning.

Commissioned comparisons. Having raters choose between responses sampled from your own SFT checkpoint gives on-policy pairs aimed at your model's actual failures. Buying RLHF comparison data compares managed collection with licensing; domain-expert preference data covers specialist raters.

AI feedback. An LLM judge labels pairs cheaply but brings its own biases and its provider's terms. Confirm the judge saw both response orders and was checked against human labels; AI feedback vs human preference data covers where judge labels suffice.

Turning approvals, edits and rejections into valid DPO pairs

Business records yield valid DPO pairs only when two complete responses answer the same input and a person with authority chose between them on quality. Draft-to-final edits are the cleanest source; a bare approval flag is usually not a pair at all.

Business recordChosenRejectedValid whenBreaks when
Supervisor-edited customer or supplier repliesApproved final as sentOriginal draftThe edit changed substanceEdits only fixed typos, signatures or placeholders
Compliance review of outbound communicationsThe approved versionThe version returned with a findingThe finding cites a rule and policy versionIt was returned for a missing attachment
QA-scored responses to the same issue typeHigher-scored responseLower-scored responseSame issue, policy and rubric versionThe responses address different facts
Contract redlinesThe clause the company acceptedThe counterparty's proposed clauseTeaching your side's drafting standardRead as quality: the rejection reflects negotiating position
Competing proposals or quotesThe selected proposalAn unselected proposalSelection criteria were recordedPrice or relationship decided it, not the text

A lone approval or rejection fits a binary method instead. For rewriting models, see draft-to-final document pairs; similar record families appear in human feedback datasets of QA scores and corrections and training data for reward models.

De-identify both sides with the same mapping, so [SUPPLIER_1] names the same entity in chosen and rejected and the only difference left is content. Keep the policy version with each pair, because a final approved under last year's policy can be the wrong answer today.

SourceX sources operational datasets from US companies, including support and sales histories, documents, and finance and legal workflows, the systems where edits and approvals are typically recorded. Buyers describe the records they need, not the companies; SourceX looks for US businesses that hold them, and every release is approved by the supplying company. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match. To scope approval-derived pairs, describe the edit or approval signal you need.

Checks to run on a sample before you buy

Ask for a sample with full metadata and run these checks before signing: a dataset card describes intent, while a sample shows whether the pairs are usable. Fix your pilot evaluation prompts before the sample arrives.

  • Schema and prompt identity. Every record has a non-empty prompt, chosen and rejected; conversational records share identical history; no record has chosen equal to rejected.
  • Trivial differences. Flag pairs whose sides differ only in whitespace, signatures, placeholders or formatting.
  • Length skew. Compute the share of pairs in which the chosen response is longer. If nearly all are, length becomes a shortcut the model can learn in place of quality.
  • Duplicates and contamination. Find exact and near-duplicate prompts and any overlap with your evaluation prompts.
  • Agreement and ties. On a double-rated subset, compute Krippendorff's alpha, where 1 means perfect agreement and 0 means agreement no better than chance [9]. Ask whether raters could mark a tie and whether tied pairs were dropped or forced into an order; inter-annotator agreement metrics explains the choice of statistic.
  • Label noise. Score pairs with reward models and read those scored strongly against their label. The Secrets of RLHF Part II study reports that, under this kind of ensemble scoring, the mean preference difference was negative for about 25% of HH-RLHF pairs, suggesting those labels may be wrong [10]. Preference data quality and noise covers cleaning methods.
  • Chosen-response quality. Read 50 random chosen responses as if they were going to customers. DPO only teaches the model to prefer chosen over rejected, so a mediocre chosen answer sets a low ceiling.
  • Source mix. Tally response-generating models and label sources; a set dominated by one generator or judge passes on its habits.
  • Pilot run. Train on the sample or a slice and compare against your SFT baseline on the held-out prompts; evaluating a fine-tuning dataset before you buy it covers pilot design.

Rights and provenance questions specific to preference data

Preference data stacks three layers of rights: the prompts, the responses (often model outputs or employee writing), and the human judgments and rationales. A license that grants "the dataset" can leave one layer unaddressed.

  • Model-written responses. Record the generating model and version for every response, and check that provider's output terms as they stood on the generation date [8]. See due diligence for purchased synthetic data.
  • Rater work product. Rationales, rewritten chosen responses and rubrics can be copyrightable work product. Contracts should assign or license them and state whether raters could use AI assistance; provenance for human-annotated and preference data covers annotator agreements.
  • Personal data on both sides. In business-derived pairs, the rejected draft can hold exactly what a supervisor removed, such as a wrong account number or an over-disclosure, so it needs the same de-identification as the final.
  • Use scope. Confirm the license covers preference tuning, release of tuned weights and any hosted fine-tuning service; see fine-tuning-only data licenses.
  • Documentation. Ask for a datasheet-style card plus machine-readable metadata. Croissant-RAI, which extends the Croissant format, lists data labeling and participatory data among its design use cases [11].

For datasets SourceX sources, rights review checks that the business owns or may share the records, and names, emails, phone numbers and account numbers are removed or replaced before delivery, with the method recorded; no de-identification method is perfect.

A request template for DPO preference data

A usable request states the behavior to change, where responses and preferences come from, the label rules, the format and the licensed uses. Copy these rows into an RFP or data request.

FieldWhat to stateWhy it matters for DPO
Target behaviorThe failure to fix, with evaluation prompts held backPairs must disagree on it
Prompt sourceReal user or business inputs versus synthetic, and the task mix; see real-world prompt setsSets what the model learns to prefer
Response originSampled from your checkpoint, written by people, or generated by named models and versionsReference-model fit and provider terms
Preference sourceRaters, domain experts, an AI judge, or recorded business decisionsSets the noise level and rights chain
Rubric and tiesRubric version, tie option, strength scale, rationalesLets you drop ambiguous pairs
AgreementShare double-rated; minimum agreement set in the pilotMakes quality an acceptance test
VolumeUnique prompts and pairs per prompt, for a pilot and then full delivery; see cost drivers for human preference dataPairs from one ranking are correlated [4]
FormatJSONL, conversational messages, field names, group IDAvoids template lock-in and split leakage
PrivacyDe-identification method, same mapping on both sidesThe difference must be content
Licensed usesPreference tuning, weight release, hosted fine-tuning, termGrant matches plan

Describe the preference data your DPO run needs

Tell SourceX the behavior you want preference tuning to change, the records that capture the decision (edits, approvals or rejections with reasons), the volume and the uses you need licensed. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Submit your preference data requirements.

Sources

  1. Rafailov, Sharma, Mitchell, Ermon, Manning, Finn, Stanford (arXiv:2305.18290; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  2. ICML 2026 (poster page), "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
  3. jsonlines.org (community format specification), "JSON Lines". https://jsonlines.org/
  4. Ouyang et al., OpenAI (arXiv:2203.02155), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  5. Lambert et al., Allen Institute for AI (arXiv:2407.18421), "Self-Directed Synthetic Dialogues and Revisions Technical Report" (2024). https://arxiv.org/pdf/2407.18421
  6. Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  7. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  8. Anthropic Help Center (provider terms guidance), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  9. Klaus Krippendorff, University of Pennsylvania (technical report), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  10. Wu et al. (arXiv:2401.06080), "Secrets of RLHF Part II: Reward Modeling" (2024). https://arxiv.org/pdf/2401.06080
  11. Jain et al., MLCommons Croissant RAI task force (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data