Fine-tuning and post-training data
Binary feedback data for KTO: using approvals, rejections and thumbs signals
Quick answer
Yes, unpaired binary labels can train a model. Kahneman-Tversky Optimization (KTO) needs only a prompt, one completion and a desirable or undesirable flag, not a chosen-versus-rejected pair for the same prompt, according to Ethayarajh and colleagues [7] (ICML 2024). That makes approvals, rejections, thumbs signals and QA pass/fail outcomes usable alignment data, provided you keep the exact text that was judged, set thresholds for implicit signals, check class balance, and audit whether the label reflects quality or just policy and workload.
By SourceX Editorial · Updated
What KTO actually consumes, and how it differs from DPO data
KTO consumes single judged outputs, while DPO consumes pairs. DPO fits the policy to a preferred and a dispreferred response for the same prompt [1], which is the comparison format behind classic RLHF pipelines: Llama 2 annotators picked the better of two outputs [2], and InstructGPT labelers ranked several outputs for the same prompt [3]. KTO instead uses a prospect-theory utility model and learns from whether each output is desirable or undesirable for its input.
The practical consequence is about collection cost. The KTO authors argue that binary signals are more abundant, cheaper and faster to collect than pairwise preferences, and they report KTO matching or exceeding preference-based methods at scales from 1B to 30B parameters, as of the ICML 2024 version. They also report KTO matching DPO while using up to 90% fewer desirable examples, which matters because real approval logs are rarely balanced. Treat those results as the paper's findings on its benchmarks, not as a guarantee for your domain.
If you already hold paired data, you do not need to choose. As of October 2026, Hugging Face TRL's KTOTrainer documentation says the trainer will split a chosen/rejected dataset into unpaired rows, labeling chosen completions true and rejected ones false. The reverse is not possible: unpaired logs cannot be turned into honest pairs unless two outputs were judged for the same prompt. For pair-specific sourcing, see our guide to preference datasets for DPO.
The KTO trainer dataset format, field by field
The minimum KTO record is three fields: prompt, completion and a boolean label. Per the TRL documentation, it accepts both the standard text format and a conversational format where prompt and completion are message lists, and it applies the chat template to conversational data automatically. Everything else in a production record exists for filtering, weighting and audit, and it should travel with the data even if the trainer ignores it.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"prompt": [
{"role": "system", "content": "You draft replies for a billing support queue."},
{"role": "user", "content": "I was charged twice for March. Can you fix it?"}
],
"completion": [
{"role": "assistant", "content": "I can see two charges dated March 3. I've opened a refund for the duplicate..."}
],
"label": true,
"signal_source": "agent_sent_without_edit",
"signal_strength": "implicit",
"edit_distance_ratio": 0.0,
"reviewer_role": "tier1_agent",
"reviewer_id_hash": "r_7f3a",
"rubric_version": "billing-qa-v4",
"policy_version": "refund-policy-2025-11",
"generator": "draft-model-2025-09",
"judged_at": "2025-12-02T14:11:09Z",
"queue": "billing",
"deidentification": "names, emails, account numbers replaced with typed tokens"
}
Three fields deserve emphasis. generator tells you whether the completion came from the model you are training, an older checkpoint or a human, which changes how on-policy the signal is. policy_version and rubric_version let you drop labels made under rules that no longer apply. signal_source is what lets you treat an explicit rejection differently from an inferred one.
Turning implicit approvals, edits and discards into labels
Implicit signals become usable labels only after you write explicit thresholds and accept that some rows will be excluded. A thumbs-down is fairly clear; a draft that an agent sent after rewriting half of it is not. The table below is a starting rule set you would tune against a manually reviewed sample.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Raw signal | Proposed label | Confidence | Notes |
|---|---|---|---|
| Explicit thumbs-up or "approve" click | true | High | Check for default-approve UI patterns |
| Sent or merged with no edit | true | Medium | Could reflect time pressure, not quality |
| Sent after light edit (e.g., under 10% character change) | true, or use edited text as a new true row | Medium | Keep both versions; the edit is a signal |
| Heavily rewritten before sending | false for original; edited version as true | Medium | Pair-like, so also useful for DPO |
| Explicit thumbs-down or "reject" with reason code | false | High | Keep the reason code |
| Discarded draft, no reason | exclude, or false at low weight | Low | Often abandoned for unrelated reasons |
| QA scorecard below pass threshold | false | Medium-high | Record rubric version and threshold |
| No interaction (timeout, closed tab) | exclude | None | Absence of a signal is not a rejection |
The edited-version rows are the most valuable output of this process. When a human rewrites a model draft, you get a desirable completion written for the same prompt plus a rejected draft, which is the closest operational data comes to a natural pair. Store the original and final text separately; a dataset that only keeps the final sent message has thrown away the judged content.
Class balance and weighting between desirable and undesirable rows
Class balance is a tuning variable in KTO, not a blocker. Approval logs in mature workflows are often heavily skewed toward accepted outputs, and the KTO paper reports that the method tolerates extreme imbalance. TRL's KTOTrainer exposes separate weights for desirable and undesirable examples so you can compensate for the ratio rather than discarding data, and its documentation discusses how to set them; check the current version of the docs for the recommended range.
TRL also notes that, in theory, a dataset should contain at least one desirable and one undesirable completion, while some users have trained on only one class; for rejected-only data it advises a conservative learning rate. In practice, report the ratio by queue, product surface and time window before training. A dataset that is 95% desirable overall can still be 60% undesirable in one queue, and that queue will dominate what the undesirable side teaches the model.
One more balance check is easy to skip: label balance per prompt type. If every refund request in the log was approved and every cancellation draft was rejected, the model may learn topic, not quality. Stratify by intent or ticket category and look for prompt types where one label is near 100%.
Label noise and bias: when approval does not mean quality
Binary labels inherit every bias of the process that produced them. Even deliberate pairwise preference labeling shows high disagreement; one analysis of RLHF reward modeling reports inter-annotator agreement typically in the 63-72% range and shows that incorrect and ambiguous labels distort reward models [4]. Benchmark test sets curated by researchers still carry measurable label errors, with one audit estimating at least 3.3% on average across ten widely used datasets [5]. Operational approvals, made under queue pressure with no rubric, should be assumed noisier until measured.
The specific failure modes to check:
- Workload bias: approval rates rise at shift end or during backlog spikes. Compare label rates by hour and queue depth.
- Policy bias: a response is rejected because it violated a policy that has since changed. Filter on
policy_version. - Reviewer drift: one reviewer rejects 40% while peers reject 10%. Normalize or reweight by reviewer, and see rater calibration for human-feedback data.
- Outcome leakage: a draft is marked bad because the customer later churned, not because the text was wrong. Exclude labels derived from downstream outcomes unless that is the target.
- Historical decision bias: in lending, claims or hiring workflows, approval may encode past discriminatory decisions. Our guide on auditing historical decision bias in operational labels covers the audit.
A practical control is a blind relabel: have two calibrated reviewers relabel a random sample of a few hundred rows against a written rubric and measure agreement with the logged signal. If agreement is low for a signal type, downweight it or drop it.
Sourcing binary feedback data you do not generate yourself
Most teams first use their own product logs, then look outside when they need domains their product does not cover. Before using thumbs data from your own users, confirm your terms and privacy notices actually allow training use; the FTC has warned that quietly changing terms to permit AI training may be unfair or deceptive [6]. For external data, the strongest candidates are companies whose existing workflows already produce judged outputs: QA scorecards on support replies, editorial approve/reject decisions, code review outcomes, and claims or document review decisions. SourceX's pages on human feedback datasets with QA scores and corrections and licensing QA scorecards describe that record family.
Illustrative example: invented to show structure; it does not describe an available dataset.
Buyer checklist for a licensed binary feedback dataset
- Is the judged text stored verbatim, including the original draft before edits?
- Who or what generated the completion (human, model and version)?
- What exactly does true mean (sent, passed QA, approved by supervisor)?
- Are rubric, policy and reviewer role recorded per row?
- What is the desirable/undesirable ratio by category and month?
- How were names, emails, phone numbers and account numbers removed or replaced, and was a sample checked?
- Does the license cover training and fine-tuning for your intended model, term and delivery?
- Has a relabel sample been run, and what was agreement with the logged signal?
For the general review process, see how to evaluate a fine-tuning dataset before you buy it, and for paired human judgments, buying RLHF comparison data. Internal sign-off typically involves legal, privacy and security; the internal approvals guide lists what each reviewer asks.
SourceX sources operational datasets from US companies on request, including support and sales histories and finance and legal workflows where approval decisions live. Datasets are not held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery; personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe the judged-output data you need on the buyer page.
Find approval and rejection data for KTO training
SourceX sources operational datasets from US companies on request and manages licensing through Find, Assess, Agree, Transact and Manage; nothing is contracted until a supplier agrees. Describe the judged outputs, labels and domains your post-training run needs, and the team will look for US businesses that hold that data. Start a buyer request at sourcex.si/buyers.
Frequently asked questions
Can I train KTO with only approved outputs?
TRL notes that some users have trained with only one class, but the expected setup includes both desirable and undesirable rows. Approved-only data is closer to supervised fine-tuning data; consider SFT on it first and reserve KTO for when you have rejections.
Should I convert my DPO pairs to KTO format?
Only if you want to mix them with unpaired data or test KTO. TRL performs the split automatically, but splitting discards the information that both responses answered the same prompt, which DPO uses directly [1].
How many binary labels do I need?
There is no fixed number; it depends on model size, domain and label noise. Our guide on how much data you need to fine-tune an LLM covers sizing, and the broader fine-tuning and post-training data hub lists related formats.
Sources
- Rafailov et al. (arXiv; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290
- Touvron et al., Meta (arXiv), "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
- Ouyang et al., OpenAI (arXiv), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- arXiv (Wang et al.), "Secrets of RLHF in Large Language Models Part II: Reward Modeling" (2024). https://arxiv.org/abs/2401.06080
- Northcutt, Athalye, Mueller (arXiv; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- arXiv (Ethayarajh et al.), "KTO: Model Alignment as Prospect Theory" (2024). https://arxiv.org/abs/2402.01306
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.