Skip to content

Code and software engineering data

Code Preference and Reward Data from Review Outcomes

Quick answer

A code preference dataset built from review outcomes pairs two versions of a change written for the same task, with a label saying which one the engineering organization accepted. The strongest signals come from real workflows: the first revision of a pull request versus the merged revision, approved versus "changes requested" reviews, abandoned PRs with a recorded reason, and reverts or hotfixes linked to incidents. Each pair must share the same base commit and task, and confounds such as business-driven abandonment must be filtered before the pairs train a reward model or DPO run.

By SourceX Editorial · Updated

Why review outcomes complement synthetic code preferences

Review outcomes carry preferences that real maintainers acted on, which synthetic judgments may not fully reproduce. Much public code preference work relies on generated data: CodeFavor, for example, trains on synthetic Commit-Instruct and Critic-Evol pairs and evaluates on CodePrefBench, 1,364 tasks spanning correctness, efficiency, security and human preference [2]. Recent reward-model research has also drawn a large share of training samples from GitHub commits that were later merged through pull requests, treating the merge as the preference signal [3].

Private review histories go further than public merges. They include the rejected side: the revision a senior reviewer blocked, the PR closed for a security concern, the change reverted after it paged on-call. Expert studies also show that preferences on readability, modularity and commenting are partly subjective [4], so labels from the organization that maintains the code can reflect its own standards more closely than generic crowd judgments; aligning models to human code preferences is an active research area [5]. For the license and data-type framing, see the owner page on licensing code reviews and pull requests and the definition of preference data.

Five review signals and what each one actually labels

Each review signal labels a different property, so treat them as separate label sources rather than one "accepted vs rejected" flag. The table below maps signals to the field evidence you should require.

Illustrative example: invented to show structure; it does not describe an available dataset.

SignalChosen sideRejected sideEvidence fields to requireMain confound
First vs merged revisionMerged head commitFirst pushed revisionPR id, base SHA, revision SHAs, review comments between themRevisions that only rebase or fix merge conflicts
Approval vs changes requestedRevision approvedRevision with CHANGES_REQUESTED review stateReview state, reviewer id (pseudonymized), timestampStyle nits weighted like defects
Rejected or abandoned PRAlternative merged fix for the same issueClosed, unmerged PRLinked issue id, close reason, closing commentAbandoned for priority, staffing or product reasons
RevertCode after the revert or follow-up fixReverted changeRevert commit, reverted SHA, reason textReverts of good code for release timing or feature flags
Incident hotfixHotfix changeCausing changePostmortem or incident id linked to the causing SHAMisattributed root cause

The first-versus-merged signal is the most plentiful and the least noisy when review comments explain each revision. Revert and incident signals are scarce but valuable for teaching a reward model about defects that passed review once. Hold them out as a separate slice so you can measure whether your model learns failure modes rather than style.

Controlling confounds before you build pairs

A preference label is only meaningful when the rejection was caused by the code itself. PRs are frequently closed because a roadmap changed, an owner left or a different team shipped first. Reverts often happen during release freezes or because a feature flag was removed, not because the code was wrong.

Ask suppliers to export reason codes where their tools record them, such as close-reason labels, revert commit messages and incident tags. Where no reason exists, filter by evidence of a defect: a failing CI run on the rejected revision, a review comment citing a bug, a linked bug ticket, or a test added in the merged revision that fails on the earlier one. Executable evidence, where the build and test environment can be reproduced, is the strongest filter and turns a soft preference into a verifiable one. Requirements for build and test reproducibility belong in your code dataset request specification.

Keeping both sides of a pair on the same task

Pairs must share the same task context, or the model learns to prefer the task rather than the solution. In practice that means the same linked issue, the same base commit and the same target files, with any rebase noise stripped out. Two PRs against different base SHAs can differ for reasons that have nothing to do with quality.

Store the prompt side as the issue text plus the repository state at the base SHA, and the response sides as diffs against that state. For agentic or repository-level reward models, include the cross-file context each side touched; the page on repository-level code context covers what to request. Deduplicate pairs across forks and vendored copies before splitting train and evaluation sets, as described in deduplicating code training data.

Weighting labels by reviewer and agreement

Not every approval carries the same information, so carry reviewer metadata forward as weights rather than discarding it. Useful fields include a pseudonymized reviewer id, the reviewer's code-owner status for the touched paths, tenure band and the number of independent approvals. A change approved by two code owners after requested changes is a stronger label than a single rubber-stamp approval on a Friday afternoon.

Where several reviewers disagree, keep the disagreement as a soft label or margin. DPO optimizes directly on chosen and rejected pairs without training a separate reward model [1], so noisy pairs flow straight into the policy; curation and pair quality strongly shape preference optimization results [6]. The same agreement data lets you calibrate an LLM judge against the organization's own reviewers before you use it to label more pairs.

A record layout that serves reward models, DPO and judges

One pair record should support reward-model training, DPO and judge calibration without re-export. A useful layout is JSON Lines with one pair per line.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "p-000412",
  "repo_ref": "repo-17",
  "language": "java",
  "base_sha": "a91f3c2",
  "task": {"issue_ref": "ISSUE-2210", "issue_text": "Retry logic double-charges on timeout"},
  "chosen": {"revision_sha": "e04b7d1", "diff": "...", "source": "merged_head"},
  "rejected": {"revision_sha": "c55a019", "diff": "...", "source": "first_revision"},
  "signal": "first_vs_merged",
  "evidence": {"ci_rejected": "fail", "ci_chosen": "pass", "review_comments": ["Idempotency key missing on retry path"]},
  "reviewers": [{"id": "r-hash-3f2", "code_owner": true, "state": "CHANGES_REQUESTED"}],
  "confound_flags": [],
  "deidentification": {"method": "pseudonymized author and reviewer ids; secrets scan on diffs"}
}

Request that the evidence and confound_flags fields are filled before delivery, not inferred later. Diffs and review text frequently contain credentials, internal hostnames and customer identifiers, so require a full-history secrets scan as covered in secrets in code datasets.

Rights and diligence questions specific to review data

Review data mixes code, comments written by employees and sometimes contractor or open source contributions, so ownership has to be checked on each layer. Confirm that the supplier owns the repositories, that contributor agreements cover employee and contractor comments, and that vendored third-party code is identified; code ownership due diligence lists the documents to ask for. Copyleft code merged into a private repository can travel into your pairs, which the page on copyleft contamination addresses.

SourceX prepares diligence materials on source, rights, preparation and allowed use for each dataset. It sources operational datasets, including engineering records, from US companies and manages the licensing process; every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names and emails are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Buyers describe the data they need on the SourceX buyer page, not the businesses that might hold it.

How to evaluate a sample before licensing

A small sample should prove that pairs are task-aligned, labeled with evidence and free of leakage. Check at least these points:

  • Every pair shares a base SHA and issue reference, and diffs apply cleanly.
  • The share of pairs with defect evidence (failing CI, bug-citing comment, linked bug) is reported per signal.
  • Close and revert reasons are present or explicitly marked missing.
  • Reviewer fields are pseudonymized consistently across the sample.
  • No pair overlaps public repositories used in your benchmarks; see code benchmark contamination.

The broader checklist is in evaluating a code dataset sample, and the code data buyer's map places this page among other code datasets. Post-training teams sourcing SFT and RL data alongside preferences can start from the post-training teams guide and human feedback and QA score datasets.

Request code preference data from real review histories

SourceX sources engineering records from US companies on request; nothing is held in stock and a request does not guarantee a match. The process runs Find, Assess, Agree, Transact and Manage, every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe the review signals, languages and evidence you need on the SourceX buyer page.

Sources

  1. arXiv (Rafailov et al., Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  2. arXiv, "Learning Code Preference via Synthetic Evolution" (2024). https://arxiv.org/pdf/2410.03837
  3. arXiv, "Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring" (2026). https://arxiv.org/pdf/2605.00754
  4. arXiv, "Subjective Code Preferences in Experts and Large Language Models" (2026). https://arxiv.org/pdf/2605.25296
  5. arXiv, "Learning to Align Human Code Preferences" (2025). https://arxiv.org/pdf/2507.20109
  6. arXiv, "When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets" (2025). https://arxiv.org/pdf/2511.10985

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data