Skip to content

Code and software engineering data

Code Review Comment Resolution Data for AI Code Reviewers

Quick answer

A code review comment dataset for training AI reviewers should pair each line-anchored human comment with the code it was written against, the commit that addressed it (or evidence it was ignored), and a category label. Raw comment dumps teach a model to talk; resolution pairs teach it which comments change code. Specify anchoring, alignment confidence, bot filtering and negatives before you license anything, and check samples against those fields first.

By SourceX Editorial · Updated

What a comment-resolution pair actually contains

A usable pair is a five-part unit: the original diff hunk, the anchored comment, the thread, the addressing change and a resolution label. Public datasets already frame the unit this way, pairing an inline review comment with the author's subsequent code change [2]. Comment generation needs the hunk and the comment. Code refinement needs the hunk, the comment and the revised code, which is exactly the part most comment dumps lack.

This page covers that aligned unit. For broader licensing of reviews and pull requests, see the owner page on licensing code reviews and pull requests for AI training; for accept/reject signals used as reward data, see code preference and reward data from review outcomes.

Anchoring: why the original diff hunk must travel with the comment

Line numbers are not stable anchors, so every comment must ship with the commit and hunk it was originally written on. On GitHub, pull request review comments attach to a portion of the unified diff and carry diff_hunk, commit_id, original_commit_id, position, original_position, line, original_line, start_line and side. When later commits touch the commented lines, commit_id moves while original_commit_id still points at the commit the reviewer actually saw, which is why an export should keep both.

The common failure mode is an export that keeps only path and line from the final PR state. The comment then appears next to code that already contains the fix, which silently turns a defect comment into nonsense and poisons refinement targets. Gerrit, GitLab and Bitbucket exports have their own patch-set or version fields; ask the supplier to map them to the same three facts: file, original revision and original hunk text.

Finding the addressing commit, and reporting how sure you are

The addressing commit is usually the next commit in the PR that modifies the anchored lines, confirmed by thread resolution, an author reply or a reviewer approval. Line overlap alone over-matches, because a later commit can touch the same lines for an unrelated reason. Resolution flags alone under-match, because many teams resolve threads without a code change ("won't fix", "done in follow-up PR").

Ask suppliers for an alignment method and a per-row confidence value, plus the share of comments aligned at each confidence level. Research on Chromium reviews modeled which feedback authors acted on from more than a million comments [1], which suggests the "was it acted on" signal is learnable but noisy. Expect a meaningful unaligned remainder and keep it labeled rather than dropped.

Categories, bots and ignored comments

Category labels let you balance the mix, because unlabeled corpora are dominated by nits and questions. A workable taxonomy is defect, security, design, readability, test, question and nit. Curation research proposes filtering review comments by type, nature and civility, and reports that comment quality filtering matters for downstream usefulness [4]. Security-focused sets such as SeRe show the value of a narrow, well-aligned slice over a large undifferentiated one [3].

Bot and linter comments (Snyk, SonarQube, CI annotations, formatter suggestions) should be removed or tagged with an author_type field. They inflate volume and teach templated phrasing rather than human judgment. Ignored or rejected comments are not waste: rows where the author declined, with the reason, are negatives that help a reviewer model stop flagging style preferences as defects. At least one public set already keeps negative examples next to comment-to-change pairs [2].

Record schema to request

Ask for one record per comment thread root, delivered as JSON Lines (UTF-8, one JSON object per line, no blank lines) [5] or Parquet, with code text stored verbatim and secrets scanned across full history.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"thread_id": "t-000412", "repo_id": "r-17", "pr_id": "pr-2291", "language": "go",
 "path": "billing/retry.go", "original_commit": "a91c...e2", "original_start_line": 84, "original_line": 91,
 "diff_hunk": "@@ -80,9 +80,14 @@ func Retry(...", "comment": "Backoff never caps; a 429 storm will spin here.",
 "author_type": "human", "reviewer_role": "maintainer", "category": "defect",
 "replies": [{"by": "author", "text": "Good catch, capping at 30s."}],
 "resolution": "addressed", "addressing_commit": "c07f...19", "alignment_method": "line_overlap+reply",
 "alignment_confidence": 0.92, "code_after": "...", "merged": true}

Allowed resolution values might be addressed, addressed_elsewhere, declined_with_reason, ignored and unknown. Keep original_* fields separate from any current-position fields, and record PII and secret handling per field.

Buyer checklist for review comment datasets

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to ask forRed flag
Anchoringoriginal_commit plus verbatim diff_hunk per commentOnly final-state path and line
AlignmentMethod, per-row confidence, share alignedEvery comment marked "addressed"
NegativesDeclined and ignored rows with reasonsUnaligned rows silently dropped
Botsauthor_type or documented removal ruleCI annotations mixed into human comments
CategoriesTaxonomy definition and label distributionLabels from an undisclosed model only
ThreadsFull reply chain and reviewer roleRoot comments without replies
Code textSecrets scan over full historyScan on HEAD only
RightsWho owns the code and who can license itContributor or client code of unclear origin
ContaminationOverlap check against public repos and benchmarksRepos mirrored on public hosts

Run the checklist on a sample using the steps in evaluating a code dataset sample before you license it. Secrets deserve their own pass, as described in scanning full git history and verifying secrets removal, because review threads often quote the credential being flagged.

Rights, people and contamination in review data

Review data contains more than code: reviewer names, handles, emails in commit metadata and sometimes customer details pasted into threads. Pseudonymize reviewer and author identities consistently, so that a role or tenure feature still works without exposing who wrote what. Code ownership questions apply to every repository in scope, and contractor or client code is the usual gap; see ownership due diligence for licensed codebases.

Private review history from a company's internal Git host is less likely to overlap public evaluation sets than GitHub-sourced sets, but forks and mirrors still leak. Check overlap before you use any of it for reviewer-model evaluation, using the methods in detecting code benchmark contamination. To see where this dataset sits among other code data types, start from the code and software engineering data buyer's map, and use the guide on how to specify a code dataset request to write languages, history depth and build requirements into the brief. For background on why review threads carry signal beyond the code itself, read what makes code reviews and pull requests valuable for AI.

How SourceX sources review comment data

SourceX sources operational datasets, including engineering records such as code review histories, from US companies, on request rather than from stock, and a request does not guarantee a match. You describe the data you need, such as languages, review platform, comment-to-commit alignment and negatives, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names and emails are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement. SourceX does not train models and does not source scraped web content. You can start a request on the SourceX buyers page, and the code review workflow overview explains what these records show.

Request code review comment resolution data

SourceX sources engineering records from US companies on request and manages the license that defines records, uses, term and delivery. Describe the anchoring, alignment and labeling you need, and nothing is contracted until a supplier agrees. Start at sourcex.si/buyers.

Frequently asked questions

Is a public code review dataset enough to train a production reviewer?

Public sets are useful for pre-training and baselines, but most come from open-source GitHub projects, carry those projects' licenses and may overlap public benchmarks. Private enterprise review history adds internal conventions, longer-lived services and domain-specific defects that open projects underrepresent.

Should we keep comments that were never addressed?

Yes, if they are labeled. Ignored and declined comments are the main source of signal for teaching a reviewer what not to say, and dropping them biases the model toward flagging everything.

What file format works best?

JSON Lines with one thread per line is easy to stream and diff, while Parquet suits large columnar analysis. Either works if code fields are stored verbatim with original encoding and line endings preserved.

Sources

  1. ACL 2018, "A dataset for identifying actionable feedback in collaborative software development" (2018). https://2018.aclweb.org/paper/220/
  2. Hugging Face (ronantakizawa), "github-codereview dataset card". https://huggingface.co/datasets/ronantakizawa/github-codereview/blob/main/README.md
  3. arXiv, "SeRe: A Security-Related Code Review Dataset Aligned with Real-World Review Activities" (2026). https://arxiv.org/pdf/2601.01042
  4. arXiv, "Balancing Usefulness and Naturalness: An LLM-based Curation Pipeline for Code Review Comments" (2026). https://arxiv.org/pdf/2607.09524
  5. jsonlines.org, "JSON Lines". https://jsonlines.org/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data