Skip to content

Code and software engineering data

Commit Histories with Change Rationale: Messages, PR Descriptions and Design Records

Quick answer

A useful commit message dataset for explanation models pairs each diff with a recorded reason for the change, not just a one-line label. Public sets such as CommitBench and MCMD supply diff-message pairs from open repositories, but the "why" usually lives in pull request descriptions, linked tickets and architecture decision records inside private engineering systems. Buyers should specify informativeness filters, merge-strategy handling, ticket and design-record linkage, and redaction rules before licensing any commit history.

By SourceX Editorial · Updated

What public commit message datasets cover, and where they stop

Public datasets are good for learning message style and short summaries but thin on rationale. CommitBench was built because earlier sets had noisy, inconsistent data [1]; its release reports 1,664,590 examples from permissively licensed GitHub repositories in Java, Python, Go, JavaScript, PHP and Ruby, with bot commits removed [2]. Microsoft Research's MCMD study observed that most prior commit message datasets came only from Java repositories [3], and ComSum treats the message as a summary of the change [5].

HQCMD goes further by splitting about 300K messages into "what" and "why" parts [4]. That split is the right target, but in open-source repositories the "why" segment is often missing, and the design discussion that motivated a change sits in mailing lists or issues the dataset never joins. For models that must justify a change to a reviewer, you need the rationale layer that enterprise teams keep in their review and tracking tools.

The rationale layer: four record types to request

Change rationale in a company codebase is spread across four linked record types, and a dataset is only as good as the joins between them. Ask for each layer explicitly rather than assuming "full history" includes it.

  • Commit messages. Subject and body, author and committer timestamps, trailers such as Co-authored-by, Reviewed-by or Fixes:, and the parent SHAs so merges and reverts can be reconstructed.
  • Pull or merge request records. Title, description body, review comments with file and line anchors, approval state, and the merge method (merge commit, squash or rebase) from GitHub, GitLab or Bitbucket APIs.
  • Tracker tickets. Jira, Linear or Azure DevOps items referenced by keys such as PAY-1423, including issue type, acceptance criteria and status transitions.
  • Design records. Architecture decision records (ADRs) in the Context / Decision / Consequences format, RFCs and design docs, ideally with links to the PRs that implemented them.

Repository-level history is covered on our proprietary codebase datasets page, and broader engineering activity on software engineering histories. This page focuses on the explanation layer that sits on top of those histories.

Message informativeness filters and how to report them

Message quality varies widely, so the supplier should apply a stated informativeness filter and report what share of commits pass it. Messages such as "fix", "wip", "address comments" or "merge branch 'main'" teach a model nothing about intent, and in some internal repositories they can be a large share of commits.

Useful filters combine several signals:

  • Minimum token length for the body, with a separate rule for the subject line.
  • A stoplist of low-content subjects and auto-generated text (dependency bots, release tooling, merge commits).
  • Presence of a causal or purpose clause ("because", "so that", "to prevent") or a linked ticket or PR that supplies one.
  • Diff-to-message consistency, for example the message names a file, function or config key actually touched.

Ask the supplier to report pass rates per repository, not a single blended number. A pass rate that differs sharply between teams usually reflects team conventions such as Conventional Commits, and it tells you which repositories carry the rationale signal.

Squash merges, rebases and where the explanation actually lives

Merge strategy decides whether rationale sits in commits or in the pull request, so the dataset must record it per repository. With squash merges, granular messages are replaced by one commit whose message is often the PR title plus a concatenation of the original messages, while the real explanation is in the PR description. With rebase-and-merge, individual commits survive but the PR description is the only place the batch is explained.

Ask which strategies each repository used and whether that changed over time. If PR records are not included, a squash-merged repository may look informative at the commit level while losing the review discussion that justified the design. Also ask for pre-squash branch commits where the hosting platform retained them, since these often show how a change evolved under review.

Linking tickets and design records without leaking identities

Ticket keys and ADR references are the joins that turn a commit history into a rationale corpus, so they must stay resolvable after de-identification. If a supplier strips PAY-1423 from messages to remove project names, the link to acceptance criteria is gone.

A safer pattern is consistent pseudonymous IDs: replace each ticket key, author handle and email with a stable token applied identically across commits, PRs, tickets and design records. Check that author emails in commit metadata are replaced as well, since they are personal data embedded in every object. Note that rewriting history to scrub content changes commit hashes, and old commits can remain reachable through existing clones and forks [6], so the delivered dataset should carry its own hash mapping rather than depend on the original repository.

Design documents carry a different risk: they can name customers, pricing, roadmap items or security architecture. Expect suppliers to exclude some documents and redact others, and ask that every redaction be marked with a typed placeholder ([CUSTOMER], [ROADMAP_ITEM]) so models do not learn from silently truncated text. Our insight on whether licensing data exposes trade secrets covers how suppliers think about this, and secrets scanning across full history is covered in secrets removal for code datasets.

Illustrative change-rationale record

A delivered record should join diff, message, review and design context under one change ID so training code never has to reconstruct links. The schema below shows the fields worth requesting; JSON Lines with diffs in unified format is a common choice.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "change_id": "chg_7f31",
  "repo_token": "repo_b2",
  "language": "Go",
  "merge_method": "squash",
  "commit": {
    "sha_token": "c_91ae",
    "subject": "Retry ledger writes on serialization failure",
    "body": "Writes failed under concurrent refunds because ... Retry up to 3 times with jitter so that ...",
    "trailers": {"Reviewed-by": "dev_044"},
    "author_token": "dev_017",
    "authored_at": "2024-03-11T14:02:00Z"
  },
  "diff_unified": "--- a/ledger/write.go\n+++ b/ledger/write.go\n...",
  "pull_request": {
    "pr_token": "pr_5530",
    "title": "Handle serialization failures in ledger writes",
    "description": "Context: ... Alternatives considered: ...",
    "review_comments": [{"path": "ledger/write.go", "line": 88, "author_token": "dev_044", "body": "Why 3 retries?"}]
  },
  "tickets": [{"key_token": "TKT_1423", "type": "bug", "acceptance_criteria": "..."}],
  "design_records": [{"adr_token": "ADR_0021", "status": "accepted", "decision": "...", "redactions": ["[CUSTOMER]"]}],
  "quality": {"informative_message": true, "rationale_source": ["pr_description", "adr"]}
}

Using rationale data for SFT and design-reasoning evaluation

Rationale-rich histories support three training and evaluation tasks: commit and PR summarization, change justification, and design-decision reasoning. For SFT, the input is the diff plus surrounding code and the target is the PR description or the "why" part of the message, as in HQCMD's what/why split [4]. For justification tasks, include review comments that challenged the change, since the author's reply is often the clearest statement of intent.

For evaluation, hold out whole repositories rather than random commits, because adjacent commits share vocabulary and leak answers. Issue-to-PR pairing is also the structure behind SWE-bench, which matched 2,294 GitHub issues to resolving pull requests [7]; private equivalents are discussed in issue-to-fix pairs from private repositories and held-out coding evaluation sets. Before licensing, review a sample using the checks in evaluating a code dataset sample.

Buyer request checklist for change-rationale data

A precise request avoids rounds of back-and-forth, so state the rationale requirements alongside the usual language and history fields. The code dataset request specification guide covers languages, history depth and build requirements; add these items for rationale data.

Illustrative example: invented to show structure; it does not describe an available dataset.

RequirementWhat to specifyWhy it matters
Informativeness filterRules used and pass rate per repositorySeparates rationale signal from "wip" noise
Merge strategyMerge, squash or rebase per repo and periodTells you whether to rely on commits or PRs
PR recordsDescriptions, review comments, approval stateHolds the explanation for squash-merged work
Ticket linkagePseudonymous keys kept consistent across objectsKeeps acceptance criteria joinable
Design recordsADRs/RFCs with links to implementing PRsProvides decision-level rationale
RedactionTyped placeholders, exclusion logAvoids silent truncation and leaks
Hash mappingToken map for rewritten SHAsHistory rewrites change hashes [6]
ProvenanceOwnership and contributor agreementsSee chain of title for AI training data

For the wider landscape of code data types, start from the code and software engineering data hub or the main AI data guide. When you are ready to scope a request with these fields, you can describe the commit and design-record data you need.

Sourcing commit histories with recorded rationale

SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details such as names and emails are removed or replaced before delivery, and every release is approved by the supplying company. Describe the commit, PR, ticket and design-record data you need at sourcex.si/buyers.

Sources

  1. arXiv (Schall, Czinczoll, de Melo), "CommitBench: A Benchmark for Commit Message Generation" (2024). https://arxiv.org/pdf/2403.05188
  2. Zenodo, "CommitBench dataset record, version 1.0.0" (2024). https://zenodo.org/record/10497442
  3. Microsoft Research, "A large-scale empirical study of commit message generation: models, datasets and evaluation". https://www.microsoft.com/en-us/research/publication/a-large-scale-empirical-study-of-commit-message-generation-models-datasets-and-evaluation/
  4. IEEE DataPort, "HQCMD dataset (IEEE paper: enhancing commit message generation by segmenting key information)". https://ieee-dataport.org/documents/hqcmd-dataset-paper-ieef-enhancing-commit-message-generation-segmenting-key-information
  5. arXiv, "ComSum: Commit Messages Summarization and Meaning Preservation" (2021). https://arxiv.org/pdf/2108.10763
  6. GitHub Docs, "Removing sensitive data from a repository". https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/removing-sensitive-data-from-a-repository
  7. arXiv (Jimenez et al., ICLR 2024), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data