Skip to content

Schemas, packaging and delivery

Delivering Code Repositories with Full History: Git Bundles, LFS and Review Metadata

Quick answer

To transfer a git repository with full history, ask the supplier for a git bundle built with git bundle create repo.bundle --all (or a bare mirror clone), plus a separate export of Git LFS objects and a JSON export of pull requests, reviews and review comments from the hosting platform's API. Git alone carries commits, trees, blobs and refs; large files and code review discussion live elsewhere and are silently lost if the delivery specification does not name them. Verify every artifact on receipt against a manifest.

By SourceX Editorial · Updated

What a full-history repository delivery actually contains

A complete repository delivery has four layers, and only the first one travels inside git. Coding-agent and code-model teams usually need all four, because commit history without review context loses the reasoning, and history without LFS content leaves broken pointer files in place of binaries and fixtures.

  • Git object database and refs. Commits, trees, blobs, annotated tags, and every branch and tag ref. This is what git bundle or git clone --mirror captures.
  • Large File Storage content. In LFS-enabled repositories, git stores small pointer files; the actual objects sit on an LFS server and must be fetched separately.
  • Platform metadata. Pull requests, reviews, inline review comments, issues, labels and linked commits are objects on GitHub, GitLab, Bitbucket or Gerrit, reached through their APIs, not through git [2][3].
  • Delivery metadata. A manifest with checksums, ref lists, commit counts, the redaction log and any old-to-new hash mapping.

If you are still deciding which repositories and history depth to request, start with the code dataset request specification; this page covers the transfer mechanics once the scope is set. For the licensing side of proprietary code, see license source code for AI training.

Git bundle vs mirror clone: choosing the history container

A git bundle is the better default for a licensed, offline handover because it is one file that can be checksummed, encrypted and verified without network access to the supplier's git server. Git's bundle command supports git bundle create backup.bundle --all packaging all refs, and git bundle verify reporting whether the bundle records a complete history. git bundle list-heads prints the refs it contains, which you should compare against the supplier's ref list.

A mirror clone (git clone --mirror) produces a bare repository with every ref, which is convenient if the supplier pushes to a git remote you control. It is also many files rather than one, so manifest and checksum work is heavier. Either way, confirm three edge cases in writing:

  • Remote-tracking and special refs. Whether refs/remotes/*, refs/notes/*, refs/pull/* (GitHub's read-only PR refs) and refs/stash are in scope. Pull request head refs matter if you want unmerged or abandoned PR branches.
  • Shallow or partial clones. A bundle made from a shallow clone has prerequisites and will not record complete history; git bundle verify exposes this.
  • Submodules. Each submodule is a separate repository and needs its own bundle; the pinned commits are gitlink entries in the superproject trees, while .gitmodules records only each submodule's path and URL.

For transfer of the resulting files, follow your normal patterns: encrypt the delivery and land it in a cross-account cloud bucket rather than an email attachment.

Exporting Git LFS objects so binaries are not lost

LFS content must be exported explicitly, because a clone or bundle carries only pointer files. Per the Git LFS manual, git lfs fetch --all downloads objects referenced by any commit reachable from the given refs, or from all refs if none are given, and is intended mainly for backup and migration. It also ignores configured lfs.fetchinclude and lfs.fetchexclude paths so nothing is skipped by local filters.

A workable supplier sequence is: mirror-clone the repository, run git lfs fetch --all inside it, then package the lfs/objects directory as a separate archive alongside the bundle. Ask for an LFS manifest listing each object ID (the SHA-256 oid from the pointer file) and size. On your side, recompute SHA-256 for each object and confirm every oid referenced in history is present.

The common failure mode is an LFS server that has already garbage-collected or lost old objects. In that case the supplier should report missing oids explicitly rather than deliver pointers that resolve to nothing.

Exporting pull requests, reviews and review comments

Pull requests and review comments are not stored in git, so they require an export from the hosting platform's API. On GitHub, reviews are exposed through the pull request reviews endpoints [2] and inline comments through the review comments endpoints [3]. GitLab merge requests and discussions, Bitbucket pull request comments and Gerrit change messages have equivalent APIs with different field names, so ask for the raw API JSON rather than a flattened CSV.

The fields that make review data usable for training are the anchoring fields. GitHub review comments include commit_id, original_commit_id, path, diff_hunk, line and original_line [3]. Keep both the original and current anchors: a comment written against one commit can be outdated after a force-push, and original_commit_id plus diff_hunk is what lets you reconstruct the exact code the reviewer saw.

Also ask for threading (in_reply_to_id), review state (approved, changes requested, commented), timestamps, thread resolution status (on GitHub this comes from the GraphQL API's review threads, not the REST comment objects), and the PR's base and head SHAs and merge commit. SWE-bench showed the value of linking real issues to the pull requests that resolved them and checking edits with tests [4]; private issue-to-PR linkage is what makes issue-to-fix pairs and agent trajectories reconstructable. Licensing for review data specifically is covered on license code reviews and pull requests.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "repo_id": "repo-007",
  "pr_number": 1482,
  "base_sha": "9f2c1e0",
  "head_sha": "a71d3b4",
  "merge_commit_sha": "c0e98aa",
  "state": "merged",
  "reviews": [
    {"review_id": 55102, "state": "CHANGES_REQUESTED", "submitted_at": "2025-03-11T14:02:09Z", "author_ref": "dev_0193"}
  ],
  "review_comments": [
    {
      "comment_id": 880341,
      "in_reply_to_id": null,
      "path": "billing/invoice.py",
      "commit_id": "a71d3b4",
      "original_commit_id": "4be0f12",
      "original_line": 212,
      "line": 218,
      "diff_hunk": "@@ -205,7 +205,9 @@ def apply_credit(...)",
      "body": "This double-counts credits when the invoice is reissued.",
      "author_ref": "dev_0412"
    }
  ],
  "linked_issues": ["ISSUE-3310"],
  "hash_map_applied": true
}

Note the pseudonymous author_ref values: reviewer and committer identities are personal data, and a consistent pseudonym preserves who-reviewed-whom structure without names or emails.

Redaction, rewritten history and the commit hash map

Any history rewrite to remove secrets or personal data changes commit hashes, so the delivery must include a mapping from old to new SHAs. GitHub's documentation states that rewriting history with git-filter-repo changes commit hashes, and that old commits can remain reachable through existing clones, forks and other channels [1]. For a buyer, the practical consequence is that every SHA in exported PR metadata, issue links and CI logs will point to commits that no longer exist in the delivered bundle.

Ask for a commit-map file (git-filter-repo writes one in its metadata directory) with old and new SHA pairs, and require the supplier to rewrite SHA fields in the API export or ship the map so you can join them yourself. Also confirm how signed commits and tags were handled, since rewriting invalidates GPG and SSH signatures. Scanning and verifying secret removal is a separate discipline covered in secrets in code datasets.

Receipt verification checklist for repository deliveries

Run verification before anything enters a training pipeline, and record results against the manifest. This builds on the general approach in dataset manifests and checksums.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckCommand or methodPass condition
Bundle integritysha256sum repo.bundle vs manifestHash matches
Complete historygit bundle verify repo.bundleNo missing prerequisites reported
Ref coveragegit bundle list-heads repo.bundle vs supplier ref listEvery agreed branch, tag and PR ref present
Object integritygit clone repo.bundle work && git -C work fsck --fullNo missing or corrupt objects
LFS completenessExtract all pointer oids from history; compare to LFS archiveZero missing oids, SHA-256 recomputed
SubmodulesCheck each superproject gitlink pin resolves in its submodule bundleAll pinned commits present
Metadata joinsJoin review commit_id/original_commit_id to bundleEvery SHA resolves, or is in the hash map
Hash mapCount old SHAs mapped vs commits rewrittenMap covers all rewritten commits
CountsCommits, PRs, comments vs manifestWithin agreed tolerance

Packaging repository metadata for code datasets

Repository-level metadata belongs in a sidecar file per repository, not scattered across READMEs. Useful fields include primary languages, build system, test command, default branch, first and last commit dates, commit count, contributor count (pseudonymous), license or ownership note, LFS usage, submodules, and whether history was rewritten. These fields let you filter for repository-level code context work and track provenance through dataset versioning.

For ongoing purchases, decide whether later drops are incremental bundles (built with a revision range such as last-delivered-tag..main, which carry prerequisites) or full re-bundles. Incremental bundles are smaller but fail verification if your copy lacks the prerequisite commits; see incremental deliveries vs full refreshes. More delivery patterns are collected in the dataset delivery hub and the wider AI data guides.

How SourceX handles proprietary code repository requests

SourceX sources operational datasets from US companies, including engineering records, and manages the commercial process, from licensing agreements to ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names and emails are removed or replaced before delivery with the method recorded, and delivery runs through private, access-controlled workflows only after an executed agreement. Teams can describe the repositories they need on the SourceX buyer page, and the proprietary code datasets page covers the data itself.

Request proprietary repositories with full history

If your coding-agent or code-model work needs private repositories with history, reviews and build context, describe the data rather than the companies you have in mind. SourceX looks for US businesses that hold it, reviews rights, and agrees allowed uses in a license before anything is transacted. Describe your code data request.

Sources

  1. GitHub Docs, "Removing sensitive data from a repository". https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/removing-sensitive-data-from-a-repository
  2. GitHub Docs, "REST API endpoints for pull request reviews". https://docs.github.com/en/rest/pulls/reviews
  3. GitHub Docs (Enterprise Server 3.22), "REST API endpoints for pull request review comments". https://docs.github.com/enterprise-server@3.22/rest/pulls/comments
  4. Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data