Software companies
Code snapshot vs full git history: what to include in a code license
By SourceX Editorial · Updated
Short answer
For most code licenses, curated git history is more useful than a bare snapshot: history linked to tickets and reviews shows AI developers how and why code changed, while a snapshot shows only the code at one moment. Full history multiplies review effort and the risk of exposing old secrets, so filter it by repository, path and date.
Key takeaways
- A snapshot shows what the code is; history shows how and why it changed, which is what coding agents learn from.
- Secrets and personal data removed from the current branch can still sit in older commits.
- Pull request reviews and issue links live on the hosting platform, not in a git clone, and need their own export.
- Curated history lets you include strong repositories in full and limit weaker ones to a snapshot.
- Files deleted years ago, such as old fixtures and configs, still sit in history, so every exclusion must apply to every commit.
What is the difference between a snapshot and full history?#
A code snapshot is the file tree at a single commit, usually the head of the main branch or a release tag. Full git history is every commit that led there: the diffs, commit messages, authors, timestamps, branches and tags. A snapshot answers what the code looks like. History answers how it got that way.
Neither format includes everything engineers think of as the code's story. Pull request descriptions, review comments, CI results and links to Jira or Linear issues live on GitHub, GitLab or the issue tracker, so a package built from git alone misses the reasoning around each change.
Packaging differs too. A snapshot can ship as a plain archive of files. History ships as a bare git repository or a git bundle, which preserves commits so the buyer can replay changes, usually alongside a separate file of platform exports.
How do snapshot, full history and curated history compare?#
Snapshot, full history and curated history differ most in what they teach a model and in how much review the supplier must do before release. The comparison below assumes a mature product repository with several years of commits.
| Factor | Snapshot | Full history | Curated history |
|---|---|---|---|
| What a buyer learns | Structure, conventions and finished code | How changes were made, reverted and refined | Change patterns from the repos and periods that matter |
| Usefulness for coding agents | Limited; no before-and-after pairs | High, especially with linked issues and reviews | High where linkage is strongest |
| Review effort | Lowest | Highest; every commit is in scope | Moderate; scoped by repo, path and date |
| Secret exposure | Only what is in current files | Anything ever committed | Reduced by date and path filters, still needs a scan |
| Personal data | Few author details | Author names and emails on every commit | Pseudonymized authors |
| Third-party code | Current vendored code only | Every library ever vendored | Excluded paths applied across all commits |
Why history linked to tickets matters#
History linked to tickets matters because it captures a complete unit of engineering work: a reported problem, the discussion, the change and the review that approved it. A commit message that cites a Jira key, a pull request that explains the approach, and a reviewer who asks for a missing test together show a model how real teams fix real defects.
That is why AI developers building coding assistants and agents look for before-and-after pairs with the reason attached. A snapshot cannot provide those pairs at all. Full history can, but only if the issue keys survive export and the review threads are pulled from the platform API.
Linkage quality varies by team and period. Repositories where engineers consistently referenced issue keys are worth far more review effort than repositories with messages like fix and wip.
Where history creates risk#
History creates risk because git remembers what the current branch has forgotten. An API key deleted from a config file still sits in the commit that added it. A test fixture with real customer records, removed years ago, is still in the object store. Commit messages may name customers, and author fields carry personal email addresses.
That changes how preparation works. A snapshot is reviewed once, file by file. History has to be filtered across every commit, branch and tag that will be delivered, and each exclusion has to reach the whole history, not just the latest version. Secret scanners such as Gitleaks, an MIT-licensed tool for finding passwords, API keys and tokens in git repositories, can cover the credential part. They do not look for customer names in commit messages or real records in old fixtures, which need their own searches and a sampled human read.
Any credential found in history is rotated before delivery, and the wider security review of what the code reveals applies to snapshots and history alike. Choosing history mainly raises the volume of material those controls must cover, and it adds the history-only risks listed below.
- Credentials committed and later deleted, which stay valid until rotated.
- Test fixtures or seed files built from real customer exports.
- Commit messages and branch names that name customers, incidents or employees.
- Author names and personal email addresses on every commit.
- Vendored libraries and large binaries removed from the tree but not from history.
How to curate a history package#
Curating a history package means deciding scope first and preparing second, so review effort lands where the value is.
- Rank repositories by product relevance and by how consistently commits reference issues.
- Set a start date for each repository, often after a major rewrite or migration.
- Exclude paths across all commits: vendored dependencies, third-party code, generated files, credentials and infrastructure configuration.
- Run a secret scan over the full history of every included repository and rotate whatever it finds.
- Replace author names and emails with stable pseudonyms so the sequence of work by one person stays visible.
- Export pull request descriptions, review comments and issue links through the platform API, keeping original IDs.
- Produce a manifest listing repositories, date ranges, excluded paths and commit counts.
Illustrative: a construction software company chooses its scope#
Illustrative: a fictional construction project management software company holds a legacy repository from its first product and a monorepo created during a platform rewrite. The legacy history is messy, with vendored libraries, large binary files and commit messages that rarely mention an issue.
The CTO offers the legacy code as a snapshot of its final release, with vendored folders removed. The monorepo is offered as full history from the rewrite onward, because most of its commits reference Jira keys and its pull requests carry substantive review threads.
The credential scan over the monorepo history comes back clean, but a search for customer names finds a seed file built from a real customer export in an early commit and a run of commit messages naming a general contractor client. History is rewritten so the seed file disappears from every commit, those messages are replaced with placeholders in the delivered copy, and the manifest records both changes. The package is smaller than everything the company holds, but every part of it is reviewed and linked.
Which format fits which repository?#
The right format depends on the repository, not the company. Most suppliers end up with a mix.
Decide per repository, then check the combination. If a legacy snapshot and a curated history overlap in time or share code, apply the same exclusion rules to both, so a file removed from one is not delivered through the other.
| Repository situation | Format to consider |
|---|---|
| Active product repo with consistent issue keys and reviews | Full or curated history with platform exports |
| Legacy repo replaced by a rewrite | Snapshot of the last release |
| Repo with heavy vendored or customer-specific code | Snapshot of first-party paths only, or exclude |
| Infrastructure, security or authentication repos | Usually exclude |
| Discontinued product with good history | Curated history, often easier to clear than live code |
How SourceX scopes code packages#
SourceX starts with a metadata inventory in the Supply step of the SourceX five-step transaction: repository names, date ranges, commit counts and whether issue keys and reviews are present. No code is shared during the fit check. The SourceX Enterprise Data Value Framework weights linked history, where a change connects to its reason and review, above standalone files.
Preparation covers secret scanning, rotation, path exclusions and pseudonymization, and the results are recorded in the SourceX Evidence Packet. Large repositories stay in the supplier's own storage or ship on encrypted drives at Delivery.
Frequently asked questions
Can we rewrite git history to remove secrets before delivery?
Yes. Tools that rewrite history can strip files or strings from every commit, and the rewritten repository becomes the delivered copy. Rewriting changes commit hashes, so issue links that cite hashes may break. Rewriting is a packaging step, not a security fix: rotate the secrets as well.
Should unmerged branches be included?
Usually not. Abandoned branches often hold experiments, half-finished work and occasionally debugging code with hard-coded credentials. Some buyers value failed approaches, but the review cost is high. Start with the main branch and merged pull requests, and add branches only on a specific request.
Should commit author identities be kept?
Replace them with consistent pseudonyms. The sequence of work by one engineer can be useful, but names and personal email addresses add privacy exposure and rarely add value. Keep the mapping table inside the company, not in the package.
How do we export pull request reviews if they are not in git?
Through the hosting platform's API. GitHub and GitLab both expose pull or merge requests, review comments and linked issues through their APIs. Export them with original IDs and timestamps so they can be joined back to the commits in the package.
Is a snapshot ever the better choice for a buyer?
Sometimes. A buyer evaluating code quality, conventions or a particular domain may only need the current code. Snapshots are also faster to clear, so a snapshot can be a sensible first package while curated history is prepared. It also lets the buyer judge fit before both sides commit to a larger history package.
Sources
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin. Source
Related resources
- InsightRecords written with AI assistance: do they lose value for licensing?
- IndustryBPO & contact centers data
- QuestionDo AI labs buy code?
- QuestionDo AI labs buy medical data?
- InsightCan you license CAD and engineering drawings to AI companies?
- InsightCan you license code reviews and pull requests to AI companies?
See if your company qualifies
A short company assessment. No data uploads are needed.