Software companies
How to find and remove secrets from git history before sharing a repo
By SourceX Editorial · Updated
Short answer
To remove secrets from git history before sharing a repo, work in four steps: scan every branch, tag and ref; rotate each live credential; rewrite a separate export copy using placeholders; then rescan a fresh clone. Rotation is the step that actually protects you. Rewriting keeps the shared copy clean.
Key takeaways
- Scan a mirror clone so every branch, tag and hidden ref is included, not just the default branch.
- Treat any credential that ever reached a commit as exposed, and rotate it before rewriting anything.
- Rewrite a separate export copy, not the repository your engineers push to every day.
- Replace secrets and sensitive strings with placeholders so commit history and diffs keep their meaning.
- Close the job with a rescan of a fresh clone and a written record of every finding.
Why do secrets survive in git history after you delete them?#
Secrets survive in git history because git stores every committed version of every file as an object, and a later commit that deletes a key only adds a new snapshot on top. Anyone who clones the repository receives the old objects too and can read them with ordinary git commands.
Sharing a repository outside the company widens that exposure. An acquirer's diligence team, an AI developer reviewing a code license or an outside contractor receives the full object store, including stale branches, old tags and anything pushed by mistake. Hosted platforms may also keep pull request references and forks that a normal clone does not show.
The four steps that follow assume the repository will leave your control: scan, rotate, rewrite and rescan. Skipping rotation is the most common and most costly mistake.
What should you scrub besides passwords and API keys?#
A repository prepared for outside eyes needs more than secret removal, because commit history also records customer names, personal details and code you may not own. This broader scrub is what separates a share-ready repository from one that merely passes a secret scanner.
Decide up front which items you will remove and which you will keep. A reviewer judging code for training value usually wants commit messages and authorship structure intact, so pseudonymize rather than delete wherever the content itself is not the problem.
- Credentials: API keys, tokens, private keys, database passwords, cloud access keys and signing certificates.
- Customer identifiers: customer names in branch names, commit messages, config files and feature flags.
- Personal data: real names, emails and phone numbers in test fixtures, seed files and committed logs.
- Data dumps: database exports, CSV extracts and support attachments committed while debugging.
- Internal infrastructure: hostnames, IP ranges and internal URLs that map your network.
- Third-party code: vendored SDKs or libraries whose license does not allow redistribution.
- Author metadata: personal email addresses in commit authorship.
Which tools find and remove secrets?#
Secret scanners and history rewriters do different jobs, and a share-ready process uses at least one of each. Scanners find candidate secrets across commits; rewriters produce a new history without them.
Running two scanners catches more than one, because each relies on different detection rules. TruffleHog's live verification sends real authentication attempts, so use it with care and only against credentials your company controls.
| Tool | Job | Notes |
|---|---|---|
| Gitleaks | Scans git repositories, files and input streams for secrets | MIT license; its default configuration had 222 detection rules as of July 2026; since May 2026 its README says it is feature complete, with future releases limited to security patches |
| TruffleHog | Scans Git and other sources such as chats, wikis, logs, object stores and filesystems | AGPL-3.0; says it classifies over 800 secret types and can log in to check whether a found secret is still live |
| git filter-repo | Rewrites history to remove paths or replace text across all commits | Suited to detailed, rule-based rewrites; run it on a fresh clone |
| BFG Repo-Cleaner | Removes files and replaces listed strings across history | Simpler for common removals, less flexible for complex rules |
| Your Git host's secret scanning | Flags secrets in pushed code and may block new ones | Features depend on your plan; most useful for prevention after the cleanup |
Steps 1 and 2: scan every ref, then rotate#
The scan step starts from a mirror clone, which copies every branch, tag and reference the host exposes. Run your scanners against the full history, export the results, and build a findings register with the secret type, file, first and last commit, and a named owner.
The rotation step follows immediately. Every live credential in the register is revoked and replaced in its source system, whether a cloud console, a payment processor or a SaaS admin page, and the owner records the date. Rotation comes before any rewrite because old clones, CI caches and laptops may already hold copies.
- Mirror-clone each repository in scope, including archived ones.
- Run two scanners across all history and merge the results.
- Triage each finding as live, already revoked, test value or false positive.
- Revoke and rotate every live finding, then confirm the old value fails.
- Search Jira, Confluence and Slack for the same values, since secrets travel.
Step 3: rewrite a copy without losing the history's value#
The rewrite step produces a new history for an export copy, and its goal is to remove sensitive content while keeping the commits, messages and diffs that make history worth sharing. Replacing a key with a placeholder such as REDACTED_API_KEY keeps the diff readable; deleting whole commits erases the story of how a bug was fixed.
Rewrite a fresh clone and treat it as a separate export repository. Commit IDs change during a rewrite, which breaks links from tickets and code reviews, so record a map of old to new commit IDs for internal use. Keep the replacement rules file out of the export, because it contains the secrets themselves. Rewriting also invalidates any signed commits, so tell reviewers that signatures in the export will not verify.
Author emails deserve their own pass. Rewriting authors with a mailmap file passed to your history rewriter turns personal addresses into stable pseudonyms, so a reviewer can still see that one engineer wrote a series of fixes without learning who that engineer is. Do not ship the mailmap file itself, because it lists the original addresses.
Step 4: rescan a fresh clone and keep the evidence#
The rescan step clones the rewritten export from scratch and runs the same scanners with the same rules. Every remaining hit is either fixed or documented as a false positive or rotated test value, with a reason.
Keep an evidence set beside the export: scan reports from before and after, the rotation log, the list of removed paths and replaced string categories, and the commit map. Reviewers ask for this record, and it protects the team if a question surfaces later.
Automated scanners miss things, especially personal data in free text. A short manual review of commit messages, test fixtures and files with unusual extensions closes much of that gap.
Illustrative: a restaurant inventory software company prepares a repo for review#
Illustrative: a fictional restaurant inventory software company is asked to share its main application repository for a code licensing evaluation. The repository holds a decade of commits, merged pull requests and links to Jira issues.
The first scan finds old cloud keys in a deploy script, a payment sandbox token, and restaurant group names in branch names and fixture files. The CTO's team rotates the keys, confirms the sandbox token was already dead, and rewrites an export copy that swaps names for tokens such as CUSTOMER_017 and pseudonymizes author emails. A vendored barcode SDK is removed because its license prohibits redistribution.
The rescan of a fresh clone is clean apart from documented test values. The evaluation team receives a repository whose commit history, review structure and issue references still make sense, along with the evidence set.
How SourceX handles code before delivery#
SourceX treats secret and sensitive-content removal as part of the Preparation step in the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Nothing is shared during the initial fit check, which relies on metadata such as repository counts, languages and years of history.
Scan results, the rotation log and the redaction rules feed the privacy record in the SourceX Evidence Packet, and the supplier approves the final export before release. Large repositories stay in the supplier's own storage or ship on encrypted drives.
Frequently asked questions
If every secret is rotated, do we still need to rewrite history?
Rotation removes the security risk, but dead credentials still reveal naming patterns, vendors and internal structure, and reviewers flag them as findings. For an external share, most teams remove them from the export copy anyway. For an internal mirror, a documented rotation may be enough.
Does deleting a repository on our Git host remove every copy?
Not necessarily. Forks, existing clones, CI caches, backups and mirrors can outlive the original repository. Check your host's documentation on forks and cached views, and assume any copy that left your control still holds the old history. That is why rotation matters more than deletion.
How do we handle Git LFS files and submodules?
Treat each as its own source. Large files stored outside the main object store need their own scan and rewrite, and a submodule points to another repository with its own history. If submodules are in scope, run the same four steps on them; if not, remove the reference from the export.
Should we rewrite the history our engineers use every day?
Usually not for an external share. Rewriting the working repository changes every commit ID, breaks open pull requests and forces every engineer to re-clone. A separate export copy gives you a clean deliverable without disrupting work, while rotation and prevention controls fix the working repository.
What should we tell an outside reviewer about the cleanup?
Share the method, never the secrets: which tools and rule sets ran, which categories of content were replaced, how authors were pseudonymized, and what the final rescan found. A short written summary with the evidence set answers most diligence questions and shows the export was prepared deliberately.
Sources
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin; its default configuration contained 222 detection rules as of July 2026. Source
- On May 21, 2026, the gitleaks README was updated to say Gitleaks is feature complete and future releases will be security patches only. Source
- TruffleHog, an AGPL-3.0 secret scanner, says it classifies over 800 secret types, can log in to confirm whether a secret is live, and scans Git, chats, wikis, logs, object stores and filesystems. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.