Skip to content

Software companies

How to remove secrets and credentials from code before licensing it

By SourceX Editorial · Updated

Short answer

To remove secrets from code before licensing it, scan the full git history rather than the current branch, then revoke and rotate every credential you find before cleaning anything. Rewriting history hides a secret but never makes it safe again. Build a separate delivery copy, rescan it, and keep the scan results with the release record.

Key takeaways

  • Scan every commit, branch and tag, because deleting a file in a later commit leaves the secret in history.
  • Rotate first and clean second: treat any credential that was ever committed as exposed.
  • A separate delivery mirror is easier to verify than an in-place rewrite of the repository your team works in.
  • Secrets also sit in issue trackers, chat exports, CI logs, test fixtures and repository config that often travel with a code package.
  • Keep scan output, rotation tickets, the final clean scan and an archive checksum as evidence for the release.

Why is deleting the file not enough?#

Deleting a file removes it from the latest commit only, while git keeps every earlier version. A database password committed in an old deploy script and removed later is still readable in that commit, in every clone, in old pull request branches and in any archive someone exported.

A licensed code package usually includes history, because commit messages, diffs and review threads are much of what makes engineering records useful for AI. That means the history itself has to be clean, not just the current tree. If history is out of scope, a snapshot of the current code is simpler, but it still needs scanning.

Where do secrets hide in a software company's records?#

Secrets collect wherever engineers needed something to work quickly, and the usual places are predictable. Start the review with the files and systems below, and add any place your own team knows it has cut corners.

  • Environment and config files: .env files, application settings and connection strings, including commented-out ones.
  • Infrastructure code: Terraform variables and state files, Kubernetes manifests, Helm values and Docker Compose files.
  • CI and deploy configuration: pipeline definitions, deploy scripts and build logs that echo variables.
  • Test fixtures and seed data: sample payloads, recorded API responses and copies of production rows.
  • Keys and certificates: private keys, signing certificates, SSH keys, Java keystores and mobile app signing files.
  • Notebooks and scripts: data analysis notebooks with saved outputs and one-off migration scripts.
  • Content outside normal history: Git LFS objects and submodules, which a plain history scan of the parent repository does not read.
  • Non-code records: Jira tickets, Slack threads and Confluence runbooks where someone pasted a token to unblock a colleague.

Step 1: scan the full history of every repository#

The first step is a history-wide scan of every repository in scope, using a mirror clone so that all branches and tags are included. A mirror clone from a Git host can also bring hidden pull request refs, which is useful for scanning but has to be decided on before delivery.

Run more than one detector, since each relies on different rules, and add custom patterns for your own token formats, internal hostnames and service account names. Scan the non-code exports that will ship with the package as well; issue trackers, chat exports and wiki pages are text files once exported.

Expect noise in the first results: example keys in documentation, dummy values in tests and tokens revoked long ago. Track every finding in one shared sheet with the repository, commit, file, secret type, whether it is live, an owner and the action taken. That sheet becomes the working record for the next three steps.

Step 1: scan the full history of every repository
Tool or methodWhat it doesWhere it fits
GitleaksOpen source detector for passwords, API keys and tokens in git history, directories and piped input; supports custom rules in a config file and a baseline so later scans show only new findings. Its README now says it is feature complete, with security patches onlyFast first pass across every repository and its history
TruffleHogOpen source scanner for Git and other sources such as chats, wikis, logs and object stores; for secret types it can classify, it can log in to confirm whether a secret is liveSorting live findings that need rotation first; live checks send real login attempts, so run them only from an approved machine
GitHub secret scanningScans the entire Git history on all branches, plus issue and pull request text; free on public repositories, while private organization repositories need GitHub Secret ProtectionOngoing alerts and a second opinion if your plan includes it
Custom rulesCompany-specific patterns for internal token formats, hostnames and customer identifiersCatching what generic rules miss
Engineer reviewManual reading of infrastructure, deploy and config pathsConfirming findings and clearing false positives

Step 2: revoke and rotate before you clean#

Rotation comes before any history rewrite, because a secret that was committed must be treated as exposed no matter how thoroughly it is removed later. Old clones, forks, CI caches and laptops may still hold it.

Work through findings in order of risk: live cloud and database credentials first, then third-party API tokens, signing keys, SSH keys and webhooks. Open a ticket for each, confirm the old value no longer works, and check access logs for use you cannot explain. If a credential belonged to a customer, for example an SFTP login stored in a test fixture, follow your incident process and the customer contract rather than handling it quietly.

Step 3: choose between rewriting and rebuilding#

Rebuilding a separate delivery mirror is usually safer than rewriting the repository your engineers use every day. An in-place rewrite changes commit IDs, breaks open pull requests and forces everyone to re-clone, all for a package that only needs to exist once.

Two widely used open source tools rewrite history, and each has a default that catches teams out. git filter-repo, which the git project now recommends over git filter-branch, expects to run in a fresh clone and stops otherwise; its replace-text option swaps listed strings across every commit, and path filters drop whole directories. BFG Repo-Cleaner, by default, does not modify the contents of the latest commit on your HEAD branch, so the current tip must be cleaned and committed first or the secret survives there.

Step 3: choose between rewriting and rebuilding
ApproachWhat it keepsEffort and riskBest when
In-place rewrite of the working repositoryFull history in the live repoDisrupts the team; forks and clones still hold old commitsThe secret must also be purged from daily use
Rewritten delivery mirrorFull history in a separate copyModerate; team workflow untouchedHistory is part of the licensed package
Current-tree snapshotLatest code onlyLowest; loses commit and review historyHistory is out of scope or too risky to clean

Step 4: verify the clean copy before release#

Verification means scanning the finished delivery copy with the same tools and rules, then explaining every remaining hit. A finding is closed only when it is removed, rotated or documented as a false positive by a named engineer.

Then check by hand. Search the delivery copy for the exact strings you rotated, open the highest-risk paths, confirm excluded directories are gone, and inspect the repository's own config, because a mirror can keep a remote URL with an embedded token. Record a SHA-256 checksum of the final archive so the reviewed package is provably the one delivered.

  • Every repository in scope was scanned as a mirror, with all branches, tags and any pull request refs that will ship.
  • Every finding is marked removed, rotated or documented as a false positive.
  • Rotated values no longer appear anywhere in the delivery copy, including LFS objects and submodules.
  • Remote URLs, credential helpers and hooks are stripped from the delivered repository config.
  • Exports from issue trackers, chat and wikis were scanned with the same rules.
  • The final scan output and archive checksum are saved with the release record.

Illustrative: a logistics software vendor cleans years of repositories#

Illustrative: a fictional transportation management software company plans to license its engineering history, covering GitHub repositories, pull request reviews and linked Jira issues. The CTO assigns the work to two senior engineers who know the deploy history and gives them the four steps above as a checklist.

The scans surface cloud keys in old deploy scripts, a Slack webhook in CI configuration and a carrier customer's SFTP password inside a test fixture. The team rotates every credential, notifies the carrier under its contract, and builds a release mirror. On that mirror they strip the fixture directory and replace the old keys across history, then rescan until nothing unexplained remains. One infrastructure repository is excluded outright because cleaning it would leave little worth licensing.

How SourceX treats credentials in code packages#

SourceX handles secrets removal in the Preparation step of the SourceX five-step transaction, alongside the removal of personal and confidential details. SourceX never hosts multi-terabyte archives: large repositories remain in the supplier's storage or move on encrypted drives, and the supplier reviews the cleaned mirror before the Approval and Delivery steps.

The scan approach, the rotation record and the final clean-scan result go into the privacy record of the SourceX Evidence Packet. Buyers see that the work was done and verified, never the secret values themselves.

Frequently asked questions

Do rotated secrets still need to be removed from history?

Yes. A dead credential can still reveal internal hostnames, naming conventions and vendor relationships, and buyers do not want credential strings in training data because models can memorize and reproduce text they were trained on. Rotation removes the security risk; removal keeps the package clean.

What about secrets in pull request comments and commit messages?

Commit messages live in git history and are covered by a history scan. Pull request comments and review threads usually live in the Git host's database, so they arrive in a separate export. Scan that export with the same rules before it joins the package.

Should the buyer be told that secrets were found?

The buyer should know that scanning, rotation and verification were done, but not what was found or where. A short description of the method and a clean final scan are normally enough, and they belong in the dataset documentation rather than in the code.

How do we handle internal hostnames and customer names that are not secrets?

Internal hostnames, IP addresses, customer names and employee emails are confidential or personal details rather than credentials. They are handled during de-identification with their own custom rules, and they need the same verification scan before release.

Can we automate the whole process?

Detection, rescanning and checksum recording automate well. Deciding what is a false positive, which repositories to exclude and whether a customer must be notified still needs an engineer and, for customer credentials, someone who knows the contract.

Sources

  • Gitleaks detects secrets such as passwords, API keys and tokens in git repos and other inputs; its README states the project is feature complete and future releases are security patches only. Source
  • TruffleHog scans Git, chats, wikis, logs, object stores and filesystems, and for each secret it can classify it can log in to confirm whether the secret is live. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify