Code and software engineering data
Secrets in Code Datasets: Scanning Full Git History and Verifying Removal
Quick answer
Removing secrets from code training data means scanning every object a delivery contains, not just the current tree: all commits, branches, tags, LFS objects, CI logs, notebooks, and issue and review text. Detected credentials should be rotated by the supplier, replaced with typed placeholders that keep code valid, and logged in a findings report. The buyer then rescans a sample with an independent tool against a residual-findings threshold written into the license before accepting the delivery.
By SourceX Editorial · Updated
This page is the buyer's acceptance standard. Supplier-side preparation is covered in how to de-identify source code for AI training, and the transfer mechanics are in delivering code repositories with full history. For the wider category map, start at the code and software engineering data hub.
Why secrets in code data are a model problem, not only a repo problem
A secret that reaches training data can become a secret the model emits, and you cannot reliably delete it afterward. Carlini et al. showed that language models memorize and reproduce verbatim training sequences, including rare strings that occur in only a few documents [3]. API keys are exactly that kind of high-entropy, low-frequency string.
Research covered by Security Boulevard found that code assistants could reproduce real credentials from public code, including keys that had since been removed from repositories [2]. HashiCorp's guidance makes the same point from the pipeline side: once a model has learned a secret, removing it can be difficult or impossible, so scanning must happen before ingestion [1]. Doppler describes the same leakage path through training and code review workflows [5].
For a buyer, three distinct harms follow. The model can leak a third party's credential, your team inherits custody of live credentials it never asked for, and a downstream red-team finding can force retraining. Each is cheaper to prevent at acceptance than to remediate later.
What a complete scan must cover in a code delivery
A complete scan covers every reachable and unreachable object in the repository plus every non-code artifact shipped with it. Deleting a .env file in a later commit does not remove it from history; the blob stays in the object database and in every clone.
Scope the scan to these surfaces:
- Git objects: all commits on all branches (
git log --all), annotated and lightweight tags, stashes (refs/stash), notes, and any dangling objects left in packfiles beforegit gc. - Large-file storage: Git LFS objects are stored outside the packfile and are easy to miss. Model checkpoints, database dumps and
.pembundles often live there. - Configuration:
docker-compose.yml, Kubernetes manifests and Helm values, Terraform*.tfvarsand state files,application.properties,.npmrc,.pypirc,settings.xmland~/.aws/credentials-style files committed by mistake. - CI and build output: GitHub Actions, GitLab CI or Jenkins logs, where
set -x, verbose curl flags or failed masking print tokens in clear text. - Notebooks:
.ipynboutputs store printed environment variables, connection strings and HTTP responses inside JSON cells, separate from the code cells a reviewer reads. - Issue, pull request and review text: engineers paste stack traces, curl commands and connection strings into tickets and review comments.
- Embedded binaries and archives: zip, tar and jar files committed into the repo need to be unpacked and scanned recursively.
If the delivery includes review metadata or commit rationale, the text channels matter as much as the code. See commit histories with change rationale and review comment resolution data for how those fields are typically structured.
How detection should work: patterns, entropy and context
Reliable detection layers known-format patterns, entropy scoring and context rules, and the supplier should report exactly which detectors and versions ran. Gitleaks uses TOML-defined regex rules combined with entropy checks, while TruffleHog walks full git history and maintains a large set of credential-type detectors. Machine-learning classifiers can reduce false positives on generic high-entropy strings such as hashes and UUIDs [4].
Each layer has a known failure mode. Pattern rules catch AWS access key IDs, GitHub tokens, Stripe keys and PEM headers well but miss internal tokens with no fixed prefix. Entropy checks find random-looking strings but flag commit SHAs, base64 test fixtures and minified assets. Context rules (a variable named password, secret or api_key, a URL with user:pass@) catch low-entropy passwords like Summer2024! that the other two miss.
Ask for the configuration file itself, not a description of it. A custom .gitleaks.toml with broad allowlists can silently exempt whole directories such as test/ or fixtures/, and those directories often contain real credentials copied from staging.
Never test whether a found secret is live
Buyers should never check whether a discovered credential works, and the supplier should treat every detected secret as compromised and rotate it. Rewriting history alone does not make the data unreachable, since old commits survive in clones, forks and cached views, and research on code assistants shows why rotation is essential [2].
Some scanners, TruffleHog among them, can verify a finding by calling the provider's API. That is a supplier-side decision about the supplier's own credentials. On the buyer side, run independent scans with verification turned off (for TruffleHog, the --no-verification flag). Authenticating with a third party's key can amount to unauthorized access regardless of intent, and it creates log evidence that your organization used the credential.
How replacements should look: typed, consistent and syntactically valid
Good replacements are typed placeholders that keep the code parseable, keep the same value consistent across files, and are distinguishable from real credentials. Blank deletion breaks code: an empty string where a connection URL was expected changes control flow, and a removed line can break YAML indentation or a JSON object.
There are two common styles, and the choice affects training. Typed tokens such as <SECRET:AWS_SECRET_ACCESS_KEY:k07> are unambiguous, never re-trigger scanners, and let you filter or mask them during tokenization. Format-preserving fakes, such as a 40-character string with the right character set, keep format validators and tests passing but will match detectors on every future rescan and teach the model the real key shape.
Whichever style is used, the same original secret should map to the same placeholder everywhere it appears, so cross-file references in repository-level context stay coherent. See repository-level code context for why cross-file consistency matters, and the redaction glossary entry for the general term.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"finding_id": "F-000214",
"repo": "billing-service",
"surface": "git_history",
"path": "config/prod.env",
"original_commit": "9c1e4d2",
"rewritten_commit": "b77f0a3",
"line": 12,
"detector": "gitleaks",
"detector_version": "8.x (exact version recorded)",
"rule_id": "aws-access-token",
"secret_type": "AWS_ACCESS_KEY_ID",
"placeholder": "<SECRET:AWS_ACCESS_KEY_ID:k07>",
"occurrences_replaced": 3,
"rotation_confirmed_by_supplier": true,
"reviewer_disposition": "true_positive"
}
A findings log in this shape gives your security team a per-secret audit trail without exposing any original value. The original value itself must never appear in the log; record only its type, location and placeholder.
Rewriting history without breaking links
Rewriting history to remove secrets changes every downstream commit hash, so the delivery should include an old-to-new hash map. The rewrite also invalidates commit and tag signatures over the changed objects. git-filter-repo writes a commit map during a rewrite, which suppliers can ship as a two-column file.
Without that map, links from issues, pull request descriptions, CI run records and changelogs to specific commits stop resolving. That breaks datasets built on issue-to-fix pairs or test-outcome labels. Ask also that signed commits be flagged as re-signed or unsigned, so provenance checks do not misread the rewrite as tampering. The data dictionary template is a good place to document the hash map file and its columns.
Acceptance testing: an independent rescan with a written threshold
Acceptance should be an independent rescan of a random sample, run with a different tool than the supplier used, judged against a residual-findings threshold agreed in the license annex. Using the same tool and ruleset only confirms that the supplier's configuration agrees with itself.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Acceptance step | What the buyer does | Pass condition (example terms to negotiate) |
|---|---|---|
| 1. Scope check | Compare refs, LFS pointers and attached logs against the delivery manifest | Every ref and artifact listed in the manifest is present and was scanned |
| 2. Independent rescan | Run a second scanner, verification off, over a random sample of repos and all history | Zero confirmed true positives of high-severity types (cloud keys, private keys, database URLs) |
| 3. Low-severity review | Triage remaining hits manually | Confirmed residual findings at or below the annex threshold per sampled volume |
| 4. Placeholder audit | Grep for placeholder tokens; check parse and build on sampled files | Placeholders consistent across files; sampled files still parse |
| 5. Text channels | Scan notebooks, CI logs, issues and review comments separately | Same thresholds as code |
| 6. Hash map check | Resolve a sample of issue-linked commits through the map | Sampled links resolve to rewritten commits |
| 7. Findings log | Reconcile log counts with your own sample | Log covers every finding type you observed |
If the sample fails, the license should say what happens next: a supplier re-scan and re-delivery, replacement of affected records, or other remedies. See remedies when a data delivery fails and evaluating a code dataset sample before you license it. For ongoing supply, apply the same rescan to every refresh, since new commits bring new secrets.
Writing secrets handling into the request
The cheapest place to fix secrets handling is the dataset request, before any supplier starts preparing data. State the surfaces in scope, the placeholder style you want, whether a hash map is required, and the acceptance threshold. The code dataset request specification guide shows where these fields fit next to language, history and build requirements.
On SourceX's buyer page you describe the data you need, not the businesses that might hold it. SourceX looks for US companies holding operational data, including engineering records, and every release is approved by the supplying company. Every dataset is rights-reviewed, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked. No method is perfect, which is why your own acceptance rescan still matters. Diligence materials covering source, rights, preparation and allowed use are prepared for each dataset. Proprietary codebases are described on the proprietary code datasets page.
Request licensed code data and set your secrets acceptance terms
SourceX sources engineering records and other operational datasets from US companies on request; categories are not inventory, and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, through private, access-controlled workflows after an executed agreement. Describe the code data and acceptance requirements you need at sourcex.si/buyers.
Sources
- HashiCorp, "Integrating secret hygiene into AI and ML workflows". https://www.hashicorp.com/blog/integrating-secret-hygiene-into-ai-and-ml-workflows
- Security Boulevard, "Yes, GitHub's Copilot can leak (real) secrets" (2023). https://securityboulevard.com/2023/10/yes-githubs-copilot-can-leak-real-secrets
- Carlini et al., USENIX Security 2021 (arXiv), "Extracting Training Data from Large Language Models" (2021). https://arxiv.org/pdf/2012.07805
- arXiv, "Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms" (2024). https://arxiv.org/html/2401.01754v1
- Doppler, "AI secret leakage: LLM training and code review" (2026). https://www.doppler.com/blog/ai-secret-leakage-llm-training-code-review
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.