Software companies
Customer data hidden in code, tests and fixtures: how to find it
By SourceX Editorial · Updated
Short answer
Customer data hides in test fixtures, seed scripts, recorded API responses, migrations and code comments in almost any codebase that grew by copying production records to reproduce bugs. Find it by searching each hotspot across the full git history, using your own customer list as the search dictionary, then replace it with synthetic data before any code is licensed.
Key takeaways
- The most effective search pattern is your own customer list: account names and email domains exported from the CRM and billing system.
- Search the full git history, not only the current branch, because a fixture deleted years ago still ships with every clone.
- Recorded API responses and seed scripts hold whole production records, while comments and branch names usually hold customer names.
- Automated PII detectors help but do not know which companies are your customers, so pair them with a dictionary search and engineer review.
- Production data found in a repository is a preparation task and may also be a security matter, so involve whoever owns incidents.
How does customer data end up in a codebase?#
Customer data usually enters a codebase through debugging. An engineer copies a failing production record into a test so the bug can be reproduced, the test passes, and the copied record stays in the repository long after anyone remembers where it came from.
Other routes are just as ordinary. Seed scripts get built from a production dump so demos look realistic, test libraries record real responses from partner APIs, migrations backfill named accounts, and log excerpts get pasted into comments. In a licensing context, each of these would hand the buyer a copy of your customers' records, which your customer contracts and privacy notice may not permit.
Hotspot checklist: where to look first#
Customer data clusters in a small number of predictable places, and checking them in order covers most of the risk quickly. The table lists each hotspot with the files it usually involves and a search approach that works there.
| Hotspot | Typical locations | What turns up | How to search |
|---|---|---|---|
| Test fixtures and factories | fixtures and testdata folders, JSON and YAML under test directories | Real names, emails, addresses and order payloads | Search for customer domains and account names from the CRM list |
| Seed and demo data | Seed scripts, demo-tenant setup, SQL and CSV files | Copied production rows; demo accounts named after real customers | Find data files by extension and flag unusually large ones |
| Recorded API responses | Cassettes, snapshots, mocks, saved HAR files | Full partner payloads, sometimes with tokens | Search for authorization headers, cookies and customer domains |
| Migrations and backfills | Migration folders and one-off scripts | Hardcoded account IDs and customer names in conditions or comments | Search for tenant IDs and production ID formats |
| Logs and notebooks | Committed .log files, test output, notebooks with saved cells | Stack traces and request bodies with personal details | Search by file type, then read outputs |
| Comments, commits and branches | Inline comments, TODOs, commit messages, branch names | Customer names and pasted ticket excerpts | Search commit messages and refs, not just files |
| Tenant config and feature flags | Per-customer config files and flag targeting rules | Customer names, account IDs and contract-specific behavior | Review any directory organized by customer |
Search patterns, and where automated scanners fit#
Search patterns work best in layers, starting with what is specific to your business and ending with generic detectors. Run each layer against every branch and the full history: git log --all -S followed by a string lists every commit that added or removed it, which catches fixtures deleted long ago.
Automated PII scanners can do part of the work, but not all of it. Presidio, an open-source MIT-licensed toolkit, combines named-entity recognition, regular expressions, rule-based logic and checksums to find personal details in text. Its own documentation warns that automated detection cannot guarantee it finds all sensitive information and that additional protections should be used.
Scanners also lack your business context. A detector can flag a person's name in a JSON fixture, but it cannot know that a company name in a migration comment is your largest customer. Secret scanners such as Gitleaks look for passwords, API keys and tokens, which matters because recorded API responses often contain both, but they do not look for customer records. Gitleaks' maintainer has also said the project is feature complete and will receive security patches only, so check a tool's status before standardizing on it. Use detectors as one layer between the dictionary search and engineer review.
Record every hit in one sheet with the repository, path, commit, pattern and decision. The same sheet later becomes the privacy record for this part of the package.
- Customer dictionary: export account names, legal names, short names and email domains from the CRM and billing system, and search for each.
- Email addresses outside test domains: match any address, then filter out documentation domains such as example.com and your own test domains.
- Production identifier formats: tenant IDs, account numbers, invoice numbers and order numbers that match live formats.
- Contact details: phone numbers, street addresses and postal codes inside fixture and seed files.
- Payment-like numbers: card-length digit strings, checked against a card checksum to cut false positives.
- Encoded payloads: base64 strings and compressed blobs in test folders, which hide records from plain-text search.
- Size outliers: any data file in a test or seed folder far larger than its neighbors.
What should you do with each finding?#
Each finding needs a decision about the delivery copy and, separately, about the working repository. Licensing preparation only requires a clean delivery copy, but production data sitting in a repository is worth fixing for its own sake.
Close the loop by rerunning the same dictionary and pattern searches against the finished delivery copy. Every remaining hit should be explained in the findings sheet, either as a false positive, such as a public company named in an open-source license file, or as a deliberate pseudonym. A delivery copy is ready only when the rerun produces nothing unexplained.
| Finding | Delivery copy | Working repository |
|---|---|---|
| Fixture copied from production | Replace with generated data that keeps the same shape and edge case | Replace going forward and note it in the security log |
| Seed script or demo tenant from a real customer | Exclude the script and its data files, including from history | Rebuild demo data from synthetic records |
| Recorded API response with real payloads | Exclude or re-record against a sandbox | Re-record and rotate any token found |
| Customer name in a comment, commit message or branch | Replace with the customer's consistent pseudonym | Usually leave as is |
| Hardcoded customer ID in a migration | Pseudonymize the ID and any name beside it | Leave; the migration has already run |
Illustrative: a freight software vendor cleans its test suite#
Illustrative: a fictional freight brokerage software vendor plans to license its code history. A dictionary search built from Salesforce accounts finds shipper names in migration comments and branch names. A pattern search finds recorded carrier API responses containing driver names and phone numbers, and a seed script that cloned a real customer's tenant for sales demos.
The CTO decides to replace the recorded responses with sandbox recordings, exclude the seed script and its data files from the delivery mirror's history, and pseudonymize customer names in comments and branch names. The licensed package keeps its full review history and carries no customer records, and the sales team stops using the cloned demo tenant.
How SourceX approaches customer data in code#
SourceX handles customer data found in code during the Preparation step of the SourceX five-step transaction, after the Rights step has agreed which repositories and record types are in scope. The supplier runs or approves each search, and nothing leaves the company during the initial fit check.
The patterns used, the hotspots checked and the treatment of each finding are written into the privacy record of the SourceX Evidence Packet, so the buyer sees how customer data was removed without seeing the data itself.
Frequently asked questions
Is masked production data in fixtures acceptable?
Sometimes, but treat it with suspicion. Masking applied by hand often misses fields such as free-text notes, addresses or IDs that can be traced back. If a fixture started as a production record, regenerating it from synthetic data is usually simpler than proving the masking was complete.
Do we need to rewrite history in our working repository?
Not for licensing. A separate delivery mirror with the affected paths removed is easier to verify and does not disrupt engineers. Rewrite the working repository only if your security team decides the data must be purged there too, which is a separate decision.
Does this matter if we only license a code snapshot?
Yes, but the search is smaller. A snapshot carries only the current files, so deleted fixtures are no longer a concern, while current fixtures, seeds and comments still are. Commit messages and branch names drop out because a snapshot carries no history.
What about images and binary files in the repository?
Screenshots, PDFs and sample documents in test folders can contain customer details that text search cannot read. List every binary file in test, seed and documentation folders, and exclude any that cannot be confirmed as synthetic or company-authored.
Who should run the searches?
An engineer who knows the codebase, working from a dictionary that operations or finance exports from the CRM and billing system. The engineer understands which folders matter, while the dictionary holder knows which names are real customers rather than test accounts.
Sources
- Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization that combines named-entity recognition, regular expressions, rule-based logic and checksums; its documentation warns there is no guarantee it will find all sensitive information and that additional protections should be employed. Source
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin; its README states it is feature complete, with future releases limited to security patches. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.