Skip to content

Software companies

Customer data hidden in code, tests and fixtures: how to find it

By SourceX Editorial · Updated

Short answer

Customer data hides in test fixtures, seed scripts, recorded API responses, migrations and code comments in almost any codebase that grew by copying production records to reproduce bugs. Find it by searching each hotspot across the full git history, using your own customer list as the search dictionary, then replace it with synthetic data before any code is licensed.

Key takeaways

  • The most effective search pattern is your own customer list: account names and email domains exported from the CRM and billing system.
  • Search the full git history, not only the current branch, because a fixture deleted years ago still ships with every clone.
  • Recorded API responses and seed scripts hold whole production records, while comments and branch names usually hold customer names.
  • Automated PII detectors help but do not know which companies are your customers, so pair them with a dictionary search and engineer review.
  • Production data found in a repository is a preparation task and may also be a security matter, so involve whoever owns incidents.

How does customer data end up in a codebase?#

Customer data usually enters a codebase through debugging. An engineer copies a failing production record into a test so the bug can be reproduced, the test passes, and the copied record stays in the repository long after anyone remembers where it came from.

Other routes are just as ordinary. Seed scripts get built from a production dump so demos look realistic, test libraries record real responses from partner APIs, migrations backfill named accounts, and log excerpts get pasted into comments. In a licensing context, each of these would hand the buyer a copy of your customers' records, which your customer contracts and privacy notice may not permit.

Hotspot checklist: where to look first#

Customer data clusters in a small number of predictable places, and checking them in order covers most of the risk quickly. The table lists each hotspot with the files it usually involves and a search approach that works there.

Hotspot checklist: where to look first
HotspotTypical locationsWhat turns upHow to search
Test fixtures and factoriesfixtures and testdata folders, JSON and YAML under test directoriesReal names, emails, addresses and order payloadsSearch for customer domains and account names from the CRM list
Seed and demo dataSeed scripts, demo-tenant setup, SQL and CSV filesCopied production rows; demo accounts named after real customersFind data files by extension and flag unusually large ones
Recorded API responsesCassettes, snapshots, mocks, saved HAR filesFull partner payloads, sometimes with tokensSearch for authorization headers, cookies and customer domains
Migrations and backfillsMigration folders and one-off scriptsHardcoded account IDs and customer names in conditions or commentsSearch for tenant IDs and production ID formats
Logs and notebooksCommitted .log files, test output, notebooks with saved cellsStack traces and request bodies with personal detailsSearch by file type, then read outputs
Comments, commits and branchesInline comments, TODOs, commit messages, branch namesCustomer names and pasted ticket excerptsSearch commit messages and refs, not just files
Tenant config and feature flagsPer-customer config files and flag targeting rulesCustomer names, account IDs and contract-specific behaviorReview any directory organized by customer

Search patterns, and where automated scanners fit#

Search patterns work best in layers, starting with what is specific to your business and ending with generic detectors. Run each layer against every branch and the full history: git log --all -S followed by a string lists every commit that added or removed it, which catches fixtures deleted long ago.

Automated PII scanners can do part of the work, but not all of it. Presidio, an open-source MIT-licensed toolkit, combines named-entity recognition, regular expressions, rule-based logic and checksums to find personal details in text. Its own documentation warns that automated detection cannot guarantee it finds all sensitive information and that additional protections should be used.

Scanners also lack your business context. A detector can flag a person's name in a JSON fixture, but it cannot know that a company name in a migration comment is your largest customer. Secret scanners such as Gitleaks look for passwords, API keys and tokens, which matters because recorded API responses often contain both, but they do not look for customer records. Gitleaks' maintainer has also said the project is feature complete and will receive security patches only, so check a tool's status before standardizing on it. Use detectors as one layer between the dictionary search and engineer review.

Record every hit in one sheet with the repository, path, commit, pattern and decision. The same sheet later becomes the privacy record for this part of the package.

  • Customer dictionary: export account names, legal names, short names and email domains from the CRM and billing system, and search for each.
  • Email addresses outside test domains: match any address, then filter out documentation domains such as example.com and your own test domains.
  • Production identifier formats: tenant IDs, account numbers, invoice numbers and order numbers that match live formats.
  • Contact details: phone numbers, street addresses and postal codes inside fixture and seed files.
  • Payment-like numbers: card-length digit strings, checked against a card checksum to cut false positives.
  • Encoded payloads: base64 strings and compressed blobs in test folders, which hide records from plain-text search.
  • Size outliers: any data file in a test or seed folder far larger than its neighbors.

What should you do with each finding?#

Each finding needs a decision about the delivery copy and, separately, about the working repository. Licensing preparation only requires a clean delivery copy, but production data sitting in a repository is worth fixing for its own sake.

Close the loop by rerunning the same dictionary and pattern searches against the finished delivery copy. Every remaining hit should be explained in the findings sheet, either as a false positive, such as a public company named in an open-source license file, or as a deliberate pseudonym. A delivery copy is ready only when the rerun produces nothing unexplained.

What should you do with each finding?
FindingDelivery copyWorking repository
Fixture copied from productionReplace with generated data that keeps the same shape and edge caseReplace going forward and note it in the security log
Seed script or demo tenant from a real customerExclude the script and its data files, including from historyRebuild demo data from synthetic records
Recorded API response with real payloadsExclude or re-record against a sandboxRe-record and rotate any token found
Customer name in a comment, commit message or branchReplace with the customer's consistent pseudonymUsually leave as is
Hardcoded customer ID in a migrationPseudonymize the ID and any name beside itLeave; the migration has already run

Illustrative: a freight software vendor cleans its test suite#

Illustrative: a fictional freight brokerage software vendor plans to license its code history. A dictionary search built from Salesforce accounts finds shipper names in migration comments and branch names. A pattern search finds recorded carrier API responses containing driver names and phone numbers, and a seed script that cloned a real customer's tenant for sales demos.

The CTO decides to replace the recorded responses with sandbox recordings, exclude the seed script and its data files from the delivery mirror's history, and pseudonymize customer names in comments and branch names. The licensed package keeps its full review history and carries no customer records, and the sales team stops using the cloned demo tenant.

How SourceX approaches customer data in code#

SourceX handles customer data found in code during the Preparation step of the SourceX five-step transaction, after the Rights step has agreed which repositories and record types are in scope. The supplier runs or approves each search, and nothing leaves the company during the initial fit check.

The patterns used, the hotspots checked and the treatment of each finding are written into the privacy record of the SourceX Evidence Packet, so the buyer sees how customer data was removed without seeing the data itself.

Frequently asked questions

Is masked production data in fixtures acceptable?

Sometimes, but treat it with suspicion. Masking applied by hand often misses fields such as free-text notes, addresses or IDs that can be traced back. If a fixture started as a production record, regenerating it from synthetic data is usually simpler than proving the masking was complete.

Do we need to rewrite history in our working repository?

Not for licensing. A separate delivery mirror with the affected paths removed is easier to verify and does not disrupt engineers. Rewrite the working repository only if your security team decides the data must be purged there too, which is a separate decision.

Does this matter if we only license a code snapshot?

Yes, but the search is smaller. A snapshot carries only the current files, so deleted fixtures are no longer a concern, while current fixtures, seeds and comments still are. Commit messages and branch names drop out because a snapshot carries no history.

What about images and binary files in the repository?

Screenshots, PDFs and sample documents in test folders can contain customer details that text search cannot read. List every binary file in test, seed and documentation folders, and exclude any that cannot be confirmed as synthetic or company-authored.

Who should run the searches?

An engineer who knows the codebase, working from a dictionary that operations or finance exports from the CRM and billing system. The engineer understands which folders matter, while the dictionary holder knows which names are real customers rather than test accounts.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization that combines named-entity recognition, regular expressions, rule-based logic and checksums; its documentation warns there is no guarantee it will find all sensitive information and that additional protections should be employed. Source
  • Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin; its README states it is feature complete, with future releases limited to security patches. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify