Skip to content

Privacy and preparation

Production data in test fixtures and seed files: personal data in your codebase

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Production data in test fixtures and seed files is real customer or employee information copied from live systems into a repository, usually to reproduce a bug. It hides in fixtures, seeds, snapshots, notebooks and SQL dumps, and survives deletion in git history. Scan the full history, replace findings with synthetic records, and exclude what cannot be cleaned.

Key takeaways

  • Deleting a fixture file removes it from the current code, not from git history, forks, CI caches or developer clones.
  • Secret scanners look for credentials; finding customer names, emails and addresses needs a separate personal data scan.
  • Your own customer list and employee directory make the most reliable search dictionary for real records.
  • Replace real values with synthetic ones that keep the same format, so tests still pass after cleanup.
  • When a repository's history cannot be cleaned, license a cleaned snapshot or leave that repository out.

How do production records end up in a codebase?#

Production records end up in a codebase through ordinary engineering shortcuts. A developer copies a failing customer payload into a test to reproduce a bug, a staging database is seeded from an old production dump, or a snapshot test captures a real API response and commits it as the expected output.

None of these steps feels like a data transfer at the time. The file is small, the repository is private and the bug gets fixed. Years later the same file sits in many commits, has been cloned by every engineer who joined since, and travels with any copy of the repository handed to an auditor, an acquirer or a data licensee.

The pattern is common in B2B software because enterprise bugs are often data-shaped. A tax calculation fails only for one customer's address format, or an import breaks on one client's spreadsheet, so the fastest fix is to test against the real thing.

Where it hides: five file families to check first#

Personal data in a repository clusters in five file families, and each needs a slightly different search. Start with these before scanning the rest of the tree.

Also check migration files that backfill data, code comments that paste an example record, and bug reports stored inside the repository. Commit metadata counts too: author names and email addresses identify employees and contractors.

Where it hides: five file families to check first
File familyTypical locationsWhat real data looks like there
Test fixturesFixtures folders, factory files with hardcoded values, YAML and JSON test dataCustomer names, emails and street addresses kept verbatim from a bug report
Seed filesDatabase seed scripts, seed SQL, demo data loadersA trimmed production export used to make staging look realistic
Snapshots and recordingsSnapshot test files, recorded HTTP cassettes, mocked webhook payloadsFull API responses with contact details, account IDs and internal notes
NotebooksJupyter notebooks in analytics, data science or support tooling foldersSaved output cells showing query results, often whole tables of customer rows
Dumps and exportsSQL dumps, CSV and spreadsheet exports, backup files, large files stored through Git LFSComplete tables of users, orders or tickets, sometimes years of records

Why deleting the file does not remove the data#

Deleting a fixture removes it from the current version of the code, but every earlier commit still contains it. Anyone who clones the repository with its history can check out an old commit and read the file, and the same is true of forks, mirrors, CI build caches and container images built from older commits.

Deletion matters for licensing because code datasets are often valued for their history. Commit sequences, code reviews and fixes linked to issues are what make engineering records useful to AI teams, so a buyer usually wants history, and history is exactly where old fixtures live.

That leaves three options for each repository: rewrite history to purge the files, deliver a prepared copy whose history has been cleaned outside the working repository, or leave the repository out of scope. Rewriting the history your team works from breaks every existing clone, so preparing a separate cleaned copy is usually the calmer route.

A scan-and-replace plan for one repository#

A scan-and-replace plan works best against a mirror of the repository with full history, never against the copy engineers push to. The steps below cover one repository; repeat them for each one in scope, including archived repositories.

Keep the two kinds of scanner distinct. Credential tools, gitleaks and TruffleHog among them, hunt for passwords, keys and tokens. The gitleaks maintainer has declared that tool finished apart from security fixes, and TruffleHog can attempt a login to test whether a found secret still works, so run that check with care. Neither is built to find a customer's home address.

Personal data detectors miss things as well. Presidio's own documentation warns that there is no guarantee it will find all sensitive information and that additional protections should be used, which is why the dictionary search and a human review of flagged files sit alongside it.

  • Inventory: list every file in every commit and flag the five file families by path and extension.
  • Dictionary search: search all history for your customer domains, account names and employee directory, which catches real records generic detectors miss.
  • Pattern scan: run a personal data detector over text files, notebook outputs and decoded fixtures for emails, phone numbers, addresses and names.
  • Secret scan: run a credential scanner as a separate pass, and rotate any live secret before removing it.
  • Triage: sort each finding as real, synthetic or public, and record who decided.
  • Replace: swap real values for synthetic ones with the same format, so field lengths, validation rules and test assertions still hold.
  • Verify: run the test suite on the cleaned copy, rescan, and record the results in the redaction log.

What to do with each finding#

Each finding needs a recorded decision, and most fall into a handful of patterns. A short decision table keeps reviewers consistent across repositories and gives counsel a clear record.

When the same real record appears in several places, replace it with the same synthetic value everywhere. Tests often join fixtures by email or account ID, and inconsistent replacements create failures that look like new bugs.

What to do with each finding
FindingActionNote for the log
Real customer contact in a fixtureReplace with a synthetic value on a reserved example domainFile, field type and replacement rule
Production dump committed to historyExclude the file from every commit in the prepared copyAsk counsel whether the past exposure needs review under your incident process
Notebook with saved query outputClear output cells and keep the codeOutputs stripped, code unchanged
Live password, key or tokenRotate first, then remove from the prepared copyRotation date and owner
Published business data, such as an office addressKeep if it carries no personal detailReason it was kept
Commit author names and emailsReplace with stable pseudonymous author IDs if identities are not neededMapping owner; mapping stays internal

Illustrative: a scheduling software company prepares two repositories#

Illustrative: a fictional company sells crew scheduling software to landscaping and snow removal contractors. It runs a Rails monolith and a newer Node service on GitHub, tracks work in Jira, and is preparing its engineering history for a licensing review.

The dictionary search, built from CRM customer domains and the HR directory, finds real crew members' phone numbers in recorded SMS webhook payloads and in a seed file built from an old staging dump. The data team's notebooks hold saved output cells with customer names.

The company keeps the Node service with full history after replacing its fixtures, because its findings were few and easy to fix. The monolith's early history is dense with seed dumps, so the company delivers a cleaned snapshot of recent history and excludes the oldest period. Engineering also adds a pre-commit scan so new fixtures are checked before they land.

How SourceX approaches personal data in code#

SourceX treats code as one record family within the Preparation step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Scoping starts with metadata such as repository names, languages, years of history and known sensitive areas, so no code is shared during the initial assessment.

During preparation, the scans, replacement rules, exclusions and test results become the privacy record in the SourceX Evidence Packet, alongside provenance, licensing rights, permitted use and release authorization. The supplier reviews that record and approves the prepared copy before anything is delivered.

Frequently asked questions

Do fixtures that look synthetic still need review?

Yes. Engineers often edit a real record slightly, such as changing a first name but keeping the real email and street address. Check fixtures that look fake against your customer and employee lists before marking them synthetic, and spot-check a sample by hand.

Is a production dump in a private repository a breach?

Not automatically. Whether it counts as an incident depends on who could access it, what it contained and which laws or contracts apply. Treat it as a finding for your security and legal leads to assess under your incident process, and remove it from any copy that leaves the company.

How do we stop production data from coming back?

Make the safe path the easy one. Provide fake-data factories for each core object, seed staging from a masked copy, add a personal data check to pre-commit hooks and CI, and include a fixture question in the pull request template. Give one person ownership of the policy.

Can engineers use AI tools to generate synthetic fixtures?

They can, with care. Ask the tool to generate data from a schema rather than from a pasted production record, and confirm your AI tool's terms do not let the vendor retain or train on prompts. Review generated fixtures before committing them.

Do buyers need commit author identities?

Often not. Many code datasets work well with stable pseudonymous author IDs, which preserve who reviewed whose code without naming anyone. If a buyer asks for identities, treat it as a scoping question for counsel, since commit metadata identifies employees and contractors.

Sources

  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source
  • On May 21, 2026, the gitleaks README was updated to state that Gitleaks is feature complete and future releases will be security patches only. Source
  • TruffleHog says that for every secret it can classify, it can log in to confirm whether the secret is live. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify