Skip to content

Code and software engineering data

Deduplicating Code Training Data: Forks, Vendored Copies and Near-Duplicates

Quick answer

Code deduplication for LLM training removes exact and near-duplicate files before tokenization so a model does not over-weight repeated code, memorize it, or score well on benchmarks it has effectively seen. Code duplicates differently from prose: forks, vendored dependencies, copied snippets, generated files and templated boilerplate dominate. The working recipe is to normalize files, hash for exact matches, run MinHash with LSH on token shingles, group results by repository, and then measure what a licensed delivery adds beyond your existing public mix.

By SourceX Editorial · Updated

Why duplication hurts code models more than the headline numbers suggest

Duplication distorts both training and evaluation, and in code the evaluation damage is often larger. Lee et al. found heavy repetition in common natural-language corpora, including one sentence repeated more than 60,000 times in C4, and reported that over 1% of unprompted model output was copied verbatim from training data; deduplication reduced memorized output and train-test overlap [1]. Their scope was natural-language text, so treat the size of the effect as a direction, not a code-specific estimate.

For code, the evaluation risk is sharper because whole files are copied verbatim across forks, mirrors and vendored directories. The failure mode is simple: a function in your test split also lives in a fork or a vendored copy in your training split. Your completion or repair benchmark then measures recall, not generalization. For the evaluation side of this problem, see code benchmark contamination from public repos, forks and mirrors.

Where code duplicates come from

Most duplicate code comes from a handful of structural sources, and each needs a different control. Knowing the source tells you whether to drop, down-weight or keep a copy.

  • Forks and mirrors. A popular repository can exist in hundreds of forks with zero or a few commits of divergence. Hash-level dedup catches the unchanged files; the lightly patched ones become near-duplicates.
  • Vendored dependencies. Directories such as vendor/, third_party/, node_modules/, and external/, plus checked-in third-party source trees, pull whole open-source projects into private repositories. In a licensed delivery these files are usually already in your public mix, and they also carry their own upstream licenses (see copyleft contamination in licensed code).
  • Copied snippets. Utility functions, Stack Overflow answers and internal "common" helpers are pasted across services. These are sub-file duplicates, so file-level hashing misses them.
  • Generated files. Protobuf and gRPC stubs (*_pb2.py, *.pb.go), OpenAPI clients, ORM migrations, minified bundles, lockfiles (package-lock.json, yarn.lock, poetry.lock) and IDE project files are high-volume, low-information text.
  • Templated boilerplate. Scaffolds from create-react-app, Spring Initializr, Cookiecutter or an internal service template produce thousands of nearly identical main files, Dockerfiles and CI configs.
  • History duplication. If a delivery includes full git history, every commit snapshot repeats most of the tree. Decide early whether you train on HEAD snapshots, diffs, or both.

Normalizing source before you hash

Normalization decides what counts as "the same code," so set it deliberately and record it. Raw byte hashes treat a reformatted file as new, which lets formatter churn, line-ending changes and license-header updates inflate token counts.

A practical normalization ladder for exact and near-dup passes:

  1. Decode to UTF-8, convert CRLF to LF, strip trailing whitespace and the byte-order mark.
  2. Remove leading license and copyright header blocks, which are themselves heavily duplicated.
  3. Strip comments and docstrings for the hashing key only (keep them in the training text).
  4. Collapse whitespace runs, or tokenize with a language-aware lexer (for example tree-sitter grammars) so formatting does not matter.
  5. Optionally canonicalize identifiers (rename locals to v1, v2) to catch renamed copies. This catches Type-2 clones but also merges genuinely different small functions, so apply it only to files above a minimum token length.

Keep two keys per file: a strict hash (after steps 1-2) for exact dedup, and a normalized token stream (after steps 3-5) for MinHash. Never train on the normalized text; it is only a matching key.

Exact, near-duplicate and repository-level passes

Run dedup as three passes that answer different questions, cheapest first. Exact hashing removes identical files, MinHash with LSH finds near-duplicates, and repository grouping decides which copy survives.

Exact pass. SHA-256 or a 128-bit hash over the strict key. This is linear in corpus size and removes unchanged forks, vendored copies and repeated generated files.

Near-duplicate pass. Build token shingles (word or lexer-token n-grams; 5-grams are a common starting point) from the normalized key, compute a MinHash signature, and use LSH banding to find candidate pairs above a Jaccard threshold. Thresholds around 0.7 to 0.85 are common starting points [2], but tune them on a labeled sample of your own code rather than copying a published setting. Public code corpora such as The Stack, filtered by repository license across 30 languages, and The Stack v2, built from the Software Heritage archive, are distributed with deduplicated variants, which makes them convenient reference indexes for this pass [2][3]. For the generic mechanics, see exact and near-duplicate detection with MinHash and LSH, and for clones that differ syntactically but not semantically, semantic deduplication with embeddings.

Repository-level pass. Group files by repository (and by organization for a single supplier's codebase) before choosing survivors. Near-duplicates across one company's own services are usually template copies, and keeping the copy inside the most complete, most-tested repository preserves cross-file context. This matters if you train on repository-level code context, where dropping a file from the middle of a repo breaks imports and call graphs. A reasonable rule: dedup files globally, but resolve each cluster by keeping the member in the repository with the most history, tests and build metadata.

Snippet pass (optional). For sub-file copies, run suffix-array or n-gram span matching on high-value languages and drop or mask repeated spans over a length threshold. Lee et al. used exact-substring matching for this in text [1]; it is expensive, so scope it to the slices you will up-weight.

Decision table: what to do with each duplicate type

The right action depends on the duplicate's source and on whether the data feeds training or evaluation. Use a table like this as a default policy and override it per slice.

Illustrative example: invented to show structure; it does not describe an available dataset.

Duplicate typeDetection signalTraining actionEval action
Unchanged fork or mirrorExact hash match, same repo name stemKeep one copy (oldest or most-starred origin)Remove from test if any copy is in train
Vendored dependencyPath under vendor/, third_party/, node_modules/; match to public indexDrop, or keep only if absent from your public mixRemove
Generated stubs and lockfilesFilename patterns, "DO NOT EDIT" markers, very low token entropyDrop or cap per repoRemove
Template boilerplateNear-dup cluster spanning many repos of one orgKeep 1 to 3 representativesRemove from test
Copied snippet inside a fileRepeated span over length thresholdMask span or down-weightFlag the test item
Renamed clone (Type-2)Near-dup on identifier-canonicalized key onlyKeep one unless functions are shortRemove
Successive commit snapshotsSame path, consecutive commitsTrain on one snapshot plus diffsSplit by time, not by file

Measuring net-new tokens in a licensed delivery

Before pricing talks, measure how much of a licensed code delivery is already in your corpus. Proprietary code often embeds open-source copies, so headline repository sizes overstate what you are buying.

Run the supplier sample through the same normalization and passes against your existing training mix plus a public reference index such as The Stack v2 dedup [3]. Report four numbers per language: exact-overlap tokens, near-duplicate tokens above your threshold, internal duplication within the delivery, and remaining net-new tokens. Then repeat the split by directory type (first-party source, tests, vendored, generated) so the conversation is about the first-party portion, which is where licensed private code differs from public, license-filtered corpora [2]. Our guides on evaluating a code dataset sample and on what drives the price of licensed code data cover how to use these figures.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "delivery_id": "sample-2026-10-a",
  "reference_indexes": ["internal_mix_v7", "the-stack-v2-dedup"],
  "normalization": {"strip_comments": true, "canonicalize_identifiers": false, "lexer": "tree-sitter"},
  "minhash": {"num_perm": 256, "shingle": "token-5gram", "jaccard_threshold": 0.8},
  "by_language": {
    "java": {"total_tokens": 410000000, "exact_overlap": 0.18, "near_dup_overlap": 0.11, "internal_dup": 0.09, "net_new": 0.62},
    "python": {"total_tokens": 95000000, "exact_overlap": 0.07, "near_dup_overlap": 0.05, "internal_dup": 0.14, "net_new": 0.74}
  },
  "by_path_class": {"first_party": 0.71, "tests": 0.12, "vendored": 0.11, "generated": 0.06}
}

Ask the supplier to describe directory conventions and build tooling so you can classify vendored and generated paths reliably. The code dataset request specification guide lists the fields to request up front.

Documenting the dedup pass for reviewers

Record the dedup configuration as dataset metadata, not as a notebook comment. Reviewers and later teams need to know the normalization, thresholds, reference indexes and survivor rule to reproduce counts or rerun against a new benchmark.

Croissant expresses dataset-level metadata and file resources in JSON-LD on schema.org, which makes it a reasonable carrier for a dedup provenance block alongside the files [4]. NIST SP 800-218A adds AI-specific secure development practices that include attention to the integrity of training and evaluation data, which is a useful frame for treating dedup manifests as controlled artifacts [5]. Store cluster IDs per file so you can later remove a whole cluster if a benchmark or a rights issue surfaces. Run secrets scanning before the near-dup pass as well, since duplicated config files spread the same credentials (see secrets removal for code datasets).

How SourceX fits code data sourcing

SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Teams scoping proprietary codebases with full git history can start from the SourceX buyer page, and the guide to evaluating data supplier quality covers broader supplier checks. The code data buyer's map and the AI data hub list related categories.

Get net-new licensed code data for your corpus

SourceX finds US businesses that hold the code and engineering records you describe, manages licensing and ongoing purchases, and releases data only with the supplying company's approval. Describe the languages, history depth and repository types you need, and the overlap you want to avoid, at sourcex.si/buyers.

Frequently asked questions

Should I deduplicate before or after splitting train and test?

Deduplicate before splitting, then split by repository or by time rather than by file. A random file-level split puts forks and template copies on both sides, which inflates benchmark scores in the same way Lee et al. found train-test overlap does for text [1].

Does identifier canonicalization remove too much?

It can. Short getters, setters and one-line wrappers collapse into the same key after renaming, so apply canonicalized matching only above a minimum token length and audit a sample of merged clusters by hand.

Is near-dedup against public code enough for decontamination?

No. It removes overlap with code you indexed, but benchmarks can leak through forks and mirrors you never ingested. Pair dedup with targeted benchmark checks and keep the cluster IDs so you can rerun them.

Sources

  1. arXiv (Lee et al.; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  2. arXiv (Kocetkov et al., BigCode), "The Stack: 3 TB of permissively licensed source code" (2022). https://export.arxiv.org/abs/2211.15533?context=cs
  3. BigCode (Hugging Face), "bigcode/the-stack-v2-dedup". https://huggingface.co/datasets/bigcode/the-stack-v2-dedup
  4. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  5. NIST, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data