Code and software engineering data
Repository-Level Code Data: Cross-File Context for Completion and Agents
Quick answer
A repository-level code completion dataset keeps whole repositories intact (imports, call sites, build files, and a pinned commit) so a model can learn to use code from other files when it completes or edits one. File-level corpora strip that structure out. To buy or build one, specify the dependency-ordered packing rule, ship build metadata and a symbol index, cap each monorepo sub-project when sampling, and draw evaluation context only from the same repository at the same commit.
By SourceX Editorial · Updated
Why file-level code data fails cross-file tasks
File-level data fails because the answer to most real completions sits in another file. The model needs the signature of a helper in utils/billing.py, a constant in a generated header, or the interface a Java class implements. RepoBench was built because earlier benchmarks were mostly single-file and left multi-file scenarios untested [1][2]. The same gap applies to pretraining: a model trained on isolated files never sees which file defines what another file uses.
The common failure modes from file-level training show up quickly in evals:
- Hallucinated APIs: the model invents a method name that sounds plausible but does not exist in the project.
- Wrong arity or types: it calls a real function with the argument order of a similar public library.
- Duplicate implementations: it rewrites a helper inline instead of importing the one that already exists.
- Broken imports: it uses paths that do not resolve under the project's package layout or build configuration.
Methods that improve repository-level completion, such as dataflow-guided retrieval, depend on the training and eval data still having intact cross-file relationships to analyze [3]. If the dataset arrives as a deduplicated bag of files with paths stripped, none of those methods can be reproduced.
What a repository-level delivery should contain
A useful delivery contains the source tree at a pinned commit plus the metadata needed to resolve references without guessing. Treat the following as the minimum manifest for each repository, whether the data feeds pretraining, completion fine-tuning or agent work.
- Pinned snapshot: repository ID, commit SHA, and the full file tree with original relative paths. For history-dependent tasks, see delivering repositories with full Git history.
- Build metadata:
compile_commands.jsonfor C and C++ (Clang's JSON compilation database: one entry per translation unit withdirectory,file, andargumentsorcommand; CMake can emit it viaCMAKE_EXPORT_COMPILE_COMMANDS);pom.xmlorbuild.gradlefor JVM projects;pyproject.tomlandrequirements*.txtfor Python;package.jsonwith the lockfile;go.modandgo.sum;Cargo.tomlandCargo.lock. - Dependency graph: a file-level edge list (importer, imported, edge type) produced by a named tool and version.
- Symbol index: definitions and references with file, line and symbol kind. SCIP or LSIF dumps from a language indexer, or ctags output, all work if the format and tool version are documented.
- Generated and vendored code flags: protobuf stubs, ORM migrations,
node_modules,vendor/and minified bundles marked so they can be excluded from loss or from context. - Per-file metadata: language, byte size, token count under your tokenizer, license header detection, and secrets-scan status.
Missing lockfiles are a common gap. Without them, a buyer cannot reconstruct which version of a third-party API the code was written against, which matters for dependency upgrade and migration pairs and for any execution-based eval.
Packing files into long-context training sequences
The packing rule decides whether related code lands in the same context window, so document it as carefully as the data itself. A reasonable default is to parse invocation relationships (Python imports, C# using, C #include), arrange files so the context each file depends on comes before it, then concatenate them into repository-level samples. Definitions then precede uses, which mirrors how a completion model sees a prefix.
Alternatives and their trade-offs:
- Topological by dependency: best match to cross-file completion, but needs cycle handling (Python and JavaScript projects often have import cycles) and a rule for disconnected components.
- Directory order: cheap and deterministic; keeps a package together but may place a consumer before its dependency.
- Retrieval-ordered: for each target file, prepend the top-k files by symbol overlap or embedding similarity. Closer to inference-time RAG, but harder to reproduce.
Position also matters at inference. Models use information at the start and end of long contexts more reliably than information in the middle [4], so a packing rule that buries the relevant definition mid-sequence can understate what the data teaches. When sequences vary widely in length, packing plus sorted batching keeps long-context fine-tuning efficient [5]; record the sequence length, separator tokens and whether file path comments are inserted between files.
Fill-in-the-middle from repositories
Fill-in-the-middle (FIM) samples built from repositories should take the prefix and suffix from one file and the extra context from other files in the same repository at the same commit. Mixing snapshots creates impossible samples: the retrieved file may define a function the target file never saw, or omit one it calls.
Practical rules for FIM construction:
- Choose spans on syntactic boundaries (function body, statement block, argument list) using a parser such as tree-sitter, rather than random character offsets alone.
- Exclude spans whose answer appears verbatim in the retrieved context, or tag them, so the model does not just learn to copy.
- Keep the target file's own imports in the prefix; they are the strongest signal for which cross-file context matters.
- Record the span selection seed so the set can be regenerated after a deletion or a rights change.
Sampling monorepos without letting one repository dominate
Monorepos need sub-project boundaries before sampling, or a single repository dominates the mix. One large monorepo can contain hundreds of services, each with its own BUILD, pom.xml or package.json, and treating it as one unit both overweights its style and produces dependency graphs too large for any context window.
Split on build-system boundaries (Bazel packages, Gradle subprojects, npm or pnpm workspaces, Go modules), then cap tokens per sub-project and per parent repository. Keep the cross-sub-project edges in the graph so a model can still learn shared-library usage, but sample targets within one sub-project at a time. Record the caps in the dataset card; JSON Lines manifests (one UTF-8 JSON object per line) [8] and a Croissant description of files and record fields [9] make the split auditable.
Repository-level evaluation without leakage
Repository-level evaluation must hold out whole repositories, or at minimum later commits, from training. Splitting at file level leaks: the model sees orders/service.py in training and is then tested on orders/handlers.py, which imports it. RepoBench separates retrieval, completion and pipeline tasks so retrieval quality and generation quality can be scored separately [1]; your internal eval should do the same.
Agent evals raise the bar further. SWE-bench builds tasks from real issues and pull requests that require changes across a codebase [6], and as of October 2026 OpenAI no longer reports SWE-bench Verified, citing contamination [7]. Public repositories leak into pretraining through forks and mirrors, which is why teams look to held-out eval sets from licensed private repositories and run benchmark contamination checks before trusting a score.
Specification for a repository-level code data request
Use a written specification so suppliers can say yes or no to concrete fields. The template below covers the repository-level items; the broader code dataset request specification covers languages, history and test requirements.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
| Unit of delivery | Whole repository at pinned commit SHA | Prevents mixed-snapshot context |
| Languages and build systems | Java 17 with Gradle; TypeScript with pnpm workspaces | Determines which metadata must ship |
| Build metadata | Lockfiles required; compile_commands.json for any C/C++ | Static resolution of cross-file references |
| Dependency graph | File-level edge list, tool name and version | Reproducible packing |
| Symbol index | SCIP dump per repository | Retrieval and eval construction |
| Packing rule | Topological by import; cycles broken by path order | Same sample regardless of who builds it |
| Monorepo handling | Split by build-system module; cap per module | Avoids one-repository dominance |
| Exclusions | Vendored, generated, minified, binary | Keeps loss on authored code |
| Eval split | Held-out repositories, not files | No cross-file leakage |
| Pre-delivery scans | Full-history secrets scan; license header report | Removes credentials and copyleft surprises |
An illustrative manifest record, one per file:
{"repo_id":"r-0412","commit":"9f3c1e7","path":"orders/handlers.py","lang":"python","tokens":1834,"imports":["orders/service.py","common/money.py"],"generated":false,"vendored":false,"subproject":"orders","secrets_scan":"clean"}
Before licensing, run your own checks on a sample using the approach in evaluating a code dataset sample, and confirm the secrets removal across full Git history and copyleft contamination results.
Where proprietary repositories fit
Private repositories add what public corpora lack: unseen code for evaluation, enterprise build systems, and internal libraries that no public model has memorized. They also bring ownership, secrets and personal-data questions that a public crawl never raises. SourceX sources operational datasets, including engineering records, from US companies on request, and handles the licensing process; it does not hold stock, and a request does not guarantee a match. See proprietary code datasets with full Git history and coding agent training data, or describe your requirements on the buyer request page.
Every dataset SourceX delivers is rights-reviewed for ownership and consents, released only with the supplying company's approval, and delivered under a license that defines the records, allowed uses, term and delivery. Personal details such as names and emails are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For the wider map of code data types, start at the code and software engineering data hub or the AI data overview.
Request repository-level code data for completion and agents
SourceX finds US businesses that hold the engineering records you describe, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees, and delivery runs through private, access-controlled workflows. Describe the repositories, languages and build metadata you need at sourcex.si/buyers.
Frequently asked questions
Is repository-level data only useful for very long context windows?
No. Even with an 8K or 16K window, repository structure lets you build retrieval-augmented samples that prepend the few most relevant cross-file snippets. The dependency graph and symbol index are what make that retrieval possible.
Should I deduplicate across repositories before packing?
Deduplicate exact and near-duplicate files, but do it after recording the original dependency edges. Removing a shared file before packing silently breaks the graph for every repository that imported it.
Do I need build metadata if I only train on Python?
Yes, in a lighter form. pyproject.toml, lockfiles and the package layout determine how imports resolve, and they are needed for any execution-based eval of completions or agent patches.
Sources
- arXiv (Liu, Xu, McAuley), "RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems" (2023). https://arxiv.org/pdf/2306.03091
- ICLR, "RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems (ICLR 2024 proceedings)" (2024). https://proceedings.iclr.cc/paper_files/paper/2024/hash/d191ba4c8923ed8fd8935b7c98658b5f-Abstract-Conference.html
- arXiv, "Dataflow-Guided Retrieval Augmentation for Repository-Level Code Completion" (2024). https://arxiv.org/pdf/2405.19782v1
- arXiv (Liu et al.), "Lost in the Middle: How Language Models Use Long Contexts" (2023). https://arxiv.org/pdf/2307.03172
- arXiv (Bai et al.), "LongAlign: A Recipe for Long Context Alignment of Large Language Models" (2024). https://arxiv.org/abs/2401.18058v1
- arXiv (Jimenez et al.), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.