Code and software engineering data
How to Specify a Code Dataset Request: Languages, History, Build and Test Requirements
Quick answer
A code dataset request should state, field by field, what a supplier must hold and prove: languages with version ranges, repository and file size bands, domains, version-control history depth, linkage between commits, issues, pull requests, reviews and CI, whether code must build and tests must pass, exclusions, train/eval split rules, delivery format and the uses you need licensed. Mark every field must-have or nice-to-have, and attach measurable acceptance metrics so suppliers can answer yes, no or partially.
By SourceX Editorial · Updated
This page is the code-specific companion to the generic guide on how to write a data request for suppliers. For the wider landscape of code data types, start at the code and software engineering data hub.
Why a generic data request fails for code
A generic request fails because code value depends on properties a file count never captures: whether it compiles, whether history is intact, and whether tests verify behavior. A supplier answering "200 repositories of Java" may be offering vendored dependencies, generated protobuf stubs and a squashed history with no tests. SWE-bench showed that useful software-engineering tasks come from pairing a real issue with the fixing patch and tests that fail before and pass after it [1]. If your request does not ask for that linkage, nobody will volunteer it.
The second failure is ambiguity about the target use. Pre-training wants breadth and deduplicated volume; fine-tuning for agents wants executable environments and trajectories; evaluation wants held-out repositories that have never been public. One request can cover several uses, but each needs its own row.
Languages, versions and size bands
Specify languages as language plus version range plus toolchain, because "Python" spans Python 2.7 code that will not run on a 3.12 interpreter. Name the build systems you can execute (Maven, Gradle, Bazel, CMake, npm/pnpm, Cargo, Go modules, Poetry or uv) and say whether you need lockfiles such as package-lock.json, poetry.lock or Cargo.lock, since missing lockfiles are a common reason old snapshots fail to resolve dependencies.
Use size bands rather than a single total: number of repositories, files per repository and lines of non-generated, non-vendored code. Ask suppliers to exclude or flag vendor/, node_modules/, third_party/, minified JavaScript and generated code, and to report counts after that exclusion. Dedup expectations belong here too, since near-duplicate removal measurably improved code model results in The Stack and reduces memorization more generally [5][6]. See deduplicating code training data for fork and vendored-copy handling.
History depth and linkage fields
Ask for history as a depth and an integrity guarantee: the number of years or commits retained, plus confirmation that history was not squashed, rewritten with git filter-repo without a log, or truncated by shallow clones. If secrets or personal data were removed by rewriting history, the request should ask for a record of the rewrite, because commit hashes change and linkage to issues may break.
Linkage is where private engineering data differs from public dumps. Specify which joins you need and the minimum linkage rate you will accept:
- Commit to issue (ticket keys such as
PROJ-1234in commit messages or branch names, resolved against Jira, Linear or GitHub Issues exports). - Pull request to commits, review comments, review state and merge outcome.
- Commit to CI run, with job status, logs and the failing test identifiers.
- Issue to fix to test, the basis for issue-to-fix pairs and review comment resolution data.
State the linkage rate as a measurable number (for example, the share of merged PRs with at least one resolvable issue key), and ask suppliers to report it on a sample rather than estimate it.
Build, test and environment requirements
Build and test requirements should be stated as pass criteria at a pinned commit, not as aspirations. "Buildable" can mean the default branch compiles today, every sampled commit compiles in a provided container, or each task instance has a Dockerfile that reproduces the environment. Pick one and say how you will verify it.
For agent and repair data, ask for task instances with a FAIL_TO_PASS set (tests that fail before the fix and pass after) and a PASS_TO_PASS set (tests that must keep passing), the convention SWE-bench popularized [1]. OpenAI's review of SWE-bench found tasks with underspecified issues and tests that rejected valid fixes [2], so ask the supplier to state how task validity was checked. Requiring strict buildability plus deep history narrows the pool of eligible repositories sharply, so mark which of the two matters more.
For test-generation use, specify coverage tooling and format (JaCoCo XML, Cobertura, coverage.py JSON, lcov), as described on the unit test generation data page. For CI-centric work, name the systems (GitHub Actions, GitLab CI, Jenkins, Buildkite) and the log retention you need, covered in CI build logs and build failure datasets.
Exclusions, secrets and license scanning
Exclusions should name what must not be in the delivery and how absence will be measured. Typical rows: credentials and keys across full history, personal data in commit metadata and comments, third-party code under copyleft licenses, code subject to export restrictions, customer data embedded in fixtures, and anything the supplier cannot show it owns.
Ask for the scanners and rules used (for example gitleaks or TruffleHog for secrets, ScanCode or a comparable tool for license detection with SPDX identifiers) and the residual-finding rate on a sample. The secrets removal guide covers full-history scanning, and copyleft contamination covers GPL and AGPL snippets. Ownership questions (contractor code, acquired codebases, open-source contributions) belong with code ownership due diligence.
If you provide a general-purpose AI model placed on the EU market, your copyright policy and training-content summary obligations under Article 53 of the EU AI Act have applied since 2 August 2025, with AI Office enforcement for new models from 2 August 2026, as of October 2026 [7]; models placed on the market before 2 August 2025 have until 2 August 2027 to comply under Article 111(3). That is a reason to ask suppliers for provenance records, not to rely on them for your compliance.
Splits, contamination controls and delivery form
Splits should be defined by repository or by time, never by random file, because files from one repository leak style, identifiers and helper functions across a random split. For evaluation sets, require repositories that were never public, since OpenAI stopped reporting SWE-bench Verified in 2026 after concluding contamination and flawed tasks undermined it [3], and SWE-Bench Pro moved toward held-out and commercial codebases for the same reason [4]. The code benchmark contamination page covers fork and mirror detection.
Delivery form should name container and schema: git bundles or mirrored bare repositories for history; JSONL for task instances and linked records (UTF-8, one JSON value per line, no byte order mark) [9]; Parquet for large file-level tables [10]. Ask for a dataset card describing contents, preparation and intended use [11], and reuse the technical delivery specification template for transfer, checksums and manifests.
Code dataset request template
The table below is a request template you can paste into an RFP or a request form, with each row marked must-have (M) or nice-to-have (N) and a measurable acceptance metric.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example requirement | M/N | Acceptance metric |
|---|---|---|---|
| Target use | Fine-tuning a repair agent; held-out eval set | M | Uses named in license draft |
| Languages and versions | Java 11–21 (Gradle/Maven), TypeScript 5.x (pnpm) | M | Share of repos matching version range |
| Size band | 50–150 repos; 20k–500k non-vendored LOC each | M | Counts reported after exclusions |
| History depth | At least 3 years, unsquashed, no undocumented rewrites | M | Commit count and rewrite log per repo |
| Linkage | PR to issue key to CI run | M | At least 60% of merged PRs linked on sample |
| Reviews | Review comments with resolution state | N | Share of PRs with one or more review comments |
| Buildability | Builds in supplied Dockerfile at task commit | M | Build rate on a 100-task sample |
| Tests | FAIL_TO_PASS and PASS_TO_PASS per task | M | Validated-task rate on sample |
| Coverage | JaCoCo XML per task commit | N | Present for share of tasks |
| Secrets | Full-history scan, rules disclosed | M | Zero confirmed live credentials in sample |
| Licensing | SPDX-tagged file inventory; no GPL/AGPL third-party code | M | License scan report attached |
| Exclusions | No customer data in fixtures; no export-restricted code | M | Supplier attestation plus sample check |
| Split | By repository; eval repos never public | M | Repo-level split manifest |
| Delivery | Git bundles plus JSONL tasks plus Parquet file table | M | Schema validation passes |
Pair the template with the sample evaluation checklist so the metrics are measured on a sample before you commit. ISO/IEC 5259-3 is useful for framing who owns each quality check, though it leaves the metrics to you [8].
How suppliers will read your request
Suppliers read a code request for three things: whether they hold the data, whether they are permitted to release it, and how much preparation the acceptance metrics imply. A request that marks every row must-have will usually return few or no matches, so rank fields and say which trade-offs you accept, such as shorter history in exchange for verified builds.
Describe the data, not the companies you think hold it. Include the uses you need licensed (pre-training, fine-tuning, evaluation, or internal benchmarking) and let licensing terms be negotiated per supplier. For a guided version of this template, the data request builder for AI teams walks through the same fields, and the proprietary code datasets page describes the kinds of engineering data involved. When the spec is ready, you can send your code data request to SourceX.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Send a code dataset request
SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, and every release is approved by the supplying company. Submit your code dataset specification.
Sources
- Jimenez et al., arXiv, "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Scale AI, arXiv, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Kocetkov et al. (BigCode), arXiv, "The Stack: 3 TB of permissively licensed source code" (2022). https://export.arxiv.org/abs/2211.15533?context=cs
- Lee et al., arXiv, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs
- Hugging Face, "Create a dataset card". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.