Code and software engineering data
Unit Test Generation Data: Focal Code, Tests and Coverage from Real Projects
Quick answer
Unit test generation data pairs each real test with the focal method or class it exercises, plus the evidence that the test is worth imitating: line and branch coverage of the focal code, a mutation score, pass stability across reruns, and the fixtures and mocks it needs to run. Public focal-test corpora are large but noisy and mostly open source. Enterprise teams that want models to write tests for internal codebases need licensed private repositories where the mapping method, coverage reports and fixture de-identification are documented per record.
By SourceX Editorial · Updated
What a usable focal-test record contains
A usable record links one test to one focal unit, states how that link was derived, and carries measured quality signals rather than just source text. Public research datasets show the range: UniTSyn collected 2.7 million focal-test pairs across five languages by resolving calls through the Language Server Protocol [1], while Go-UT-Bench is far smaller at 5,264 Go pairs but records the repository, file paths and commit hash for each entry [2]. For a commercial dataset, ask for both scale and that level of provenance.
The minimum fields are the focal method signature and body, the enclosing class or module context, the test method, test-file imports and setup (@BeforeEach, setUp, pytest fixtures, TestMain), the build target, and the commit SHA where both existed together. Add coverage, mutation and flakiness fields, described below, and you have a record that can serve as both an SFT target and a reward signal.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "utg-000412",
"repo_alias": "repo_17",
"commit_sha": "9f3c2e1",
"language": "java",
"build_system": "gradle",
"focal": {
"path": "billing/src/main/java/InvoiceCalculator.java",
"signature": "BigDecimal applyDiscount(Invoice inv, DiscountRule rule)",
"class_context_tokens": 1840
},
"test": {
"path": "billing/src/test/java/InvoiceCalculatorTest.java",
"method": "appliesTieredDiscountAboveThreshold",
"framework": "junit5",
"mocks": ["mockito", "internal:FakeLedgerClient"]
},
"mapping": { "method": "coverage", "confidence": 0.93, "also_matched_by": ["naming"] },
"coverage": { "focal_line_pct": 88.0, "focal_branch_pct": 75.0, "tool": "jacoco" },
"mutation": { "tool": "pitest", "killed": 14, "survived": 3, "no_coverage": 1 },
"stability": { "reruns": 20, "failures": 0, "uses_network": false, "uses_wall_clock": false },
"fixture_deid": { "method": "pseudonymized", "fields": ["customer_name", "email", "account_no"] }
}
How the focal mapping was derived matters
The mapping method determines label noise, so require suppliers to report it per record. Three methods are common: naming conventions (FooTest.testBar maps to Foo.bar), static call-graph or LSP resolution from the test body [1], and dynamic mapping from per-test coverage, which records what the test actually executed. Naming is cheap but breaks on helper-heavy suites; call graphs miss reflection and dependency injection; coverage-based mapping is the most faithful but requires a working build.
Noise is not a theoretical concern. A data-quality study of Methods2Test (about 781,000 pairs) and Atlas (about 188,000 pairs) found that noisy focal-test data degrades test-generation models, which supports filtering for quality before scaling for volume [3]. When two methods disagree, keep both labels and a confidence value rather than silently picking one.
Coverage and mutation scores turn tests into ranked targets
Per-test coverage and mutation results let you train on strong tests and use weak ones only as negatives. Line coverage alone rewards tests that execute code without checking it; mutation testing (PIT for Java, Stryker for JavaScript and C#, mutmut for Python, go-mutesting for Go) measures whether assertions actually catch injected faults. Ask for killed, survived and no-coverage counts per test, not a project-level percentage.
Coverage should be per test, scoped to the focal unit, and produced by a named tool and version (JaCoCo, coverage.py, Istanbul/nyc, go test -coverprofile, gcov/llvm-cov). A suite-level LCOV or Cobertura report cannot tell you which test covered which branch. If a supplier can only provide suite-level coverage, treat the dataset as unranked and budget for re-running tests yourself.
These signals also make good reward functions for RL or rejection sampling: compile success, pass on the original code, focal coverage gain and mutants killed. That only works if the dataset ships enough build context to execute tests, which overlaps with the environment requirements in SWE task environments with tests.
Flaky and environment-dependent tests poison both training and evaluation
Flag flaky tests before training, because a model that learns from them learns to write nondeterministic tests. Ask for rerun counts and failure counts per test, and for flags on tests that touch the network, wall-clock time, random seeds without fixing them, the filesystem, ordering of hash maps, or shared state between tests. OpenAI's write-up on SWE-bench Verified named overly specific unit tests and unreliable development-environment setup among the problems it found in the original benchmark [5].
Quarantined tests (@Disabled, @pytest.mark.skip, t.Skip(), CI retry annotations) are useful signal: keep them with their status rather than dropping them. They can serve as negative examples or as a held-out flakiness-detection task.
Fixtures, mocks and the enterprise gap
Enterprise test suites contain what public data lacks: internal mocking frameworks, service stubs, contract-test fakes, golden files and seeded databases that reflect how a company actually isolates dependencies. Ask suppliers to include these helper modules and the test-utility packages they import, or your model will learn to call fakes it has never seen. This is the main reason to license private code instead of relying on mined open-source pairs; see open code datasets vs licensed private code.
Fixtures are also where real customer data hides. JSON fixtures, SQL seed scripts, VCR or WireMock cassettes and snapshot files are often copied from production. Require the same de-identification standard for test data as for any record: replace names, emails, phone numbers and account numbers, record the method, and check a sample.
If fixtures hold protected health information from a HIPAA covered entity or business associate, de-identify them via Safe Harbor or Expert Determination [8]. Credentials in fixtures and CI config are a separate problem covered in secrets removal for code datasets.
Buyer checklist for a unit test generation dataset
Use this checklist when writing a request or reviewing a sample. It complements the broader code dataset request specification.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | What to ask for | Red flag |
|---|---|---|
| Focal mapping | Method per record (naming, call graph, coverage) and confidence | One unexplained mapping for all records |
| Coverage | Per-test, focal-scoped line and branch, tool and version | Suite-level percentage only |
| Mutation | Killed, survived, no-coverage per test; operator set | Mutation "score" without counts |
| Stability | Rerun count, failures, network/time/random flags | No rerun data |
| Build context | Build files, lockfiles, toolchain versions, test helpers | Test files without imports or helpers |
| Provenance | Repo alias, commit SHA, file paths per record [2] | Snapshots with no history |
| Fixtures | De-identification method and sample check | Raw cassettes or seed dumps |
| Contamination | Overlap check against public benchmarks and mirrors | No dedup or overlap report |
| Rights | Ownership and consent review, allowed uses in license | "Internal code, no review needed" |
Contamination, duplication and synthetic test data
Deduplicate before splitting, because copied test helpers and generated tests create near-duplicates that inflate evaluation scores. Near-duplicate examples are common in training corpora and increase verbatim memorization [7]; in test suites this shows up as vendored test utilities, parameterized tests expanded into many files and snapshot tests. Split by repository, not by record, so a held-out evaluation does not share a codebase with training. See deduplicating code training data.
Public benchmarks lose value once they leak into training: OpenAI stopped reporting SWE-bench Verified in 2026, citing contamination [6]. Private enterprise tests are less exposed, but check for public forks and mirrors using the approach in code benchmark contamination. Synthetic generation is an active alternative, including file-level test data synthesized with chain-of-thought and self-debugging [4]; the trade-offs are covered in synthetic vs licensed code data.
How SourceX sources unit test data
SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. You describe the data you need, such as Java services with JUnit 5 suites and per-test coverage, and SourceX looks for US businesses that hold it. Each dataset is rights-reviewed for ownership and consents, personal details in records are removed or replaced with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval.
Related owner pages cover proprietary codebases with full git history and software engineering histories from private repos; the code data buyer's map and the AI data hub cover adjacent needs. You can describe your test-generation data needs at any point.
Request unit test generation data
SourceX sources engineering records from US companies on request and manages licensing, with every release approved by the supplying company. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe the focal-test data you need on the SourceX buyers page.
Frequently asked questions
Is coverage-based mapping always better than naming-based mapping?
It is more faithful to what the test executed, but integration-style tests touch many methods, so coverage mapping needs a rule for picking the focal unit, such as the method with the highest share of covered lines in the class under test. Keep naming matches as a secondary label to catch disagreements.
Can tests without mutation data still be useful?
Yes, for SFT on style and framework usage, but you cannot rank them by fault-detection strength. If the dataset ships buildable projects, you can run PIT, Stryker or mutmut yourself and add the scores.
Should failing tests be included?
Include them with status labels. Tests that fail on the paired commit are useful negatives or repair targets, but they should never appear as positive SFT targets.
Sources
- arXiv, "UniTSyn: A Large-Scale Dataset Capable of Enhancing the Prowess of Large Language Models for Program Testing" (2024). https://arxiv.org/pdf/2402.03396
- arXiv, "Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go" (2025). https://arxiv.org/pdf/2511.10868
- arXiv, "Less is More: On the Importance of Data Quality for Unit Test Generation" (2025). https://arxiv.org/pdf/2502.14212
- arXiv, "Synthesizing File-Level Data for Unit Test Generation with Chain-of-Thoughts via Self-Debugging" (2026). https://arxiv.org/pdf/2602.03181
- OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv (ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.