Code and software engineering data
Vulnerability Fix Data for Secure Code Models
Quick answer
A vulnerability fix dataset pairs code before and after a security fix, labeled with the weakness type (usually a CWE ID), severity and the finding that exposed it. Public sets such as CVEfixes, Vul4J+ and SAP's curated Java fixes cover CVE-linked open-source commits [1][3][4]. Teams training detection or auto-fix models often also need private fixes: internal SAST, pen-test and bug-bounty findings with minimal patches and regression tests. Specify those fields, release rules and verification evidence before you evaluate any supplier.
By SourceX Editorial · Updated
What public vulnerability fix datasets already cover
Public datasets give you scale on disclosed open-source CVEs, but almost nothing on the private, proprietary code where many enterprise vulnerabilities are found and fixed. CVEfixes v1.0.8 links 12,107 fixing commits in 4,249 projects to 11,873 CVEs across 272 CWE types, and it can be regenerated from the NVD [1]. Its collection is automatic: the tool pulls CVE records, follows reference links to repositories and extracts the vulnerable and fixed code [2].
Curated sets trade size for precision. SAP's dataset maps 624 vulnerabilities in 205 Java projects to 1,282 fix commits, reviewed by hand, and was used to train classifiers that flag security-relevant commits [4]. Vul4J+ adds test oracles so a generated Java patch can be checked by execution rather than by diff similarity [3].
The gaps matter for buyers. CVE-linked data skews toward popular libraries, toward weaknesses that earned a CVE and toward languages with strong open-source ecosystems. It carries no internal finding source, no triage notes and little of the business-logic, authorization and configuration flaws that pen testers find in proprietary services. That is the case for licensing private fixes, covered alongside general history data on the software engineering datasets from private repos page.
Fields a usable vulnerability fix record needs
A usable record carries the pre-fix code, the minimal fix, the weakness label and the evidence that triggered and verified the change. Without the trigger and the verification, you have a diff with a guess attached. The schema below is a starting point for a request.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"fix_id": "vf-000184",
"repo_alias": "svc-billing-api",
"language": "Java",
"pre_fix_commit": "a1c9e0f",
"fix_commit": "7d42b13",
"files_changed": ["src/main/java/invoice/InvoiceController.java"],
"minimal_security_hunks": [1, 3],
"unrelated_hunks_removed": true,
"cwe_id": "CWE-639",
"cwe_mapping_basis": "analyst assigned, second-reviewer confirmed",
"severity": {"scheme": "CVSS v3.1", "score": 8.1, "internal_rating": "high"},
"finding_source": "pen_test",
"finding_summary": "Invoice ID accepted from path without tenant ownership check",
"detector_output": null,
"regression_test": "InvoiceControllerTenantIsolationTest",
"test_fails_pre_fix": true,
"test_passes_post_fix": true,
"deployed_all_environments": true,
"disclosure_window_closed": true,
"redactions": ["internal hostnames", "service account names"]
}
Three fields do most of the work. finding_source (SAST, DAST, pen test, bug bounty, incident, code review) lets you balance tool-detectable and human-found weaknesses. minimal_security_hunks separates the security change from refactors bundled into the same commit. The test pair (test_fails_pre_fix, test_passes_post_fix) turns a record into an executable evaluation case, the same idea Vul4J+ applies to public Java fixes [3].
For repository context beyond a single file, see repository-level code context. For the broader issue-and-patch shape, the issue-to-fix pairs guide covers how to build and validate pairs from private trackers.
Label noise: the main failure mode in fix pairs
Label noise is the most common defect in vulnerability fix data, and it usually comes from commits that mix the security change with unrelated edits. A fix commit that also bumps dependencies, renames variables and reformats files teaches a model that all of those changes are "the fix." Automated collection from CVE references inherits whatever the linked commit contains [2], which is why curated sets invest in manual review [4].
Weakness labels add a second layer of noise. CWE assignments in NVD-derived data can be coarse (a parent category rather than the specific weakness) or reflect the reporter's view rather than the root cause. Even well-known benchmarks carry measurable label error: an audit of ten widely used test sets estimated an average error rate of at least 3.3% [5]. Confident learning, implemented in the cleanlab package, is one practical way to rank suspect labels for re-review [6].
Ask suppliers to state how each record's minimal change was isolated, who assigned the CWE and whether a second reviewer checked it. Then audit a sample yourself, as described in evaluating a code dataset sample.
Release rules for private security fixes
Only fixes that are deployed everywhere and past any disclosure window should leave the supplier. A pre-fix snapshot of a still-unpatched service can serve as an exploit map, so suppliers should not release it. Expect the supplying company's security team to sign off on each batch, and expect redaction of internal hostnames, service names and infrastructure details.
Secrets are a related risk. A vulnerability fix often removes a hard-coded credential, so the pre-fix side may contain a live key. Require full-history secret scanning and proof of rotation, as covered in secrets in code datasets. Ownership of the code and any third-party components should be confirmed through code ownership due diligence.
Decision table: which fix data fits which model goal
Match the data to the model you are building, because detection, repair and secure generation need different record shapes and different evaluation evidence.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Model goal | Minimum record | Must-have labels | Verification evidence | Public data enough? |
|---|---|---|---|---|
| Security-commit classifier | Commit diff plus message | Security vs non-security | Manual review sample | Often, for open-source signals [4] |
| Function-level vulnerability detection | Pre-fix function, post-fix function | CWE, severity | Second-reviewer CWE check | Partly; skews to CVE-linked libraries [1] |
| Automated repair (SFT) | Pre-fix file, minimal fix, finding text | CWE, finding source | Failing-then-passing test | Limited outside Java test-oracle sets [3] |
| Repair evaluation | Runnable environment and test | CWE, difficulty | Held-out, never public | Rarely; public fixes risk contamination |
| Secure code generation | Prompt-style task plus fixed solution | CWE avoided | Static analysis on outputs | Needs private business-logic cases |
For evaluation sets, contamination is the deciding factor: public CVE fixes appear in mirrors, forks and advisories that models have likely seen. Read code benchmark contamination before treating any public fix as held-out, and see SWE task environments with tests for runnable evaluation design.
Request checklist for vulnerability fix data
A strong request describes the data and its verification, not the suppliers you hope to find. Use this list when drafting a request or comparing offers.
- Languages, frameworks and service types (APIs, mobile, infrastructure-as-code).
- Target CWE coverage, including business-logic and authorization classes, and how mapping is done.
- Finding sources wanted, with rough proportions (SAST, pen test, bug bounty, incident).
- Minimal-change isolation method and the share of records with unrelated hunks removed.
- Pre-fix and post-fix tests, and whether tests run in a provided build environment.
- Release rules: deployed everywhere, disclosure window closed, supplier security sign-off.
- Redaction scope and how redaction is recorded so it does not break compilation.
- Secrets scan of full history and confirmation that exposed credentials were rotated.
- Delivery format (JSONL records plus repository snapshots or patches) and transfer method; see dataset delivery formats.
- Supplier security posture for handling pre-fix code, per security review of a training data supplier.
The code dataset request specification guide covers the general fields; the code and software engineering datasets map shows adjacent data types.
How SourceX approaches private fix data
SourceX sources operational datasets, including engineering records, from US companies and manages the commercial process through licensing and ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match. You describe the data you need, and SourceX looks for US businesses that hold it; every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, allowed uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery happens through private, access-controlled workflows only after an executed agreement and supplier approval. Security teams can also see how SourceX frames demand on the cybersecurity services buyers page, or start from the buyer overview.
Source private vulnerability fix pairs for your model
If public CVE-linked fixes do not cover your languages, weakness classes or held-out evaluation needs, describe the vulnerability fix data you need, including labels, finding sources and tests. SourceX will look for US companies that hold it, assess data and licensing permissions, and nothing is contracted until a supplier agrees. Describe your vulnerability fix data request.
Sources
- Zenodo, "CVEfixes dataset, version v1.0.8" (2024). https://zenodo.org/records/13118970
- GitHub (inksong), "CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software". https://github.com/inksong/cvefixes
- Zenodo, "Vul4J+ (Version 2.0)" (2024). https://zenodo.org/records/13752193
- arXiv (ar5iv), "A Manually-Curated Dataset of Fixes to Vulnerabilities of Open-Source Software" (2019). https://ar5iv.labs.arxiv.org/html/1902.02595
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Northcutt, Jiang, Chuang), "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/abs/1911.00068v4
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.