Skip to content

Code and software engineering data

What Drives the Price of Licensed Code Data

Quick answer

There is no public price list for proprietary code data, and quotes for similar-looking corpora can differ widely. Price follows a handful of measurable drivers: how many tokens survive deduplication against public code, language scarcity, history depth and linkage to issues and reviews, whether the code builds, how strong the tests are, how many validated tasks you get, and which uses, term and exclusivity you buy. Budget per usable unit, not per repository or per gigabyte.

By SourceX Editorial · Updated

Why list prices for code data do not exist

Code data is priced deal by deal because every corpus has a different overlap with public code and a different preparation burden. A seller quoting "per repository" and another quoting "per million tokens" are not quoting the same thing until both are converted to what your pipeline will actually keep.

The public baseline matters. The original release of The Stack alone contained 3.1 TB of permissively licensed code in 30 languages, with an opt-out process for developers [1]; later versions are larger. Anything a proprietary seller offers is valued against what you could already train on for free, which is why the general pricing drivers for licensed enterprise data only partly apply to code.

Net-new tokens after deduplication set the floor

The first price driver is how much of the corpus is genuinely new once vendored dependencies, forks and copied snippets are removed. Private monorepos often carry vendored copies of open-source libraries, generated protobuf or OpenAPI clients, minified JavaScript and lockfiles; all of that inflates raw size without adding value. Run exact and near-duplicate matching (MinHash or similar) against your existing public corpora before you accept a token count, as described in deduplicating code training data.

Language and domain scarcity then scale the value of what remains. Python and TypeScript are abundant in open corpora; COBOL, PL/I, RPG, Ada, embedded C for specific microcontrollers, PLC ladder logic and proprietary DSLs are not. A smaller corpus in a scarce language can reasonably command more per token than a large web-stack codebase.

History depth and linkage add value per commit

Linked history is often worth more than the code snapshot itself. A single HEAD checkout supports pre-training on code; full Git history joined to issues, pull requests, review comments and CI results supports issue-to-patch tasks, review modeling and preference data. Ask how links are preserved: commit SHAs referenced in PR metadata, issue keys in commit messages (for example Jira keys), and CI run IDs tied to commits.

Linkage quality varies. Squash-merge workflows erase intermediate commits, migrations from SVN or Perforce truncate history, and ticket systems are often retired separately from the repository. Price the portion of commits with a resolvable issue and review thread, not the raw commit count.

Buildability and test strength drive agent-grade pricing

Code that builds and has meaningful tests is priced as a different product from code that only parses. The SWE-bench design shows why: each of its 2,294 tasks pairs a real issue with a pull request and checks resolution with tests [2]. When annotators audited it, they found overly specific unit tests, underspecified issues and unreliable environment setup, and kept a 500-task verified subset [3]; OpenAI has since stopped reporting that subset, citing contamination [10], which is another reason unpublished tasks carry value. That attrition is a cost someone pays, and a seller who has already paid it will price it in.

Ask suppliers for build success rate on a clean machine, the share of tasks with fail-to-pass tests, flakiness across repeated runs, and coverage of the changed lines. Weak tests make a task cheap to produce but expensive to trust.

Environment building is the hidden line item

Reproducible execution environments are frequently the largest preparation cost for coding-agent data. The SWE-bench harness runs every task in its own Docker container built from layered base, environment and instance images [4], and SWE-Gym's 2,438 Python instances each ship a codebase with a working runtime, unit tests and a task description [5]. Pinning dependency versions, mirroring private package registries, stubbing internal services and handling licensed toolchains all take engineering time per repository.

For priced quotes, separate three costs: the code license, the preparation work (secrets removal, license scanning, redaction, environment builds) and ongoing maintenance as base images and package mirrors rot. Secrets scanning across full history is its own effort; see secrets removal for code datasets.

Validated tasks are priced per task, not per token

When you buy coding-agent training or evaluation tasks, the unit is a validated task, and the per-task price depends on how it was produced. Machine-generated trajectories can be cheap: one web-agent paper reports about $0.28 per successful synthetic trajectory against $0.55 for an earlier method [7]. Tasks mined from real private engineering history with human-verified tests and working environments sit at the other end, and agent trajectory cost drivers explain the human-labor side.

Contamination resistance adds value for evaluation. Public permissively licensed repositories are likely already in pre-training corpora, which is why SWE-Bench Pro keeps part of its tasks private [6]. Private code that has never been published is valuable as a held-out evaluation set precisely because it is unseen; check this using the methods in code benchmark contamination.

Use rights, term and exclusivity change the price of the same files

The same repository can carry very different prices depending on the rights you license. Evaluation-only use, fine-tuning, and pre-training differ in how much value transfers into your weights, and sellers price accordingly. Term, the right to retain trained models after expiry, sublicensing to affiliates and exclusivity in a field or time window each move the number.

Rights diligence is also a cost the price must absorb. The U.S. Copyright Office concluded that many acts in AI training implicate copyright, in a report still marked pre-publication as of October 2026 [8]. Providers placing general-purpose AI models on the EU market must also publish a training-content summary under AI Act Article 53(1)(d) using the Commission template [9], so you need provenance records from the seller either way. Ownership questions (contractor code, acquired codebases, copyleft snippets) are covered in code ownership due diligence and copyleft contamination.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

A quote normalization worksheet for code data

Normalize every offer to a cost per usable unit before you negotiate. Convert raw size into post-dedup tokens, then into the units your use case consumes: tokens for pre-training, linked change records for fine-tuning, validated tasks for agents and evals.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldQuote A (snapshot corpus)Quote B (linked history)How to verify
Quoted unitper repositoryper validated taskContract schedule
Raw size1,200 repos, 40B tokens300 repos, 9,000 tasksManifest file
Overlap with public code after MinHash dedup55%10%Your dedup run on sample
Net-new tokens18Bnot the unitRecount on sample
LanguagesJava, TypeScriptCOBOL, JavaLinguist-style breakdown
History and linkageHEAD onlyFull history, PR and issue linksJoin rate on sample
Build success on clean machinenot offered92% of tasksRerun 50 tasks
Fail-to-pass tests, flake ratenoneyes, 3% flakyTriple-run sample
Licensed usespre-trainingfine-tuning and evaluationLicense grant clause
Term, retention, exclusivity2 years, no exclusivity3 years, field-limited exclusivityLicense schedule
Preparation includedsecrets scan of HEADfull-history secrets scan, environmentsSeller's preparation notes
Normalized unit for comparisonprice / net-new tokensprice / validated taskYour worksheet

Quote A's headline size shrinks by more than half after dedup and offers no agent value; Quote B is smaller but every unit is usable for its stated purpose. The vendor quote comparison method and per-token pricing audit extend this worksheet.

Negotiation levers procurement teams actually use

You lower code data cost mainly by narrowing scope and rights, not by haggling over a headline. Practical levers:

  • Request a sample first and run your own dedup, build and test checks; see evaluating code dataset samples.
  • Buy evaluation-only rights before training rights when the goal is a held-out benchmark.
  • Drop exclusivity unless a competitor gaining the same data is a real risk.
  • Specify languages, history depth and test requirements up front with a code dataset request specification so you do not pay for repositories you will discard.
  • Agree on who pays for environment builds and how broken tasks are replaced.

Teams on a tight budget can find more tactics in data licensing for AI startups on a budget, and the wider code data buyer's map shows which data types fit which training stage.

Where SourceX fits in code data sourcing

SourceX sources operational datasets, including engineering records, from US companies on request and manages the commercial process, including licensing agreements. It does not publish prices; terms are agreed per deal, and nothing is contracted until the supplying company agrees. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Describe your requirement on the SourceX buyer page or review proprietary code datasets with full Git history.

Budgeting for licensed code data with SourceX

If you are budgeting for proprietary code, describe the languages, history, build and test needs and intended uses rather than naming companies. SourceX looks for US businesses that hold matching data, but a request does not guarantee a match, and every release is approved by the supplier. Start a code data request with SourceX.

Sources

  1. Kocetkov et al. (BigCode), "The Stack: 3 TB of permissively licensed source code" (2022). https://export.arxiv.org/abs/2211.15533?context=cs
  2. Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
  3. OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  4. SWE-bench project, "SWE-bench evaluation harness reference". https://swebench.com/SWE-bench/reference/harness
  5. Pan, Wang, Neubig et al., "Training Software Engineering Agents and Verifiers with SWE-Gym" (2025). https://arxiv.org/pdf/2412.21139v2
  6. Scale AI, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  7. arXiv, "Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents" (2025). https://arxiv.org/pdf/2502.11357
  8. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  9. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  10. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data