Proprietary codebases with full version history for AI training
A proprietary codebase dataset is a company's private source code delivered as whole repositories with their version history: commits, branches, tags and files, along with build configuration, tests and internal documentation. SourceX sources codebases from established software companies and engineering teams, confirms the partner owns the code, identifies open-source and other third-party components, removes secrets and personal data, and screens for export-controlled code before licensing.
Dataset manifest
Sourced to your spec- What it is
- Complete private repositories with full version history, build files, tests and docs
- Typical systems
- GitHub Enterprise, GitLab, Bitbucket, Azure Repos, Perforce, Subversion
- Typical history
- Varies by partner; older history may come from a pre-git system
- Modality
- Source code with commit history, build and test configuration and internal docs
- Delivery formats
- Agreed per order; git bundles or mirrors, or file-level JSONL or Parquet
- Preparation
- Secrets removed from all history; authors pseudonymized; third-party code removed or flagged
- Licensing
- Partner confirms ownership; third-party code stays under its own license terms
- Availability
- Usually scoped to selected repositories rather than a whole code estate; not guaranteed
What a delivery contains
Fields vary by source system and are fixed per order. A typical delivery includes:
| Field | Type | What it holds |
|---|---|---|
| repo_id | string | Pseudonymous repository identifier, used across files, commits and the manifest. |
| summary | string | What the software does, written without naming the partner, its products or its customers. |
| provenance | object | How the code came to be owned, such as built in-house, built by contractors under assignment, or acquired. |
| vcs | object | Version control system, conversion from any older system, first and last commit dates and the bundle reference. |
| languages | object | Share of code by language, including legacy and domain-specific languages. |
| build | object | Build systems, package managers, lockfiles, CI definitions and whether the code builds offline. |
| tests | object | Test frameworks, where tests live and how to run them. |
| docs | array | READMEs, architecture decision records, API specifications and other in-repo documentation. |
| commits | array | Full commit graph with parents, pseudonymous authors, timestamps, messages and diffs. |
| files | array | Per-file records with path, language, content or blob reference, and generated or vendored flags. |
| third_party | array | Detected open-source and commercial components with their license and the action taken. |
| secrets_scan | object | Scope of the secret scan, what was found by category and how it was replaced. |
| export_screen | object | Result of screening for encryption implementations and code built for controlled uses. |
| exclusions | array | Paths removed before delivery, such as customer data, licensed assets or client-owned modules, with the reason. |
Example record
{
"repo_id": "repo_8d02",
"summary": "Warehouse management backend and handheld scanner app",
"provenance": { "origin": "built_in_house", "contractor_code": "assigned_by_contract",
"acquired": false },
"vcs": { "type": "git", "converted_from": "subversion", "first_commit": "2011-04-18",
"last_commit": "2025-11-30", "bundle": "bundles/repo_8d02.bundle" },
"languages": { "java": 0.61, "kotlin": 0.17, "sql": 0.09, "typescript": 0.08, "other": 0.05 },
"build": { "systems": ["gradle", "npm"], "lockfiles": true, "ci": "jenkins",
"builds_offline": "partial",
"missing_dependencies": ["[INTERNAL_PACKAGE_REGISTRY]/scanner-sdk"] },
"tests": { "frameworks": ["junit5", "jest"], "runs_in_container": true },
"docs": ["README.md", "docs/adr/", "api/openapi.yaml"],
"third_party": [
{ "path": "vendor/barcode-decoder/", "license": "LGPL-2.1", "action": "removed_listed_in_manifest" },
{ "path": "src/main/java/[ORG]/util/StringUtils.java", "match": "public_snippet",
"license": "Apache-2.0", "action": "flagged" }
],
"secrets_scan": { "scope": "all_commits_and_tags",
"categories": ["api_key", "db_password", "private_key"],
"action": "replaced_with_placeholder", "history_rewritten": true },
"export_screen": { "crypto_implementations": false, "controlled_use": false, "status": "screened" },
"exclusions": [
{ "path": "db/fixtures/prod_snapshot.sql", "reason": "customer_personal_data" },
{ "path": "assets/fonts/", "reason": "commercial_license" }
],
"commits": [
{ "sha": "3c9e5a1", "parents": ["a07b2d4"], "author": "dev_0b3",
"committed_at": "2019-06-11T14:03:27Z",
"message": "Pick path: route around aisles flagged as blocked by scanners ([TICKET_ID])",
"diff_ref": "diffs/3c9e5a1.patch" }
],
"files": [
{ "path": "src/main/java/[ORG]/wms/picking/PickPathPlanner.java", "language": "java",
"generated": false, "vendored": false, "last_commit": "3c9e5a1",
"blob_ref": "blobs/9a/1f3e7c" }
]
}Synthetic record for illustration. Field names, structure and format are agreed per order.
What AI teams use it for
Train on real application code
Business logic, internal frameworks, database migrations and infrastructure configuration show how long-lived commercial systems are actually built and maintained.
Cover scarce languages and domains
Legacy and industry-specific code, such as COBOL, RPG, ABAP and PLC programs, is scarce in public repositories and mostly lives inside the companies that still run it.
Learn from real migrations
Full history captures framework upgrades, language-version moves and monolith splits as they were carried out, giving before-and-after pairs for modernization models.
Evaluate repository-level understanding
Questions that can only be answered by tracing calls across many files and internal libraries test long-context and navigation skills on code that was never published.
Ground code search and documentation
Code paired with its design records, API specifications and READMEs supports retrieval, code search and documentation generation inside a real codebase.
Use-case guides: Coding agents, Private evaluation sets
What makes this data valuable
Clear chain of title
Documented ownership for every contributor group and every acquired component.
Unbroken history
History carried over from older version control systems, with authors mapped consistently.
Builds and tests run
Lockfiles, CI definitions and a container that builds the code without internal services.
Mostly original code
A low share of vendored and copied third-party code relative to code the partner wrote.
In-repo design records
Architecture decision records and specs that explain why the code is shaped the way it is.
Still in production
Code that is still deployed and maintained, so recent history reflects current practice, not an abandoned project.
What private code adds to a training mix
Public repositories over-represent certain kinds of software: libraries, developer tools, tutorials, side projects and the early life of new projects. The code most companies actually run looks different. It is application code built around business rules, held together by internal frameworks and years of migrations, with configuration, database schemas and deployment scripts that never appear in a public repository. Full history adds the time dimension: why a module was split, which abstractions were abandoned, how a team moved off a deprecated framework and what broke along the way, recorded in commit messages written for colleagues rather than for the public.
That history is also a record of conventions. Naming, error handling, logging and test styles are enforced by the same team over many years, so a single codebase shows consistent engineering judgment in a way that a scrape of unrelated repositories cannot.
Why codebases are licensed piece by piece
Partners rarely license an entire code estate. A company's repositories usually include code it owns outright alongside client deliverables, acquired products with incomplete paperwork, vendored dependencies, modules subject to export controls and code that is simply too sensitive to share. Scoping therefore starts from the buyer's spec and narrows to specific repositories, products or time ranges whose ownership is documented.
Preparation then removes what should not travel, and records it. Secrets are purged from every commit, which rewrites history. Customer data in fixtures, commercial assets and third-party code are removed or flagged. Generated code, minified bundles, binaries and duplicated forks are excluded or marked, since they add volume without teaching much. Every removal is listed in the manifest, so you know, for example, that a build will need a stub for a vendor SDK that could not be delivered.
The result is narrower than the partner's full codebase but cleaner to license: each delivered file has a known owner, a known license position and a known history.
What to check before licensing
- Trace ownership of every part of the codebase. Employees' work usually belongs to the employer, contractors' code usually needs a written assignment, and acquired code depends on what the acquisition agreement transferred.
- Separate code built for clients. Agencies and consultancies often assign deliverables to the client and keep only their own background tools, so client-owned modules are excluded unless the client authorizes licensing.
- Run software composition analysis across the full history, not only the latest snapshot, to find vendored libraries, copied snippets and copyleft-licensed files.
- Look for commercial third-party code the partner cannot redistribute, such as vendor SDKs, licensed fonts and generated client libraries under proprietary terms.
- Screen for export-controlled code, such as encryption implementations and software built for defense or other controlled uses, before it crosses a border or is shared with foreign nationals.
- Check what has ever been public, such as open-sourced modules, published SDKs and code samples in public docs, since those parts may already be in training corpora.
- Search history and fixtures for secrets, customer data and employee personal data, and confirm the partner has rotated any live credentials found.
- Agree permitted use precisely, including training versus evaluation only, confidentiality, verbatim reproduction in model outputs and deletion at the end of the term.
How licensing works through SourceX
- 1
Define
Send the domain, modality, volume, format, timeline and permitted use you need.
- 2
Source
SourceX identifies businesses that hold matching data and are open to licensing it.
- 3
Qualify
Fit, rights and quality are checked, and you review samples before committing.
- 4
License
Scope, permitted use, exclusivity, price and obligations are agreed in writing.
- 5
Deliver
Approved data is prepared, de-identified where required and transferred securely.
Questions buyers ask
Who owns code written by contractors or for clients?
It depends on the contracts, which is why ownership is checked repository by repository. Code written by employees usually belongs to the employer. Contractors' code usually needs a written assignment, and agencies often assign deliverables to their clients while keeping background tools. SourceX asks the partner to confirm its rights to each repository in scope, and client-owned code is excluded unless the client authorizes licensing.
What happens to open-source code inside a proprietary repository?
It is identified and either removed or delivered under its own license. The partner can license only what it owns; open-source components remain under their original terms, and copyleft licenses such as the GPL attach conditions to redistribution. Public code may also already be in pretraining data, so flagging it per file lets you exclude it from evaluation sets.
Is encryption or defense-related code a problem?
It can be. Software that implements encryption, or was built for defense, aerospace or other controlled uses, may be subject to export controls such as the US Export Administration Regulations or ITAR. Delivering it to another country, and in some cases to foreign nationals, can require a classification or a license, so such code is screened during qualification and usually excluded or scoped separately.
Can the source company stay anonymous?
Only partly. Package namespaces, copyright headers, internal hostnames, product names and comments all point to the company. They can be rewritten, but code is hard to fully anonymize without breaking builds, so the license can include confidentiality obligations and a commitment not to identify or contact the partner.
Do I get the full version history or only a snapshot?
Full history where it exists, typically as git bundles or mirrors, with snapshots at chosen tags if you prefer. Older history may have been converted from Subversion, Perforce or another system, which can lose branch structure or author mapping. Squash-merge workflows leave one commit per change on the main branch, and removing secrets rewrites history, so commit hashes differ from the partner's originals.
Which languages and kinds of code can I request?
Any that partners hold and can license. Specify languages, frameworks, domains such as embedded, scientific, enterprise or data engineering, and the age and size of the repositories you want. Requests for legacy languages, industrial control code or in-house domain-specific languages depend on finding companies that still maintain such systems and are willing to license them, so they can take longer to match.
Related datasets
- Software engineering histories (issues, PRs, reviews)
Issues linked to commits, pull requests, code review, CI runs, deploys and incidents
- IT service management and incident histories
Incidents, problems, changes and requests with work notes, CI links and outcomes
- CAD and PCB engineering files with revision history
Native CAD and ECAD design files with revisions, change orders and BOMs
- Enterprise document archives
A company's working files with folders, versions and sharing metadata
Evaluating this data for procurement?
Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.
Request dataset diligenceTell us what your models need
Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.
Updated 3 October 2026. Own data like this? See how companies license it to AI developers.