Skip to content

Software companies

Data license vs software license: why licensing code for AI training is different

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

A data license for AI training lets a developer use your code as training or evaluation material, not run, resell or redistribute it, which is what EULAs and source code licenses govern. The grant is usually non-exclusive, limited in term and field of use, and specific about trained models and outputs. Prove ownership and scrub secrets before signing.

Key takeaways

  • An AI training license grants use of code as material to learn from, not as software to deploy.
  • Ownership comes first: employee and contractor assignments, acquired code and third-party code all need checking.
  • Most code training licenses are non-exclusive; exclusivity limits future deals and needs a clear reason.
  • Secrets and customer data must be removed from the full commit history, not just from the latest branch.
  • Terms on trained models and outputs matter more than in software licenses, because models outlive the data license.

How do the four license types compare?#

The four license types differ in what they grant, which use they permit, how long they last and how they treat what the licensee produces. A EULA lets a customer run your software; an AI training license lets a developer learn from your code.

The last row is the one most teams have never negotiated. It looks like a software license on paper but behaves like a data license, which changes which clauses matter.

How do the four license types compare?
LicenseWhat is grantedPermitted useTypical termOutputs
EULA or SaaS termsRight to use compiled software or a hosted serviceRun the product for internal business purposesSubscription or perpetualCustomer owns its business outputs
Source code licenseAccess to source, often for a customer or through escrowMaintain, modify or support the licensed productTied to the commercial relationshipModifications often owned by or licensed back to the vendor
Open source licenseBroad public rights to use, modify and redistributeAny use within the license conditionsGenerally irrevocable for distributed copiesDerivative works may carry copyleft obligations
AI training data licenseRight to use code and its history as training or evaluation dataTrain or test models; no redistribution or deployment of the codeFixed term, often with deletion dutiesTrained models and outputs addressed by specific clauses

What does an AI training license actually grant?#

An AI training license grants the developer the right to copy, process and learn from the licensed material for a defined purpose, usually model training, fine-tuning or evaluation. It does not give the right to ship your code as a product, publish it or sublicense it to others.

Deletion at term end usually applies to the licensed material and working copies, while models trained during the term often survive. Agree on that point explicitly, because it is the clause most likely to be misunderstood later.

  • Field of use: which kinds of models or products the material may be used for.
  • Copies and derived datasets: which processing copies the licensee may make, and where they may be stored.
  • Restrictions: no redistribution, no deployment of the code as software, no attempts to identify people in the data.
  • Term and deletion: when the license ends and what must be deleted, separating raw data from trained models.
  • Outputs: whether the licensee must take steps against models reproducing licensed code verbatim.
  • Confidentiality and audit: how the licensee protects the material and how compliance can be checked.

Why is ownership the first question?#

Ownership is the first question because you can license only what you own or have the right to sublicense. A software company's repository usually mixes code written by employees, contractors, acquired teams and outside projects, and each source has a different chain of title.

Gaps can often be closed with confirmatory assignments or by excluding affected folders. This is general information, not legal advice; counsel should review the chain of title before any grant.

Why is ownership the first question?
Code sourceWhat to checkCommon gap
EmployeesInvention assignment agreementsEarly hires who joined before agreements were standard
Contractors and agenciesWritten IP assignment in the statement of workAssignment missing, or only a license granted
Acquired codePurchase agreement and IP schedulesAssets transferred without full assignment records
Customer-funded developmentOwnership terms in the development agreementCustomer owns custom modules
Open source and vendored codeInbound licenses and noticesCopyleft code mixed with proprietary files
AI-assisted codeTool terms and internal policyUnclear provenance of generated sections

What must be removed from code before licensing?#

Secrets, customer data and personal details must be removed from the full commit history before code is licensed, not only from the latest version. A key deleted from the main branch years ago still sits in older commits.

Scanners help, and their status matters. Gitleaks, an MIT-licensed tool, detects passwords, API keys and tokens in git repositories; its README was updated in May 2026 to say it is feature complete and future releases will be security patches only. TruffleHog, an AGPL-3.0 scanner, also covers chats, wikis and logs and can test whether a secret it classifies is live, which calls for care when scanning.

Beyond secrets, check fixtures, seed files and test data for copies of customer records, commit metadata for personal emails, and comments for internal hostnames or customer names. Rotate any credential found, even one that appears expired.

Should a code training license be exclusive?#

A code training license is usually non-exclusive, and exclusivity should require a specific reason. Non-exclusive terms let you license the same history to other developers, reuse it yourself and publish parts later if you choose.

When a licensee insists, exclusivity can be narrowed to a field of use, a time period or a type of model rather than all uses. Any exclusivity will surface in diligence if the company is sold, so record it clearly and keep the clause short.

Who needs to approve a code license?#

A code license usually needs approval from the CEO and an authorized signer, and sometimes from the board, investors or lenders. Investor agreements may require consent for licensing material IP outside the ordinary course, and credit agreements may treat IP as collateral.

Check these before a term sheet is discussed, not after. A consent requirement found late can delay signing or force changes to scope.

  • Board: whether the license falls outside the ordinary course of business.
  • Investors: protective provisions covering IP transfers or exclusive licenses.
  • Lenders: security interests in IP and covenants on asset dispositions.
  • Customers: development agreements that give them rights in specific modules.

Illustrative: a restaurant point-of-sale software company#

Illustrative: a fictional software company builds point-of-sale and kitchen display software for restaurants. Before a proposed AI training license, its CTO reviews the repository and finds code from an early agency, a vendored payment terminal SDK and test fixtures containing real menu and order data from a pilot customer.

Counsel obtains a confirmatory assignment from the agency. The payment terminal SDK is excluded because it belongs to the hardware vendor, and the fixtures are replaced with synthetic data. The team scans the full history, rotates the old keys it finds and pseudonymizes commit authors.

The final grant is non-exclusive, limited to training and evaluating coding models, prohibits redistribution and requires deletion of the licensed material at term end. It also preserves the company's right to publish parts of the code later.

How SourceX structures code licenses#

SourceX structures code licenses as data licenses: licensed, not sold, with the company keeping ownership. In the Rights step of the SourceX five-step transaction, ownership, third-party code and customer material are reviewed before anything is prepared for delivery.

The SourceX Evidence Packet records provenance, licensing rights, permitted use, the privacy record and release authorization for the package. Large repositories stay in the company's own storage or ship on encrypted drives.

Frequently asked questions

Can a licensee ship our code inside its own product?

Not under a typical AI training license. The grant covers using the material to train or evaluate models, and redistribution or deployment of the code is expressly excluded. If a licensee wants to run or ship your code, that is a software license with different terms and should be negotiated separately.

Can we license code that customers already hold under a source code license?

Usually yes, if their license is non-exclusive and does not restrict your own use. Check for exclusivity, confidentiality obligations and customer ownership of modules built for them. Customer-specific code is often excluded from the training package to keep the scope clean.

What happens to trained models when the license ends?

That depends on the contract. Many licenses require deletion of the licensed material and working copies while allowing models trained during the term to remain in use. Some add limits on further training. Agree on this clause explicitly rather than relying on general deletion language.

Should the commit history be part of the package?

For coding-agent developers, usually yes. Commits, pull requests, review comments and linked issues show how code changed and why, which is often more useful than a final snapshot. Including history also increases the scanning and de-identification work, so plan for it.

Do outside contributors to our public repositories affect a license?

They can. Contributions accepted under a contributor license agreement may be licensable, while contributions without one remain under the project's open source license. Separate community-contributed code from proprietary code and review each source before including it.

Sources

  • Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin. Source
  • On May 21, 2026, the gitleaks README was updated to state that Gitleaks is feature complete, that future releases will be security patches only, and that the maintainer is shifting focus to Betterleaks. Source
  • TruffleHog, an AGPL-3.0 open-source secret scanner, can log in to confirm whether a secret it classifies is live, and scans sources including Git, chats, wikis, logs, object stores and filesystems. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify