Software companies
What permitted uses should a code license allow: training, evaluation or RL environments?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
A code license for AI should name each permitted use separately, because training, evaluation and reinforcement learning environments carry different exposure. Evaluation is usually the narrowest grant, training the broadest because what a model learns cannot be deleted, and RL environments sit between, since agents run the code as tasks. Grant only uses you have reviewed.
Key takeaways
- A grant for any AI purpose is too broad; list training, evaluation and RL environment use separately.
- Training is the hardest use to reverse, because deletion clauses reach copies of the code but not what a model learned.
- Evaluation sets lose their value if published, so confidentiality and no-release terms matter most there.
- RL environments execute your code, so dependencies, build scripts and any live credentials need extra review.
- Exclusivity can be granted per use, per field and per period rather than across every use.
Why should a code license split permitted use by purpose?#
A code license should split permitted use by purpose because training, evaluation and RL environment work do different things to the licensed code and leave different traces. A single grant for any AI purpose hands the licensee the widest of those uses by default.
The distinction matters most at the end of the term. Copies of code can be deleted and the deletion certified. What a model learned during training cannot be removed the same way, so a training grant behaves closer to a permanent grant than its stated term suggests.
Naming uses separately helps the licensee too. Its own governance teams need to know which models may contain the code and which may not, and a clear use schedule answers that question without a dispute over what any AI purpose was meant to cover.
Training, evaluation and RL environments compared#
The main uses differ in what the licensee does with the code, what the supplier is exposed to, and which restrictions are typical. The table adds two adjacent uses that often appear in the same term sheet.
Value tends to track exposure. Training rights are often the most valuable to a licensee and the hardest to unwind, while evaluation-only rights are narrower and easier to contain.
| Use | What the licensee does | Main exposure for the supplier | Typical restrictions |
|---|---|---|---|
| Pretraining or fine-tuning | Feeds code and history into model training | Patterns persist in model weights after the term ends | Reproduction safeguards, model scope limits, no resale of the dataset |
| Evaluation | Holds code and tasks out as a private test set | A leak would expose the code and make the set useless | No training on the set, no publication, access limited to named teams |
| RL environments | Turns repositories and issues into executable tasks, with tests as rewards | Code runs repeatedly, so secrets and dependencies matter | Sandboxed execution, no network access to supplier systems, deletion of environments at term end |
| Synthetic data generation | Uses code as seed material for new examples | Derived data may outlive the license | Ownership and deletion terms for derived data |
| Internal research only | Exploratory analysis without production models | Lowest, if work stays internal | No deployment, publication only with consent |
Which clauses keep each use narrow?#
Narrow use is built from several clauses working together; the permitted use definition alone is not enough. A tight definition paired with loose derived-data or affiliate terms still leaks.
Read the clauses as a set. If the definition says evaluation only but the derived-data clause lets the licensee keep everything it generates, the licensee can train on generated material that mirrors your code.
- Permitted use definition that names each use and excludes everything else.
- Model scope: internal models only or publicly released models, and whether fine-tuned variants count.
- Derived data: who owns synthetic examples, labels and environments built from the code, and whether they must be deleted.
- Confidentiality and no-release terms for evaluation sets and environments.
- Reproduction safeguards: reasonable measures to stop a model from emitting the licensed code verbatim.
- Affiliates and sublicensing: whether group companies or contractors may use the code.
- Deletion and certification at term end, with plain language on what deletion cannot reach.
- Audit or attestation rights confirming use stayed within scope.
How does each use change the rights review?#
Each permitted use changes the rights review because it changes how the code is copied, run and retained. Ownership, open-source licenses and secrets are checked for every use, but the depth differs.
Training raises the bar on open-source review, since copyleft components in the training set may carry their own conditions. Evaluation calls for the same ownership checks but usually involves smaller, curated sets that are easier to review by hand.
RL environments add execution. Build scripts, dependency manifests and test suites become part of the deliverable, and a live credential in the history could be used, not merely read. TruffleHog, for example, says it can log in to confirm whether a detected secret is still live, which is the kind of check that belongs before environments leave the company.
How should exclusivity and term be set per use?#
Exclusivity in a code license can be granted per use, per field and per period rather than across everything. A supplier might grant one licensee exclusive evaluation rights in a narrow domain for a set period while keeping training rights non-exclusive.
Each counterproposal in the table trades a little of the licensee's flexibility for a lot of the supplier's control, which is usually an easier negotiation than a flat refusal.
| If the licensee asks for | Consider instead | Why |
|---|---|---|
| Exclusive rights for all AI uses | Exclusivity for one use or field | Keeps other licensing options open |
| Perpetual training rights | A fixed term with clear survival language | Makes continuing obligations visible to future acquirers |
| Rights for all affiliates | Named entities, with consent for others | Keeps track of who holds the code |
| Use in publicly released models | Internal models first, release by amendment | Lets the supplier review once uses are clear |
Illustrative: a routing software company splits its grant#
Illustrative: a fictional route optimization software company that serves regional distributors holds many years of monorepo history, Jira issues and review threads. A model developer asks for rights to use all of it for any AI purpose.
The general counsel splits the request. Evaluation and RL environment use are granted for the repositories holding the dispatch and integration services, with sandboxed execution and deletion of environments at term end. Training rights are granted non-exclusively for the same repositories but exclude the core routing engine, which the company treats as a trade secret.
Before delivery, the CTO runs secret scanning across the full history, rotates old credentials found in deployment scripts, and removes a vendored library under a copyleft license. The final permitted use schedule names each use, each repository and the model scope in plain terms.
How SourceX documents permitted use#
In the SourceX Evidence Packet, permitted use sits beside provenance, licensing rights, the privacy record and release authorization as one of five recorded elements. Each use, repository and restriction is written down before Approval, so the supplier signs off on exactly what the licensee may do.
Industry standards point the same way. The Data & Trust Alliance Data Provenance Standards include metadata elements for license to use and intended data use in their Use group, which reflects a wider expectation that datasets carry their permitted purpose with them.
Frequently asked questions
Is an evaluation-only license worth offering?
It can be, because private evaluation sets are hard for model developers to build from public code that may already sit in training data. Evaluation-only rights also keep exposure lower. Whether it is worthwhile depends on the licensee's need and the code's distinctiveness, which become clear only once a buyer engages.
Can we revoke training rights if the licensee breaches the contract?
You can usually terminate the license and require deletion of copies, but a model already trained on the code cannot simply forget it. That is why breach remedies, model scope limits and reproduction safeguards matter more for training grants than for other uses.
Does an RL environment license include our tests?
Usually yes, because tests often serve as the reward signal. Tests can encode business rules and fixtures copied from production, so review them as carefully as the code. State in the license whether tests, fixtures and build scripts are included.
Should a code license for AI mirror open-source license terms?
No. Open-source licenses address distribution and modification of software, not model training or evaluation. A data license for code should define AI-specific permitted uses directly, and counsel should check how any open-source components in the set interact with those terms.
Who should negotiate the permitted use schedule on our side?
Usually the general counsel or outside counsel, working with the CTO. Counsel drafts the clauses, but the CTO knows which repositories hold trade secrets, which contain vendored code, and which would be risky to execute. Schedules drafted without engineering input tend to grant either too much or too little.
Sources
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.