Software companies
Can you license your engineers' AI coding assistant logs?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
AI coding assistant logs can sometimes be licensed, but only the parts your company controls and your AI vendor's terms allow. Engineers' prompts, your code and test results are usually company records; model outputs may carry vendor restrictions, often including limits on building competing models. Read those terms before scoping anything.
Key takeaways
- Split a coding session log into prompts, context, model replies, tool calls, diffs and test results before judging rights.
- AI vendor terms, not only employment agreements, decide what you can do with model outputs.
- Session logs often capture credentials, customer records and proprietary code, so they need scanning and review before any license.
- Where logs exist depends on tool configuration and retention settings, and many teams hold less than they assume.
What does a coding assistant session log contain?#
A coding assistant session log is the record of an engineer and an AI tool working on a task together: the request, the files and terminal output the tool read, the model's plan and replies, the edits and commands it ran, and whether the engineer kept, changed or rejected the result. In agent tools, these step-by-step records are often called trajectories.
Those records live in more places than most CTOs expect. Some tools keep session files on each developer's machine, enterprise plans may log activity in an admin console, companies that route traffic through an internal gateway hold their own copy, and the final result lands in commits and pull requests. Retention also differs by access path. GitHub's Copilot Trust Center FAQ, for example, says that for Business and Enterprise customers IDE chat and code-completion inputs and outputs are not retained by default, while Copilot coding agent session logs are kept for the life of the account. Check the current vendor documentation for each tool rather than assuming one rule covers every log.
The value of a log comes from linkage. A session that starts from a Jira issue and ends in a reviewed, merged pull request with passing tests shows a real task, real steps and a verified outcome. A pile of chat transcripts with no link to what shipped shows much less.
Who controls each part of the log?#
Control of a coding session log splits by part, because each part came from a different source. Your engineers wrote the prompts, your repository supplied the context, the vendor's model produced the replies, and your systems produced the tool output.
The pattern is simple to state: your side of the session and your code are easier to clear, and the model's side depends on terms you did not write.
| Log part | Typical starting position | What to check |
|---|---|---|
| Engineer prompts and task descriptions | Company work product under employment and contractor agreements | Contractor IP assignment; personal remarks |
| Repository context read by the tool | Your code, plus open source and third-party files | Open source licenses, customer-owned code, vendored SDKs |
| Model replies and generated code | Often assigned to the customer, with use restrictions | Output ownership and competing-model clauses |
| Tool calls and terminal output | Output of your own systems | Credentials, customer records, internal hostnames |
| Merged diffs, commits and test results | Company code once accepted | Same review as any code license |
| Accept, edit and reject signals | Your telemetry; the vendor may hold a copy | Where it is stored and the vendor's own rights |
What do AI vendor terms say about outputs?#
AI vendor terms decide what you can do with model replies, even when they say you own them. Commercial terms from AI providers commonly address two separate points: who owns outputs, and what the customer may not use outputs for. Restrictions on using outputs to develop competing models are a common example, and licensing logs to an AI developer for training is exactly that kind of use.
Read the business terms that actually applied, not the consumer version, and check whether they changed during the period the logs cover. Terms incorporated by link can be updated, so keep dated copies. If engineers used personal accounts, the company may have no rights in those sessions at all, because the contract was between the vendor and the individual.
Also check what the vendor may do with the same sessions. GitHub's Terms of Service, for instance, grant GitHub a license to use AI-feature inputs and outputs to train models, with an opt-out in account settings, but exclude customers under a GitHub Customer Agreement or volume licensing agreement. Which contract applied affects any exclusivity you could offer, so confirm it with counsel against the live text.
The decision rule: if the terms restrict using outputs for model development, either exclude model-generated turns or get written permission from the vendor before including them.
Which parts of a session can realistically be licensed?#
The realistically licensable part of a session is usually the human side plus the verified result. A task description paired with the merged code, the review discussion and the test outcome is still useful material, even with the model's turns removed.
This is general information, not legal advice. Employment agreements, vendor terms and customer contracts differ, so scope is assessed deal by deal with counsel.
- Usually in scope after review: engineer-written prompts and task descriptions, merged diffs, review comments on the resulting pull request, CI and test results, and linked issue history.
- Conditional: model-generated explanations and code, depending on vendor terms or written permission.
- Usually out of scope: sessions from personal accounts, sessions that touched customer-owned code or customer data, and sessions with credentials that cannot be cleanly removed.
- Out until reviewed: sessions in repositories covered by client contracts or restrictive third-party source licenses.
What has to be cleaned before a log leaves your control?#
Coding session logs need heavier cleaning than source code, because terminal output and debugging steps capture whatever was on screen. Environment variables, connection strings, API tokens and sample customer records show up in tool output far more often than in committed code.
Scan every log with a secret scanner and treat any hit as a credential to rotate, not just text to redact. TruffleHog, an open-source scanner, says it scans sources including Git, chats, wikis, logs and filesystems, and that it can check whether a found secret is still live. That check sends real login attempts, so run it with care.
Then remove personal data such as names and email addresses in commit metadata, customer records in fixtures or debug output, internal hostnames and IP addresses, and remarks engineers made about colleagues or customers. Tell engineers how session logs may be used before the review starts.
Illustrative: a route-planning software company audits its agent logs#
Illustrative: a fictional route-planning software company for regional carriers wants to know whether its coding agent sessions could form part of a licensing package. The CTO finds three sources: an enterprise plan whose activity is logged through an internal gateway, local session files on developer laptops, and a team that used personal accounts during an early trial.
Counsel reads the enterprise terms and finds a restriction on using outputs to build competing models. The company excludes the personal-account sessions entirely, removes model turns from the rest, and keeps the engineer's task, the merged diff, the review thread, the CI result and the linked Jira issue. A secret scan finds live staging keys in terminal output, which are rotated before anything else happens.
The result is narrower than the CTO first imagined but clean enough to describe in a fit check, and the company now requires company accounts for every AI tool its engineers use.
How SourceX scopes coding assistant logs#
SourceX scopes coding assistant logs as one record family within the Supply step of the SourceX five-step transaction, alongside issues, code reviews and CI history. The first fit check asks which tools were used, on which plans, and where logs are stored; no logs are shared at that stage.
In the Rights step, AI vendor terms are read together with employment agreements and customer contracts, and each log part is marked in, out or conditional. Preparation covers secret scanning and personal data removal. The SourceX Evidence Packet then lists, for exactly the parts the company approves, their provenance, the licensing rights and permitted use, the privacy record and the release authorization. SourceX's own rights in a deidentified dataset are set out in the signed supplier agreement.
Frequently asked questions
Do engineers need to agree before their sessions are licensed?
Employment agreements usually assign work product to the company, but that is not the whole question. Some state privacy laws may apply to employee information, and logs can contain personal remarks. Most companies tell engineers how session logs may be used, exclude personal content, and review the plan with counsel before scoping.
Does a zero-retention setting mean we have no logs?
Not necessarily. Zero retention usually describes what the vendor keeps, not what exists on your side. Local session files, gateway logs, CI records and the resulting commits and pull requests may still hold most of the session. Check each location before concluding the history is gone.
Are agent trajectories more useful than plain code?
They show different things. Plain code shows the result; a trajectory shows the steps, tool calls, failed attempts and corrections that led to it. AI developers building coding agents look for real tasks with verified outcomes, so a trajectory linked to a merged, tested change is the stronger record. Value is known only once a buyer engages.
Can we license logs from an assistant we no longer use?
Possibly. The terms in effect when the sessions were created usually govern, and some clauses survive termination. Check whether the vendor deleted its copies and whether your local or gateway copies are complete, and keep dated copies of the terms with the logs.
What about open source code that appears in a session?
Open source code keeps its license wherever it appears, including inside a log. Note which components appear, keep their notices, and exclude files under licenses your counsel considers incompatible with the planned use.
Sources
- TruffleHog, an AGPL-3.0 open-source secret scanner from Truffle Security, says it can log in to confirm whether a classified secret is live, and it scans sources including Git, chats, wikis, logs, object stores and filesystems. Source
- For Copilot Business and Enterprise, IDE chat and code-completion inputs and outputs are not retained by default, and Copilot coding agent session logs are retained for the life of the account. Source
- GitHub's Terms of Service (Section J) grant GitHub a license to use AI-feature inputs and outputs to train AI models, with an opt-out in account settings; customers under a GitHub Customer Agreement or volume licensing agreement are excluded. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.