Agent, workflow and domain-reasoning data
License terms for agent data: environments, replay, derived tasks and benchmarks
Quick answer
A license for agent workflow data needs more than a training grant. Agent builders also turn records into sandbox environments, re-execute recorded trajectories, generate synthetic task variants and publish benchmark scores, and each of those is a separate use that a "train machine learning models" clause may not cover. Before signing, name every one of these uses in the grant, define who owns the derivatives, and decide what survives termination in models, environments and task banks.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a training-only grant breaks agent programs
A training-only grant breaks agent programs because most of the value in workflow data is realized outside the gradient step. A typical pipeline loads ticket histories or ERP audit trails into a mock system, replays human action sequences against it to check that they still reach the recorded end state, mutates those tasks into thousands of variants for reinforcement learning, and holds some out as an evaluation set whose scores end up in a model card. If the license says only "use the Data to train models," the supplier can argue that the environment, the replayed episodes and the published scores are unlicensed copies or derivative works.
Dataset licensing is already error-prone without that extra complexity. An audit of more than 1,800 text datasets found that popular hosting sites omitted license information for over 70% of datasets and miscategorized licenses at error rates above 50% [1]. Agent data multiplies the problem, because one source record can end up as a seed row, a trajectory, a grader fixture and a leaderboard task at once. For the clause-by-clause baseline, see AI data licensing agreements: every clause explained and AI data license terms: use, exclusivity, deletion; this page covers only the grants that are specific to agents.
The six grants an agent data license should name
An agent data license should name six uses explicitly: training, environment construction, replay, derived task generation, evaluation, and publication of results. Each one has a distinct failure mode if left implied, and suppliers often agree to some and refuse others, so separating them makes negotiation easier.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Grant | What it covers | Typical failure if missing | Drafting question for counsel |
|---|---|---|---|
| Training | SFT, preference and RL updates on records and trajectories | Usually covered; disputes arise over RL "rollouts" generated from the data | Does "training" include on-policy RL that uses the data as environment state? |
| Environment construction | Loading records into mock CRMs, ITSM queues, spreadsheets or databases; building state snapshots | Environment treated as an unlicensed copy or a competing product | May we host the data inside a sandbox, and may contractors access it? |
| Replay | Re-executing recorded action sequences, screenshots and API calls against the environment | Replay logs treated as new copies outside the license | Are replayed episodes and their logs "Licensed Data" or our derived output? |
| Derived tasks and variants | Synthetic tasks, paraphrases, perturbed states and new goals generated from source records | Supplier claims ownership of or royalties on the task bank | Who owns derived tasks, and do they inherit use limits? |
| Evaluation | Internal held-out sets, regression suites, red-team scenarios | Eval use read as outside a "training" grant | Is evaluation a permitted use, including for third-party models? |
| Publication | Scores, aggregate statistics, example tasks in papers, model cards or leaderboards | Publishing an example task discloses confidential source content | What may be published: scores only, redacted examples, or full tasks? |
Treat the table as a checklist for your own draft, not a market standard. The terms suppliers accept vary by deal, data type and the sensitivity of the systems the records came from.
Environment construction and RL rights: seeding, cloning and state snapshots
Environment rights let you load licensed records into a running system that an agent can act on, and they should be granted expressly. Public research environments show what this involves: WebArena ships fully functional self-hosted websites in four domains [2], OSWorld provides real desktop environments that support task setup, execution-based evaluation and interactive learning [3], and CRMArena-Pro populates CRM object environments with synthetic enterprise data [4]. Licensed workflow data replaces or enriches that synthetic state, which means the data persists inside infrastructure for as long as the environment runs.
The grant should cover hosting the data in your sandbox, snapshotting and restoring state, sharing the environment with named contractors and evaluation vendors, and running unlimited RL episodes against it. It should also say whether an environment built from the data may be used to train models for third parties, which matters if you are an AI data or evaluation company passing rights down to lab customers. For the data side of this, see seed data and state snapshots for enterprise agent sandboxes and RL environments from business workflows.
A separate question is the software itself. Cloning the screens, schemas or behavior of a third-party SaaS product raises issues that a data license from the product's customer cannot settle, because that customer does not own the vendor's interface; rights questions when replicating third-party software covers that boundary.
Replay rights: re-executing recorded trajectories
Replay rights cover re-running a recorded sequence of actions against an environment and keeping the resulting logs, and they matter because replay creates new artifacts that contain source content. A replayed computer-use episode produces fresh screenshots, accessibility trees, DOM snapshots and tool-call outputs, each of which can reproduce customer names, ticket text or account numbers from the original record. If the license is silent, those logs sit in an unclear category between licensed data and your own output.
Draft the replay clause to say that replayed episodes, grader outputs and failure traces are derived materials subject to the same confidentiality and use limits as the source, and that you may retain them for debugging and regression testing. If the source data was de-identified, require that replay environments use only the de-identified version, so that a replay cannot reintroduce original identifiers through cached state. For the step-record fields that replay depends on, see computer-use trajectory data: what every step record must contain.
Derived synthetic tasks and variants: who owns the task bank
Derived-task rights decide whether you own the synthetic tasks you generate from licensed records, and the default answer should be written down rather than inferred. Methods such as OS-Genesis derive new tasks and trajectories after the fact from GUI interactions, through reverse task synthesis [5], and teams routinely paraphrase goals, perturb starting states and swap entities to turn one licensed workflow into hundreds of training tasks. If the supplier later claims those tasks are derivative works, your task bank becomes encumbered.
Three positions are common in drafts. The buyer owns all derived tasks outright; the buyer owns them but they inherit the source license's use limits and confidentiality; or derived tasks that reproduce source content verbatim remain licensed data while abstracted tasks belong to the buyer. Whichever you choose, define a test for "reproduces source content" (for example, verbatim spans above a set length, or presence of any original identifier) so the line can be audited. The trade-offs between licensed, commissioned and generated data are covered in licensed workflow records, commissioned demonstrations or synthetic trajectories.
Benchmark and result publication: scores, examples and contamination
Publication rights decide what you may say publicly about results produced with licensed data, and most disputes come from example tasks rather than scores. Benchmarks such as tau-bench pair simulated users and APIs with domain policy documents and databases [6], and a paper or model card built on a private equivalent will want to show sample tasks, policies and transcripts. Each shown example can disclose a supplier's internal procedures or customer content.
Separate three tiers in the clause: aggregate scores and statistics (usually acceptable), redacted or paraphrased examples (often acceptable with supplier review), and release of the task set itself (rarely acceptable for operational data). Add a non-release obligation for held-out tasks, because public exposure ends their value as an evaluation set. SWE-bench Verified illustrates the lifecycle: it was built as a 500-sample human-verified subset to fix underspecified tasks and unreliable environment setup [7], and as of October 2026 OpenAI has stopped reporting it, citing contamination [8]. Private evaluation sets drawn from licensed data are useful precisely because they are not public, as discussed in agent evaluation task suites.
Termination: what survives in models, environments and task banks
Termination clauses for agent data must address four artifact types separately: raw records, trained model weights, running environments, and derived task banks. A deletion obligation drafted for raw files can unintentionally require you to tear down a sandbox, purge a regression suite or, in the worst reading, retrain a model.
Common buyer asks are that trained weights survive termination, that derived tasks meeting the "abstracted" test survive, that environments containing source records are destroyed or reseeded with synthetic state within a stated period, and that certification of deletion covers backups and replay caches. Where personal data is involved, the European Data Protection Board's Opinion 28/2024 addresses when a model can be considered anonymous and how unlawful processing during development can affect later use of the model [12][13], so model survival is not purely a contract question. The ownership and consent issues around your own production traces are covered in using agent deployment logs for training.
Disclosure duties that the license must support
Your license must let you meet training-data disclosure duties, which apply to environments and derived tasks as much as to raw training sets. California AB 2013 (Civil Code Section 3111) requires developers of generative AI systems made available to Californians to post documentation about training data, covering systems released since January 1, 2022, with postings due on or before January 1, 2026 [11]. Under the EU AI Act, Article 53 obligations for general-purpose AI model providers include a copyright compliance policy and a public summary of training content [9], for which the Commission published a template on 24 July 2025 [10].
A confidentiality clause that forbids naming the data's source type, time range or domain can collide with those duties. Negotiate a carve-out permitting disclosures required by law at the level of detail those templates require, without naming the supplier unless the law demands it.
Agent data license rider: a working checklist
Use this rider as a drafting aid alongside your main agreement. It lists the points that agent programs most often leave undefined.
Illustrative example: invented to show structure; it does not describe an available dataset.
agent_data_rider:
licensed_data: "Ticket and case histories, 2022-2025, de-identified; fields listed in Schedule A"
permitted_uses:
training: [sft, preference, rl_on_policy, rl_offline]
environment_construction: { hosted_by: licensee, contractors_allowed: named_only }
replay: { logs_status: derived_materials, retention: "term + debugging period" }
derived_tasks: { owner: licensee, verbatim_test: "no original identifiers; no span > N tokens" }
evaluation: { internal: true, third_party_models: "requires consent" }
publication: { scores: true, redacted_examples: "supplier review", task_release: false }
survives_termination: [model_weights, abstracted_derived_tasks]
destroyed_on_termination: [raw_records, seeded_environments, replay_caches]
regulatory_disclosure_carve_out: [ca_ab_2013, eu_ai_act_art_53]
downstream_customers: "may receive models; may not receive raw records or environments"
Questions to put to a supplier before signing
Ask questions that test whether the supplier can actually grant what you need, not only whether it will. Ownership and consents come first: confirm that the supplying company holds the records and that its customer contracts and privacy notices permit licensing for AI development.
- Which systems did the records come from (Zendesk, ServiceNow, Salesforce, SAP), and does any vendor contract restrict export or secondary use?
- How were personal details removed or replaced, and was a sample checked after transformation?
- Are any fields subject to sector rules, such as health data that would need HIPAA de-identification?
- Will the supplier review example tasks before publication, and within what process?
- Can the supplier approve environment hosting by your named evaluation contractors?
- Does the supplier have its own deletion obligations to upstream customers that would flow down to you?
If you are still defining the data itself, start from writing an agent data specification and the category overview on enterprise workflow datasets and agent trajectories, then return to the license terms. Buyers who want a sourcing partner to run this process can describe the workflow data they need on the SourceX buyers page.
Licensing workflow data for agent training with SourceX
SourceX sources operational datasets, including support and sales histories, engineering records and finance and legal workflows, from US companies on request, and manages the process from assessing data and licensing permissions through agreeing pricing and allowed uses in a license. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, and nothing is contracted until the supplying company agrees. To start, describe the agent data and uses you need to license.
For more on this cluster, see the AI agent training data hub and the AI data hub.
Sources
- Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Zhou, Xu et al., Carnegie Mellon University (arXiv), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
- Xie et al., XLANG Lab (arXiv), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- Salesforce AI Research (arXiv), "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions" (2025). https://arxiv.org/pdf/2505.18878
- arXiv, "OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis" (2024). https://arxiv.org/pdf/2412.19723
- Sierra Research (arXiv), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- CMS (summary of EDPB Opinion 28/2024), "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.