Skip to content

Agent, workflow and domain-reasoning data

Using agent deployment logs for training: ownership, consent and contract terms

Quick answer

Agent interaction logs usually do not have a single owner. The customer that deploys the agent normally controls the content inside them (prompts, retrieved documents, tool outputs, end-user messages), while the vendor may hold rights to service and usage telemetry, but only as the contract defines it. A vendor can train on customer logs when the agreement grants a specific model-improvement right, privacy law allows the processing, and the vendor's public commitments match. Without that grant, assume no training right.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What an agent deployment log actually contains

An agent log is a bundle of different data classes with different owners, so ownership has to be decided field by field. A single trace from a support or back-office agent typically mixes the customer's confidential business records, personal data about the customer's own end users and employees, the vendor's prompts and tool schemas, and machine telemetry. Hosted platforms make this concrete: Google's Dialogflow documentation warns that user information may reside in an agent's interaction logs stored by the platform, and that logged data persists until the agent or conversations are deleted [8].

Attribution fields are not ownership fields. Google Workspace audit events can carry an agentOwner value identifying who owns the agent that took an action [9], which answers "who is accountable for this action," not "who may reuse this record." Product teams often conflate the two when they read their own observability schema.

A useful split for contract drafting:

  • Customer content: user turns, uploaded files, retrieved passages from the customer's knowledge base, CRM or ticket records read by tools, tool responses, and final agent outputs.
  • Human review data: approvals, rejections, edits, escalation notes and QA scores written by the customer's reviewers. These are employee-authored records about customer content (see employee-authored records in training data).
  • Vendor materials: system prompts, tool definitions, routing policies and model weights.
  • Service and usage data: latency, token counts, error codes, tool-call success flags, model version and session metadata, ideally stripped of content.

Who owns AI agent interaction logs by default

By default the deploying customer controls its content and the vendor controls its own materials; anything else depends on the contract. Most enterprise SaaS and AI agreements say the customer retains all rights in customer data and grant the vendor a limited license to process it to provide the service. The same pattern applies to SaaS tools generally, as covered in who owns data in a SaaS tool.

Under the GDPR, the vendor running a deployed agent is usually a processor that may act only on the controller's documented instructions under Article 28(3)(a) [3]. A processor that decides on its own to use customer logs to train a general model is determining purposes and means, and Article 28(10) treats it as a controller for that processing [3]. That shift brings its own lawful-basis, transparency and data-subject-rights obligations.

Under the CCPA, a "service provider" processes personal information on behalf of a business under a written contract that restricts other uses [6]. The CCPA regulations have allowed a service provider to use personal information internally to build or improve the quality of its services, provided it does not use that information to perform services for another person, according to law-firm analysis of the regulation text [7]. Whether training a cross-customer foundation model counts as "improving the service" is exactly where counsel disagrees, so write the right down rather than relying on the default.

When a vendor can train on customer agent logs

A vendor can train on customer agent logs only when four conditions hold together: a contractual grant, a lawful basis, matching public commitments, and a preparation method the customer accepted. Missing any one of them is the common failure.

  1. Contractual grant. The customer agreement or DPA must name model training or model improvement as a permitted purpose, with scope (which models, which data classes) stated.
  2. Lawful basis for end-user data. The customer's end users never signed the vendor's terms. In the EU, EDPB Opinion 28/2024 sets out how legitimate interest may be assessed for AI development and warns that unlawfully processed training data can affect later use of the model [5].
  3. Consistent promises. In 2024 staff posts, the FTC said that model-as-a-service companies that break promises not to use customer data for undisclosed purposes, such as training, may face liability [1]. Staff also wrote that quietly or retroactively changing terms to permit AI training may be unfair or deceptive [2].
  4. Accepted preparation. If the vendor relies on de-identification, the method should meet the applicable standard. The CCPA's definition of deidentified information requires reasonable measures plus a public commitment not to re-identify and contractual flow-down to recipients [6]. GDPR Recital 26 only takes anonymous information out of scope when identification is no longer reasonably likely [4].

Human review logs from agent deployments

Human review logs are the most valuable part of a deployment for training and the most legally exposed. Approve/reject labels, reviewer edits to agent drafts and escalation rationales are close to preference and correction data, which is why vendors want them. They are also authored by the customer's employees, often contain free-text notes about named end users, and can reveal internal policy.

Treat reviewer identity as personal data and the review policy as customer confidential information. Before using review data, confirm whether the customer's workforce notices cover AI training, pseudonymize reviewer IDs, and keep the policy document that defined "correct" with the labels. A label without its policy is hard to use for evaluation in the policy-following style of benchmarks such as tau-bench, which pairs tasks with domain policy documents and database state [11]. The policy-following service agent data and agent-to-human handoff data pages cover how those pairings are structured.

Agent pilot data rights

Pilot and proof-of-concept agreements are where training rights are most often lost or granted by accident. Pilots run on short order forms or click-through evaluation terms that rarely mention model improvement, so the default "customer owns its data, vendor processes to provide the service" applies. When the pilot converts, the vendor discovers it cannot use the most informative traces it collected.

Settle three points before go-live: whether pilot logs may be retained after the pilot ends, whether they may be used for evaluation only or also for training, and what happens to derived artifacts (fine-tuned adapters, reward models, eval sets) if the customer does not convert. The license terms for agent workflow data page covers derived tasks and benchmarks in more depth.

Model-improvement clause checklist

A workable model-improvement clause is specific about data classes, purposes, preparation and exit. Use the checklist below as a review aid when drafting or negotiating.

Illustrative example: invented to show structure; it does not describe an available dataset.

Clause elementWhat to specifyFailure mode if missing
Opt-in mechanismSeparate signature, order-form checkbox or admin toggle; default offTraining right buried in general license; FTC deception risk [1][2]
Data classes coveredContent, review data, telemetry listed separately"Usage data" read to include full transcripts
Purpose scopeCustomer-specific model, vendor's shared model, or evaluation onlyCustomer's data improves competitors' agents
PreparationRedaction or de-identification standard, sample check, record of methodNames and account numbers in training corpus
Third-party dataCustomer warrants notices or lawful basis for end users and reviewersEnd users never told; no lawful basis
Regulated dataExclude PHI, payment card data and children's data unless separately addressedHIPAA, PCI DSS or COPPA exposure
Retention and deletionRetention period for raw logs; whether trained weights must be retrainedDeletion request cannot be honored for weights
Derived artifactsOwnership of eval sets, reward models, adapters built from logsDispute over benchmarks after termination
TerminationWhether rights survive; what data is purgedTraining use continues after churn
Disclosure dutiesData needed for EU AI Act Article 53 training-content summary, if a GPAI model is trained [10]Provider cannot describe its training sources

Logs versus licensed operational records

Deployment logs and licensed historical records solve different problems. Logs capture your agent's own behavior on live traffic, which makes them strong for regression evaluation and failure analysis but biased toward whatever the current agent already attempts. Licensed operational records, such as ticket and case histories reconstructed as trajectories, capture how humans did the work before any agent existed, under a license negotiated specifically for training.

Many teams use both: logs, under a narrow evaluation right, to measure the agent, and separately licensed records to teach it new workflows. The agent training data hub compares the main sources, and licensed, commissioned or synthetic trajectories covers the trade-offs. General license terms such as use scope, exclusivity and deletion are explained in AI data license terms.

When customer contracts do not permit training, licensing operational records directly from companies that hold them is one alternative. SourceX sources operational datasets such as support and sales histories, engineering records and finance and legal workflows from US companies on request, and every release is approved by the supplying company; buyers can describe the records they need.

Source agent training data with clear rights

If your deployment contracts do not support training, licensed operational records are a separate route. SourceX looks for US businesses that hold the data you describe, reviews ownership and consents, and delivers under a license that defines records, uses, term and delivery; nothing is contracted until a supplier agrees. Describe the agent data you need licensed.

Frequently asked questions

Does anonymizing agent logs remove the need for a training right?

Not on its own. Anonymization may take data outside the GDPR if identification is no longer reasonably likely [4], but customer confidentiality and the contract's purpose limits still apply to non-personal business content.

Are telemetry fields like latency and tool error codes free to use?

Often yes, if the contract defines them as service or usage data and they are stripped of content. Problems start when "usage data" definitions sweep in prompts, tool payloads or reviewer comments.

Can a vendor add a training right at renewal?

It can propose one, but FTC staff warned in 2024 that quietly or retroactively broadening data use through terms changes may be unfair or deceptive [2]. Use an affirmative, documented opt-in and apply it prospectively.

Sources

  1. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  2. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  3. gdpr-text.com, "GDPR Article 28: Processor" (2016). https://gdpr-text.com/en/read/article-28/
  4. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  5. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  6. California Privacy Protection Agency, "California Consumer Privacy Act of 2018 (statute text)". https://cppa.ca.gov/regulations/pdf/ccpa_statute.pdf
  7. Cleary Gottlieb, "California AG proposes modified CCPA regulations" (2020). https://clearygottlieb.com/-/media/files/alert-memos-2020/california-ag-proposes-modified-ccpa-regulations.pdf
  8. Google Cloud, "Interaction logging (Dialogflow ES)". https://docs.cloud.google.com/dialogflow/es/docs/interaction-logging
  9. Google for Developers, "AgentAttributionInfo (Admin SDK Reports API)". https://developers.google.com/workspace/admin/reports/reference/rest/v1/AgentAttributionInfo
  10. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  11. arXiv (Yao et al.), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data