Skip to content

Provenance, rights and permitted use

Copyright in Operational Business Records: What It Means for Training-Data Rights

Quick answer

Some business records are copyrighted, but much of what AI teams buy is not. Under US law, facts carry no copyright, and neither do business procedures or methods; free-text expression does, such as a support agent's long reply, an engineering design doc or a contract memo, and so can an original selection or arrangement. Structured fields, timestamps, amounts and log lines usually carry thin protection or none. In practice, contracts, confidentiality and access control decide who may license operational data, so your license should say which right it actually conveys.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Copyright protects original expression, and Feist holds that facts are not original because nobody authors them. The Court required independent creation plus "a modicum of creativity." It found an alphabetical telephone listing unprotectable even though compiling it took real effort, because the "sweat of the brow" theory does not create copyright. Section 102(b) adds a second limit. Ideas, procedures, processes, systems and methods of operation stay unprotected even when a protected work describes them under 17 U.S.C. 102(b).

For operational records, this produces a split. An order's amount, sku, created_at and status, a ticket's priority and resolution_code, and a server log's level, latency_ms and stack frame are facts or machine-generated values. The workflow the records reveal, such as an escalation ladder or a reconciliation procedure, is a method. Neither is copyright subject matter. The narrative text a person writes into a description, comment or body field can be.

The amount of copyright in a record roughly tracks how much human-written prose it contains. The table below is a working triage for counsel, not a legal ruling, and close cases depend on the actual text. For a full labeling scheme, see classifying training data by copyright status.

Record typeTypical fieldsCopyright exposureWhat usually controls use
Transaction and ledger rowsIDs, amounts, dates, GL codes, counterparty keysLittle or none: facts and codesContract, confidentiality, financial privacy rules
System and application logstimestamp, level, service, message template, trace IDThin: templated messages, machine outputContract, security policy, trade secret
Workflow and RPA logs (job, queue item, exception events)state transitions, exception type, retriesThin: process facts; the process itself is a methodContract, trade secret
Support ticketssubject, category, priority, agent replies, customer messagesMixed: fields thin, long agent and customer prose can be protectedContract, customer terms, privacy law
Canned replies and macrostemplate text, variablesPossibly protected as text, but often third-party or vendor-suppliedAuthorship check, vendor terms
Email threads and chat transcriptsheaders, bodies, attachmentsBodies often protected; authors include outsidersContract, privacy, authorship of each party
Internal documents (design docs, SOPs, memos, playbooks)prose, diagrams, tablesUsually protectedEmployer ownership, contract
Filled formsform template plus entered valuesTemplate may be protected; entered facts are notTemplate owner terms, privacy
Curated knowledge bases and annotation guidesselected and arranged entries, editorial summariesSelection, arrangement and summaries can be protected [1]Copyright plus contract

Two rows need special care. Canned replies often come from a help-desk vendor's starter library or a consultant, so the supplier may not own them; see templates and boilerplate in business records. Customer-written ticket text is authored by people outside the supplier, which raises the third-party questions covered below.

A dataset can be protected as a compilation only where the selection, coordination or arrangement of its contents is original, and that protection covers the creative choices, not the underlying facts. Feist found that an alphabetical directory failed this test. So did including every subscriber, which is the obvious completeness choice. Most operational exports follow the same pattern: every ticket in a date range, sorted by created_at, with the schema the help-desk product imposes.

Editorial layers are different. In September 2026 the Third Circuit affirmed that Westlaw headnotes, the editorial summaries attached to judicial opinions, were original enough to be copyrightable. It also affirmed that copying them to build a competing non-generative legal search tool was not fair use [1]. Business-record analogs include curated resolution articles, analyst-written case summaries and internal annotation of records. If a supplier's dataset includes these layers, treat them as copyrighted expression the supplier must own or have rights to license. Do not treat them as "just data."

Outside the US, the analysis can change. The EU and UK also recognize a separate database right based on substantial investment rather than originality, so non-US records may carry protection that Feist would not give them. Ask counsel in that jurisdiction instead of relying on the US baseline.

Who owns the protected parts: employees, customers and vendors

Where records do contain copyrightable text, the next question is who authored it. For text employees write within the scope of their jobs, the employer is treated as the author under the work-made-for-hire rule and owns the copyright unless the parties agree otherwise in a signed written instrument [2]. That covers agent replies, engineering docs and internal memos written by staff. It does not automatically cover contractors, offshore support vendors working under their own master services agreements, or customers.

Customer-authored messages in tickets, emails and community forums are usually owned by the customer. Whatever the supplier can do with them comes from its customer terms of service or contracts. Check whether those terms grant a license broad enough to cover disclosure to a third party for model training. That question often turns on the customer agreement or data processing agreement more than on copyright; see customer contracts and DPAs for training use. The employee side is covered in detail in employee-authored records in training data, and the owner-level question in who owns enterprise data.

For the factual core of business records, contract law and trade secret law do the controlling. A supplier's pricing history, defect logs or underwriting outcomes can qualify as trade secrets. That requires the supplier to take reasonable measures to keep the information secret and for it to derive independent economic value from not being generally known [3]. That protection depends on confidentiality being maintained, which is why NDAs, access control and delivery controls matter to a deal's legal footing, not only its security.

Contract terms bind only the parties to them. Commentary on AI terms-of-use restrictions makes the general point: where material lacks copyright, use restrictions rest on contract, and contract cannot bind strangers the way a property right can [4]. For a buyer, the license is the main source of permission. Equally important, the supplier's own contracts can limit what it may grant. Those include customer agreements, vendor terms for the help-desk or CRM platform that stores the records, and confidentiality obligations owed to counterparties. A record with no copyright can still be off-limits.

Copyright still matters for training-data compliance in two places. In the EU, Article 53(1)(c) of the AI Act requires general-purpose AI model providers to maintain a policy to comply with Union copyright law [5]. That makes it useful for a provenance file to separate copyrighted expression from factual data. In the UK, the text and data analysis exception in CDPA s29A covers non-commercial research only, so commercial training on protected text generally needs a license or another legal basis [6].

What a provenance claim should state about rights

A provenance claim should name the legal basis for each component instead of saying the data is "owned" or "licensed" in general. A useful claim answers four questions. Which components contain protected expression? Who authored them? On what basis does the supplier license them: authorship, work for hire, assignment or a third-party grant? And which contracts restrict the factual remainder? That structure also feeds chain-of-title documentation and per-record permitted-use metadata.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset: support_tickets_2023_2025   # invented example
components:
  - name: structured_fields
    fields: [ticket_id, created_at, priority, category, resolution_code, csat]
    copyright_basis: none_claimed      # facts and codes
    controlling_instruments: [supplier_customer_tos_v7, nda_buyer_supplier]
  - name: agent_replies
    fields: [reply_body]
    copyright_basis: work_made_for_hire
    authors: employees_in_scope
    exclusions: [replies_by_outsourced_vendor_before_2024-03]
  - name: customer_messages
    fields: [message_body]
    copyright_basis: third_party_authored
    supplier_permission: customer_tos_section_license_to_content
    open_question: "Does ToS license extend to disclosure for model training?"
  - name: macros
    fields: [macro_text]
    copyright_basis: mixed
    note: "Vendor starter macros removed; in-house macros retained"
  - name: kb_article_summaries
    copyright_basis: employer_owned_editorial
    risk_note: "Original editorial layer; license explicitly"
derived_artifacts:
  pii_redaction: "method recorded; sample checked"

Derived artifacts follow the same logic. A redacted ticket, a translated email or a model-written summary has its own rights trail back to the source text; see tracing rights in derived records.

Use this before you sign, and keep the answers in the deal file.

  1. Map free text. List every field holding human-written prose (body, description, notes, attachments) and estimate its share of tokens. That share is your copyright exposure.
  2. Identify authors per field. Employees, contractors, outsourced vendors, customers, counterparties. Request the vendor and contractor agreements for any field they wrote.
  3. Find third-party templates. Flag help-desk starter macros, purchased SOP libraries and form templates. Remove or license them.
  4. Find editorial layers. Look for curated summaries, KB articles and annotation produced by staff. Those layers are where compilation and headnote-style claims arise [1].
  5. Read the supplier's restricting contracts. Check customer ToS and DPAs, platform terms of the system of record, and NDAs with counterparties named in the records.
  6. Separate the warranty from the permission. Ask for representations scoped to each component, not a single "we own all rights" clause.
  7. Mark jurisdiction. Note records created or hosted in the EU or UK, where database rights may apply, and whether the model will be placed on the EU market, which brings the AI Act copyright policy duty for general-purpose model providers [5].
  8. Record the result. Put the component-level basis into your training-data register and data rights attestation.

Personal data is a separate track. A ticket with no copyright can still contain names, account numbers and health details, and privacy law applies regardless of copyright status. The privacy buyer's guide covers de-identification.

How much the copyright analysis matters depends on what the model reproduces. For supervised fine-tuning on instruction-response pairs built from business records, the training targets are often the agent-written replies. These are the most expressive and the most likely to be protected, and their verbatim reproduction is the clearest risk. For evaluation sets, the record's outcome fields serve as ground truth. They are usually factual, so the binding constraint is typically the contract's permitted-use list rather than copyright.

For agent training on workflow logs and decision records, the valuable signal is process: which tool ran next, which exception fired, which approval was needed. Section 102(b) leaves the process unprotected, but trade secret status and confidentiality terms remain [3]. In every case, the license should name the use (SFT, evaluation, agent training) explicitly. Copyright silence is not a permission.

For a broader view of how the Copyright Office's position on training bears on licensing, see what the Copyright Office's AI training report means for licensing. The full topic map is in the provenance hub.

Sourcing operational records with a clear rights basis

Because control over operational records comes mostly from contract, buyers need a supplier willing to document authorship and restrictions per component. SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents and finance and legal workflows, on request rather than from stock, so a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Describe the data you need on the buyer request page, and see do you own your business data for the supplier-side view.

Request business records with a documented rights basis

SourceX finds US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted, with every release approved by the supplying company. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked. Describe the records you need at sourcex.si/buyers.

Sources

  1. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  2. Office of the Law Revision Counsel, U.S. House of Representatives, "17 U.S.C. 201: Ownership of copyright". https://uscode.house.gov/view.xhtml?req=%28title%3A17+section%3A201%28b%29+edition%3Aprelim%29
  3. U.S. Government Publishing Office (GovInfo), "18 U.S.C. 1839: Definitions" (2021 edition). https://www.govinfo.gov/content/pkg/USCODE-2021-title18/html/USCODE-2021-title18-partI-chap90-sec1839.htm
  4. SpicyIP, "Discussing Lemley and Henderson's The Mirage of Artificial Intelligence Terms of Use Restrictions" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
  5. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  6. UK Intellectual Property Office (GOV.UK), "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data