Skip to content

Provenance, rights and permitted use

Third-Party Content Inside Licensed Corpora: Attachments, Quoted Text and Stock Media

Quick answer

A supplier can license only what it owns or is authorized to license, yet its email archives, shared drives and slide libraries are full of material written by others: inbound attachments, vendor proposals, analyst reports, stock photos, syndicated articles and quoted passages. Buyers should expect this content, detect it with file metadata, sender domains, credit lines and hash matching, then decide per class whether to exclude it, clear it, or keep it under a documented risk decision.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a clean chain of title still leaves embedded third-party rights

Chain of title proves the supplier's rights in its own records; it says nothing about the documents other people sent it. Under US law a transfer of copyright ownership needs a signed writing [2], so a company that received a consultant's PDF or bought a stock image holds a copy and perhaps a use license, not the copyright. Practitioner guidance on AI licensing flags exactly this gap: datasets can include stock or syndicated items the vendor had no right to sublicense [1]. Read this page alongside the chain of title guide, which covers the supplier's own rights.

The stakes are not hypothetical. The US Copyright Office's Part 3 report on generative AI training, still a pre-publication version as of October 2026, treats the training use of copyrighted works as a fact-specific fair use question rather than a safe harbor [3]. In September 2026 the Third Circuit affirmed that copying Westlaw headnotes to build ROSS's non-generative legal research tool was not fair use [4]. Neither outcome tells you how a court would treat your model, which is why buyers manage embedded content at the dataset level instead of relying on a defense.

Where third-party content hides in operational corpora

Third-party content clusters in predictable places, and each place has a recognizable signature in the file or message metadata. The categories below cover most of what turns up in email archive licenses, document repositories, knowledge bases and presentation deck datasets.

  • Inbound attachments. MIME parts on messages whose From domain is external: vendor quotes, customer RFPs, signed contracts on the counterparty's paper, résumés.
  • Forwarded and quoted threads. Text below -----Original Message----- or > quote markers, often authored by people outside the supplier, sometimes including whole newsletters.
  • Analyst and research reports. PDFs from subscription research services, usually with a cover page, a copyright line and a "licensed to" or "for internal use" footer.
  • Stock media in decks and pages. Images inside ppt/media/ of a .pptx package, or embedded in .docx and HTML knowledge base articles, many still carrying IPTC or XMP credit and copyright fields.
  • Templates, fonts and icon sets. Licensed slide masters, embedded font files and icon libraries whose terms cover the company's use, not redistribution.
  • Syndicated news clippings. Media-monitoring digests and pasted articles in wikis and Slack exports.
  • Third-party software artifacts. Vendor documentation, release notes and screenshots of other products in support tickets (see screen recordings and third-party content).

Confidential material from counterparties is a separate risk with different tests: NDAs and trade secrets rather than copyright. Route it through third-party confidential information screening rather than this workflow.

Detection signals buyers can ask suppliers to run

The most reliable detection combines provenance metadata with content matching, because each signal alone misses a category. Ask for these checks before delivery and for the counts they produce, so the residual risk is measured instead of assumed.

SignalWhat it catchesTypical fields or toolsKnown failure mode
Sender and author domainInbound attachments, forwarded documentsEmail From, Reply-To, Received headers; Office docProps/core.xml creator and lastModifiedByInternal staff re-saving external files resets author fields
Quote and forward markersQuoted third-party textIn-Reply-To, References, > prefixes, "Forwarded message" separatorsInline replies interleave authors in one body
Embedded media metadataStock imagery, agency photosXMP dc:rights, IPTC Credit Line, Copyright Notice, C2PA manifests and the CAWG training-mining assertion [7]Metadata is stripped by screenshots, compression and slide export
Visible credit and watermark textStock previews, analyst reportsOCR over images and PDF footers for "©", agency names, "licensed to"Cropped credits; OCR misses on small fonts
Perceptual hash matchingReused stock and agency imagespHash or PDQ hashes compared against a reference set the buyer or supplier is entitled to useHeavy edits and composites escape matching
Near-duplicate text matchingSyndicated articles, boilerplate reportsMinHash or shingle overlap against known publisher text or a sample of external sourcesParaphrased or translated copies
Font and template inventoryEmbedded licensed fonts and mastersppt/fonts/, word/fonts/, theme XML, font embedding flagsFonts subset at export with no name retained

Where media files carry provenance manifests, the C2PA Content Credentials guide and the CAWG do-not-train assertion explain how to read them. A manifest that reserves training use is a strong exclusion signal even when the file sits in a supplier's own folder.

Exclude, clear or accept: a decision rule per content class

The default for low-value embedded items is exclusion, because clearing rights item by item rarely pays for itself. Clearing makes sense when the third-party content is the reason you want the corpus, for example a knowledge base built on vendor documentation. Accepting residual items should be a written risk decision with a measured rate, not a silent default.

Illustrative example: invented to show structure; it does not describe an available dataset.

Content classDefault actionClear instead whenRecord in manifest
Inbound attachments from external domainsExclude file, keep message body if supplier-authoredThe attachment type is core to the use case and the counterparty can grant rightsexclusion_reason=external_attachment, count by MIME type
Quoted text from external sendersStrip quoted blocks below the first external markerThread context is essential and quotes are short and incidentalquote_stripping=v2, share of messages affected
Analyst and research reportsExcludeA publisher or collective license covers AI use [6]Publisher, license reference or exclusion
Stock images in decks and articlesReplace with placeholder token or drop image, keep slide textSupplier's stock license is confirmed to permit the transfer (rare) [5]Hash list of removed images
Licensed fonts and templatesRemove embedded font files; keep rendered textNot usually worth clearingFonts removed, template IDs
Syndicated news clippingsExcludeA publisher license covers the useSource domains excluded

Stock licenses are the clearest case for exclusion. A typical stock end-user license grants rights to the licensee for defined uses and restricts sublicensing [5], so a supplier generally cannot pass those images on, and an image-text pair built from a stock photo inherits that problem. For multimodal records built from several rightsholders, see licensing multimodal records.

Recording exclusions so the attestation stays accurate

Every exclusion and clearance should appear in the delivery manifest, because the supplier's rights attestation is only true for the corpus as filtered. If the attestation says "supplier owns or is authorized to license all included content" while 4% of files are unfiltered vendor PDFs, the warranty and the data disagree. Pair the manifest with the data rights attestation template and log the rules in your training data use register.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "msg-000184233",
  "source_system": "exchange_online_export",
  "authored_by_supplier": true,
  "third_party_screen": {
    "external_attachments_removed": 2,
    "removed_mime_types": ["application/pdf", "image/jpeg"],
    "quoted_external_text_stripped": true,
    "stock_image_hash_matches": 0,
    "rule_version": "tp-screen-2026.10"
  },
  "residual_risk_class": "low",
  "notes": "Vendor quote PDF and inline logo excluded; supplier-authored reply retained"
}

Keep removed items out of the delivery entirely rather than flagging them in place. When originals ship with extracted text, apply the same exclusions to both layers; original files with sidecar metadata describes how the layers relate.

How intended use changes the risk: pre-training, SFT and RAG

The same embedded item carries different exposure depending on whether it trains weights or is retrieved verbatim at inference. In pre-training, rare third-party items are diluted but can still be memorized when duplicated across a corpus, which is why deduplicating syndicated text matters. In SFT, a small number of examples carries heavy weight, so a vendor proposal reused as a target answer is a concentrated risk.

RAG is the most exposed case because retrieved passages are reproduced to users, often with attribution that names the original author. A license to train is not a license to display; check grounding versus training licenses before indexing a corpus that still holds analyst reports or news clippings.

EU obligations that make third-party screening a documentation duty

For providers of general-purpose AI models in the EU, screening embedded content is part of a required copyright policy, not only good hygiene. Article 53(1)(c) of the AI Act requires a policy to comply with Union copyright law, including identifying Article 4(3) rights reservations [8], and for signatories the voluntary GPAI Code of Practice copyright chapter turns that into a written policy they keep up to date and implement [9]. As of October 2026, AI Office enforcement powers apply to new models from 2 August 2026, and the training-content summary template dated 24 July 2025 asks providers to describe their data sources [10].

A supplier corpus that silently includes publisher content with an Article 4 opt-out undermines that policy. Ask for the screening rules and counts in writing so they can feed your own documentation.

Questions to put to a supplier before signing

Buyers get better answers by asking for counts and rule definitions than by asking whether third-party content exists. Use these alongside the training data due diligence checklist.

  1. Which signals identify external authorship in each source system, and what share of files or messages did each rule remove?
  2. Are quoted and forwarded external passages stripped, and at which marker?
  3. Which media were hash-matched, against what reference set, and with what threshold?
  4. Are embedded fonts, templates and icon sets removed from delivered originals?
  5. Does any content rely on a third-party license (stock, research, news, collective license), and does that license permit the transfer you are buying?
  6. How is a later-discovered third-party item reported and removed from delivered data?

For knowledge base article licenses in particular, ask what share of articles reproduce vendor documentation, since that is where supplier-authored and third-party text mix most.

SourceX sources operational datasets from US companies on request and rights-reviews every dataset for ownership and consents before it is delivered under a license that defines records, uses, term and delivery. Buyers can describe the corpus they need on the buyers page, and the diligence materials prepared per dataset cover source, rights, preparation and allowed use. More context sits in the provenance hub and the AI data guides.

Sourcing licensed training data with third-party content in mind

SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and nothing is contracted until the supplying company agrees and approves the release. Sourcing is on request, so a request does not guarantee a match. Describe the records, use and exclusions you need at sourcex.si/buyers.

Sources

  1. Osborne Clarke, "Session 4: AI Licensing (13 November 2024)" (2024). https://osborneclarke.com/system/files/documents/24/11/21/Session-4---13-Nov---AI-Licensing%28157063266.2%29.pdf
  2. Legal Information Institute, Cornell Law School, "17 U.S. Code § 204 - Execution of transfers of copyright ownership". https://law.cornell.edu/uscode/text/17/204
  3. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf?amp=&stream=top
  4. IPWatchdog, "Third Circuit Affirms Revised Fair Use Ruling Against ROSS' AI Legal Research Platform in Sealed Opinion" (2026). https://ipwatchdog.com/2026/09/30/third-circuit-affirms-revised-fair-use-ruling-against-ross-ai-legal-research-platform-in-sealed-opinion/
  5. Storyblocks, "Storyblocks Individual License Agreement" (2024). https://www.storyblocks.com/license
  6. Copyright Clearance Center (Business Wire), "CCC Launching New AI Content Re-Use Rights for U.S. Academic Customers and Transactional Licensing Capabilities for AI" (2026). https://www.businesswire.com/news/home/20260303630177/en/CCC-Launching-New-AI-Content-Re-Use-Rights-for-U.S.-Academic-Customers-and-Transactional-Licensing-Capabilities-for-AI
  7. IPTC, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
  8. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  9. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  10. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data