Provenance, rights and permitted use
Third-Party Content Inside Licensed Corpora: Attachments, Quoted Text and Stock Media
Quick answer
A supplier can license only what it owns or is authorized to license, yet its email archives, shared drives and slide libraries are full of material written by others: inbound attachments, vendor proposals, analyst reports, stock photos, syndicated articles and quoted passages. Buyers should expect this content, detect it with file metadata, sender domains, credit lines and hash matching, then decide per class whether to exclude it, clear it, or keep it under a documented risk decision.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a clean chain of title still leaves embedded third-party rights
Chain of title proves the supplier's rights in its own records; it says nothing about the documents other people sent it. Under US law a transfer of copyright ownership needs a signed writing [2], so a company that received a consultant's PDF or bought a stock image holds a copy and perhaps a use license, not the copyright. Practitioner guidance on AI licensing flags exactly this gap: datasets can include stock or syndicated items the vendor had no right to sublicense [1]. Read this page alongside the chain of title guide, which covers the supplier's own rights.
The stakes are not hypothetical. The US Copyright Office's Part 3 report on generative AI training, still a pre-publication version as of October 2026, treats the training use of copyrighted works as a fact-specific fair use question rather than a safe harbor [3]. In September 2026 the Third Circuit affirmed that copying Westlaw headnotes to build ROSS's non-generative legal research tool was not fair use [4]. Neither outcome tells you how a court would treat your model, which is why buyers manage embedded content at the dataset level instead of relying on a defense.
Where third-party content hides in operational corpora
Third-party content clusters in predictable places, and each place has a recognizable signature in the file or message metadata. The categories below cover most of what turns up in email archive licenses, document repositories, knowledge bases and presentation deck datasets.
- Inbound attachments. MIME parts on messages whose
Fromdomain is external: vendor quotes, customer RFPs, signed contracts on the counterparty's paper, résumés. - Forwarded and quoted threads. Text below
-----Original Message-----or>quote markers, often authored by people outside the supplier, sometimes including whole newsletters. - Analyst and research reports. PDFs from subscription research services, usually with a cover page, a copyright line and a "licensed to" or "for internal use" footer.
- Stock media in decks and pages. Images inside
ppt/media/of a.pptxpackage, or embedded in.docxand HTML knowledge base articles, many still carrying IPTC or XMP credit and copyright fields. - Templates, fonts and icon sets. Licensed slide masters, embedded font files and icon libraries whose terms cover the company's use, not redistribution.
- Syndicated news clippings. Media-monitoring digests and pasted articles in wikis and Slack exports.
- Third-party software artifacts. Vendor documentation, release notes and screenshots of other products in support tickets (see screen recordings and third-party content).
Confidential material from counterparties is a separate risk with different tests: NDAs and trade secrets rather than copyright. Route it through third-party confidential information screening rather than this workflow.
Detection signals buyers can ask suppliers to run
The most reliable detection combines provenance metadata with content matching, because each signal alone misses a category. Ask for these checks before delivery and for the counts they produce, so the residual risk is measured instead of assumed.
| Signal | What it catches | Typical fields or tools | Known failure mode |
|---|---|---|---|
| Sender and author domain | Inbound attachments, forwarded documents | Email From, Reply-To, Received headers; Office docProps/core.xml creator and lastModifiedBy | Internal staff re-saving external files resets author fields |
| Quote and forward markers | Quoted third-party text | In-Reply-To, References, > prefixes, "Forwarded message" separators | Inline replies interleave authors in one body |
| Embedded media metadata | Stock imagery, agency photos | XMP dc:rights, IPTC Credit Line, Copyright Notice, C2PA manifests and the CAWG training-mining assertion [7] | Metadata is stripped by screenshots, compression and slide export |
| Visible credit and watermark text | Stock previews, analyst reports | OCR over images and PDF footers for "©", agency names, "licensed to" | Cropped credits; OCR misses on small fonts |
| Perceptual hash matching | Reused stock and agency images | pHash or PDQ hashes compared against a reference set the buyer or supplier is entitled to use | Heavy edits and composites escape matching |
| Near-duplicate text matching | Syndicated articles, boilerplate reports | MinHash or shingle overlap against known publisher text or a sample of external sources | Paraphrased or translated copies |
| Font and template inventory | Embedded licensed fonts and masters | ppt/fonts/, word/fonts/, theme XML, font embedding flags | Fonts subset at export with no name retained |
Where media files carry provenance manifests, the C2PA Content Credentials guide and the CAWG do-not-train assertion explain how to read them. A manifest that reserves training use is a strong exclusion signal even when the file sits in a supplier's own folder.
Exclude, clear or accept: a decision rule per content class
The default for low-value embedded items is exclusion, because clearing rights item by item rarely pays for itself. Clearing makes sense when the third-party content is the reason you want the corpus, for example a knowledge base built on vendor documentation. Accepting residual items should be a written risk decision with a measured rate, not a silent default.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Content class | Default action | Clear instead when | Record in manifest |
|---|---|---|---|
| Inbound attachments from external domains | Exclude file, keep message body if supplier-authored | The attachment type is core to the use case and the counterparty can grant rights | exclusion_reason=external_attachment, count by MIME type |
| Quoted text from external senders | Strip quoted blocks below the first external marker | Thread context is essential and quotes are short and incidental | quote_stripping=v2, share of messages affected |
| Analyst and research reports | Exclude | A publisher or collective license covers AI use [6] | Publisher, license reference or exclusion |
| Stock images in decks and articles | Replace with placeholder token or drop image, keep slide text | Supplier's stock license is confirmed to permit the transfer (rare) [5] | Hash list of removed images |
| Licensed fonts and templates | Remove embedded font files; keep rendered text | Not usually worth clearing | Fonts removed, template IDs |
| Syndicated news clippings | Exclude | A publisher license covers the use | Source domains excluded |
Stock licenses are the clearest case for exclusion. A typical stock end-user license grants rights to the licensee for defined uses and restricts sublicensing [5], so a supplier generally cannot pass those images on, and an image-text pair built from a stock photo inherits that problem. For multimodal records built from several rightsholders, see licensing multimodal records.
Recording exclusions so the attestation stays accurate
Every exclusion and clearance should appear in the delivery manifest, because the supplier's rights attestation is only true for the corpus as filtered. If the attestation says "supplier owns or is authorized to license all included content" while 4% of files are unfiltered vendor PDFs, the warranty and the data disagree. Pair the manifest with the data rights attestation template and log the rules in your training data use register.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "msg-000184233",
"source_system": "exchange_online_export",
"authored_by_supplier": true,
"third_party_screen": {
"external_attachments_removed": 2,
"removed_mime_types": ["application/pdf", "image/jpeg"],
"quoted_external_text_stripped": true,
"stock_image_hash_matches": 0,
"rule_version": "tp-screen-2026.10"
},
"residual_risk_class": "low",
"notes": "Vendor quote PDF and inline logo excluded; supplier-authored reply retained"
}
Keep removed items out of the delivery entirely rather than flagging them in place. When originals ship with extracted text, apply the same exclusions to both layers; original files with sidecar metadata describes how the layers relate.
How intended use changes the risk: pre-training, SFT and RAG
The same embedded item carries different exposure depending on whether it trains weights or is retrieved verbatim at inference. In pre-training, rare third-party items are diluted but can still be memorized when duplicated across a corpus, which is why deduplicating syndicated text matters. In SFT, a small number of examples carries heavy weight, so a vendor proposal reused as a target answer is a concentrated risk.
RAG is the most exposed case because retrieved passages are reproduced to users, often with attribution that names the original author. A license to train is not a license to display; check grounding versus training licenses before indexing a corpus that still holds analyst reports or news clippings.
EU obligations that make third-party screening a documentation duty
For providers of general-purpose AI models in the EU, screening embedded content is part of a required copyright policy, not only good hygiene. Article 53(1)(c) of the AI Act requires a policy to comply with Union copyright law, including identifying Article 4(3) rights reservations [8], and for signatories the voluntary GPAI Code of Practice copyright chapter turns that into a written policy they keep up to date and implement [9]. As of October 2026, AI Office enforcement powers apply to new models from 2 August 2026, and the training-content summary template dated 24 July 2025 asks providers to describe their data sources [10].
A supplier corpus that silently includes publisher content with an Article 4 opt-out undermines that policy. Ask for the screening rules and counts in writing so they can feed your own documentation.
Questions to put to a supplier before signing
Buyers get better answers by asking for counts and rule definitions than by asking whether third-party content exists. Use these alongside the training data due diligence checklist.
- Which signals identify external authorship in each source system, and what share of files or messages did each rule remove?
- Are quoted and forwarded external passages stripped, and at which marker?
- Which media were hash-matched, against what reference set, and with what threshold?
- Are embedded fonts, templates and icon sets removed from delivered originals?
- Does any content rely on a third-party license (stock, research, news, collective license), and does that license permit the transfer you are buying?
- How is a later-discovered third-party item reported and removed from delivered data?
For knowledge base article licenses in particular, ask what share of articles reproduce vendor documentation, since that is where supplier-authored and third-party text mix most.
SourceX sources operational datasets from US companies on request and rights-reviews every dataset for ownership and consents before it is delivered under a license that defines records, uses, term and delivery. Buyers can describe the corpus they need on the buyers page, and the diligence materials prepared per dataset cover source, rights, preparation and allowed use. More context sits in the provenance hub and the AI data guides.
Sourcing licensed training data with third-party content in mind
SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and nothing is contracted until the supplying company agrees and approves the release. Sourcing is on request, so a request does not guarantee a match. Describe the records, use and exclusions you need at sourcex.si/buyers.
Sources
- Osborne Clarke, "Session 4: AI Licensing (13 November 2024)" (2024). https://osborneclarke.com/system/files/documents/24/11/21/Session-4---13-Nov---AI-Licensing%28157063266.2%29.pdf
- Legal Information Institute, Cornell Law School, "17 U.S. Code § 204 - Execution of transfers of copyright ownership". https://law.cornell.edu/uscode/text/17/204
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf?amp=&stream=top
- IPWatchdog, "Third Circuit Affirms Revised Fair Use Ruling Against ROSS' AI Legal Research Platform in Sealed Opinion" (2026). https://ipwatchdog.com/2026/09/30/third-circuit-affirms-revised-fair-use-ruling-against-ross-ai-legal-research-platform-in-sealed-opinion/
- Storyblocks, "Storyblocks Individual License Agreement" (2024). https://www.storyblocks.com/license
- Copyright Clearance Center (Business Wire), "CCC Launching New AI Content Re-Use Rights for U.S. Academic Customers and Transactional Licensing Capabilities for AI" (2026). https://www.businesswire.com/news/home/20260303630177/en/CCC-Launching-New-AI-Content-Re-Use-Rights-for-U.S.-Academic-Customers-and-Transactional-Licensing-Capabilities-for-AI
- IPTC, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.