Skip to content

AI data market

Third-party content hiding in your archives: what to exclude before licensing

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Third-party content in company archives is any material your company holds but did not author or cannot relicense: purchased reports, books, industry standards, vendor manuals, customer files and pirated copies. Exclude it before licensing. The working rule is simple: if your company did not write it, keep it out until a written license says otherwise.

Key takeaways

  • Holding a copy of a file is not the same as holding the right to license it.
  • Shared drives, email attachments, wikis and ticket attachments hold most third-party material.
  • Purchased research, standards, books and vendor manuals are excluded by default, even when the company paid for them.
  • Pirated or improperly obtained files should be removed, logged and escalated to counsel, not quietly skipped.
  • A written exclusion log lets the buyer see what was removed and why.

What counts as third-party content in a company archive?#

Third-party content in a company archive is any file or passage your company holds but did not author, or authored under terms that leave control with someone else. A subscription, purchase or download usually grants permission to read and use material internally. It rarely grants a right to pass that material to an AI developer under a new license.

That distinction decides what can go into a dataset. A license grants rights, and a company cannot grant rights it never had. Buyers now ask suppliers where each record family came from and what permits its use, so an unexamined reference folder tends to become a problem at exactly the moment a deal is close.

  • Purchased market, analyst and benchmarking reports.
  • Books, ebooks and textbooks, including scanned chapters.
  • Industry standards, building codes and certification study guides.
  • Vendor manuals, spec sheets, catalogs and product documentation.
  • Customer and client files: drawings, data and uploads received to do the work.
  • Third-party training courses, webinar recordings and conference decks.
  • News articles, newsletters and paywalled content saved as PDFs.
  • Vendored libraries, SDKs and open-source code inside your repositories.
  • Stock images, fonts and design templates.
  • Pirated or improperly obtained copies of any of the above.

Where does third-party material usually hide?#

Third-party material usually hides in the places employees treat as personal libraries: shared drives, email, wikis and ticket attachments. These locations collect years of downloads that nobody catalogued, and folder names rarely say whether a file was bought, borrowed or written in house.

Structured systems are usually cleaner. Records created inside a CRM, an ERP or a ticketing tool are mostly typed by your own staff, although the attachments stored inside those systems carry the same risk as any shared drive.

Where does third-party material usually hide?
LocationWhat tends to turn upHow to spot it
Shared drives (Google Drive, SharePoint, Box)Reference folders of purchased reports, standards and ebooksFolder names such as Library or Reference, publisher names in PDF properties, copyright pages
Email archivesNewsletters, vendor quotes with spec sheets, conference decksPublisher and vendor sender domains, bulk-mail headers, attachment types
Wikis (Confluence, Notion)Pasted articles and copied vendor documentationLong passages with outside formatting and no internal author
Help desk (Zendesk, Intercom)Customer uploads and screenshots of other vendors' softwareAttachment fields, file types, customer-side authorship
Code repositories (GitHub, GitLab)Vendored libraries, SDKs and vendor sample codeLicense files, third-party directories, package manifests
Project foldersClient-supplied drawings, surveys and datasetsTransmittal logs, client names in paths, external authors
Learning management systemsPurchased courses and certification prep materialCourse author fields and vendor branding

Which items should be excluded by default?#

Purchased reports, books, standards, vendor manuals and third-party courses should be excluded by default, even when the company paid for them. The purchase bought a use right, typically internal and often limited to named users, and those terms rarely allow redistribution or AI training.

The useful line runs between the third-party item and your company's own work around it. A technician's note explaining why a manual's procedure failed on a particular unit is your record; the manual is not. Keep the note, drop the manual, and refer to it by title rather than including its text.

Which items should be excluded by default?
Content typeDefault treatmentWhat could change it
Purchased research and analyst reportsExcludeWritten permission from the publisher, which is uncommon
Books, ebooks and textbooksExcludeNothing within the company's control
Industry standards and codesExcludeWritten permission from the standards body
Vendor manuals and spec sheetsExcludeWritten vendor permission; your own service notes stay in scope
Customer and client filesCarve out pending reviewContract terms and client consent, reviewed by counsel
Third-party courses and webinarsExcludeNothing; your own training material is assessed separately
Vendored and open-source codeExclude from the code packageLicense review of each component
Pirated or improperly obtained copiesRemove, log and escalateNothing

Why pirated material needs separate handling#

Pirated material needs separate handling because the problem is how it was obtained, not only how it would be used. Typical finds include ebook PDFs from shadow libraries, course videos recorded from one employee's subscription and shared with a team, paywalled reports forwarded far beyond the licensed seats, and cracked software installers.

Copyright disputes over AI training have drawn attention to how material was acquired, not only how it was used. The class settlement in Bartz v. Anthropic, which covers books on a works list drawn from LibGen and PiLiMi downloads, requires the original torrented files and any copies derived from them to be destroyed. Buyers have taken note and now screen suppliers for acquisition taint. One folder of downloaded books in a shared drive can cast doubt on a whole package if the buyer's review finds it before yours does.

Remove such files from the licensing scope, log what was found and where, and raise it with counsel as a separate internal matter. Avoid bulk deletion in a hurry: retention obligations or a legal hold may apply, and destroying files carries its own risk.

How do you screen an archive without reading every file?#

An archive screen works from metadata first and human review second. File names, folder paths, author fields, sender domains and document properties identify most third-party material before anyone opens a file, and they let you set rules at folder level instead of file by file.

  • Export a file inventory with paths, types, authors and created dates for each system in scope.
  • Flag known markers: ISBNs, copyright pages, licensed-to watermarks, standard document numbering and publisher names in file properties.
  • Flag files whose author field sits outside your email domain or is blank.
  • Sample flagged folders by hand and turn what you find into folder-level include or exclude rules.
  • Run a license scanner over code repositories to separate vendored and open-source components from code your team wrote.
  • Record each exclusion rule, its reason and its reviewer, and keep the list with the dataset documentation.

What buyers expect the exclusion record to show#

Buyers expect the exclusion record to show what was removed, by which rule and on whose decision. Expect false positives along the way: your own white papers carry copyright notices, and vendor logos appear on internal templates. The reviewer's job is to decide, not to trust a flag.

Provenance metadata standards point the same way. The Data & Trust Alliance's Data Provenance Standards list license to use and copyright, patent and trademark status among the elements of their Use group, the metadata meant to travel with a dataset. A clean exclusion log makes those fields easy to complete.

Illustrative: a regional distributor clears its shared drive#

Illustrative: a fictional industrial distributor prepares to license customer service notes and quote revisions from its CRM, along with its Epicor order history. A metadata pass over its SharePoint shows a Supplier Library folder full of manufacturer catalogs and spec sheets, a Market Intel folder of purchased industry reports, and years of vendor newsletters in Outlook.

The COO and outside counsel set folder rules. Catalogs, spec sheets and purchased reports are excluded. Customer emails stay in scope only as the distributor's own replies, after privacy preparation, with customer-attached drawings removed. In a former sales manager's personal share, the review finds a set of downloaded business books; they are removed from scope and logged for counsel.

The resulting package is smaller and easier to defend. The exclusion log goes into the dataset documentation, so the buyer's review starts with answers instead of questions about where the files came from.

How SourceX treats third-party content#

SourceX treats third-party content as a Rights question within the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The fit check runs on metadata only, so nothing is shared while record families and their likely third-party exposure are mapped.

Exclusions and their reasons are recorded in the SourceX Evidence Packet alongside provenance, licensing rights, permitted use, the privacy record and release authorization. The supplier approves the final scope before anything is delivered, and the company keeps ownership of everything it licenses.

Frequently asked questions

Can we include our own notes on a purchased report?

Notes written in your own words are generally your company's work, so they can usually stay in scope. Notes that reproduce long passages or tables from the report carry the publisher's restrictions with them. Keep the commentary, drop the copied material, and have counsel review borderline items such as internal digests built mainly from one purchased source.

Are quoted customer replies in email threads third-party content?

Quoted replies are written by people outside your company, but they are usually handled through privacy preparation and contract review rather than a copyright exclusion. The question is whether customer contracts and notices permit the use and whether personal and confidential details are removed. Files the customer attached are a separate question and are often carved out.

What about open-source code in our repositories?

Open-source licenses differ. Some allow reuse with attribution, while others attach conditions to distribution that may not suit a data license. Most code packages exclude vendored third-party directories and include only code your team wrote, with the license scan results recorded so the buyer can see how the line was drawn.

Should we delete the third-party files we find?

Not automatically. Excluding a file from a license scope is different from deleting it from your systems. Retention schedules, legal holds and the original purchase terms may require you to keep or destroy material in a particular way, so decide on deletion separately, with counsel, after the licensing scope is set.

Does AI-generated text in our archives count as third-party content?

Not in the copyright sense this checklist covers, but it raises a different question for buyers: whether a record reflects human work. Drafts produced with AI tools and pasted into wikis or tickets can be flagged during the same metadata pass and documented, so the buyer knows which material was written by people.

Sources

  • The Use group of the Data & Trust Alliance Data Provenance Standards includes elements for confidentiality classification, consent documentation location, privacy-enhancing technologies applied, allowed and excluded processing and storage geographies, license to use, intended data use, and copyright, patent and trademark status. Source
  • The Bartz settlement class covers works on a Works List drawn from the LibGen and PiLiMi versions Anthropic downloaded, and Anthropic must destroy the original torrented files and any copies derived from them. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify