Skip to content

AI data market

Pirated vs lawfully acquired training data: why acquisition now matters

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Pirated training data lawsuits made acquisition a legal issue of its own: courts have treated how a developer obtained copies separately from whether training on them was fair use. Material downloaded from unauthorized sources carries the highest exposure, while records licensed from their owner with documented rights carry the lowest. Document acquisition before you license anything.

Key takeaways

  • Courts have separated the question of how copies were obtained from the question of how they were used.
  • In Bartz v. Anthropic, pirated library copies were not fair use even though training was, and those claims settled for $1.5 billion.
  • Records licensed from their owner with documented rights avoid the acquisition question instead of defending it.
  • Operating company archives often contain third-party material, such as purchased reports and client files, that must be carved out.
  • Acquisition records are easiest to keep at the moment data arrives and hardest to rebuild later.

Why does acquisition now matter for training data?#

Acquisition now matters because US courts have begun treating how a developer obtained training material as its own question, separate from whether training on that material could be fair use. A developer can win an argument about use and still face liability over copies it should never have held.

The US Copyright Office's May 2025 pre-publication report on generative AI training pointed the same way, saying commercial use of vast troves of works to produce competing content, especially through illegal access, goes beyond established fair use boundaries. That split moved sourcing from a back-office detail to a legal and commercial priority. Buyers now ask where every dataset came from, and suppliers who can answer with documents are simply easier to work with than those who answer with assurances.

What the book piracy cases showed#

The book piracy cases showed the difference between buying a copy and downloading one from an unauthorized library. In Bartz v. Anthropic, Judge William Alsup of the Northern District of California held on June 23, 2025 that training on books was fair use and that scanning purchased print books was fair use too, but that downloading pirated copies from sources including Books3, Library Genesis and Pirate Library Mirror to build a permanent central library was not, and he left that claim for trial.

The piracy claims then settled. Anthropic agreed to pay $1.5 billion, roughly $3,000 per work for about 482,460 works, and the court granted final approval on July 20, 2026. The settlement also requires destruction of the original torrented files and copies derived from them. The lesson commentators drew: a transformative-use defense does not clean up material obtained unlawfully, and an acquisition problem can cost far more than licensing would have.

The details of any settlement, including who may claim and on what terms, are set by the court's orders. Read those documents directly rather than relying on summaries or headlines.

How data was obtained and how exposed it is#

Acquisition routes fall along a range of exposure, from material taken from pirate sites to records licensed directly from the company that holds them. The table gives general starting positions; actual exposure depends on the facts and is assessed with counsel.

The bottom row is where most operational records belong. Support tickets, job histories and code reviews licensed by the company that created them avoid the acquisition question altogether, provided the company can show it had the right to license them.

How data was obtained and how exposed it is
Acquisition routeTypical legal questionsRelative exposureEvidence that helps
Downloaded from pirate or shadow librariesUnauthorized copying, possible willfulnessHighestNothing cures the source
Scraped from behind logins or paywallsSite terms, anti-circumvention and computer access lawsHighLittle, because access itself is the issue
Scraped from public web pagesFair use, site terms and crawler signalsContestedCrawl logs and honored opt-outs
Purchased copies used internallyFair use of lawfully owned copiesModerateReceipts and digitization records
Openly licensed or public domainLicense conditions such as attribution or share-alikeLower if conditions are metLicense record for each source
Licensed from the owner with documented rightsContract scope and complianceLowestSigned license and provenance file

How AI companies acquire data legally#

AI developers acquire data lawfully through several routes, and each carries its own documentation burden. None of them is free of legal questions, but all of them produce a record of permission that a pirated copy never has.

Each route needs a paper trail. The Data Provenance Initiative's first audit covered 44 data collections spanning more than 1,800 fine-tuning text datasets and documented their sources, licenses and creators, the kind of record that is far easier to keep at acquisition than to reconstruct later.

  • Licensing directly from rights holders such as publishers, platforms and operating companies.
  • Commissioning new work from contractors and experts under agreements that assign rights.
  • Using openly licensed material while honoring attribution, share-alike and other conditions.
  • Using public domain works after confirming their status.
  • Purchasing copies where counsel has assessed the intended use.
  • Using data from the developer's own products under terms that permit it.

Your own archive may hold acquired material too#

An operating company's archive can hold acquired material even when the company wrote most of its records itself. Shared drives and inboxes collect paywalled industry reports, purchased market studies, standards documents, client deliverables, vendor manuals and contact lists bought from data vendors.

None of that belongs in a licensed package unless the company holds rights to license it, which is rare. A rights review should search for it by source, file type and folder, then exclude it or document why it may be included.

Your own archive may hold acquired material too
MaterialWhere it tends to sitUsual treatment
Paywalled reports and purchased researchSharePoint, Google Drive, email attachmentsExclude
Industry standards and codesEngineering libraries and project foldersExclude
Client deliverables and drawingsProject folders, Procore, BluebeamExclude unless the contract allows
Purchased contact or lead listsCRM imports in Salesforce or HubSpotExclude; also a privacy question
Vendor manuals and catalogsService libraries and ERP attachmentsUsually exclude; check vendor terms
Staff-written notes about any of the aboveTickets, project reviews and playbooksOften licensable after review

What buyers now ask about acquisition#

Buyers now ask suppliers to show how each record family was created or obtained, and to warrant that no unlicensed third-party material is included. The questions are short, but they need documentary answers rather than a general assurance.

Provenance standards point the same way. The Data & Trust Alliance's Data Provenance Standards organize dataset metadata into Source, Provenance and Use groups, which gives buyers and suppliers a shared vocabulary for describing where data came from and what it may be used for.

Illustrative: a consulting firm clears its shared drive#

Illustrative: a fictional management consulting firm wants to license proposals, project reviews and internal playbooks stored in SharePoint and its CRM. A first scan of the drive shows that this material is mixed with purchased industry studies, paywalled research and client deliverables.

The firm's counsel treats all third-party material as out of scope from the start. The team excludes files by source and folder, flags client deliverables by project code, and keeps the consultants' own lessons-learned notes after client names are removed.

The resulting package is smaller than the drive but clean on acquisition. Everything in it was written by the firm's staff under employment terms that leave the work with the firm, and the file records how each excluded category was found.

How SourceX documents acquisition#

Acquisition is one of the first things SourceX checks. Third-party material found during the Rights step of the SourceX five-step transaction, Supply, Rights, Preparation, Approval and Delivery, is excluded or documented before preparation begins, and each approved package carries a SourceX Evidence Packet covering provenance, licensing rights, permitted use, the privacy record and release authorization.

The supplier approves each step, and records are licensed, not sold outright. SourceX works only with records the supplier created or holds rights to, and what it may then do with a deidentified dataset is defined in the signed agreement.

Frequently asked questions

Does a settlement make pirated data lawful to use?

Not necessarily. A settlement resolves claims between specific parties on terms the court approves; it does not by itself license future use of the same material. What a settlement permits or requires, including any destruction of copies, depends on its terms, so read the settlement documents rather than assuming.

If a developer bought a copy of a book, can it train on that copy?

Owning a lawful copy helped in Bartz v. Anthropic, where scanning purchased print books and training on them was held fair use on its facts. That is not a general rule, and other factors, especially market harm, still apply. Developers that want certainty license the material instead.

Can a buyer be liable for data a supplier obtained unlawfully?

Possibly, which is why buyers ask suppliers to prove acquisition and to give warranties and indemnities. A supplier licensing its own records with documented rights removes most of that concern. Where third-party material is mixed in, exclude it or document the rights that cover it.

Is scraped public web data the same as pirated data?

No. Pirated material usually comes from sources distributing copies without permission, while public web pages were openly accessible. Scraping still raises questions about site terms, crawler signals and fair use, so it sits in a contested middle position rather than at either end of the range.

Do these cases affect records that were never published?

Only indirectly. Private operational records are not available from public sources in the usual sense, and if someone took them without permission, other laws such as trade secret and computer access laws would apply. The main effect is that buyers value suppliers who can document lawful acquisition.

Sources

  • The Data Provenance Initiative released a first audit covering 44 data collections that span more than 1,800 fine-tuning text-to-text datasets, documenting their sources, licenses, creators and other metadata. Source
  • The Data & Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups: Source, Provenance and Use. Source
  • On June 23, 2025, Judge William Alsup held in Bartz v. Anthropic that using books to train large language models was exceedingly transformative and therefore fair use. Source
  • Judge Alsup held that Anthropic's purchase and scanning of print books, with the originals discarded, was fair use. Source
  • Judge Alsup ruled that downloading pirated books from sources including Books3, Library Genesis and Pirate Library Mirror to build a permanent central library was not fair use and left that claim for trial. Source
  • Anthropic agreed to pay $1.5 billion to settle Bartz v. Anthropic, covering about 482,460 pirated books at roughly $3,000 per work; the court granted final approval on July 20, 2026. Source
  • Under the Bartz settlement, Anthropic must destroy the original torrented files and any copies derived from them. Source
  • The Copyright Office's Part 3 pre-publication report states that commercial use of vast troves of copyrighted works to produce competing content, especially through illegal access, goes beyond established fair use boundaries. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify