Text and language data
Licensing Scholarly Journal Content for LLM Training
Quick answer
To license scientific papers for AI training, you negotiate with whoever controls full-text rights: the publisher (for subscription journals it owns), a hosting aggregator such as JSTOR (for content it distributes on publishers' behalf), or a society that retains copyright. The usable corpus is then shaped by author opt-in or opt-out terms, article version, and retraction status. Budget as much effort for per-article rights metadata as for price, because those fields determine what you can actually train on and defend later.
By SourceX Editorial · Updated
This guide covers paid publisher and aggregator licensing. Free open-access subsets and their CC license tiers are covered in open-access scientific full text for training; the broader pre-training picture sits in the text and language data hub.
Who actually holds training rights in a journal article
The rights holder is whoever the author agreement assigned copyright or an exclusive license to, and that varies article by article. Under a traditional copyright transfer agreement, the publisher holds the rights; under an exclusive license to publish, the author keeps copyright but the publisher controls reuse; under CC BY or CC BY-NC open access, terms flow from the license, not a negotiation. Society journals published by a commercial house often split ownership between the society and its publishing partner, so check whether the publishing agreement lets the partner sublicense content for AI use.
Aggregators add a further layer. A journal hosted on JSTOR has reported that the platform is making hosted content available to AI developers, with individual authors able to opt out of inclusion [2]. That means the aggregator can deliver a multi-publisher corpus under one contract, but only for publishers that have elected in, and minus any author exclusions.
The U.S. Copyright Office's pre-publication Part 3 report (still pre-publication as of October 2026) concludes that many acts involved in training, such as copying works into datasets, may implicate the reproduction right, and it treats fair use as a case-by-case question [6]. That uncertainty is a common reason to license paywalled text rather than rely on an exception. Outside the US, the UK's section 29A text and data analysis exception remains limited to non-commercial research, so it does not cover commercial model training [9].
Publisher, aggregator and society deals compared
Each route trades corpus breadth against rights certainty. Deal trackers such as the Ithaka S+R Generative AI Licensing Agreement Tracker catalog agreements by publisher, purchaser, deal type and reported size, and are the fastest way to benchmark what peers have signed [1]. Press coverage has reported that publishers are earning millions from licensing papers for AI training, though most individual terms remain undisclosed [4].
| Route | Typical scope | Rights certainty | Main failure mode |
|---|---|---|---|
| Large commercial publisher | Own journal portfolio, back file plus current issues | High for owned titles | Society titles and author-retained CC articles carved out late |
| Aggregator program | Many publishers' hosted content | Medium; depends on each publisher's election | Opted-out authors and withdrawn publishers change the corpus mid-term |
| Society or university press | Single field, deep back file | High, often with author opt-in | Small volume; opt-in yields a fraction of the list |
| Independent academic publisher with opt-in | Books and journals | High for opted-in works | Revenue share and consent tracking per work [3] |
Book-side deals show the same pattern. One trade-book agreement was reported to require author opt-in and to pay a per-title fee split between author and publisher [5]. Scholarly publishers that follow this model, such as Boydell & Brewer, will not license works for LLM training without the author's affirmative agreement and share revenue under contract [3]. If you are also sourcing monographs, compare with book corpus licensing routes.
Handling author opt-in and opt-out
Author consent mechanics decide your effective corpus, so specify them in the contract rather than accept them as a list delivered later. An opt-out model starts with the full catalog and removes exclusions; an opt-in model starts empty and adds consenting authors. Opt-in corpora are smaller but carry a cleaner consent record for each work.
Ask for four things in writing:
- Consent granularity. Is consent recorded per author, per article, or per co-author? Multi-author papers need a rule for when one co-author opts out.
- Timing. Is the opt-out window closed before delivery, or can authors opt out during the term? The second case needs a removal feed.
- Removal obligations. What happens to removed articles already in a training run, a tokenized shard, or a retrieval index? Distinguish training from grounding rights using grounding license vs training license.
- Evidence. A machine-readable consent flag and date per DOI, not a PDF list of names.
Metadata that makes a scholarly corpus usable
A licensed scholarly corpus is only as defensible as its per-article metadata, and the DOI is the join key for everything else. Require each record to carry the DOI, the article version, the license or rights basis, the retraction and correction status, and the consent flag. An audit of 1,800+ text datasets reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites, so do not assume a corpus-level license statement is accurate at article level [7].
Version matters because the accepted manuscript and the version of record can differ in text, figures and corrections. NISO's Journal Article Versions practice names stages from Author's Original through Accepted Manuscript, Proof, Version of Record and Corrected Version of Record. Contracts often grant rights only to one version, so record which one you received.
Retraction status should be refreshed, not captured once. Crossref now distributes the openly licensed Retraction Watch retraction data alongside its DOI metadata, so you can poll status by DOI. A missing flag does not prove an article stands, because publisher-supplied retraction metadata is inconsistent; check both sources and re-poll on a schedule. Pair this with factual accuracy checks for licensed text.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doi": "10.0000/example.2024.0117",
"journal_issn": "0000-0000",
"publisher_rights_holder": "Example Society (published by Example Press)",
"article_version": "VoR",
"version_standard": "NISO RP-8-2008",
"rights_basis": "copyright_transfer",
"author_consent": {"model": "opt_out", "status": "included", "as_of": "2026-09-30"},
"retraction_status": {"source": "crossref_rest_api", "state": "none", "checked": "2026-10-01"},
"corrections": ["10.0000/example.2024.0117.corr1"],
"full_text_format": "JATS XML",
"permitted_uses": ["pretraining"],
"excluded_components": ["third_party_figures", "supplementary_data"]
}
Describe the delivery itself with a dataset-level manifest. Croissant, a schema.org-based JSON-LD format from MLCommons, records dataset metadata, file resources and record structure, which lets your ingestion pipeline validate shards against the contract [8].
Format and content components to scope
Request JATS XML for journals and BITS XML for books where available, because PDF extraction loses section structure, equations and reference lists and introduces hyphenation and column-order errors. Specify whether you receive abstracts only, full text, references, figure captions, tables and supplementary files. Third-party figures and reproduced images are often excluded from the publisher's grant even when the article text is included.
Decide how you will treat retracted, corrected and expression-of-concern items before the data arrives. A common pattern is to exclude retracted articles from pre-training, keep corrected versions of record with the correction notice attached, and log every exclusion by DOI so the decision can be reproduced. Also screen any supplemental crawl against the licensed DOIs to catch copies from shadow libraries, using pirated-source screening.
Clauses to negotiate before signing
The clauses that matter most are scope, removal, and model outputs. Use this checklist on any scholarly license:
Illustrative example: invented to show structure; it does not describe an available dataset.
| Clause | Question to settle | Why it matters |
|---|---|---|
| Permitted uses | Pre-training, fine-tuning, evaluation, retrieval grounding, or all? | Training and retrieval are priced and risk-assessed differently |
| Corpus definition | Titles, years, versions and components, by ISSN and DOI list | Prevents silent carve-outs at delivery |
| Updates | Are new issues and corrections delivered as a feed? | Keeps the corpus current; see recurring text feeds |
| Opt-out handling | Removal duty for future runs vs trained weights | Weights cannot easily forget; agree what is required |
| Attribution | Any citation or display obligations in outputs | Affects product design for answer engines |
| Term and survival | What survives expiry for models already trained | Defines whether a model can stay in production |
| Audit | Can you verify delivered DOIs against the agreed list? | Catches gaps and duplicates |
| Retractions | Who notifies whom, and how fast | Keeps known-bad science out of later runs |
Standardized model licenses for AI use of scholarly content have been proposed in industry discussions, but as of October 2026 most terms are still negotiated deal by deal. Compare drafts against general licensed text corpus planning.
Where operational research content fits
Peer-reviewed journals are not the only source of scientific reasoning text. Companies produce technical reports, test protocols, lab notebooks and research memos that never reach a journal. SourceX sources operational datasets from US companies, including engineering records and documents, on request rather than from stock, and every release is approved by the supplying company. Its scope does not include scraped web content, and a request does not guarantee a match. See the research deliverables page for how that content is described, or tell SourceX what research text you need.
Each dataset SourceX delivers is rights-reviewed for ownership and consents and comes under a license that defines records, uses, term and delivery. Personal details such as names and emails are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
License scientific and research text through SourceX
If your scientific corpus needs proprietary research documents or engineering records from US companies alongside publisher licenses, describe the data rather than the businesses that might hold it. SourceX looks for suppliers, assesses data and licensing permissions, and agrees pricing and allowed uses in a license; nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.
Sources
- Ithaka S+R, "Generative AI Licensing Agreement Tracker". https://sr.ithaka.org/our-work/generative-ai-licensing-agreement-tracker/
- Trinity College Dublin, Hermathena, "Hermathena and JSTOR". https://www.tcd.ie/classics/assets/pdf/hermathena/Hermathena_JSTOR.pdf
- Boydell & Brewer, "Boydell & Brewer statement on AI licensing". https://boydellandbrewer.com/?p=35458
- InfoDocket, "Report: Publishers Are Selling Papers to Train AIs and Making Millions of Dollars" (2024). https://www.infodocket.com/2024/12/09/report-publishers-are-selling-papers-to-train-ais-and-making-millions-of-dollars/
- EMARKETER, "Microsoft and HarperCollins sign AI licensing deal, but author opt-in still required" (2024). https://www.emarketer.com/content/microsoft-harpercollins-sign-ai-licensing-deal--author-opt-in-still-required
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Akhtar et al. (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988, section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.