Skip to content

Text and language data

Open-Access Scientific Full Text for Training: License Tiers and Bulk Access

Quick answer

Only part of open-access science is usable for commercial training. In PubMed Central, the Open Access Subset is the portion with machine-readable reuse licenses, and the license on each article governs [1]. Articles under CC0, CC BY, CC BY-SA or CC BY-ND form the commercial-use tier; NC-licensed and custom-licensed articles do not [2]. On arXiv, authors choose a license per paper, so filter by license metadata before treating any full text as commercially usable.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why "open access" is not a training license

Open access describes free reading, not a uniform grant of reuse rights, so a commercial pre-training team has to resolve rights article by article. PMC as a whole contains many articles that are free to read but that may not be text-mined or redistributed; the OA Subset exists precisely because only some articles carry licenses that allow broader reuse [1]. The practical consequence: the user, not NCBI, carries copyright compliance, and the license statement in each article is the operative document.

The common failure mode is aggregation drift. A team downloads a "PMC corpus" from a mirror or a hub, the license field gets flattened to a single value, and NC or custom-licensed articles end up in a commercial training mix. Dataset audits show how often this happens: the Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [5]. Treat any repackaged scientific corpus as unverified until you can trace each record back to its source license.

For the broader question of how public sources differ from licensed proprietary data, see public vs proprietary data for AI.

The PMC OA license tiers and what each permits

PMC OA content is distributed in three license groupings, and only the commercial grouping is a candidate for commercial model training [2]. The grouping is a starting filter, not the final answer, because the license text on each article is what binds [1].

Tier (bulk grouping)Licenses includedCommercial training candidateMain obligations and caveats
Commercial use allowedCC0, CC BY, CC BY-SA, CC BY-NDYes, subject to per-article checkAttribution for BY variants; ShareAlike terms on BY-SA; BY-ND bars distributing adapted versions of the article itself
Non-commercialCC BY-NC, CC BY-NC-SA, CC BY-NC-NDNo, without a separate license from the rights holderUsable for non-commercial research in some jurisdictions; never mix into a commercial training set
OtherNo machine-readable CC license, or custom termsOnly after reading the specific termsCustom publisher licenses vary; some tools skip this grouping entirely

Two details matter for engineers. First, figures, tables and supplementary files can carry different rights from the article text, often with third-party credit lines, so text-only extraction is the cleaner path. Second, publisher text-and-data-mining policies can add constraints on outputs: one publisher allows commercial text mining but may limit public reproduction of text to short snippets [3]. If your model or retrieval product surfaces passages verbatim, map those limits into your output filters.

Bulk access channels for PMC and what to pull

Use NCBI's documented bulk services rather than crawling article pages; NLM distributes the subset through bulk packages, cloud datasets split by license, and BioC XML and JSON [1]. Check NCBI's current terms for which automated retrieval routes are permitted, since crawling HTML pages is both a terms risk and a quality problem.

Pick the channel by what you need downstream:

  • Bulk packages (cloud or FTP): compressed packages of JATS XML or plain text grouped by license tier, plus file lists that map each PMCID to its license. Best for pre-training snapshots [1].
  • BioC API (XML or JSON): passage-segmented text with offsets, convenient for domain adaptation, entity annotation and RAG chunking [1].
  • OAI-PMH: incremental harvesting of new and updated records, useful for keeping a corpus current.
  • E-Utilities: targeted lookups by query or ID, better for building eval sets than whole-corpus pulls.

Keep the per-article license string and the file list version with every record. Articles can be retracted, corrected or re-licensed between snapshots, and a clean lineage record is what lets you answer a takedown or a disclosure request later.

arXiv full text: per-paper licenses and the distribution default

arXiv full text is not a single licensed corpus; each submission carries the license its authors selected at submission. Most papers use arXiv's default license, which is a distribution grant to arXiv rather than a reuse grant to third parties, so treat it as not covering commercial training or redistribution. Confirm the current wording on arXiv's license page before relying on this.

The CC options change the picture. CC BY 4.0 allows reuse, adaptation and redistribution with attribution, while NC and ND variants carry the same limits they do in PMC. Different versions of the same paper can carry different licenses, so key your filter to the version you actually ingest.

For bulk work, harvest metadata through arXiv's OAI-PMH interface, which exposes a license field per record, and review arXiv's API terms of use before any harvesting. A practical pipeline harvests metadata, keeps only versions whose license URL is CC0, CC BY or CC BY-SA, and then fetches full text for that set alone. Expect the commercially usable share to be a minority of arXiv; treat the default-license papers as read-only references, not training text.

Filtering checklist for an open-access scientific corpus

A defensible corpus is built from a per-record license manifest, not from a dataset name. Use the checklist below as a gate before any open-access scientific text enters a commercial training mix.

Illustrative example: invented to show structure; it does not describe an available dataset.

StepCheckPass conditionRecord in manifest
1Source channelPulled via a documented bulk service (PMC bulk packages, cloud, BioC; arXiv OAI-PMH) [1]channel, snapshot date, file list version
2License stringMatches an allow list (CC0, CC BY, CC BY-SA; CC BY-ND only if no adapted article is redistributed)license URL as published, not a normalized label
3Version matchLicense belongs to the exact version ingested (arXiv v1 vs v2)version ID
4Embedded contentFigures, tables and supplements excluded or separately clearedextraction scope
5Publisher TDM termsNo output reproduction limits violated (for example a snippet cap) [3]policy reference
6Retractions and correctionsRetracted items removed or flaggedretraction status, check date
7AttributionCitation metadata (DOI, authors, license) retained for BY obligationsDOI, PMCID or arXiv ID
8DeduplicationSame paper across PMC, arXiv and publisher mirrors resolved to one license decisioncanonical ID, chosen source

An illustrative manifest record:

{
  "record_id": "pmc:PMC0000000",
  "source": "PMC OA bulk package, commercial-use tier",
  "snapshot": "2026-10-01",
  "license_url": "https://creativecommons.org/licenses/by/4.0/",
  "license_tier": "commercial",
  "version": "published",
  "scope": "body_text_only",
  "attribution": {"doi": "10.0000/example", "first_author": "Example"},
  "retraction_checked": "2026-10-01",
  "decision": "include"
}

Step 8 is where corpora usually break. The same paper often appears as an arXiv preprint under the default license and as a CC BY journal article in PMC, and the two copies carry different rights. Pick the copy whose license you have verified and drop the other.

Disclosure and jurisdiction checks that touch open-access corpora

Open licenses settle copyright permission, but they do not remove disclosure duties or change how exceptions apply. In the UK, the CDPA s29A text and data analysis exception covers only non-commercial research, so a commercial lab cannot rely on it to use NC-licensed papers [6]; see the UK text and data mining guide for detail as of October 2026.

In California, AB 2013 requires developers of generative AI systems made available to Californians to post documentation about training data, with postings due 1 January 2026 [7]. A per-record license manifest like the one above makes those disclosures accurate rather than estimated. For CC compatibility questions across mixed sources, use the open data license compatibility matrix, and for a broader review method see auditing open dataset licenses.

Where open-access science stops and licensed text begins

Open-access science gives a strong, citable domain baseline, but it covers published findings, not the working language of operations. Openly licensed pre-training corpora such as the Common Pile already include research papers, so a commercial-tier PMC and arXiv slice is largely a baseline that other teams also have [4]. See openly licensed text corpora for commercial training for what those corpora cover.

What OA papers rarely contain is the text produced inside organizations: protocol deviations, lab notebooks, quality investigations, support tickets, engineering records and finance or legal workflows. Paywalled journal content is a separate negotiation covered in licensing scholarly journal content. For proprietary operational text, SourceX sources datasets from US companies on request, rather than holding stock, and every release is approved by the supplying company; you can describe the data you need to SourceX. The text data hub maps the other options.

Sourcing scientific and operational text beyond open access

Once the open-access baseline is filtered, the gap is usually domain text that never gets published. SourceX sources operational datasets from US companies, such as engineering records, documents, support histories and finance or legal workflows, and manages licensing and ongoing purchases. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery; a request does not guarantee a match, so start by describing the text data you need at SourceX for buyers.

Frequently asked questions

Is the whole PMC Open Access Subset usable for commercial training?

No. Only the commercial-use grouping (CC0, CC BY, CC BY-SA, CC BY-ND) is a candidate, and the license statement on each article still governs [1][2]. Non-commercial and custom-licensed articles need separate permission.

Can I crawl PMC article pages instead of using the bulk packages?

Avoid it. NCBI documents specific bulk and API routes for automated access, and its terms restrict other systematic retrieval; check the current terms. Bulk packages and BioC output also give cleaner, structured text with the license attached [1].

Does a CC BY-ND license block training?

ND bars distributing adapted versions of the work. Whether training is an adaptation is a legal question to resolve with counsel; many teams include BY-ND text only where no article content is redistributed in modified form.

Sources

  1. National Library of Medicine via Data.gov, "PubMed Central Open Access Subset (PMC OA)". https://catalog.data.gov/dataset/pubmed-central-open-access-subset-pmc-oa
  2. CASRAI, "PMC open access (tag archive)". https://casrai.org/wp/tag/pmc-open-access/
  3. Rockefeller University Press, "Journal of Cell Biology editorial on text and data mining (jcb.201303016)" (2013). https://rupress.org/jcb/article-pdf/201/1/7/689215/jcb_201303016.pdf
  4. Kandpal et al., "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  5. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  6. UK Intellectual Property Office (GOV.UK), "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  7. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data