Skip to content

Text and language data

News Archive Text for AI Training: Depth, Metadata and Training Rights

Quick answer

News archive data for AI training is licensed, not scraped: you negotiate with a publisher, wire service or syndicator for a defined slice of its back catalog, for a defined term, with training rights spelled out separately from display or grounding rights. The variables that decide value are archive depth (how many years, and how complete), per-article metadata (dates, sections, corrections, wire versus staff), carve-outs for content the publisher does not own, and what happens to trained weights when the term ends.

By SourceX Editorial · Updated

Public deal reporting shows the shape of the market. The 2023 AP and OpenAI agreement licensed part of AP's text archive going back to 1985, reportedly for two years, with undisclosed financial terms [1][3]. A later deal reported for Time covered archives from the past 101 years and added citation and link-back duties in model responses [2]. Commentators count many such deals with major news organizations since mid-2023 [4], while publishers increasingly push usage-based "grounding" deals instead of one-time training payments [5].

What news archives add that web crawls do not

News archives add dated, edited, fact-checked prose with a reliable temporal axis, which is what continued pre-training for factual and temporal knowledge needs. A crawl captures whatever was live on the crawl date, often paywall teasers, syndicated duplicates and boilerplate, and it rarely preserves the original publication timestamp or later corrections. A licensed archive can deliver the full text of every version, keyed to a stable article ID.

That matters for three model behaviors. Temporal grounding improves when each document carries a trustworthy published_at, so you can build time-sliced training mixes or knowledge cutoffs that are actually true. Factuality improves when retractions and corrections are linked to the article they amend rather than silently overwritten. Deduplication improves when wire stories are flagged, because one AP or Reuters item can appear verbatim across hundreds of member and subscriber sites.

For the wider context on why proprietary text beats crawls, see proprietary text data beyond web crawls and the cluster hub on licensed text datasets for LLM training.

Archive depth and term: the two numbers that set the deal

Depth and term are the first two numbers to pin down, because they determine both the token yield and your exposure when the license lapses. The AP deal paired roughly four decades of depth with a two-year term [1], which illustrates a common asymmetry: deep history, short permission.

Depth is rarely uniform. Pre-digital decades are often OCR from microfilm or print scans, with column-merge errors, hyphenation breaks and missing bylines; the born-digital era (roughly from the late 1990s for most publishers) is cleaner but may have gaps from CMS migrations. Ask for a per-year article count and a sample from each era before you price on "since 1985".

Term needs a weights clause, not just a data clause. When the license ends, you need to know whether you must delete the raw corpus, whether models already trained may keep shipping, and whether checkpoints derived during the term are affected. Read pre-training data license rights for the grant language to request.

Metadata fields to require per article

Require metadata at the article level, because most downstream filtering, dedup and audit work depends on it. Publishers usually hold these fields in their CMS or in NewsML-G2 or NITF exports; ask for them as a JSON or Parquet sidecar keyed to the text file.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "article_id": "pub-2009-04-17-000812",
  "version": 3,
  "published_at": "2009-04-17T14:05:00Z",
  "updated_at": "2009-04-18T09:12:00Z",
  "section": "Business",
  "headline": "Regional lender posts first quarterly loss",
  "byline": "Staff",
  "content_origin": "staff",
  "wire_source": null,
  "rights_status": "owned",
  "correction_of": null,
  "has_correction": true,
  "correction_text": "An earlier version misstated the loss figure.",
  "language": "en-US",
  "source_format": "born_digital",
  "ocr_confidence": null,
  "word_count": 642,
  "embedded_media_excluded": true
}

The fields that most often go missing are content_origin, rights_status and the correction linkage. Without them you cannot apply carve-outs or prefer final versions, and the publisher cannot later prove what you were allowed to use. The full field list for text corpora is on text corpus metadata fields, and the rights encoding pattern is in the permitted-use metadata schema.

Wire, syndicated and freelance content: the carve-out problem

A publisher can license only what it owns or has sublicensable rights to, so wire copy, syndicated columns and freelance work often need to be excluded or separately cleared. A typical regional daily's archive mixes staff reporting with AP or Reuters wire, syndicated opinion columns, freelance features and reader letters, each with different underlying rights.

Practical checks before signing:

  • Ask what share of each year's articles are wire, syndicated or freelance, and how the publisher identifies them (byline patterns, a CMS source field, or slug conventions).
  • Confirm whether freelance contracts from earlier decades granted electronic or derivative rights at all; many pre-2000 contracts did not contemplate either.
  • Require photos, graphics, embedded tweets and video transcripts to be stripped or flagged unless separately cleared.
  • Ask for a warranty-backed exclusion list rather than a best-effort filter.

The same issue arises in any corpus that quotes or embeds others' work; see third-party content inside licensed corpora.

Training rights versus display and grounding rights

Training rights and display rights are separate grants, and a news archive license for pre-training should say which it covers. Display-side terms such as citation, attribution and link-back obligations appeared in reported deals [2], and publishers increasingly price live grounding by usage rather than as a lump sum [5]. Those terms belong to retrieval products, not to weights.

If you need both, negotiate them as separate schedules with separate terms. Grounding and vector-index questions are covered on embedding and vector index rights for licensed content. For news archive training, the clean grant usually names: copying and transformation for training, internal evaluation, derivative model weights, and the outputs of those models, with explicit exclusions for redistribution of the text itself.

Direct publisher deals versus aggregated licensing

Direct deals give you the deepest archive and cleanest provenance for one title; aggregated licensing gives breadth across many titles with less negotiation. Syndicators market full-text articles with machine-readable metadata for LLM training, sourced from many publishers [6], though their rights coverage and metadata depth are self-reported and need checking. Collective and blanket schemes are a third route, discussed on collective licenses for AI training.

Illustrative example: invented to show structure; it does not describe an available dataset.

RouteStrengthTypical gapAsk before signing
Direct publisherDeep back catalog, full CMS metadata, corrections linkageOne voice and region; long negotiationPer-year counts, carve-out list, weights clause
Wire serviceBroad topical coverage, consistent styleHeavy duplication with member sitesStory ID for dedup against other sources
Syndicator or aggregatorMany titles under one agreementUneven metadata; upstream rights vary by titlePer-title chain of rights, opt-outs by title
Trade and B2B titlesDense domain vocabularySmaller volumeSee trade publication archives

Transparency duties that touch news training data

If you place a general-purpose AI model on the EU market or make a generative AI system available to Californians, you will have to describe training data publicly, so license terms should let you name the source category. Article 53(1)(c) of the EU AI Act requires GPAI providers to maintain a copyright policy that honors rights reservations, and 53(1)(d) requires a public summary of training content [7]; the Commission's template for that summary is dated 24 July 2025 [8].

California AB 2013 requires developers of generative AI systems to post documentation about training datasets, including their sources and whether they include copyrighted material; that documentation was due by 1 January 2026 [9]. A confidentiality clause that forbids you from naming the publisher or describing the archive can collide with these duties, so agree disclosure language up front. Keep the license, the carve-out list and the metadata manifest together as your audit record.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Quality checks before a news archive enters the training mix

Run quality checks per era and per content origin, because OCR decades and wire-heavy years fail differently. A short acceptance checklist:

  • Sample at least a few hundred articles from each decade and measure OCR character error against a hand-corrected subset.
  • Count near-duplicates (MinHash or SimHash) within the archive and against your existing web crawl; wire copy drives most overlap.
  • Verify that published_at is the original publication time, not the digitization or CMS-migration timestamp.
  • Confirm that corrected articles carry the final version and that retracted items are flagged or removed.
  • Check that paywall prompts, newsletter sign-up boxes and related-article widgets were stripped.

Factual-accuracy screening for licensed knowledge text is covered on factual accuracy checks for licensed text. If you need ongoing updates rather than a static archive, see recurring text data feeds.

Request template for a news archive license

Describe the data you need, not a named publisher, and you will get comparable answers from different holders.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Content: English-language US local and regional news, staff-written only, 1995 to present.
  • Use: continued pre-training and internal evaluation; no display, retrieval or redistribution of text.
  • Depth: per-year article counts and a 200-article sample per decade.
  • Metadata: article_id, version, published_at, updated_at, section, byline, content_origin, rights_status, correction linkage, source format.
  • Exclusions: wire, syndicated, freelance without electronic rights, images, embedded third-party media.
  • Delivery: JSONL or Parquet text plus sidecar metadata through an access-controlled transfer.
  • Term questions: raw-data deletion at end of term, status of trained weights, disclosure wording for training-data summaries.

How SourceX fits news archive sourcing

SourceX sources operational datasets from US companies on request and manages licensing; it does not hold news archives in stock, and a request does not guarantee a match. Its focus is business data such as support and sales histories, engineering records, documents, and finance and legal workflows, so a classic newspaper archive is a narrower fit than domain text held by operating businesses. SourceX does not source scraped web content.

When a US business does hold described text, every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with supplier approval for every release. Buyers in publishing can also review the media and publishing buyer page and data licensing rules for media and publishing companies, and can describe a dataset request to SourceX.

License news archive text for training

SourceX sources datasets from US companies on request and manages the process from finding and assessing a holder through license agreement, transaction and ongoing management. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Tell SourceX what text data you need.

Frequently asked questions

Is a historical news archive enough for pre-training on its own?

No. Even a deep archive is small relative to pre-training budgets, so news text is usually a high-quality slice in a broader mix or a continued pre-training set for temporal and factual knowledge. See licensed text corpora for LLM pre-training for volume planning.

Should OCR-era articles be included?

Include them only after measuring error rates by decade. Heavily garbled pages teach the model noise; a common approach is to keep high-confidence OCR, exclude the rest, and record the threshold in your data card.

Does a training license let us quote articles in product answers?

Not unless the license says so. Quoting, citing or linking is a display right with its own terms [2][5], and should be negotiated separately from the training grant.

Sources

  1. CTV News (Associated Press), "ChatGPT maker OpenAI signs deal with Associated Press to license news stories" (2023). https://www.ctvnews.ca/sci-tech/article/chatgpt-maker-openai-signs-deal-with-associated-press-to-license-news-stories/
  2. The Morning Context, "OpenAI and Time strike content licensing deal" (2024). https://themorningcontext.com/internet/openai-and-time-strike-content-licensing-deal?type=short
  3. Penningtons Manches Cooper, "Associated Press and Open AI: the first news sharing and technology partnership" (2023). https://www.penningtonslaw.com/insights/associated-press-and-open-ai-the-first-news-sharing-and-technology-partnership/
  4. ProMarket, "The false hope of content licensing at internet scale" (2025). https://www.promarket.org/2025/11/19/the-false-hope-of-content-licensing-at-internet-scale/
  5. Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?". https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
  6. SyndiGate, "Fully licensed premium datasets for training LLMs and AI applications". https://www.syndigate.info/fully-licensed-premium-datasets-for-training-llms-and-ai-applications/
  7. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  8. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  9. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data