Skip to content

Text and language data

Recurring Text Data Feeds for Continual Model Updates

Quick answer

An ongoing text data feed for LLM training is a licensed supply of newly created text delivered on a fixed cadence, with a date stamp on every record, a delta format that separates inserts from corrections and deletions, and a license term that sets how long new text keeps arriving. Buyers should specify freshness in terms of temporal coverage and knowledge cutoff, not just delivery frequency, and should hold out evaluation text dated after each training cutoff.

By SourceX Editorial · Updated

What a recurring feed buys that an archive does not

A recurring feed buys temporal coverage: text written after your last training run, which is what moves a model's knowledge cutoff forward. An archive purchase, such as the deep collections covered in news archive text for AI training, gives you depth up to one date and nothing after it. Continual pre-training and knowledge-refresh runs need the opposite shape, a thin but steady slice of new documents with reliable creation dates.

The market already prices these as separate structures. One corpus licensor notes that while search indexing is welcome, large-scale AI training requires a separate license [1]. Platform deals have been reported as giving access to "real-time, structured" content [2], while a widely reported 2023 news archive license was announced as a time-bounded agreement [3]. The term is the hidden variable: once it ends, new text stops even if the archive license survives.

Choosing a delivery cadence for text

Match cadence to how fast the knowledge in the text goes stale and how often you actually retrain, not to what the supplier can technically push. A daily feed into a model retrained twice a year buys storage, not freshness.

  • Real-time or daily (API or streaming): news, support tickets, forum threads and other text used for retrieval grounding or fast knowledge updates. Compare the mechanics in API and streaming feeds vs batch files.
  • Monthly: operational text such as case notes, engineering tickets or sales call summaries feeding domain adaptation runs.
  • Quarterly: reference text, documentation and long-form writing feeding scheduled continual pre-training.
  • Annual: stable professional corpora where the license renewal itself is the refresh point; see can I license data every year.

For how other buyers set this, read how often AI buyers want fresh data.

Date stamps and knowledge cutoffs

Every record needs at least two timestamps, because the date a document was written and the date it was delivered answer different questions. created_at (authored or published) defines what your model's knowledge cutoff actually is; ingested_at or delivered_at defines what was in a given batch. A feed that only stamps delivery dates will quietly mix back-catalog text into "fresh" batches, and your cutoff claim becomes unverifiable.

Ask for a third field, modified_at, for documents edited after first publication, plus the supplier's timezone convention (ISO 8601 with offset, or UTC). If you publish a cutoff date for a model, you need these fields to defend it. California's AB 2013 requires developers of generative AI systems offered to Californians to post documentation about training data, including the time period during which data was collected, with postings due on or before 1 January 2026 (status as of October 2026) [9]. Per-record dates make that disclosure a query rather than an estimate.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Delta formats: inserts, corrections and deletions

A usable text feed distinguishes three operations per record: insert, update and delete. Without an explicit operation code, a corrected news story or an edited support article arrives as a second, near-duplicate document, and your deduplication pass has to guess which version is authoritative.

Common delivery shapes:

  • Change files in Parquet [5] or JSON Lines, partitioned by delivery date, each row carrying op (insert, update, delete), a stable doc_id and a monotonic version.
  • Shared tables over Delta Sharing [4], where the supplier maintains the table and you read new versions; the protocol is open and documented in the Delta Lake project, which reduces lock-in to one cloud.
  • Full snapshots plus a manifest, simpler to reason about but expensive at scale; the tradeoffs are in incremental vs full-refresh deliveries.

Describe each batch with a machine-readable manifest. Croissant, a JSON-LD vocabulary built on schema.org, can describe dataset-level metadata and the files in each release [6], which lets your loaders validate a batch before it touches a training mix. When the schema itself changes mid-term, follow schema evolution across recurring deliveries.

Handling corrections and deletions after training

Corrections and deletions are the text-specific hard problem, because a retracted article or a removed post may already sit inside a trained checkpoint. Decide up front what each tombstone obliges you to do: purge from the raw store, exclude from the next training mix, rebuild retrieval indexes, or all three.

Common failure modes buyers report:

  • Silent overwrite: the supplier replaces text in place with no version bump, so you cannot tell which copy trained which model.
  • Orphaned derivatives: a deleted document survives in your tokenized shards, embeddings and deduplicated mixes because lineage stops at the raw file.
  • Late tombstones: deletions arrive in a quarterly batch while the corrected text arrived daily, leaving a window where both versions coexist.

Keep a lineage table mapping doc_id to every shard, index and training run that consumed it. The mechanics are covered in propagating deletions and corrections through recurring deliveries; the commercial side, including what happens to delivered text when a term ends, belongs in your ongoing data supply agreement.

Keeping evaluations clean as the feed grows

Hold out evaluation text dated after each training cutoff, or every new batch risks leaking your test set into training. Temporal leakage, where future data informs a model judged on an earlier point in time, is a well-documented failure in time-ordered data, and the standard fix is a time-based split rather than a random one [7].

The stakes are concrete: OpenAI stopped reporting SWE-bench Verified scores after concluding the benchmark was contaminated, so gains increasingly reflected training-time exposure rather than capability [8]. For a recurring text feed, the practical rule is to freeze an evaluation window per release (for example, documents with created_at in the month after the cutoff), deduplicate it against all prior batches with near-duplicate hashing such as MinHash, and never let that window flow into the next training mix.

Feed specification template

Use this as the attachment to an RFP or term sheet so suppliers quote against the same structure. Pair it with SLAs for recurring data deliveries and SourceX's guide to data supplier SLAs.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
Text typeEnglish support tickets and resolution notesDefines domain and register
Temporal coverage per batchDocuments created in the prior calendar monthMoves the knowledge cutoff predictably
CadenceMonthly batch, by the 10thMatches retraining schedule
Delivery shapeParquet change files, partitioned by delivered_atLoader-ready, columnar
Required fieldsdoc_id, version, op, created_at, modified_at, delivered_at, language, textSupports dedup, lineage, cutoff claims
Correctionsop=update with incremented versionNo silent overwrite
Deletionsop=delete tombstone within the next batchEnables purge from shards and indexes
ManifestCroissant JSON-LD with row counts and checksumsBatch validation before ingest
Personal dataNames, emails, phones, account numbers removed or replaced; method recordedPrivacy review sign-off
Acceptance checkSample per batch against field completeness and date validitySee acceptance sampling
Term and renewalStated term; renewal triggers a fresh rights reviewBounds how long new text arrives

Sourcing recurring text from operating companies

Much of the freshest professional text, such as support conversations, engineering records, and finance and legal workflows, is produced inside operating businesses rather than published on the web. That text continues to be generated every day, which makes it a natural fit for a recurring feed, and it is covered in more depth in proprietary text data beyond web crawls and operational free-text notes.

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Buyers describe the data, not the businesses; SourceX looks for US companies that hold it, and every release is approved by the supplying company. A request does not guarantee a match, and nothing is contracted until a supplier agrees. You can describe a recurring text requirement to SourceX using the template above as a starting point. For the wider landscape of licensed corpora, start at the text and language data hub or the AI data hub.

Set up a recurring text feed with SourceX

SourceX sources operational text from US companies on request and runs the process from Find and Assess through Agree, Transact and Manage, so ongoing purchases stay under a license that defines records, uses, term and delivery. Every dataset is rights-reviewed, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe the text feed you need.

Frequently asked questions

Is a recurring feed the same as continual pre-training data?

Not exactly. A feed is the supply mechanism; continual pre-training is one way to consume it. The same monthly text batches can also refresh a retrieval index or a fine-tuning mix, and each use may need different rights in the license.

What happens to delivered text when the feed term ends?

That depends entirely on the license. Some reported news licenses carried a fixed term [3], so ask explicitly whether text delivered during the term may stay in trained models, in raw storage and in retrieval indexes after expiry.

How do I prove a model's knowledge cutoff from a feed?

Query the maximum createdat across every batch included in the training mix, excluding any held-out evaluation window. If records only carry delivery dates, you cannot prove it.

Sources

  1. Source Library, "Licensing" (2026). https://sourcelibrary.org/licensing
  2. TechCrunch, "OpenAI inks deal to train AI on Reddit data" (2024). https://techcrunch.com/2024/05/16/openai-inks-deal-to-train-ai-on-reddit-data
  3. CTV News, "ChatGPT maker OpenAI signs deal with Associated Press to license news stories" (2023). https://www.ctvnews.ca/sci-tech/article/chatgpt-maker-openai-signs-deal-with-associated-press-to-license-news-stories/
  4. Delta Lake project, "Read Delta Sharing Tables". https://docs.delta.io/delta-sharing/
  5. The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs
  6. MLCommons (arXiv:2403.19546), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  7. temporalcv documentation, "Why Time Series Is Different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
  8. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  9. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data