Skip to content

Retrieval, RAG and grounding data

Excerpt display rights and verbatim limits for AI answers

Quick answer

An AI answer may show only what your license's display grant covers, and display is a separate right from training or retrieval. Map every element your interface renders (generated summary, short snippet, direct quote, full text, image, logo, link) to a named tier in the contract. Express verbatim limits in measurable units such as characters, tokens or percent of source per answer. Then enforce them in the output pipeline with length caps, quote detection and per-source display policies.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why display is its own grant, separate from training and retrieval

Display rights govern what a user actually sees in a live answer, which is why they are negotiated separately from the right to train on or index the content. Industry coverage of recent publisher deals describes display rights as the permission for a chatbot to present summaries, quotes, logos and links during live answers, distinct from training rights [1]. One licensing vendor warns that a license for training but not output display can be incompatible with chat products that quote, paraphrase or cite [2].

The commercial market has moved in the same direction. Publishers are shifting from one-time training payments toward usage-based grounding deals, where the system fetches content at query time [4]. If you already hold a training license, read the grounding license vs training license comparison before assuming it extends to your answer surface; the broader term map sits in the RAG content licensing hub.

Three separate permissions usually stack in an answer engine: the right to ingest and index (including embeddings), the right to pass content to a model at inference time, and the right to render content or derivatives to an end user. Indexing is covered in embedding and vector index rights. This page covers only the third.

The display tiers you can license, from summary to full article

Display rights are often sold as tiers, and the cheapest tier may not cover what a quoting chat product renders. One vendor markets licenses in tiers from summarization use cases to full-article display [3]; treat that as a market signal, not a standard taxonomy. The table below is a working vocabulary for drafting your own schedule.

Illustrative example: invented to show structure; it does not describe an available dataset.

TierWhat the user seesTypical measurable limit to negotiateMain risk if under-licensed
T0 Reference onlyTitle, publisher name, linkNo body text; headline length onlyLow; still check logo use
T1 SummaryModel-written paraphraseNo contiguous verbatim run above N tokensSummary that substitutes for the article
T2 SnippetShort verbatim excerpt plus linkMax characters per excerpt and per answerExcerpts chained across turns
T3 QuoteAttributed verbatim passagesMax quote length, max percent of source per sessionLong passages from paywalled pieces
T4 Full textEntire article or documentAuthenticated users only, per-view meteringRedistribution outside the product
Media add-onImages, thumbnails, charts, logosSeparate grant per asset class and credit linePhoto and logo rights held by third parties

Two points trip up buyers. First, a publisher may own the text but not the photographs, wire images or embedded charts in an article, so an image tier needs its own chain of title, the same problem described for rights layers in video clips. Second, logos and masthead marks raise trademark questions that a copyright license may not address; name logo display explicitly.

How to write verbatim limits in measurable units

A verbatim limit is enforceable only when it names a unit, a scope and a measurement method. Verbatim excerpt limits are a negotiated term [2], so the drafting burden is on you to make them testable. "Brief excerpts" or "reasonable quotation" leaves the dispute to whoever audits the logs.

Specify each of these in the display schedule:

  • Unit. Characters (Unicode code points, not bytes), whitespace-delimited words, or tokens under a named tokenizer. Tokens drift between model versions, so characters or words are easier to audit.
  • Matching rule. What counts as verbatim: exact match after normalizing whitespace, case and punctuation, or a fuzzy threshold such as longest common substring of 40 or more characters.
  • Scope. Per excerpt, per answer, per conversation or session, and per source document per user per day. Session scope matters because a user can ask "continue" until the article is reconstructed.
  • Aggregation. Percent of the source document that may be displayed cumulatively, measured against the canonical text the licensor delivers.
  • Exceptions. Headlines, short factual data such as a price or date, and code snippets in technical documentation often need explicit carve-outs.
  • Paywall rule. Whether content behind the licensor's paywall can appear at T2 or above, and whether display requires the end user to be a subscriber.

Answers that substitute for the source article

Substitution, not length alone, is a central display risk: an answer that gives the user everything they would have clicked through for. Litigation against a RAG-powered answer engine is ongoing in 2026, and its outcome is unsettled [5]. A law-firm summary of the Copyright Office's Part 3 report says it treats RAG as involving reproduction of works [6]; that report remains a pre-publication version as of October 2026 [7].

You cannot contract away substitution with a character cap, because a dense paraphrase of a scoop or a recipe can substitute with zero verbatim overlap. Practical controls include a maximum share of an answer drawn from a single source, mandatory link-out for time-sensitive news within an embargo window, and a rule that T1 summaries of paywalled articles stay at headline-plus-key-point depth. Attribution placement is covered in citation, attribution and link-back terms. Whether RAG itself is fair use is a separate legal question for counsel, not something this page answers.

For EU-facing general-purpose model providers, the AI Office's Code of Practice Copyright chapter is a reference point: it gives signatories a way to show compliance with Article 53(1)(c) through a copyright policy and addresses the risk of infringing outputs [8]. It is voluntary and does not by itself settle display questions under a specific license.

Output controls that enforce display limits in production

The contract sets the limit; the answer pipeline has to make violations rare and detectable. Build controls at four points: retrieval, generation, post-processing and logging.

  • Per-source display policy at retrieval. Attach a display_tier, max_excerpt_chars, paywalled flag and image_rights flag to every chunk in the index. The orchestrator reads these before the chunk reaches the prompt, so a T1 source is passed with an instruction to paraphrase only.
  • Prompt-level constraints. System prompts that forbid quoting T0 and T1 sources help, but models ignore them often enough that they cannot be the only control.
  • Post-generation quote detection. Compare the draft answer against retrieved chunks with n-gram or suffix-array matching, then truncate, paraphrase or drop spans above the per-source cap. Run it on streaming output in windows, or buffer before display.
  • Session accumulators. Track cumulative characters displayed per source document per user so "continue" requests stop at the aggregate limit.
  • Media gating. Render images, thumbnails and logos only from an allowlist of assets with recorded rights, never from whatever the page HTML linked.
  • Audit logs. Store answer ID, source IDs, displayed character counts and matched spans. These logs also feed usage reporting and metering and any per-display fee.

Illustrative example: invented to show structure; it does not describe an available dataset.

source_policy:
  source_id: pub-0412
  display_tier: T2_snippet
  max_excerpt_chars: 300
  max_chars_per_answer: 600
  max_pct_document_per_user_day: 10
  match_rule: "normalized exact; LCS >= 40 chars counts as verbatim"
  paywalled: true
  paywalled_display: "T1_summary_only"
  images: { allowed: false }
  logo: { allowed: true, asset_id: "logo-0412-v2" }
  link_back: required
  log_fields: [answer_id, chunk_ids, verbatim_chars, matched_spans]

Procurement checklist for display rights

Run this checklist against every source before it goes live in a customer-facing answer. Internal-only deployments often need narrower terms; see internal vs customer-facing RAG licensing.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Inventory every render element your UI produces: answer text, inline quotes, cards, image carousels, favicons, logos, preview panes and export or share features.
  2. Map each element to a tier and confirm the license names that tier for that surface (web, mobile, API, partner embeds).
  3. Confirm the field of use covers answer display, not only research or internal search.
  4. Write verbatim limits with unit, matching rule, scope and aggregation.
  5. Confirm separate chain of title for images, charts and logos.
  6. Agree the paywall rule and how entitlement is checked.
  7. Align caching and retention of displayed text with caching and retention limits.
  8. Agree how display is metered if pricing is per display; see usage-based pricing for RAG content.
  9. Define what happens to displayed and cached content at termination, alongside vector index deletion.
  10. Run a red-team test of "continue" and "quote the whole thing" prompts before launch and keep the results.

For the general vocabulary of use, term and deletion clauses, see AI data license terms explained and the full clause list in RAG content license terms.

Where operational and documentation content fits

Not every source in an answer engine is news or reference publishing. Enterprise assistants often ground answers in product manuals, support histories and internal documents licensed from the companies that hold them, and the same tiering applies: decide whether answers may quote a procedure verbatim or only paraphrase it. If that is the content you need, SourceX's buyer intake is where you describe it; SourceX sources operational datasets from US companies on request rather than from stock, and does not source scraped web content. See also product manuals and service documentation as RAG corpora.

License the display rights your AI answers need

SourceX sources operational datasets from US companies on request, including support histories, engineering records and documents, and manages licensing and ongoing purchases. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, and nothing is contracted until the supplier agrees. Describe the content and display use you need at sourcex.si/buyers.

Frequently asked questions

Does a grounding license automatically include display?

No. Retrieval and display are distinct permissions, and a license can allow the model to read content at inference time while restricting what is shown verbatim [2]. Check the schedule for an explicit display tier.

Is paraphrasing always safer than quoting?

Not always. A paraphrase can avoid verbatim overlap and still substitute for the source, so pair summary rights with share-of-answer and link-out rules.

Should we measure limits in tokens?

Prefer characters or words for the contract and convert internally. Token counts change with the tokenizer, which makes audits across model upgrades harder to reconcile.

Sources

  1. Everything PR, "The publishers who took the deal". https://everything-pr.com/the-publishers-who-took-the-deal
  2. Newstex, "Editorial content licensing for AI training and grounding". https://www.newstex.com/blog/editorial-content-licensing-for-ai-training-and-grounding
  3. TollBit, "Licensed RAG". https://tollbit.com/licensed-rag/
  4. Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?". https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
  5. Blck Alpaca, "Perplexity AI copyright lawsuit: key insights 2026" (2026). https://blckalpaca.at/en/blog/perplexity-ai-copyright-lawsuit-key-insights-2026
  6. Jenner & Block, "US Copyright Office Releases Pre-Publication Version of Report on Copyright Issues in Generative AI Training" (2025). https://www.jenner.com/en/news-insights/client-alerts/us-copyright-office-releases-pre-publication-version-of-report-on-copyright-issues-in-generative-ai-training
  7. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  8. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data