Retrieval, RAG and grounding data
Embedding and vector index rights for licensed content
Quick answer
Yes, plan on it. Creating embeddings copies licensed text into a model's input, and most RAG indexes also store chunk text as a payload, so an embedding pipeline makes reproductions that a read or display license may not cover. No settled US authority says whether a vector is a derivative work, so do not rely on contract silence. Get explicit grants to embed, re-embed with new models, index, store payload text, retain after term and restrict sharing.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a content license must name embedding and indexing
A content license should name embedding and indexing because each step in a RAG pipeline copies the work, and the default grant language in most content deals was written for display or internal reading. The U.S. Copyright Office's Part 3 report, still a pre-publication version as of October 2026, walks through the acts involved in building and deploying generative AI systems and treats copying at several stages as implicating the reproduction right [1]. A law-firm summary of that report reads it as saying retrieval-augmented generation involves reproduction of the works retrieved, which is the practical reason to get the embedding and indexing grant in writing [2].
Content licensors in the grounding market already split rights this way. One editorial syndicator notes that a training license does not authorize live grounding and a grounding license does not grant training rights, with rules for how long content may be retained and when it must be deleted [3]. If your agreement only says "access and use for internal purposes", counsel on the other side can argue that a persistent vector index of the full corpus is outside it.
For the cluster overview, start at the RAG content licensing buyer's guide. For the broader clause set, see RAG content license terms and the distinction drawn in grounding license vs training license.
Are embeddings derivative works? What is actually unsettled
Whether an embedding is a derivative work, a non-expressive fact about the text, or something else has not, as far as we found, been squarely decided by a US court as of October 2026, so treat the question as open and contract around it. The arguments cut both ways. A dense vector from a model such as a 768- or 1,024-dimension sentence encoder does not look like the text, yet research shows it can carry most of it.
Embedding inversion is the strongest reason licensors treat vectors like the text itself. Published research on inversion attacks, such as the Vec2Text line of work, reports that short passages and personal names can be reconstructed from some text embeddings, with risk highest for short chunks. Treat that as a reason licensors and privacy reviewers may handle a vector index with the same care as the source text, and ask your security team to assess it for your specific model and chunk size.
Two practical consequences follow. First, expect licensors to define "Licensed Content" to include "any embedding, vector, index or other representation derived from it", which pulls vectors into deletion and confidentiality obligations. Second, if you are licensing content that contains personal data, the vector store belongs in your data map and your privacy review; the de-identified data guide covers that side.
The rights to separate in an embedding clause
An embedding clause should separate at least seven rights, because licensors price and restrict them independently. Bundling them as "use for RAG" leaves the gaps a later dispute will find. The definition of embeddings themselves is on the embeddings glossary entry; the list below is about who may do what with them.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Right | What to ask for | Common licensor position | Failure mode if silent |
|---|---|---|---|
| Embed | Create vectors from all licensed content with any embedding model you choose | Allowed for licensed purpose only | Licensor argues indexing exceeds "access" |
| Re-embed | Regenerate vectors when you change model, dimension or chunking, without new fees | Sometimes capped or priced per refresh | Locked to an old model; migration needs a new deal |
| Store payload | Keep chunk text, title, URL and section IDs alongside vectors | Limited to excerpt length or display rules | Cannot show citations or snippets from the index |
| Derived artifacts | Summaries, keyword indexes (BM25), knowledge-graph triples | Often treated as "adaptations" | Hybrid search fields fall outside the grant |
| Retention | Keep vectors for a defined tail period, or purge on term end | Purge within a set window | Vectors outlive the license with no rule |
| Sharing | Host in your managed vector DB (Pinecone, Weaviate, S3 Vectors) and expose results to end users | No transfer of vectors to third parties | Hosting or multi-tenant use argued as redistribution |
| Training use | Use vectors or retrieval logs to fine-tune retrievers or rerankers | Excluded unless separately licensed | Retriever training becomes model training |
The last row is where grounding and training licenses collide. Fine-tuning an embedding model on licensed passages, or on query-passage pairs harvested from your logs, is training; see training data for domain-specific embedding models before you assume a RAG grant covers it.
The embedding model license is a separate question
The license on the embedding model is separate from your right to embed the content, and you need both. An open-weight model released under MIT or Apache 2.0 lets you run it commercially, while a proprietary embedding API has its own terms of service, and neither says anything about the rights in the text you feed it [4].
Watch two interactions. If you embed through a hosted API, the licensed text leaves your environment, which may trigger the licensor's restrictions on sending content to third-party processors; sending licensed content to third-party model APIs covers sublicense and processor language. If the model provider's terms allow it to retain inputs, ask whether that is a disclosure the content license prohibits.
Retention: vectors do not disappear when the source does
Vectors persist independently of source documents, so deleting a licensed document from your content store does not remove its embeddings from the index [5]. Vector stores treat deletion as its own operation keyed by vector ID; Amazon S3 Vectors, for example, documents a separate delete call against the vector index [6]. Unless your pipeline keeps a mapping from each chunk ID to its source document and license, you cannot purge precisely.
Before signing, decide what "retain after term" means for you. Options include a full purge within a set number of days, keeping vectors but dropping payload text, or keeping vectors only for content already surfaced to users. Any of these needs metadata on every record: source ID, license ID, ingest date, embedding model and version. The operational steps are on removing licensed content from vector indexes when a license ends, and cache limits are on caching and retention limits for licensed RAG content.
Buying pre-computed embeddings instead of raw content
Buying pre-computed vectors makes sense when you only need similarity search over a corpus you never display, but it is usually the weaker deal for a RAG product. Some data vendors already sell a distinct "RAG / Vector Index" license tier next to model-training and derivative-model tiers, and prohibit raw redistribution [7], so the market does treat vectors as their own licensable product.
Vector-only deliveries carry three costs. They lock you to the vendor's embedding model, because you cannot mix vectors from different models in one index or re-embed without the text. They cannot ground a cited answer, since the model needs passage text to quote and the user needs a source to click. They also inherit the inversion risk above, so the licensor may still impose text-level confidentiality on them.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Factor | Buy pre-computed vectors | License text and embed yourself |
|---|---|---|
| Model choice | Vendor's model and dimension, fixed | Your model; re-embed when you upgrade |
| Citations and snippets | Not possible without payload | Possible, subject to excerpt display terms |
| Hybrid search (BM25 plus dense) | Dense only unless text supplied | Both |
| Retention exposure | Vectors only | Text and vectors; larger purge scope |
| Best fit | Dedup, clustering, recommendation signals | Customer-facing RAG answers |
If you do buy vectors, ask for the model name and version, normalization, chunking rules, and the chunk-to-source map, and negotiate a right to receive re-embedded vectors when the vendor changes models.
Embedding rights request checklist
Use this checklist when you send a licensor or data source a request for embedding and index rights; it keeps the ask specific and comparable across offers. For price structure, pair it with usage-based pricing for RAG content, and for general term definitions, see AI data license terms explained.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Licensed purpose: internal search, customer-facing answers, or both (internal vs customer-facing RAG).
- Grant: "reproduce, chunk, embed, index and store" licensed content, with any embedding model you select.
- Re-embedding: unlimited re-embedding for model, dimension or chunking changes during the term.
- Payload: chunk text length, metadata fields (title, canonical URL, section, publication date) and display limits.
- Derived artifacts: summaries, sparse indexes and extracted entities are inside or outside the grant.
- Hosting: named vector DB or cloud, region, and whether a managed service counts as a third party.
- Retention: purge window at term end, what survives (query logs, aggregate metrics) and certification of deletion.
- Exclusions: no training of generative or embedding models unless separately granted.
- Reporting: retrieval counts per document if pricing is usage-based (usage reporting and metering).
How SourceX fits embedding-rights requests
SourceX sources operational datasets from US companies, such as support histories, engineering records, documents and finance or legal workflows, and manages licensing and ongoing purchases. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, so state up front that you intend to embed and index the data when you describe it. Data is sourced on request rather than held in stock, and a request does not guarantee a match. You can describe the content you need to embed as a buyer.
Licensing embedding rights for your RAG content
SourceX looks for US businesses that hold the operational data you describe, and every release is approved by the supplying company. It handles the Find, Assess, Agree, Transact and Manage steps, and nothing is contracted until a supplier agrees. Tell SourceX what content you need to embed and index.
Frequently asked questions
Can I keep embeddings after a content license ends?
Only if the license says so. Vectors stay in the index after the source document is deleted [5], and licensors often define licensed content to include derived vectors, which puts them inside the deletion obligation; negotiate a specific retention or purge rule rather than assuming vectors are yours.
Does an open-source embedding model license cover the content I embed?
No. MIT or Apache 2.0 covers your use of the model weights, not your rights in the text you run through it [4].
Is an embeddings-only license safer than a text license?
It reduces display rights you need but not confidentiality risk, because published inversion research reports that short passages and names can sometimes be reconstructed from vectors.
Sources
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Jenner & Block, "US Copyright Office Releases \"Pre-Publication Version\" of Report on Copyright Issues in Generative AI Training" (2025). https://www.jenner.com/en/news-insights/client-alerts/us-copyright-office-releases-pre-publication-version-of-report-on-copyright-issues-in-generative-ai-training
- Newstex, "Editorial content licensing for AI training and grounding". https://www.newstex.com/blog/editorial-content-licensing-for-ai-training-and-grounding
- Zilliz, "What are the licensing considerations for different embedding models?". https://zilliz.com/ai-faq/what-are-the-licensing-considerations-for-different-embedding-models
- Twig, "Erasure from Vector Index". https://help.twig.so/rag-scenarios-and-solutions/privacy/right-to-erasure
- Amazon Web Services, "Deleting vectors from a vector index". https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-vectors-delete.html
- CSIMarket, "AI licensing". https://csimarket.com/help/ai-licensing/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.