Skip to content

Text and language data

Licensing Review Text Data for AI Training

Quick answer

Licensing customer review text for AI training means clearing three separate layers of rights: the reviewer who wrote the text, the platform or merchant whose terms govern it, and any repackager who compiled it. A dataset card that says "commercial use allowed" only speaks for the last layer. Buy review data with a documented chain from collection to delivery, a license that names training as a permitted use, and recorded filtering for fake, incentivized and personal content.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Who holds rights in a customer review

Rights in review text are usually split, so no single party can grant everything by default. The reviewer typically authors the text and keeps copyright; the platform or merchant usually receives a license through its terms of service; and an aggregator or dataset publisher holds, at most, rights in its own selection, labels and formatting.

In the US, the Consumer Review Fairness Act limits form-contract terms that penalize consumers for reviews or require them to hand over intellectual property in their review content, which is why many merchant terms take a non-exclusive license rather than ownership [6]. That license may or may not be sublicensable, and it may be scoped to display, marketing or "operating the service" rather than model training. A buyer therefore needs to see the actual reviewer-facing terms in force when the reviews were collected, not the current version.

Platforms add a further layer: their terms of service and API agreements often restrict bulk extraction and machine-learning use regardless of what the reviewer granted. A dataset compiled from platform pages inherits those restrictions even when the compiler publishes it under a permissive license.

Why public review datasets fail commercial review

Most public review sets are unsuitable for commercial training because their licenses are non-commercial, missing or contradicted by their sources. Licenses for similar-looking review sets range from CC0 to non-commercial: one Kaggle Amazon review set is released under CC0 1.0 with binary labels [1], while a widely used review polarity benchmark listing carries a non-commercial license [2].

Mixed-source sets are the hardest case. One 20,000-review sentiment set combines marketplace, local review site and SaaS review platform content under an unclear license [3]; the repackager's license cannot override platform or reviewer rights in that material. This matches the wider pattern found by the Data Provenance Initiative, which audited more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular hosting sites [5].

Research exceptions do not close the gap. The UK text and data analysis exception in CDPA s29A covers computational analysis for non-commercial research only [8], and as of October 2026 it remains non-commercial only. A model trained under a research exception does not become clean for a commercial product later.

First-party review data from merchants and brands

First-party reviews collected by a merchant on its own site or app are often the cleanest commercial source, provided the merchant's reviewer terms and privacy notices support the use. The merchant controls collection, can show the terms version each review was submitted under, and usually holds the linked order, product and rating metadata that makes the text useful for aspect modeling.

Check three things before relying on a merchant grant. First, the reviewer terms must give the merchant a license broad enough to sublicense for training. Second, the merchant's privacy policy and any public statements must not promise that customer content will stay out of AI training; the FTC has warned that companies that break promises not to use customer data for training can be liable under laws it enforces [7]. Third, reviews collected through a syndication or review-management vendor may be governed by that vendor's agreement, which can restrict reuse outside the vendor's network.

First-party data also brings operational context that public sets lack: verified-purchase flags, moderation decisions, merchant responses and return or support outcomes. If you want the downstream conversation, see order-support conversations with order state; for the products themselves, see licensed product catalogs and descriptions.

Fields and structure that make review text trainable

Review data is most useful for sentiment and aspect models when each record keeps its rating, product context, timestamps and moderation status alongside the text. Star ratings alone are a noisy proxy label: a three-star review often contains strong positive and negative aspects, so aspect-based sentiment needs span-level annotation or at least a product taxonomy to map aspects against.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "review_id": "r_000184213",
  "source_channel": "merchant_site",
  "terms_version": "reviewer-tos-2024-03",
  "collected_at": "2025-06-14T18:22:07Z",
  "product_id": "SKU-44812",
  "product_category": "Home > Kitchen > Blenders",
  "rating": 3,
  "verified_purchase": true,
  "incentivized": false,
  "moderation_status": "published",
  "language": "en-US",
  "text": "Crushes ice well but the lid seal leaked after two weeks. [PRODUCT_NAME] support replaced it.",
  "pii_method": "names and order numbers replaced with typed placeholders",
  "aspects": [
    {"span": "Crushes ice well", "aspect": "performance", "polarity": "positive"},
    {"span": "lid seal leaked", "aspect": "durability", "polarity": "negative"},
    {"span": "support replaced it", "aspect": "customer_service", "polarity": "positive"}
  ]
}

Ask any annotation supplier for the aspect schema, annotator guidelines, inter-annotator agreement and how conflicting spans were adjudicated. Document all of this in a data card covering upstream sources, collection and annotation methods and intended use [9].

Filtering fake, incentivized and personal content

Review text needs documented filtering before it trains anything, because fake reviews, incentivized reviews and embedded personal details each distort models in different ways. Fake and coordinated reviews teach a sentiment model the vocabulary of manipulation as if it were genuine opinion; incentivized reviews skew polarity upward; duplicate and templated reviews inflate apparent volume and leak across train and test splits.

Ask suppliers how each was handled and keep the answer with the dataset:

  • Fake or coordinated reviews: detection method (platform moderation flags, burst detection by product and time, reviewer-graph signals, near-duplicate hashing) and whether removed records are excluded or flagged.
  • Incentivized reviews: whether disclosures such as "received free product" are captured as a field rather than stripped.
  • Personal details: reviewers name staff, include order numbers, phone numbers and sometimes health details. Record the redaction or replacement method and the sample check performed; see the de-identification guide for AI training data.
  • Authorship: post-2023 reviews may be machine-generated; see verifying human-written text before you buy.
  • Dedup and contamination: exact and near-duplicate removal, and checks against public benchmarks built from the same platforms, as covered in training data quality assessment.

Rights checklist for a review data license

A review data license should name the source of every record, the rights each upstream party granted and the training uses permitted. Use this checklist when comparing offers from marketplaces, where terms vary by seller [4], or from direct suppliers.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to requestRed flag
Source channelsPer-record source_channel and platform name"Aggregated from the web" with no breakdown
Reviewer termsTerms versions in force at collection, with sublicensing languageOnly the current terms, or none
Platform termsConfirmation that collection did not breach platform ToS or API termsScraped from third-party review sites
Privacy commitmentsPrivacy notice text for the collection periodNotice says customer content is not used for AI
Permitted usesTraining, fine-tuning, evaluation and commercial deployment named explicitly"Research" or "internal analytics" only
Derived labelsWho owns ratings-derived and human aspect labelsLabels licensed separately or unclear
Personal dataRedaction method, fields covered, sample check result"Anonymized" with no method
Filtering recordFake, incentivized and duplicate handlingNo record of what was removed
DeliveryFormat (JSONL or Parquet), schema, splitsSingle CSV with mixed encodings

For deciding between a license and an assignment of the compiled set, see buying data outright vs licensing it. Review text sits close to community posts, so compare the rights model with licensing forum and community content.

Where review text ends and survey verbatims begin

Review text is unsolicited, public-facing and tied to a product listing, while survey verbatims are solicited answers collected under a research consent. That difference changes the rights analysis: survey programs typically carry respondent consents and panel terms, while reviews depend on reviewer terms and platform rules. If your corpus is mostly NPS comments or open-ended survey answers, start with licensed survey and research data. For the wider landscape of licensed language data, start at the text and language data hub or the AI data buyer hub.

SourceX sources operational datasets from US companies and manages the licensing process; every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. It does not source scraped web content. You can describe the review data you need to SourceX.

License first-party review text for training

SourceX looks for US businesses that hold the review data you describe, and every release is approved by the supplying company. Personal details are removed or replaced before delivery with the method recorded, and data is sourced on request, so a request does not guarantee a match. Describe your review data requirements to SourceX.

Frequently asked questions

Can I train on a review dataset labeled CC0 or MIT on Hugging Face?

Only if the publisher had the right to apply that license. A permissive license on a compiled set does not bind the platform or reviewers whose text it contains, so check where the records came from before relying on it [3][5].

Are star ratings enough as sentiment labels?

They work for coarse polarity but miss mixed reviews. Aspect-based models need span or aspect annotations, a defined aspect schema and agreement metrics.

Does a merchant own the reviews on its own site?

Usually the merchant holds a license under its reviewer terms rather than ownership. Whether that license extends to sublicensing for AI training depends on the terms version each review was submitted under [6].

Sources

  1. Baselight, "Amazon product reviews (Kaggle mirror)". https://baselight.app/u/kaggle/dataset/thedevastator_amazon_product_reviews
  2. HyperAI, "Review polarity dataset listing". https://hyper.ai/en/datasets/5481
  3. Hugging Face, "ecommerce-reviews-sentiment". https://huggingface.co/datasets/Jermelle/ecommerce-reviews-sentiment
  4. Datarade, "Product Review Dataset search". https://datarade.ai/search/products/product-review-dataset
  5. Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  6. Kyle E. Mitchell, "Consumer Review Fairness Act" (2021). https://writing.kemitchell.com/2021/11/18/Consumer-Review-Fairness-Act
  7. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  8. UK Intellectual Property Office (GOV.UK), "Copyright, Designs and Patents Act 1988, section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  9. Pushkarna, Zaldivar, Kjartansson (Google Research), arXiv, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data