Skip to content

Provenance, rights and permitted use

User-Generated Content in Training Data: When a Platform Can License Its Users' Posts

Quick answer

A platform can license its users' posts for AI training only as far as its user terms let it. On most forums, review sites and Q&A communities, users keep copyright and grant the platform a license. The deal works only if that license is sublicensable, broad enough to reach machine learning, and in force when each post was written. Counsel should test the grant, its version history, any opt-out, and third-party material embedded in posts before relying on a platform-level license.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Who owns a forum post, a review or a Q&A answer

The author usually owns the copyright, and the platform holds a license, not title. Typical user agreements say something like "you retain ownership of content you submit" followed by a grant to the platform that is worldwide, royalty-free, non-exclusive and, in the strongest versions, sublicensable and transferable. That structure is the same one buyers meet in business software, covered in who owns data in a SaaS tool, but here the "customers" are millions of individuals who never signed a negotiated contract.

Three consequences follow for a buyer. First, the platform cannot give you more than it received, so your license is derivative of the user grant. Second, any gap in the grant is a gap in your chain of title, which is why the chain of title documents for a UGC corpus must include the user terms themselves. Third, some communities publish user contributions under an open license; Stack Exchange contributions, for example, carry a Creative Commons attribution-share-alike license, while the company restricts bulk dumps and sells commercial AI access separately [9]. In that case the open license, not the platform contract, defines the floor of what anyone may do, including its attribution and share-alike conditions.

Reading the grant clause: the five words that decide the deal

The grant clause decides the deal, and five terms carry most of the weight: sublicensable, transferable, purpose, media and duration. A grant "to operate, promote and improve the Services" is purpose-limited and is the weakest basis for licensing posts to an outside AI developer. A grant "to use, reproduce, modify, create derivative works, distribute, and sublicense through multiple tiers" is far stronger, especially if it adds "for any purpose" or names machine learning expressly.

Watch for these failure modes:

  • No sublicense right. Without "sublicensable," the platform can use content itself but may not authorize a third party to train on it. The sublicense glossary entry covers why tiered sublicensing matters.
  • Purpose tied to the service. "To provide the Services to you" may support the platform's own features but strains when stretched to an external model.
  • Derivative-work rights missing. Training produces weights and outputs; a grant limited to "display and distribute" does not clearly reach them.
  • Termination on deletion. Many grants end when a user deletes content or an account, sometimes with a carve-out for copies already sublicensed. Your agreement must say how deletions propagate after delivery.
  • Moral rights and attribution. Open-license communities may require attribution that a pre-training corpus cannot practically preserve.

Copyright analysis of training itself is still unsettled. The U.S. Copyright Office's Part 3 report concludes that many acts in training implicate the reproduction right and discusses the growth of voluntary licensing, but as of October 2026 it remains a pre-publication version [4]. A clean contractual license is the buyer's way to avoid depending on a fair use outcome.

Terms versions: which rules applied when each post was made

Each post is governed by the terms in force when it was submitted, unless a later change validly reached it. A community that is fifteen years old may have six or more terms versions, and the AI-training language often appears only in the latest. Ask the platform for a dated archive of every user agreement and privacy policy version, the change notices sent, and the acceptance mechanism (clickwrap, browsewrap, email notice with continued use).

Retroactive changes carry regulatory risk as well as contract risk. FTC staff warned in February 2024 that adopting more permissive data practices, such as using data for AI training, through a quiet or retroactive terms change could be unfair or deceptive [2]. A month earlier, the same office said companies may be liable for breaking promises not to use data for undisclosed purposes [3]. If old privacy policies said "we will not share your content with third parties for purposes other than operating the site," a later license to an AI developer can contradict a commitment users relied on.

The practical fix is record-level dating. Each post should carry a created_at timestamp and a terms_version_id so you can filter to posts made under adequate terms, which is the same approach described in record-level provenance.

User notices, opt-outs and how excluded content is removed

An opt-out is only as good as the pipeline that enforces it. Several platforms have added settings that let users exclude their content from AI training or third-party data licensing. Counsel should ask when the setting launched, whether it applies retroactively to past posts, whether it is account-level or per-post, and how an exclusion reaches copies already delivered to licensees.

Request evidence rather than descriptions: the notice text and dates, opt-out rates by month, and the filter logic in the export job (for example, a join on a training_opt_out flag at export time plus a deletion feed for later changes). The consent and notice records guide lists the artifacts to request. Separately, an audit of web content found that terms of service and robots.txt signals restrict AI reuse inconsistently and increasingly [1]; a platform's crawler rules are not a substitute for its user grant, but conflicting signals are a diligence flag.

If you place a general-purpose model on the EU market, Article 53(1)(c) of the AI Act requires a copyright policy that identifies and respects text and data mining reservations, an obligation that has applied since 2 August 2025 [5]. The GPAI Code of Practice copyright chapter gives signatories a documented way to show that policy [6]. A licensed UGC corpus should therefore carry its reservation handling with it, as described in the EU TDM opt-out guide.

Embedded third-party content inside posts

Users routinely post material they do not own, and the platform's grant cannot cover it. Forum threads quote paywalled news articles in full, review sites host photos that include other people's work, and Q&A answers paste code under incompatible licenses. A user can grant only rights the user holds.

Screen for embedded content before training: long quotations matched against news and book corpora, image attachments with external EXIF or watermark signals, code blocks scanned for license headers, and URLs pointing to copyrighted media. Decide per class whether to strip, truncate or retain with a documented basis. The third-party content guide covers attachment and quoted-text handling in more depth.

Usernames, signatures, profile fields and personal details inside posts are a privacy question rather than a copyright one and belong in your privacy review. One rights-adjacent point does belong here: if the platform allowed users under 13, the amended COPPA Rule, with a compliance date of 22 April 2026, governs how children's personal information may be used [8], so ask how minors' accounts are identified and excluded.

UGC platform rights review: a decision table for counsel

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to requestGreenAmberRed
Grant scopeEvery user agreement version, with datesSublicensable, any purpose or names MLSublicensable, purpose tied to "Services"Not sublicensable
Derivative worksGrant clause text"Modify, create derivative works"SilentDisplay/distribute only
Terms coverageterms_version_id per postAll posts under adequate termsOlder posts need filteringNo version mapping possible
Change noticeNotices sent, acceptance methodProspective change, affirmative noticeEmail notice plus continued useSilent retroactive change [2]
Prior promisesPrivacy policy archiveNo contrary statementAmbiguous "won't share" languageExplicit "never share for other purposes" [3]
Opt-outSetting launch date, flag, deletion feedFlag enforced at export plus ongoing feedFlag at export onlyNo opt-out or unenforced
Open licenseLicense text applied to contributionsNone, or terms reconcile with CCCC BY with attribution planShare-alike or non-commercial unresolved
Embedded contentScreening reportQuotes and media screened and handledPartial screeningNone
MinorsAge gate and exclusion methodUnder-13 accounts excludedSelf-declared age onlyUnknown

Illustrative record fields that make this table testable at delivery:

{
  "post_id": "p_000184223",
  "created_at": "2019-03-14T09:22:11Z",
  "terms_version_id": "tos_2018_05",
  "grant_sublicensable": true,
  "training_opt_out": false,
  "content_license": "platform_terms",
  "embedded_media_flag": false,
  "quoted_text_ratio": 0.04,
  "deleted_at": null
}

Contract terms to add when the platform is the licensor

The platform's agreement with you should convert diligence findings into obligations. Common buyer asks include a representation that the platform holds sufficient rights under its user terms for the licensed uses, a schedule of terms versions relied on, a covenant to deliver ongoing deletion and opt-out feeds, and an indemnity scoped to user-grant defects. Specify whether the license covers pre-training, supervised fine-tuning, retrieval and evaluation separately, since a RAG index that displays post text raises different questions than weights trained on it.

Disclosure obligations also flow downstream. California AB 2013 required developers of generative AI systems made available to Californians to post training-data documentation by 1 January 2026, including sources and whether data was purchased or licensed [7], so keep the platform's name, data types and date ranges in a form you can publish. Market deals between community platforms and AI developers show these arrangements are negotiated at the platform level [10]; the forum and community content licensing guide and the guide to user-generated text via official APIs cover commercial tiers and access mechanics.

How this differs from business-generated community data

Community data held by a business for its own customers raises a different question than a public forum. A software company's customer support community, a product-review module inside a B2B tool, or an internal Q&A system is governed by customer contracts and employee policies, not consumer terms, and the SaaS platform terms guide explains that analysis. SourceX sources operational datasets, such as support and sales histories, from US companies, and does not source scraped web content. Where a supplier's records include user-written text, every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.

Buyers can describe the user-written data they need on the SourceX buyer page without naming specific businesses. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match. For the wider framework, return to the provenance hub or the AI data guides.

Request community and review data with rights you can verify

SourceX looks for US businesses that hold the data you describe, assesses the data and its licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Personal details are removed or replaced before delivery, the method is recorded, and a sample is checked, though no method is perfect. Describe your dataset requirements to SourceX.

Sources

  1. arXiv (Data Provenance Initiative), "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  2. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  3. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  4. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  5. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  6. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  7. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  8. Federal Trade Commission, Federal Register, "Children's Online Privacy Protection Rule (Final Rule amendments), 90 FR 16918" (2025). https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule
  9. DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
  10. TechCrunch, "OpenAI inks deal to train AI on Reddit data" (2024). https://techcrunch.com/2024/05/16/openai-inks-deal-to-train-ai-on-reddit-data

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data