Skip to content

Regulation and governance for data buyers

Model Cards That Reference Licensed Training Data: What to Disclose and What to Withhold

Quick answer

The training data section of a model card should describe licensed data by category, domain, time span, volume range, acquisition basis and processing, and should state the data-driven limitations the model inherits. It should not name suppliers, quote license terms or reveal record-level detail unless the license allows it. Write it last, from the same supplier record that feeds your California AB 2013 page, your EU training-content summary and any Colorado developer documentation, so that no two public statements disagree.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why the model card is now a compliance document

A model card is no longer only a research courtesy, because its training-data claims can now be compared with statutory disclosures and legal commentary treats it as developer documentation. Mitchell et al. proposed model cards in 2018 as short documents shipped with a released model, covering intended use, evaluation conditions and the data behind the results; dataset cards and datasheets for datasets grew up as their companions [1][2]. Commentary on Colorado's AI law lists model cards and dataset cards among the forms developer documentation can take [10].

That status cuts both ways. A card that says "trained on proprietary data" while your AB 2013 page says datasets were "purchased or licensed" invites the question of which statement is accurate [3][4]. Treat every sentence in the training data section as a disclosure that someone may compare against another one.

The disclosures your card has to agree with

As of October 2026, three public or semi-public disclosures overlap with the model card's training data section, and each asks for slightly different granularity.

  • California AB 2013. Civil Code section 3111 requires developers of generative AI systems made available to Californians to post documentation about training data, including a high-level summary of datasets, their sources or owners, approximate data point counts, data types, whether datasets include copyrighted material or personal information, whether they were purchased or licensed, the collection time period, and processing applied [3]. Disclosures were due 1 January 2026, and the model card should not contradict them [4]. Our California AB 2013 disclosure guide covers the field list in detail.
  • EU public summary of training content. Article 53(1)(d) of the AI Act requires general-purpose AI model providers to publish a summary using the AI Office template dated 24 July 2025 [5][6]. The template separates data by source type, including publicly available datasets, private datasets obtained from third parties, crawled data, user data and synthetic data; check the published DOC for the exact fields before mirroring them [5]. See completing the EU summary for licensed and private datasets.
  • Colorado SB26-189. Signed 14 May 2026, it replaces SB 24-205, and from 1 January 2027 developers of automated decision-making technology used in consequential decisions must give deployers documentation that covers intended uses, categories of training data and known limitations [9]. The model card is a natural vehicle for this.

For high-risk systems in the EU, Annex IV technical documentation adds a non-public layer: datasheets describing provenance, scope, main characteristics, how data was obtained and selected, labeling procedures and cleaning methods [7]. Article 10 sets the governance practices behind them; as of October 2026 its application dates were reportedly moved by Regulation (EU) 2026/1744 to 2 December 2027 for Annex III systems [8]. The disclosure requirements comparison maps all of these to one supplier record.

What to disclose about licensed datasets

Disclose the facts a downstream user needs to judge fitness for use, at the level of the dataset category rather than the contract. For licensed operational data, that usually means six attributes.

  1. Category and domain. "English-language B2B software support tickets with agent resolutions" is useful. "Proprietary enterprise text" is not.
  2. Acquisition basis. State that the data was licensed from data holders, as distinct from crawled, user-contributed or synthetic. AB 2013 and the EU template both ask for this split [3][5].
  3. Time span. Give the collection window, for example 2019 to 2024, and note if collection is ongoing, as AB 2013 requires [3].
  4. Volume range. Use ranges such as "1 to 5 million tickets" or token ranges. AB 2013 accepts general ranges, and ranges avoid revealing negotiated record counts [3].
  5. Processing. Describe de-identification, deduplication, filtering, language detection and any labeling, at method level ("direct identifiers replaced with typed placeholders; residual review on a sample").
  6. Personal information and IP status. State whether the licensed sets contained personal information before processing and whether they include copyrighted works, as AB 2013 asks [3].

Keep the dataset-level depth for the dataset card. Hugging Face's dataset card workflow and the datasheets question set give a better home for composition, collection process and preprocessing detail [1][2], and our guide to dataset cards for licensed enterprise data explains how to split content between the two. Link the dataset card from the model card when it is public; reference it by internal ID when it is not.

What to withhold, and how to withhold it consistently

Withhold anything that identifies a counterparty, discloses commercial terms or lets a reader reconstruct records, unless the license expressly permits publication. Withholding is safer when it is structured, because an ad hoc omission in one document and a disclosure in another is the most common inconsistency.

Typical withheld items are the supplier's legal name, the license price and term, exclusivity, exact record counts, field-level schemas that reveal a supplier's internal systems, and examples drawn from the licensed data. AB 2013 asks for "sources or owners" of datasets [3]; describe owners by type ("US mid-market software companies") where the license prohibits naming them, and have counsel confirm that description meets the statute for your system. If you must name a source, get written consent in the license and record it.

Never paste a verbatim training example from licensed data into a card. Even de-identified support transcripts carry product names, internal ticket IDs and writing style that can identify a supplier. Write invented examples and label them.

Illustrative example: invented to show structure; it does not describe an available dataset.

Card fieldDiscloseGeneralizeWithhold unless license allows
Data categorySupport tickets, sales call notesn/an/a
Data ownern/a"US B2B software companies (3 to 10)"Company names
Acquisition basis"Licensed from data holders"n/aLicense terms, price, exclusivity
Time span2019 to 2024, not ongoingn/an/a
Volumen/a"1M to 5M records; 2B to 4B tokens"Exact counts per supplier
Personal information"Present before processing; removed or replaced"Method familyRe-identification test results with record IDs
IP status"Includes copyrighted text licensed for training"n/aScope-of-use clauses
ProcessingDedup, PII replacement, language filterThresholdsSupplier-specific redaction rules
Known limitationsCoverage, geography, time gapsn/an/a

Writing known limitations that come from the data

The limitations subsection should name the coverage gaps the licensed data creates, because that is what deployers act on and what Colorado's documentation duty asks about [9]. The original model card format already included caveats and recommendations; licensed data makes them more concrete.

Useful limitation statements are specific and testable:

  • Geography and language. "US and Canadian English only; no EU customer interactions."
  • Time. "No tickets after December 2024; product names and policies after that date are unknown to the model."
  • Supplier concentration. "Most records come from a small number of companies in two verticals; expect lower accuracy on hardware and healthcare support."
  • Processing artifacts. "Names and account numbers were replaced with placeholders such as [CUSTOMER_NAME]; the model may emit these tokens."
  • Label provenance. "Resolution labels reflect the originating agents' closure codes, not independent review."

If you claim the licensed data adds coverage beyond the open web, test it first; checking licensed data for overlap with public web corpora describes the method.

Versioning the card for fine-tunes and data refreshes

Update the training data section for every fine-tune and every data refresh, and version it alongside the model weights. A fine-tune on newly licensed data can change the AB 2013 obligation for a substantially modified system and, in the EU, can make the fine-tuner a provider in its own right; see when fine-tuning creates provider or developer duties.

A practical scheme keeps one changelog row per data event: dataset ID, version, date added, date removed, and which public documents were regenerated. When a license ends and data must be deleted, record that the dataset was used for model version 1.2 and removed before 1.3, rather than silently dropping it from the card. Article 53 documentation for GPAI providers must be kept up to date [6], so a card that drifts from your technical file creates an inconsistency a reviewer can find.

Illustrative example: invented to show structure; it does not describe an available dataset.

training_data:
  card_version: 1.3
  model_version: support-assist-1.3
  sources:
    - id: LIC-SUPPORT-02
      basis: licensed_from_data_holder
      category: b2b_software_support_tickets
      owner_description: "US B2B software companies"
      language: en-US
      collection_window: "2019-01/2024-12"
      ongoing: false
      volume_range: "1M-5M records"
      personal_info_before_processing: true
      processing: [dedup_minhash, pii_typed_placeholders, sample_review]
      copyrighted_material: true
      dataset_card: internal://cards/LIC-SUPPORT-02/v2
  removed_since_previous: [LIC-SALES-01]
  regenerated_documents: [ab2013_page, eu_summary, colorado_dev_doc]
  known_limitations:
    - "No interactions after 2024-12"
    - "Placeholder tokens may appear in outputs"

A review checklist before publishing the card

Run the card through a consistency and confidentiality review before release, with the dataset owner, counsel and the person who maintains your regulatory disclosures in the room.

  • Every dataset in the card appears in the AB 2013 page and EU summary with the same basis, time span and volume range [3][5].
  • No supplier name, price, term or exact count appears unless the license permits it in writing.
  • The license's permitted-use clause covers the use the card describes, including fine-tuning and evaluation.
  • Personal-information statements match the de-identification method actually recorded.
  • Known limitations name geography, language, time and supplier concentration.
  • No example in the card is taken from licensed data.
  • The card version, model version and changelog match the technical documentation held for regulators [6][7].

The compliance hub collects the related guides, and the dataset card glossary entry defines the companion artifact.

How SourceX supports licensed data documentation

SourceX sources operational datasets from US companies on request, and every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your model card and regulatory disclosures a documented starting point. Buyers can describe the data they need, knowing that a request does not guarantee a match and that every release is approved by the supplying company.

Request licensed data with documentation you can disclose from

If your next model card will reference licensed support, sales, engineering, document or finance data, start by describing the data rather than the businesses that might hold it. SourceX manages the process from Find and Assess through Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe your training data requirement to SourceX.

Sources

  1. arXiv (Gebru et al.), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
  2. Hugging Face, "Create a dataset card (Datasets library documentation)" (v2.19.0). https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
  3. California Legislative Information, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  4. Conventus Law, "US: California's AB 2013 requires generative AI data disclosure by January 1, 2026". https://conventuslaw.com/report/us-californias-ab-2013-requires-generative-ai-data-disclosure-by-january-1-2026/
  5. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  6. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  7. European Commission, AI Act Service Desk, "AI Act Annex IV: Technical documentation". https://ai-act-service-desk.ec.europa.eu/en/ai-act/annex-4
  8. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  9. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  10. CASRAI, "Colorado AI Act: What It Requires". https://casrai.org/guides/colorado-ai-act-explained

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data