Skip to content

Data licensing for AI training

Market data licenses and AI: non-display, derived data and model-training use

Quick answer

Most exchange and vendor market data licenses were written for screens and trading systems, not model training. Feeding prices, quotes or order-book data into a training pipeline is usually machine (non-display) use. As of October 2026, at least one exchange, Nasdaq, publishes an AI-specific data policy [1], and vendors say training, inference and agents need explicit terms [4]. Before training, confirm four things in writing: the category of use, derived-data treatment, output redistribution, and how usage is reported and audited.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a standard market data license rarely covers model training

A standard subscriber or vendor agreement grants rights by category of use, and training is usually not one of them. Exchange policies separate display use (a person viewing data on a terminal) from non-display use (a machine consuming it), and they bill, report and audit the two separately. Vendors now say openly that standard licenses cover display and human consumption, while training, inference and agents need explicit terms [4]. One vendor describes licensing as the main hurdle for using financial data in AI [5].

The rights problem is mostly contractual, not copyright. Raw prices are largely factual, but your redistribution agreement or subscriber agreement still binds you, whatever the copyright analysis of training in the Copyright Office's Part 3 report (still a pre-publication version as of October 2026) [6]. Treat the license as the controlling document, and read every exhibit and policy it incorporates by reference.

How exchanges define non-display use, and why AI falls inside it

Non-display use generally means any machine or automated access that is not solely in support of a display to a natural person. Under that definition, a training job, feature pipeline or retrieval tool is machine access. Some exchange policies list automated or black-box trading systems as examples of non-display use. Read silence on AI as unresolved, not as permission.

The category has widened over time. Exchanges have moved from targeting algorithmic or high-frequency trading toward treating most automated uses as non-display, including order generation, price referencing, smart order routing, risk management, surveillance and portfolio management. A model that ingests a feed to produce forecasts or signals sits naturally inside that list.

Nasdaq publishes a separate Data AI Policy alongside its data agreements [1]. Expect that pattern to spread: an AI-specific policy layered on top of the master agreement, the non-display policy and the fee schedule. Your team has to reconcile all of them.

Watch these failure modes in practice:

  • Historical files bought "for back-testing." Back-testing rights do not imply training rights. Many historical products carry their own terms.
  • Delayed or end-of-day data. Lighter fees for delayed data do not remove non-display or derived-data rules.
  • Inference-time retrieval. A RAG or agent tool that pulls live quotes into a prompt is machine access, even when a person reads the answer.
  • Shared research environments. One licensed feed landing in a data lake used by several model teams can create several non-display uses, each counted separately.

Derived data: when a feature, a signal or a model weight counts

Derived data clauses decide whether what you build from licensed data belongs to you, and most were drafted before anyone asked about model weights. Exchange policies commonly treat values computed from their data, such as indices, VWAPs, analytics or signals, as derived data. They often add conditions: the original values must not be recoverable, and redistributing the derived values may need its own license.

Three questions matter for AI teams:

  1. Are engineered features derived data? Rolling returns, implied volatility surfaces and order-book imbalance features are classic derived data. Check whether they may leave the licensed environment.
  2. Are trained weights derived data? Few policies say so explicitly. If the clause is silent, get written confirmation either way before you ship weights to another legal entity, a cloud tenant you do not control, or a customer.
  3. Can the original data be reconstructed? Models that memorize price series, or that let a user query a tick-by-tick history, can defeat the "not reverse-engineerable" condition that many derived-data carve-outs rely on.

Successor models, distillation and fine-tuned variants raise the same questions one level up. See derivative and successor model rights and rights to synthetic data generated from licensed data.

Model outputs and redistribution limits

Model outputs that reveal or approximate licensed prices are redistribution, and redistribution is the most heavily policed right in market data. A financial LLM that answers "where did XYZ close yesterday" is displaying exchange data to a person. A forecast product sold to clients may be derived data distributed externally. Each needs a matching right.

Map every output surface before you negotiate:

  • Internal research and trading only. Usually the easiest case, but still non-display for the training and inference systems.
  • Client-facing analytics or chat. Likely needs display entitlements per user, plus derived-data redistribution rights.
  • Model or API sold to third parties. Treat it as an external distribution product. Expect vendor-of-record obligations, per-customer reporting and possibly a redistributor agreement.

Field-of-use limits often sit next to these clauses. See drafting field-of-use restrictions and prohibited-use clauses in AI data licenses.

Usage declarations, entitlements and audits

Market data licensing assumes you can prove what you used, so an AI pipeline needs the same entitlement controls as a trading floor. Exchanges and vendors typically require periodic usage declarations by user, device or application, plus an entitlement or permissioning system that logs who and what received each feed. SIX, for example, publishes an audit code of practice for its market data license agreements [2]. Vendors pitching AI-ready market data emphasize entitlement enforcement and audit trails at the data layer [4].

Training pipelines break the usual counting units. A single ingestion job may feed dozens of experiments, and a checkpoint can outlive the subscription. Before an audit, be ready to show:

  • which datasets and feeds entered which training runs (dataset IDs, date ranges, symbols, venues);
  • which models, checkpoints and features descend from them;
  • where those models are deployed and who can query them;
  • what happens to weights and caches after termination.

For negotiating audit scope, frequency and lookback, see audit and usage-reporting rights. For vendor diligence, the FISD Alternative Data Council's questionnaire includes generative AI questions you can reuse [3].

AI rider checklist for an exchange or vendor market data license

Use a short rider that names each AI activity, rather than relying on a general non-display grant. The table below maps common activities to the license questions to resolve.

Illustrative example: invented to show structure; it does not describe an available dataset.

AI activityLikely license categoryQuestions to put in the rider
Pretraining or fine-tuning on historical tick dataNon-display; historical product termsIs training an expressly permitted purpose? Does it survive termination?
Feature store built from real-time feedNon-display; derived dataCan features be stored, shared across teams, or moved to another entity?
RAG or agent tool calling live quotesNon-display at retrieval; display for the end userPer-user display fees? Caching limits? Delay requirements?
Forecasts or signals sold to clientsDerived data redistributionIs a redistributor or vendor agreement needed? Attribution?
Model weights shared with an affiliate or customerDerived data (often unaddressed)Are weights derived data? Reconstruction prohibition?
Evaluation benchmark built from pricesNon-display; derived dataCan the benchmark be published or shared with model vendors?

Rider request template, written to the licensor:

We request written confirmation that the licensed data may be used for (a) training and fine-tuning of machine learning models, (b) automated inference and retrieval, and (c) creation of model weights and features. Please confirm whether trained weights and engineered features are derived data, which outputs may be displayed or distributed externally, the reporting unit and frequency for AI usage, and the treatment of trained models after termination.

Scope also matters: a per-model grant and an enterprise grant price and audit very differently. See per-model vs enterprise-wide licenses. If you buy from a consolidator, trace the exchange terms that pass through to you, using the method in tracing upstream licenses in aggregated datasets.

Where exchange market data ends and operational finance data begins

Exchange and vendor market data comes from licensors with standard policies, while much of the most useful data for financial AI sits inside companies' own operations. Buyers training forecasting models or financial LLMs often also need records that no exchange sells: finance and legal workflow documents, support and sales histories, and engineering records. That data is licensed one supplier at a time, not under an exchange policy.

SourceX sources operational datasets from US companies on request, including finance and legal workflows and support and sales histories, and manages the licensing process and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery. The method is recorded and a sample is checked, but no method is perfect. For transaction-level records, see licensing financial transaction data for AI training, or describe the operational data you need.

Get licensed operational finance data for your models

SourceX sources operational datasets from US companies on request, including finance and legal workflows, and supports AI teams wherever they are based. It assesses data and licensing permissions first, and nothing is contracted until a supplier agrees. Tell SourceX what data you need.

Related: AI training data licensing hub · all AI data guides

Frequently asked questions

Is delayed or end-of-day data free to use for training?

Not by default. Delayed and end-of-day data often carry lower fees, but non-display, derived-data and redistribution terms usually still apply. Check the specific product's policy and ask for written training rights.

Does a non-display license automatically cover training a model?

No. Non-display fees are often structured around trading, routing or risk systems. Nasdaq now publishes an AI-specific data policy [1], but the scope of permitted purposes, derived-data treatment and post-termination rights still needs to be confirmed in writing.

What happens to a trained model when the market data license ends?

It depends on the contract, and many agreements are silent on weights. Negotiate explicit survival for trained models, or accept a deletion or retraining obligation knowingly, and record which checkpoints used the data.

Sources

  1. Nasdaq (NasdaqTrader), "Data AI Policy". https://nasdaqtrader.com/content/AdministrationSupport/AgreementsData/Data_AI_Policy.pdf
  2. SIX Group, "SIX Exfeed MDLA Audit Code of Practice". https://www.six-group.com/dam/download/market-data/exfeed/agreements-mdla/six-exfeed-mdla-audit-code-of-practice.pdf
  3. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire with GenAI Questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  4. QUODD, "Market data for AI". https://solutions.quodd.com/market-data-for-ai
  5. Intrinio, "Financial data licensing for AI: common roadblocks and solutions". https://intrinio.com/blog/financial-data-licensing-for-ai-common-roadblocks-solutions
  6. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data