Industry-specific operational data
Investment research notes and analyst work product as AI training data: rights and MNPI limits
Quick answer
Investment research notes are licensable AI data only when the holder owns the work product, can strip or embargo anything touching material nonpublic information (MNPI), and excludes third-party content it cannot redistribute. Buy-side memos, investment committee (IC) records and analyst models add what public filings corpora lack: a thesis, a decision and an outcome. Before licensing, confirm authorship, carve out sell-side research and expert-network transcripts, and set a recency embargo on live positions.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How analyst work product differs from public filings corpora
Analyst work product records judgment, while filings corpora record disclosure. EDGAR-CORPUS, for example, packages 10-K annual reports from 1993 to 2020 split into items, as raw text with no analyst annotations [1]. Most open financial LLM evaluation, including Fin-RATE and SECQUE, is built over the same public SEC filings [2][4], and even FinanceBench publishes only an open-source subset while the full question set is licensed [3].
That leaves a gap for research copilots. A model trained on 10-K Item 7 text learns how companies describe themselves; it does not learn how an analyst turns that text into a variant view, sizes a position, or admits a thesis broke. Private notes supply those steps, which is why they matter for supervised fine-tuning (SFT), retrieval-augmented generation (RAG) over a firm's own history, and evaluation of financial reasoning against what actually happened.
For broader context on how labs approach finance data, see do AI labs buy financial data and the finance buyers overview.
Who holds rights in research notes, memos and models
The firm that employs the analyst usually holds rights in internally created notes, but not in everything those notes contain. Treat each research archive as layered content with different owners.
- Firm-authored work product. Initiation notes, earnings recaps, IC memos, pre-mortems and post-mortems written by employees are typically the firm's work product, subject to employment agreements and fund documents.
- Third-party sell-side research. Broker reports pasted or attached into notes are generally licensed to the reader for internal use. Training or redistribution needs explicit rights from the publishing bank; assume the firm cannot grant them.
- Expert-network and channel-check transcripts. These sit under the network's terms and often carry their own confidentiality and compliance screens.
- Licensed data embedded in models. Excel models frequently pull estimates, prices or fundamentals from terminal or vendor feeds through add-ins; those cells belong to the vendor's license, not the firm's.
- Portfolio-company materials. Board decks, management models and data-room documents received under NDA are confidential to the counterparty.
The Third Circuit's September 2026 affirmance in Thomson Reuters v. ROSS (as of October 2026) is a caution for the second and fourth layers: copying a publisher's editorial headnotes to train a non-generative legal-research tool was held not to be fair use, and the headnotes were found sufficiently original to be protected [7]. Analyst summaries and ratings rationales from a third-party publisher can raise similar issues, and vendor estimate tables are usually restricted by contract even where the underlying facts are not protected. Teams placing general-purpose models on the EU market also need a copyright compliance policy under AI Act Article 53(1)(c), so provenance of each layer has to be documented, not assumed [8].
Where MNPI risk enters a research archive
MNPI risk sits in specific record types and time windows, not evenly across an archive. Investment advisers must maintain written policies reasonably designed to prevent misuse of MNPI under Advisers Act Section 204A, and buyers' counsel typically anchors data-vendor diligence on how a supplier sourced data and handled MNPI and personal information [5]. Expect a supplier's compliance team to apply the same Section 204A policies to any export of research records.
High-risk sources inside a research archive:
- Private-side materials. Notes written while the firm was wall-crossed on a private placement, credit amendment or PIPE, or while it held a board seat.
- Expert calls with current or former employees of covered companies, including notes flagged by compliance.
- Restricted- and watch-list periods. Any note dated while the issuer was on the firm's restricted list.
- Management meetings where selective disclosure may have occurred.
- Live positions and pipeline ideas. Not MNPI about the issuer, but commercially sensitive information about the fund's own trading, which licensors usually embargo.
The practical control is to join the archive to the supplier's compliance records before export: restricted-list history, wall-crossing logs and expert-call approvals. A note that cannot be matched to a clean compliance status should be excluded rather than redacted, because MNPI is a property of the information, not of a name field.
What a training-ready research record contains
A useful record links the thesis, the evidence, the decision and the later outcome, so a model can learn reasoning and an evaluator can score it. Notes without the decision and outcome are commentary; with them, they become decision records with rationale (see decision records with rationale).
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "ic-2023-0412-07",
"doc_type": "ic_memo",
"authored_by_role": "senior_analyst",
"issuer_ref": "PSEUDO-ISSUER-311",
"sector_gics": "Industrials / Machinery",
"memo_date": "2023-04-12",
"thesis": "Aftermarket parts mix shift lifts margins 200-300 bp over 8 quarters.",
"catalysts": ["Q3 pricing reset", "dealer inventory normalization"],
"valuation": {"method": "EV/EBITDA", "base_multiple": 11.0, "target_upside_pct": 24},
"key_risks": ["OEM volume downturn", "steel input costs"],
"decision": {"action": "initiate_long", "size_bps_nav": 150, "vote": "approved_4_1"},
"dissent_summary": "One member questioned dealer destocking duration.",
"model_file_ref": "model-ic-2023-0412-07.xlsx (vendor-sourced cells removed)",
"outcome_review": {"review_date": "2024-10-15", "thesis_status": "partially_correct", "return_vs_benchmark_pct": 6.1},
"compliance": {"restricted_list_at_date": false, "wall_crossed": false, "expert_calls_cited": 0},
"third_party_content_removed": ["sell_side_excerpt", "terminal_estimates"],
"personal_data_method": "names of employees and contacts replaced with role tokens"
}
For copilots that draft memos, the pair of thesis/valuation inputs with the decision and dissent_summary gives SFT and preference signal. For evaluation, outcome_review lets a team grade predictions against realized results rather than against another model. Analyst spreadsheets need separate handling; the spreadsheet and financial model datasets page covers formula and structure questions.
Diligence questions before licensing research data
Diligence for research data should confirm authorship, third-party carve-outs, MNPI screening and personal-data handling, record by record. The FISD Alternative Data Council's data-provider due diligence questionnaire, whose 2024 edition adds generative AI questions, is a familiar starting point for investment-firm counsel [6]; adapt it in the reverse direction when the investment firm is the supplier.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | What to ask for | Red flag |
|---|---|---|
| Authorship | Confirmation that included documents were created by the supplier's employees | Archive mixes shared-drive files of unknown origin |
| Third-party research | Method for detecting and removing broker PDFs, excerpts and ratings tables | "Removed obvious attachments" with no text-level scan |
| Vendor data in models | List of add-in or feed sources and how those cells were stripped | Models delivered with live links to a data terminal |
| MNPI screening | Join to restricted-list, wall-cross and expert-call logs, with excluded counts | Screening by keyword only |
| Recency | Embargo cutoff relative to delivery, applied to live and recently exited positions | Notes on current holdings included |
| Fund and client terms | Confirmation that LPAs, side letters and client agreements do not restrict use of research | No one has read the side letters |
| Personal data | Names, emails and phone numbers of employees, executives and experts replaced, with method recorded | Expert names retained in call notes |
| Documentation | A dataset card covering sources, collection and annotation methods and intended use [9] | No written description of what was excluded |
License terms specific to investment research
Research licenses need terms that reflect market sensitivity, not just copyright. Generic text-license templates rarely address these points, so raise them early.
- Embargo period. A fixed lag between a note's date and its inclusion, so no record reflects current positioning.
- No-trading and no-reidentification covenants. The buyer agrees not to use records to infer the supplier's holdings or to reconstruct issuer identities that were pseudonymized.
- Output restrictions. Whether a trained model may reproduce a memo verbatim, and how retrieval systems cite source records.
- Exclusions schedule. An exhibit listing excluded categories (sell-side research, expert transcripts, private-side materials) so both parties can audit the delivery.
- Evaluation holdout. If records serve as a held-out financial-reasoning benchmark, terms that keep them out of any training set.
Related industry pages cover adjacent record types: wealth client-service and advisor notes, credit decision records with adverse action reasons and tax research memos with authority citations. For when to build, buy or synthesize instead, see build, buy or synthesize training data, and the industry-specific operational data hub maps the rest of the cluster.
How SourceX approaches research and finance work product
SourceX sources operational datasets from US companies on request, including documents and finance workflows, and manages the licensing and ongoing purchases; it does not hold inventory, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails and phone numbers are removed or replaced before delivery, the method is recorded, and a sample is checked, though no method is perfect. Buyers can describe the research records they need without naming specific firms.
License investment research data for your AI team
SourceX sources operational datasets, including documents and finance workflows, from US companies for AI teams wherever based. Each dataset is rights-reviewed, approved by the supplying company and delivered under a license, and nothing is contracted until a supplier agrees. Start a buyer request at SourceX.
Frequently asked questions
Can sell-side equity research be used to train a model?
Usually not under a standard research subscription, which typically licenses reading for internal use. Training or redistribution needs explicit rights from the publishing bank or research provider, so treat broker content inside buy-side notes as an exclusion unless separately licensed.
Is an old investment memo still MNPI?
Information stops being MNPI once it is public or no longer material, but a memo can still hold confidential portfolio-company material or reveal the fund's positioning. That is why record-level compliance checks plus an embargo are safer than a blanket age cutoff.
Should issuer names be removed from research notes?
It depends on the use. Evaluation of financial reasoning often needs the real issuer to check facts against filings, while drafting copilots can learn from pseudonymized issuers. Decide in the license and record the method.
Sources
- Hugging Face (c3po-ai), "EDGAR-CORPUS (dataset card)". https://huggingface.co/datasets/c3po-ai/edgar-corpus
- arXiv, "Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings" (2026). https://arxiv.org/pdf/2602.07294
- Patronus AI, "FinanceBench documentation". https://docs.patronus.ai/docs/financebench-1
- arXiv, "SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities" (2025). https://arxiv.org/pdf/2504.04596
- Lowenstein Sandler, "Key considerations for alternative data and AI vendors to investment firms". https://www.lowenstein.com/media/iyrpwxij/key-considerations-for-alternative-data-and-ai-vendors-to-investment-firms.pdf
- FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (with GenAI questions)" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.