Skip to content

Evaluation and benchmarking datasets

Financial analysis evaluation datasets: questions, answers and evidence strings

Quick answer

A financial LLM evaluation dataset pairs analyst-style questions with gold answers, the exact evidence string that supports each answer, and the source document and page it came from. Public benchmarks such as FinanceBench cover 10-K, 10-Q, 8-K and earnings-report questions for listed companies, but finance copilots also face private documents such as management accounts, board packs, credit memos and models. A useful eval set therefore mixes question types, ties every answer to evidence, and is scored both end-to-end and with gold context.

By SourceX Editorial · Updated

What a financial QA evaluation record must contain

Each record needs four core parts: a question, a gold answer, an evidence string quoted from the source, and a pointer to the source document. FinanceBench uses exactly this structure across 10,231 questions about public companies, with answers grounded in filings [1]. The evidence string is what makes the set auditable: a reviewer can check whether a model's answer matches the gold value and whether the model cited the same passage.

For production copilots, extend the core with fields that let you slice results. Record the document type (10-K, 10-Q, 8-K, earnings call transcript, management accounts), the fiscal period, the reporting basis (GAAP, non-GAAP, IFRS), the unit and scale (USD thousands vs millions), and whether the answer requires a calculation. Numeric answers should store a tolerance and a canonical unit so that "$1.2bn" and "1,200 (USD millions)" grade as equal.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "qid": "fa-000412",
  "question_category": "ratio_analysis",
  "question": "What was the company's FY2024 operating margin, and how did it change from FY2023?",
  "gold_answer": {"value": 0.142, "unit": "ratio", "tolerance": 0.002,
                   "text": "14.2%, up from 12.9% in FY2023"},
  "calculation": "operating_income / total_revenue for each fiscal year",
  "evidence": [
    {"doc_id": "doc-7731", "doc_type": "annual_report", "page": 48,
     "evidence_string": "Operating income 412.6 ... Total revenue 2,905.3", "scale": "USD millions"},
    {"doc_id": "doc-7731", "doc_type": "annual_report", "page": 48,
     "evidence_string": "Operating income 351.0 ... Total revenue 2,721.0", "scale": "USD millions"}
  ],
  "requires_multi_doc": false,
  "answerable": true,
  "source_visibility": "private",
  "annotator_id": "rater-cfa-03",
  "adjudicated": true
}

If you are writing the full specification, the field, slice and acceptance decisions are covered in writing an evaluation dataset specification, and the triple format itself is detailed in question-answer-citation triples for RAG evaluation.

Which analyst question categories to cover

Cover the question categories analysts actually ask, not just single-fact lookups. SECQUE argues that most financial LLM benchmarks test isolated tasks and organizes its questions around comparison and trend analysis, ratio analysis, risk factors and analyst insights [3]. That taxonomy is a good starting allocation for a finance copilot eval.

A practical set usually spans these slices:

  • Extraction: a single reported figure (total debt, diluted EPS, segment revenue) from one table.
  • Calculation: ratios and derived metrics (gross margin, net debt to EBITDA, days sales outstanding) that require two or more figures.
  • Comparison and trend: period-over-period or peer comparisons across filings or quarters.
  • Risk and narrative: questions over Item 1A risk factors, MD&A or footnotes where the answer is text, not a number.
  • Judgment: analyst-insight questions with rubric-graded answers rather than exact match.
  • Unanswerable: questions whose answer is not in the documents, to test refusal instead of fabrication.

Allocate more examples to rare but costly cases such as restatements, segment reclassifications and fiscal years that do not end in December. The stratified evaluation sets guide covers allocation, and unanswerable questions in RAG evaluation covers the abstention slice.

Why retrieval setup changes your scores

Score every model twice: once end-to-end through your retrieval pipeline and once with the gold evidence supplied as context. FinanceBench showed how large the gap can be: one tested configuration using a shared vector store answered incorrectly or refused 81% of a 150-case sample [1]. Configurations handed the relevant evidence pages did far better, which points to retrieval over long filings as a major bottleneck.

The split tells you where to invest. A low end-to-end score with a high gold-context score points to chunking, table parsing or retrieval ranking; a low score in both modes points to the model or the prompt. Log retrieval recall against the evidence pages in each record so failures can be attributed, and keep table-heavy pages as a separate slice because financial statements break naive text chunkers. Related failure modes for superseded filings and amended 10-K/A documents are covered in evaluating RAG on versioned, outdated and conflicting documents.

Public filings vs private financial documents

Public filings give you scale and reproducibility, but they are the most likely content to have appeared in pretraining corpora. Test data that leaks into newer models' training sets contaminates a benchmark and can make it obsolete quickly [4]. EDGAR filings and popular earnings transcripts are widely mirrored, so high scores on them may partly reflect memorization rather than analysis.

Private financial documents test what a deployed copilot actually sees: monthly management accounts, budget-vs-actual workbooks, lender compliance certificates, internal close checklists and board materials. These documents use company-specific account names, non-standard layouts and spreadsheet logic that no public benchmark represents. They also carry confidentiality and personal-data risk, so names, account numbers and counterparties typically need to be removed or replaced before an eval set leaves the company.

A strong design uses both: a public-filing slice for comparability with published results and a private slice as the held-out decision set. The trade-offs are covered in private evaluation sets vs public benchmarks and contamination-resistant evaluation design.

FinanceBench alternatives and when to license or commission

Start with open benchmarks for smoke tests, then license or commission data for decisions. FinanceBench publishes only a subset openly, and the full set is licensed from the vendor [2]. SECQUE adds analyst-style categories over SEC filings [3]. Both are public-company benchmarks, so neither tests private documents, your own chart of accounts or your users' question distribution.

Illustrative example: invented to show structure; it does not describe an available dataset.

OptionDocumentsContamination riskRealism for copilotsBest use
Open benchmark subsetPublic filingsHighLow to mediumSmoke tests, regression checks
Licensed full benchmarkPublic filingsMedium to highMediumComparable published results
Golden set from your own recordsYour internal documentsLowHighRelease decisions on your deployment
Commissioned set from licensed company documentsPrivate documents from data holdersLowHighCoverage of document types you lack

Building from your own records is covered in golden datasets from real business records, and the commissioning process in how to commission a custom LLM evaluation dataset. If you also need training data rather than eval data, see financial services LLM fine-tuning data; for agentic accounting workflows rather than analyst QA, see finance and accounting AI training data.

Annotation and acceptance checks for finance QA

Gold answers in finance need qualified annotators and a second reviewer, because one wrong sign or scale error flips a grade. Widely used test sets carry an estimated average label error rate of at least 3.3%, enough to change model rankings [5]. In financial QA the common errors are predictable: thousands vs millions, fiscal vs calendar periods, GAAP vs adjusted figures, and restated vs originally reported values.

Use this acceptance checklist before relying on a delivered set:

  • Every record has at least one evidence string that appears verbatim in the referenced document and page.
  • Numeric answers recompute from the evidence within the stated tolerance.
  • Units, scale and currency are explicit on every numeric answer.
  • At least two qualified reviewers agree on judgment-question rubrics, with adjudication logged.
  • Unanswerable questions are confirmed absent from the full document set, not just the cited pages.
  • Personal and counterparty identifiers in private documents are removed or replaced, with the method recorded.
  • The license states that evaluation use is permitted and whether results may be published.

For expert sourcing, see domain-expert raters for LLM evaluation; for delivery checks, see gold-label audits and adjudication. Spreadsheet-native source material is covered on spreadsheets and financial models.

Where SourceX fits

SourceX sources operational datasets from US companies, including finance and legal workflows and documents, and manages the licensing process with the supplying company. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplier. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe the private financial documents your eval set needs by type, period and question slices, not by company name.

For the wider map of benchmarks and licensed test data, start at the LLM evaluation datasets hub or the AI data hub.

Request private financial documents for your evaluation set

Describe the document types, periods and question categories your finance copilot must handle. SourceX looks for US businesses that hold matching data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered through a private, access-controlled workflow. Start a buyer request.

Sources

  1. arXiv (Islam et al., Patronus AI, Contextual AI, Stanford), "FinanceBench: A New Benchmark for Financial Question Answering" (2023). https://arxiv.org/abs/2311.11944v1
  2. Patronus AI, "FinanceBench documentation". https://docs.patronus.ai/docs/financebench-1
  3. arXiv, "SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities" (2025). https://arxiv.org/pdf/2504.04596
  4. arXiv (ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  5. arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data