Fine-tuning and post-training data
Expert worked solutions for quantitative and professional reasoning
Quick answer
A worked solutions dataset pairs a clearly stated applied problem, such as a loan amortization, a beam deflection check or a depreciation schedule, with an expert's step-by-step derivation and a final value that can be checked mechanically. Public sets like GSM8K [8] and FinQA [9] show the format, but production finance, engineering and tax reasoning lives in business spreadsheets, calculation packages and review memos. Buyers should specify the problem statement, the step format, the verifiable answer and the masking method before sourcing.
By SourceX Editorial · Updated
Why public math sets do not cover applied professional reasoning
Public benchmarks teach the shape of a worked solution but not the conventions professionals actually use. Grade-school sets such as GSM8K [8] pair word problems with natural-language solutions and a numeric answer, which made them a standard anchor for training verifiers and outcome rewards. Financial sets such as FinQA [9] moved closer to practice by attaching an explicit arithmetic program to questions over earnings reports, but they still draw on public filings rather than the working files where analysts actually compute.
Applied work adds what those sets leave out. A real engineering calculation cites a code clause, carries units through every line and states a load combination. A real tax computation applies a specific rule to a specific fact pattern and records the judgment calls. SpreadsheetBench made the same point for spreadsheet work by building its 912 questions from real Excel forum problems rather than synthesized queries and simplified files [1].
The gap matters most when your target model must handle the following:
- Domain conventions: day-count bases (30/360 vs actual/365), sign conventions in cash-flow models, ASCE 7 load combinations, ACI 318 strength reduction factors, MACRS recovery periods.
- Units and rounding policy stated per step, not only at the end.
- Intermediate assumptions that an expert flags ("assumes mid-year convention"), which models routinely drop.
- Multi-source inputs: a schedule, a contract term and a prior-year figure combined in one derivation.
For general human-written versus model-generated traces, see the companion guide on reasoning trace datasets; this page covers applied calculation problems specifically.
Where expert worked solutions already exist inside companies
Most applied worked solutions already exist as operational records, just not in prompt-response form. The richest sources are files where an expert computed something for a decision and someone else reviewed it.
- Financial models and workpapers: valuation models, budget-to-actual reconciliations, loan pricing sheets and close workpapers. Formulas encode each step, and cell references preserve the dependency graph. The owner pages on spreadsheet and financial model datasets and licensing spreadsheets and financial models cover that data type in depth.
- Engineering calculation packages: structural, HVAC, electrical load and pressure-vessel calcs, often as Mathcad, Excel or PDF packages with a checker's initials.
- Tax and accounting computations: provision workpapers, depreciation schedules, transfer-pricing benchmarks and amended-return reconciliations.
- Pricing, underwriting and estimating records: quote builds, rate calculations and cost estimates where a final number was approved.
These files rarely state the question. A depreciation schedule shows the answer to "what is the year-3 deduction under this method," but nobody wrote that sentence down. Restating the problem is part of the preparation, and it is where quality is won or lost.
Converting spreadsheets and calc packages into solution records
The conversion keeps the expert's logic while adding the problem statement the original file lacked. A practical pipeline looks like this:
- Parse the workbook, not the rendered values. Read formulas with a library such as openpyxl so you capture
=PMT(B3/12,B4*12,-B2)rather than only the cached result. The formula chain becomes the step sequence. - Trace precedents to a target cell. Pick an output cell an expert would be asked about, walk its precedent graph, and order the steps topologically.
- Restate the problem in words. A domain expert writes the question, the given inputs and the units, using only information visible in the file. This step needs a reviewer, because restatements can leak the answer or omit a necessary input.
- Render steps in natural language plus formula. Keep both: the prose step for SFT and the executable expression for verification.
- Recompute the final value. Re-execute the formulas in a clean environment and confirm the result matches the cached value; mismatches often reveal stale links, circular references or manual overrides.
- Mask or perturb confidential figures. Replace client names and account numbers, and where the figures themselves are sensitive, perturb the inputs and recompute every downstream step so the solution stays internally consistent.
Perturbation is the step most often done badly. Scaling revenue by a random factor without recomputing ratios produces solutions whose steps no longer agree with their answer, and that inconsistency teaches the model to be wrong confidently.
Record schema for SFT and RLVR use
One record should serve both supervised fine-tuning and reinforcement learning with verifiable rewards. Tulu 3 describes RLVR as rewarding outputs that can be checked against a verifiable answer [2], so the final value must be machine-checkable, with tolerance and unit stated.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "ws-fin-000412",
"domain": "corporate_finance",
"subdomain": "lease_accounting",
"problem": "A lessee signs a 5-year lease with annual payments of 24,000 paid in arrears and an incremental borrowing rate of 6%. Compute the initial lease liability.",
"given": [
{"name": "payment", "value": 24000, "unit": "USD"},
{"name": "term_years", "value": 5},
{"name": "rate", "value": 0.06}
],
"steps": [
{"n": 1, "text": "Payments are in arrears, so use an ordinary annuity.", "expr": null},
{"n": 2, "text": "Annuity factor = (1 - 1.06^-5) / 0.06", "expr": "(1-(1.06)**-5)/0.06", "value": 4.2124},
{"n": 3, "text": "Lease liability = payment x factor", "expr": "24000*4.2124", "value": 101097}
],
"final_answer": {"value": 101097, "unit": "USD", "tolerance_abs": 5},
"assumptions": ["no initial direct costs", "no residual value guarantee"],
"provenance": {"source_type": "lease_workpaper", "author_role": "senior accountant", "reviewed": true},
"masking": {"method": "inputs_perturbed_and_recomputed", "fields": ["payment", "lessor_name"]},
"split": "train"
}
Note the separation between steps (what SFT imitates) and final_answer with tolerance_abs (what an RLVR verifier checks). Records whose answer is a judgment, such as "is this expense deductible," belong in a rubric-scored set instead; see rubric-based rewards data.
Confidentiality, masking and financial privacy rules
Applied worked solutions often originate in client files, so masking is a design decision, not a cleanup task. Direct identifiers (names, account numbers, addresses) should be removed or replaced, while commercially sensitive numbers usually need perturbation with full recomputation.
Financial records add regulatory constraints. Under the Gramm-Leach-Bliley Act privacy rule, a business that receives nonpublic personal information from a financial institution generally steps into that institution's shoes for any further disclosure [5]. Regulation P sets the limits at 12 CFR 1016.11: information received under an exception may be used only to carry out the purpose for which it was received, and information received outside an exception may be redisclosed only as the originating institution itself could [6]. Ask suppliers whether any record derives from customer financial data and how it was de-identified before release.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Acceptance tests before you train
Acceptance should test correctness, independence and leakage, not just volume. A practitioner study of LLM training data found that many models do not document how their datasets were built [4], so require a data card that states authorship, review and masking per record.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | How to run it | Typical failure it catches |
|---|---|---|
| Recompute rate | Execute every expr and compare to final_answer within tolerance | Perturbed inputs not propagated; stale cached values |
| Problem sufficiency | Second expert solves from problem and given only | Restatement omits an input present only in the source file |
| Answer leakage | Search problem text for the final value or an intermediate | Question reveals the result |
| Independent re-solve agreement | Two experts re-solve a sample; measure agreement with Krippendorff's alpha [7] | Ambiguous conventions (day count, rounding) |
| Near-duplicate scan | MinHash or suffix-array dedup across problems and against your eval sets [3] | Same template with changed numbers inflating counts |
| Domain coverage | Tabulate subdomain and step counts | Long tail of one easy calculation type |
| Masking audit | Sample records for residual names, account numbers and real client figures | Identifiers left in comments or sheet names |
Template inflation deserves particular attention. Deduplication research found common corpora full of near-duplicates and showed that removing them reduces verbatim memorization [3]; a hundred depreciation schedules with different asset costs are one problem type, not a hundred. For a broader pre-purchase review, use the fine-tuning dataset evaluation checklist, and for checking who wrote the solutions, see verifying expert annotator qualifications.
Writing a sourcing request for applied worked solutions
A good request describes the problem types and record format, not the companies you imagine hold them. Specify:
- Domains and subdomains: for example, lease accounting, project finance debt sizing, steel connection design, state apportionment.
- Source artifacts acceptable: native workbooks with formulas, calc packages, reviewed workpapers.
- Step format: prose plus executable expression, units per step, rounding policy.
- Verifiability: share of records needing a numeric final answer with tolerance, versus rubric-scored judgments.
- Review evidence: author role, reviewer sign-off, re-solve sample size.
- Masking: identifier removal plus perturbation-with-recompute for sensitive figures, documented per record.
- Intended use: SFT, RLVR or held-out evaluation, so the license and splits match.
For how worked solutions fit with other post-training data, start from the fine-tuning and post-training data hub and the guide to RLVR datasets with prompts, reference answers and verifiers. When you are ready to describe a specific need, submit a buyer request to SourceX.
Sourcing expert worked solutions through SourceX
SourceX sources operational datasets, including finance workflows, engineering records and documents, from US companies on request, rights-reviews each dataset and delivers it under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery with the method recorded, and nothing is contracted until the supplying company agrees. Describe the worked-solution data you need.
Sources
- arXiv (NeurIPS 2024 Datasets and Benchmarks), "SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation" (2024). https://arxiv.org/html/2406.14991v2
- arXiv (Allen Institute for AI), "Tulu 3: Pushing Frontiers in Open Language Model Post-Training" (2024). https://arxiv.org/pdf/2411.15124
- arXiv (Lee et al., ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- ASE 2024 (conf.researchr.org), "What Makes a High-Quality Training Dataset for Large Language Models: A Practitioners' Perspective" (2024). https://conf.researchr.org/details/ase-2024/ase-2024-research/53/What-Makes-a-High-Quality-Training-Dataset-for-Large-Language-Models-A-Practitioners
- Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
- Consumer Financial Protection Bureau, "12 CFR 1016.11 Limits on redisclosure and reuse of information". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- University of Pennsylvania, Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- arXiv (Cobbe et al., 2021), "Training Verifiers to Solve Math Word Problems". https://arxiv.org/abs/2110.14168
- arXiv (Chen et al., 2021), "FinQA: A Dataset of Numerical Reasoning over Financial Data". https://arxiv.org/abs/2109.00122
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.