Evaluation and benchmarking datasets
Accounting agent evaluation: reconciliation and GL coding test sets with reviewer sign-offs
Quick answer
To evaluate an accounting AI agent, score it against records where a reviewer has already approved the final answer: the posted journal entry, the cleared bank match, the signed-off reconciliation. Build test cases from GL coding, bank and AP matching, accruals and close checklists. Grade the ledger end state, not the agent's narrative. Keep the preparer's first attempt and the reviewer's correction as separate fields, and split by period so the agent never sees the answer month.
By SourceX Editorial · Updated
Why analyst-QA finance benchmarks miss bookkeeping work
Most public finance benchmarks test question answering over filings, not the operational tasks a bookkeeping agent performs. SECQUE, a 2025 benchmark, is typical: it scores models on expert-written analysis questions over SEC filings [1], not on posting or matching entries. Reading a 10-K and answering "what was Q3 gross margin" says little about whether an agent can code a $1,412.50 vendor bill to the right expense account and department.
Agent benchmarks for business software help with method but not with content. CRMArena-Pro, for example, evaluates agents in a seeded CRM environment built on synthetic enterprise data [3]. Synthetic ledgers rarely reproduce the messy parts of real books: vendor names that drift across invoices, split payments, truncated bank descriptors, and chart-of-accounts changes mid-year. For the general argument, see when synthetic evaluation data misleads.
Which accounting tasks belong in the test set
A useful suite covers four task families, each with a different input record and a different ground truth. Treat this as a working hypothesis to adapt to your agent's scope, not a fixed taxonomy.
| Task family | Input the agent sees | Ground truth to source | Primary metric |
|---|---|---|---|
| Transaction / GL coding | Bill, card transaction or bank line; chart of accounts; vendor history | Approved account, department/class, tax code, memo | Exact match on account + dimension; cost-weighted error |
| Bank reconciliation | Bank statement lines (BAI2, MT940 or camt.053) plus open GL items | Reviewed match pairs, including one-to-many and unmatched items | Match precision/recall; unreconciled difference in currency |
| AP / three-way matching | PO, receipt, invoice | Approved match or exception code (price, quantity, duplicate) | Exception detection recall at fixed false-positive rate |
| Accruals and close | Trial balance, prior-period entries, close checklist | Posted adjusting entries and checklist sign-offs | Ledger end-state diff; checklist completion accuracy |
Bank inputs matter more than teams expect. Statements reach accounting systems as BAI2, SWIFT MT940 or ISO 20022 camt.053 files, with different field structures and reference data. If your production traffic is camt.053 with end-to-end IDs and your test set is all MT940, matching scores will not transfer. The invoice reconciliation workflow page describes what these records typically contain.
The reviewed entry is the label, not the preparer's first draft
Ground truth for accounting evaluation should be the entry a reviewer approved, because the preparer's first attempt carries the very errors you want the agent to avoid. Many ERP and close tools record both: a draft or "prepared by" state and an approved or "reviewed by" state with timestamp and user. Ask for both, plus any reversing or correcting entry posted later in the period.
That gives you three useful signals per case. The approved entry is the target. The diff between preparer and reviewer marks the hard cases where human experts disagreed or erred. A later reversal shows when even the approved answer was wrong, which you should either relabel or exclude.
Label noise is not a minor issue. Audits of widely used ML test sets found an average label error rate of at least 3.3%, enough to reorder model rankings [4]. In bookkeeping, miscodings that slip past review tend to cluster in low-materiality accounts, so stratify review effort toward accounts where a wrong answer costs money. The same approach applies to audited codes in medical coding evaluation.
Grade the ledger end state after the agent acts
Score an accounting agent by comparing the books after its actions with the approved end state, rather than grading its explanation. τ-bench applies this pattern to customer-service agents: it compares the database state after a conversation with an annotated goal state, under written domain policies [2]. The same idea maps to a ledger, though the mapping is an inference, not something the benchmark tests.
In practice, load the opening trial balance and open items into a sandbox ledger, let the agent post entries and matches, then diff against the reviewed close. Report, at minimum:
- Account-level diff: accounts whose ending balance differs from the approved balance, and by how much.
- Entry-level match: each approved entry found, missing or extra, keyed on date, amount, account and dimension.
- Reconciliation residual: unreconciled difference per bank account, in currency.
- Policy violations: postings to locked periods, entries above an approval threshold without escalation, edits to closed vendor records.
End-state grading tolerates different but equivalent paths, such as two entries instead of one compound entry. It still catches the failure that matters: a balanced but wrong ledger. For environment design, see agent evaluation task suites with state-based grading.
A test case record with sign-off fields
Each test case should carry the inputs, the approved answer, the review trail and the split metadata in one record. The schema below shows the fields worth requesting from a data supplier.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"case_id": "glc-2025-11-00417",
"task_family": "gl_coding",
"entity": "ENTITY_A",
"period": "2025-11",
"input": {
"source_doc": "vendor_bill",
"vendor": "VENDOR_0193",
"amount": 1412.50,
"currency": "USD",
"line_text": "Annual license renewal - analytics seats",
"chart_of_accounts_version": "coa_v7"
},
"preparer_entry": {"account": "6100 Software Subscriptions", "department": "Ops"},
"approved_entry": {"account": "1450 Prepaid Expenses", "department": "Data", "amortize_months": 12},
"review": {"status": "approved", "reviewer_role": "controller", "reviewed_at": "2025-12-04", "changed_fields": ["account", "department"]},
"later_reversal": false,
"materiality_band": "low",
"split": "test_2025H2"
}
The changed_fields list lets you build a hard subset in one query. The chart_of_accounts_version prevents grading an agent against an account that did not exist in its prompt.
Splitting by period to prevent leakage
Split accounting evaluation data by fiscal period and entity, never randomly by transaction. Vendor-to-account mappings repeat month after month, so a random split puts near-duplicates of each test case in the training or few-shot pool. Process-mining researchers found the same problem in business event logs: random splits leak information across cases, and temporal, case-level splits reduce it [6].
Hold out the most recent closed periods as the test window, and keep at least one entity entirely unseen to test transfer to a new chart of accounts. Refresh the window as new periods close, so public or vendor models are less likely to have trained on it [5]. More patterns are covered in contamination-resistant evaluation design.
De-identifying counterparties without breaking matching
Accounting records are full of counterparties, employee names and account numbers, and de-identification must keep the joins that matching depends on. Replace each vendor, customer and employee with a consistent pseudonym across bills, payments and bank lines. Otherwise a reconciliation test becomes unsolvable or trivially easy.
Mask bank and card numbers to a stable token, keep amounts and dates exact unless the supplier requires perturbation, and scrub free-text memo fields where people paste names and emails. If records relate to California consumers, check whether the result meets the CCPA definition of deidentified information, which also places conditions on the business holding it [7]. See de-identifying evaluation data without breaking the test.
How SourceX approaches accounting evaluation data
SourceX sources operational datasets from US companies, including finance and legal workflow records, and manages the licensing process. Records are sourced on request rather than held in stock, so a request does not guarantee a match. You describe the data you need, not the businesses that hold it, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery.
Before delivery, personal details such as names, emails, phones and account numbers are removed or replaced. The method is recorded and a sample is checked, though no method is perfect. Diligence materials on source, rights, preparation and allowed use are prepared per dataset. To scope reviewed GL, reconciliation or close records for an agent eval, start with the SourceX buyer intake. Related pages include finance and accounting AI training data, accounting reconciliation and close datasets, the month-end close workflow and evaluation datasets built from real business work. For the wider cluster, see the LLM evaluation datasets buyer's map and building a golden evaluation set from business records.
Sourcing an accounting agent evaluation dataset
Describe the task families, ledger systems, bank formats and periods you need to test, and SourceX looks for US businesses that hold matching reviewed records. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set in a license per deal. Describe your accounting evaluation data needs.
Sources
- arXiv, "SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities" (2025). https://arxiv.org/pdf/2504.04596
- arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- arXiv (Salesforce AI Research), "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions" (2025). https://arxiv.org/pdf/2505.18878
- arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- arXiv (Weytjens and De Weerdt), "Creating unbiased public benchmark datasets with data leakage prevention for predictive process monitoring" (2021). https://export.arxiv.org/abs/2107.01905
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.