Skip to content

Tables, time series and transactional data

Journal Entry Data for Accounting Anomaly Detection and Audit Analytics

Quick answer

A useful journal entry anomaly detection dataset is a line-level general-ledger extract from a real company, with posting metadata that auditors actually test (posting user, entry source, document and posting dates, approvals, reversal links) plus labels whose origin is documented: reviewed exceptions, audit adjustments and reversals. Public data rarely offers this, so teams usually need to license it from the companies that hold it. Specify the population, field list, label provenance and period-aware splits before you compare suppliers.

By SourceX Editorial · Updated

Why public ledger data rarely works for journal entry testing models

Public datasets almost never contain real journal entry detail, because general-ledger records carry confidential financial information that companies have no reason to release. SAP's release of SALT, anonymized linked business tables from a real ERP system, was notable precisely because real linked ERP data is so rarely published [2]. Even SALT targets sales-document autocompletion, not posting-level audit metadata.

Labeled anomalies are scarcer still. Even in adjacent e-commerce fraud detection work, researchers have judged existing open datasets insufficient and had to prepare a combined training set [3], and synthetic ledgers tend to encode the generator's idea of fraud rather than the patterns a reviewer actually flags. If your model will be sold to auditors or controllers, it needs to be trained and evaluated on entries that came out of real close processes. See the structured data buyer's guide for how this fits the wider tabular category.

What auditors test, and therefore what your data must carry

Your dataset needs the attributes that journal entry testing actually turns on, because buyers will judge model output against how an auditor would reason. Under PCAOB AS 2401 and ISA 240, auditors design journal entry procedures around the risk of management override of controls. In practice that means selecting entries against documented fraud-risk criteria (manual versus automated, unusual users, seldom-used accounts, period-end and post-closing entries, thin descriptions) and confirming that the journal entry population is complete.

That translates directly into data requirements. A model cannot learn "unusual posting user" without a stable user identifier and the user's role, and it cannot learn "manual top-side entry after close" without an entry-source flag and both document and posting dates. It also cannot be evaluated on completeness unless the extract ties back to the trial balance. Treat these as must-have fields, not enrichments.

Field specification for a journal entry extract

Request line-level detail with header attributes repeated or joinable, plus a chart of accounts and period trial balances for reconciliation. The AICPA's voluntary Audit Data Standards, including a General Ledger Standard, are a common reference for file and field definitions; aligning to it makes multi-company data easier to combine, even if each supplier's ERP (SAP S/4HANA, Oracle, NetSuite, Dynamics 365, Sage Intacct) names fields differently.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy the model needs it
je_id / line_noJE-2025-0412883 / 3Groups lines into balanced entries
company_code, ledger1000, 0L (leading ledger)Separates entities and parallel ledgers
gl_account, account_type6105, expenseRare-account and unusual-pairing signals
debit_amount / credit_amount, currency0.00 / 48,500.00 USDMagnitude, round-number and threshold patterns
document_date, posting_date, entry_timestamp2025-12-31, 2026-01-09, 2026-01-09T22:14ZBackdating, post-close and off-hours signals
fiscal_period12 (plus special periods 13-16)Period-end and adjustment-period behavior
posting_user_id (pseudonymized), user_roleU-7781, controllerUnusual-user and segregation-of-duties signals
source / doc_typemanual, SA (GL document) vs. automated subledgerManual versus automated split
approver_id, approval_timestampU-2034, 2026-01-10T08:02ZSelf-approval and missing-approval checks
reversal_of / reversed_byJE-2025-0412101Reversal chains and weak labels
line_description"Accrual true-up per CFO"Text signals; also a privacy risk
exception_label, label_source, label_dateadjusted, audit_adjustment, 2026-02-20Supervision with documented provenance

Ask for the chart of accounts with account descriptions and a mapping column, because each company's chart is its own. For cross-company training, map accounts to a standard taxonomy (for example, a common financial statement line hierarchy) and keep the original account so the mapping can be audited. The ERP data guide and the tabular license specification page cover grain, keys and history in more depth.

Building ledger anomaly labels you can defend

Most journal entry data has no clean fraud label, so treat each label as a weak signal and record exactly how it was produced. Confirmed fraud is extremely rare in any single ledger, and a model trained on a handful of cases will memorize them. Weak labels give you volume, but each type carries its own bias.

Illustrative example: invented to show structure; it does not describe an available dataset.

Label sourceWhat it indicatesKnown bias or failure mode
Auditor-proposed adjustment linked to entryEntry was materially wrongReflects the auditor's sample, not the population
Controller-reviewed exception (close checklist)Entry was flagged and investigatedDisposition codes vary; "cleared" is not "normal"
Reversal within N daysEntry was later undoneMany reversals are routine accrual reversals
Rule hit from prior JE testing scriptsEntry met a risk criterionModel learns the rule, not the risk
Internal audit or investigation findingConfirmed issueVery rare; often unshareable

Ask the supplier for a disposition field (adjusted, cleared, unresolved) rather than a binary flag, plus the date the label was assigned. Exclude routine auto-reversing accruals from the reversal label, or the model will learn period-end mechanics instead of risk. Document all of this in a data statement or dataset card so downstream evaluators understand what "anomalous" means in your set [5]. For transaction-level fraud labels from payments rather than ledgers, see fraud-labeled transaction data; for approval-chain records, see approval and rejection records.

Splitting by period to avoid leakage

Split journal entries by time, and keep whole close cycles together, because ledgers are dominated by month-end and year-end seasonality. Random splits on time-ordered data leak future information into training [1]; in ledgers that leak is worse, since a reversal posted in January reveals the label of a December entry. Hold out complete later periods, including at least one fiscal year-end, and keep a company-level holdout if you claim cross-company generalization.

Watch for label leakage too. If label_date falls after your prediction cutoff, or if fields like reversed_by or post-audit document types are present at training time, the model sees the answer. The target leakage audit guide gives a column-by-column procedure that applies directly here.

Confidentiality, pseudonymization and MNPI timing

Ledger data is commercially sensitive even after personal details are removed, so review confidentiality and identifiability together. Posting user IDs, vendor and customer names in line text, employee reimbursement entries and bank account numbers all need pseudonymization or removal. Identifiability is a risk judgment, and the motivated intruder test is a practical way to frame it: could someone with public filings and industry knowledge work out the company or individuals from amounts, dates and account patterns [4]?

For public issuers, unreleased period results can be material nonpublic information, so ask whether the extract covers periods that have not yet been reported and agree on a release lag with the supplier's counsel. Consider scaling or bucketing amounts, shifting dates consistently within an entity, and removing free-text descriptions you do not need. Each transform affects which anomaly signals survive, so test model performance on the transformed data, not the raw extract.

Buyer checklist before signing

Confirm population, labels and permissions in writing before any data moves, because gaps in any of the three can make the data unusable for anomaly work.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Population: Which entities, ledgers and periods? Does the extract reconcile to period trial balances, and are special periods included?
  • Fields: Are posting user, user role, entry source, document and posting dates, approver and reversal links all present and non-null at stated rates?
  • Labels: Which label sources exist, how was each produced, and what disposition codes are used?
  • Mapping: Is the chart of accounts supplied with descriptions, and who maintains the mapping to your taxonomy?
  • Privacy and confidentiality: What was pseudonymized or removed, by what method, and was a sample checked?
  • Timing: Are any periods unreported, and what release lag applies?
  • Rights: Does the supplying company own the data and authorize the stated AI uses, term and delivery method?

How SourceX approaches journal entry requests

SourceX sources operational datasets from US companies on request, including finance and legal workflow records, and manages licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match; you describe the data, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details such as names and account numbers removed or replaced before delivery, the method recorded and a sample checked (no method is perfect). You can start a buyer request with the field list above.

For adjacent owner pages, see accounting reconciliation datasets, what month-end close records show and finance operations AI. If you are working on system-log anomalies instead of ledger anomalies, see production log data for anomaly detection.

Request journal entry anomaly detection data

Describe the entities, periods, fields and label sources you need, and SourceX will look for US companies that hold matching ledger data and manage the process from assessment to license. Pricing and allowed uses are agreed per deal, and nothing is contracted until a supplier agrees. Describe the journal entry data you need.

Sources

  1. temporalcv documentation, "Why Time Series Is Different". https://temporalcv.readthedocs.io/en/latest/guide/why_time_series_is_different.html
  2. SAP News Center, "SAP SALT: Real ERP Dataset for Enterprise AI Research" (2025). https://news.sap.com/2025/04/sap-salt-real-erp-dataset-enterprise-ai-research/
  3. Information Technology and Mathematical Modelling journal, "Methodology of dataset preparation for training e-commerce fraud detection models". https://journals.nmetau.edu.ua/index.php/itmm/en/article/view/2468
  4. Information Commissioner's Office (ICO), "How do we ensure anonymisation is effective?" (2026). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  5. University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data