Industry-specific operational data
SOX and internal control testing workpapers for controls-testing AI
Quick answer
SOX control testing data for AI is the documented trail of an ICFR testing cycle: risk and control matrix (RCM) rows, test attributes and procedures, sample selections, evidence references, exceptions, and deficiency evaluations graded as control deficiency, significant deficiency or material weakness. Buyers training controls-testing or workpaper-review agents need these artifacts linked end to end, sourced from internal audit or SOX program teams with company consent, and de-identified at the employee and system-credential level before delivery.
By SourceX Editorial · Updated
What a SOX control-testing workpaper actually contains
A usable workpaper set is a chain of linked records, not a folder of PDFs. Under Section 404, management evaluates internal control over financial reporting (ICFR), and SEC interpretive guidance from 2007 describes that evaluation as top-down and risk-based [5], starting from financial reporting risks and working down to the controls that address them. Most US programs map controls to the COSO 2013 Internal Control–Integrated Framework, and integrated audits by the external auditor follow PCAOB AS 2201.
The records a model can learn from usually include:
- RCM rows: process, sub-process, risk statement, financial statement assertion (existence, completeness, accuracy, valuation, presentation), control ID, control description, control type (preventive or detective), nature (manual, automated, IT-dependent manual), frequency, key-control flag and owner role.
- Test of design and walkthrough notes: narrative of how the control operates, the system or report it relies on, and whether the design addresses the risk.
- Test of operating effectiveness scripts: numbered attributes (for example "approver is not the requester", "approval dated before posting"), population source, sample size rule and procedure per attribute.
- Sample selections: population definition, completeness check, selection method and the selected item keys.
- Evidence references: tickmarked screenshots, system reports, ticket exports and sign-offs linked to each sample item and attribute.
- Exceptions and conclusions: attribute failed, root-cause note, compensating control considered, and the tester's and reviewer's conclusions.
- Deficiency evaluations: severity rating, aggregation with other deficiencies, and remediation plan with retest results.
PCAOB AS 1215 sets the bar that makes these records valuable for training: documentation must let an experienced auditor with no previous connection to the engagement understand the work performed, who did and reviewed it, and the conclusions reached [1]. Workpapers written to that standard are close to self-explanatory labeled examples.
Why public data cannot cover SOX and ITGC testing
Real control-testing records are effectively absent from public datasets because they are confidential company records. Even general enterprise transaction data is scarce publicly, since privacy, confidentiality and commercial interests make cleansed company datasets hard to procure [2]. Control workpapers are more sensitive still: they describe where a company's controls are weak.
Synthetic RCMs from templates teach vocabulary but not judgment. What models miss without real data is the distribution of messy cases: populations that fail completeness checks, evidence that is ambiguous rather than clearly passing or failing, exceptions later cleared by a compensating control, and reviewer notes that send a test back.
ITGC testing evidence: the hardest and most useful slice
IT general controls (ITGCs) are where controls-testing agents earn their keep, because the evidence is high-volume and structured. Typical ITGC domains are access to programs and data, program change management, computer operations, and program development, tested over in-scope systems such as SAP, Oracle E-Business Suite, NetSuite, Workday and the databases and identity layers beneath them.
Representative test types and their evidence:
- User access provisioning: access request tickets (ServiceNow, Jira) matched to role assignments and approvals.
- User access reviews: periodic user listings with reviewer sign-off and evidence that removals were actioned.
- Terminated-user removal: HR termination reports joined to account disable dates.
- Privileged access and segregation of duties: SoD conflict reports and firefighter or break-glass logs.
- Change management: change tickets, approvals, test evidence and deployment logs, including the completeness check that every production change has a ticket.
- Job monitoring and backups: scheduler failure logs and resolution tickets.
The failure mode buyers most often hit is evidence that cannot be reattached to the test. If screenshots are separated from attribute-level tickmarks, or user listings lack the report parameters and extraction timestamp, the record no longer shows why the tester concluded what they did. Require evidence-to-attribute linkage in the specification.
Deficiency evaluation data and severity labels
Deficiency evaluations are the highest-value labels in the set because they encode judgment. PCAOB standards (AS 2201 and AS 1305) distinguish a control deficiency, a significant deficiency (less severe than a material weakness but merits attention from those overseeing financial reporting) and a material weakness (a reasonable possibility that a material misstatement will not be prevented or detected on a timely basis). Significant deficiencies and material weaknesses must be communicated in writing to management and the audit committee, and AS 1305 points auditors to the AS 2201 evaluation guidance.
For a deficiency-classification model, ask for the reasoning inputs alongside the label: the magnitude considered, the likelihood assessment, the accounts and assertions affected, compensating controls evaluated, and whether deficiencies were aggregated. A severity label without those fields teaches pattern matching on wording rather than evaluation.
Expect a skewed label distribution. Most tests pass, most exceptions resolve as control deficiencies, and material weaknesses are rare, so plan oversampling or targeted collection for the severe classes and hold out a balanced evaluation set (see LLM evaluation datasets).
Illustrative record: one linked test of operating effectiveness
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"control_id": "P2P-07",
"process": "Procure to pay",
"risk": "Invalid or unapproved vendor invoices are paid",
"assertions": ["existence", "accuracy"],
"control_type": "preventive",
"nature": "IT-dependent manual",
"frequency": "per transaction",
"key_control": true,
"test": {
"type": "operating_effectiveness",
"period": "FY-Q1..Q3",
"population": {"source": "ERP payment run report", "row_count": 18422, "completeness_check": "reconciled to GL cash disbursements"},
"sample": {"method": "random", "size": 40},
"attributes": [
{"id": "A", "text": "Invoice approved by authorized role per DOA matrix"},
{"id": "B", "text": "Approver differs from invoice creator"},
{"id": "C", "text": "Approval timestamp precedes payment run"}
]
},
"results": [
{"sample_ref": "S-0013", "A": "pass", "B": "fail", "C": "pass", "evidence_refs": ["EV-0013-1", "EV-0013-2"]}
],
"exception": {"count": 1, "root_cause": "temporary role assignment during system migration", "compensating_control": "monthly vendor payment review P2P-11"},
"deficiency": {"severity": "control_deficiency", "aggregated_with": [], "rationale_present": true},
"people": {"tester": "TESTER_04", "reviewer": "REVIEWER_01", "control_owner": "ROLE_AP_MANAGER"},
"remediation": {"plan": "remove temporary role", "retest_result": "pass"}
}
Note what is replaced: tester, reviewer and owner names become stable pseudonymous tokens, and evidence is referenced by ID so screenshots can be redacted separately.
Who holds licensable workpapers and what limits them
Internal audit departments, SOX program offices and consulting firms that run outsourced or co-sourced testing hold the most licensable sets, because the company that owns the controls can consent to release. External-auditor workpapers sit under professional confidentiality obligations to the audit client and firm retention rules, so treat them as a separate and harder rights question; external-audit judgment data such as risk assessments and review notes is outside this page's scope.
Practical limits to plan for:
- Company consent is required because the records describe the company's own controls and systems.
- Evidence attachments carry identities and secrets: user listings, approver names, email headers, hostnames, IP addresses and occasionally credentials or API keys visible in screenshots.
- Templates vary by firm and year, so RCM fields and attribute wording need normalization across sources.
- License terms are often missing or wrong in shared datasets generally, with one audit of 1,800+ text datasets finding license omissions above 70% on popular hosting sites [3]; for confidential workpapers, insist on a written license from the data owner.
Where workpapers include personal information of California residents, including employees, note that the CCPA treats information as deidentified only when it cannot reasonably be linked to a particular consumer and the business takes reasonable measures, commits publicly not to reidentify, and contractually binds recipients [4]. That contractual element belongs in your license, not just your pipeline.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Buyer specification checklist for control-testing data
A precise request separates usable sets from document dumps. Use this as a request template:
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to specify | Why it matters |
|---|---|---|
| Scope | SOX 404 business-process controls, ITGCs, or both; framework mapping (COSO 2013 components) | Determines label vocabulary |
| Systems | ERP, identity and ticketing systems in scope | Evidence formats differ by system |
| Linkage | RCM row to test to sample to evidence to exception to deficiency | Without links, no reasoning chain |
| Labels | Pass or fail per attribute; severity tier; reviewer overrides | Supervision targets |
| Rationale | Free-text conclusions and review notes kept | Teaches judgment, not keywords |
| Volume | Controls, test instances, periods and companies | Diversity beats one large program |
| Formats | Excel RCMs, PDF or image evidence, CSV extracts, tool exports (AuditBoard, Workiva and similar) | Parsing cost |
| De-identification | Names, emails, user IDs, hostnames, credentials replaced; method recorded | Privacy and security review |
| Holdout | Periods or companies reserved for evaluation | Prevents leakage into eval |
| Rights | Owner consent, allowed uses, term, delivery | Approval by counsel |
Before committing to volume, request a sample and score it against this table (see how to request a training data sample and how to evaluate a fine-tuning dataset before you buy). After delivery, evidence repositories should stay behind role-based access, as covered in access controls for licensed training data.
How this differs from reconciliation and remediation data
Control-testing workpapers record whether a control worked; reconciliation records record the accounting work itself. If your model needs account reconciliations, close checklists or variance explanations, see accounting reconciliations and close work. If it needs the remediation trail after a finding, see compliance remediation records. Buyers by industry are summarized for accounting and compliance consulting, and the industry-specific operational data hub maps neighboring categories such as SOC alert triage decisions.
How SourceX handles requests for control-testing workpapers
SourceX sources operational datasets from US companies, including finance and legal workflows and documents, on request rather than from stock, so describing the data does not guarantee a match. Buyers describe the data, not the businesses; SourceX looks for US companies that hold it, and each release is approved by the supplying company. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with names, emails, account numbers and similar details removed or replaced, the method recorded and a sample checked, though no method is perfect. You can describe the workpapers your agent needs at any point in your scoping.
Request SOX control testing data for AI
Tell SourceX which controls, systems, labels and linkage your controls-testing model needs. The process runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Start a buyer request.
Sources
- AS 1215: Audit Documentation. https://pcaobus.org/oversight/standards/auditing-standards/details/AS1215
- SAP advancing enterprise AI research with first real ERP dataset. https://silicon-saxony.de/en/sap-advancing-enterprise-ai-research-with-first-real-erp-dataset/
- A large-scale audit of dataset licensing and attribution in AI. https://www.nature.com/articles/s42256-024-00878-8
- California Civil Code section 1798.140 (CCPA definitions). https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Commission Guidance Regarding Management's Report on Internal Control Over Financial Reporting. https://www.sec.gov/files/rules/interp/2007/33-8810.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.