Skip to content

Code and software engineering data

COBOL and Mainframe Code Datasets for LLM Training

Quick answer

A COBOL dataset for LLM training is only useful if it is a complete application estate: every program with the copybooks it includes, the JCL that runs it, the CICS BMS maps and DB2 DDL or IMS DBDs it touches, and enough run documentation to explain behavior. Public COBOL is scarce, so most teams need licensed estates from companies that ran mainframes. Specify completeness, encoding, ownership exclusions and an evaluation plan before reviewing any sample.

By SourceX Editorial · Updated

Why public COBOL will not carry a modernization model

Public COBOL is too thin and too clean to train a model that explains or converts real production estates. Recent research builds COBOL training sets from documentation-derived synthetic question-answer pairs plus curated source validated by a compiler, precisely because real corpora are lacking [1]. Mainframe-modernization models such as XMainframe show the demand, and their authors had to assemble dedicated mainframe training data to get there [2].

Community efforts are trying to close the gap. The Open Mainframe Project's Zorse effort collects permissively licensed COBOL, including code from decommissioned systems [3], and other researchers generate deliberately degraded synthetic COBOL to mimic decades of patching [5]. Synthetic rot is a useful supplement, but it cannot reproduce the naming conventions, dead paragraphs, REDEFINES tricks and business rules of a 30-year-old claims or ledger system. That is what licensed estates from banks, insurers and other operators add (see the finance buyer page and insurance buyer page).

The completeness rule: a program is not a compile unit

A COBOL program cannot be parsed, let alone understood, without the copybooks pulled in by its COPY statements. Record layouts, 88-level condition names and shared working storage usually live in copy members, which the compiler resolves from SYSLIB concatenations of PDS or PDSE libraries, so the dataset must also record which library each member came from. A dataset of .cbl files without their copy libraries yields truncated programs and hallucinated field names.

Batch behavior lives outside the program too. JCL job streams define step order, DD statements map logical files to datasets, and PROCs and scheduler definitions decide what runs when. Online programs depend on CICS BMS mapsets for screens, and data access depends on DB2 DDL or IMS DBDs and PSBs. Require every compile unit with its includes, plus the job streams and database definitions it references, and ask for a manifest that proves closure.

Preprocessed code and the EXEC blocks

Embedded EXEC SQL and EXEC CICS blocks are not standard COBOL, so a buyer needs the preprocessing context to train on them correctly. Current IBM Enterprise COBOL toolchains can handle EXEC CICS with an integrated translator and EXEC SQL with the Db2 coprocessor, which emits a DBRM, while older builds ran separate precompile steps; check the compiler and Db2 documentation for the versions in the estate. The same estate may therefore contain programs prepared both ways, with different JCL compile procedures.

Ask whether compile listings or preprocessed output can be supplied alongside source. Listings show resolved copybooks, compiler options and diagnostics, which makes them strong supervision for explanation tasks. DBRMs or bind information connect SQL statements to the DB2 objects they hit, which matters for data-lineage questions a documentation model will be asked.

Encoding, columns and sequence numbers

Mainframe source is stored in EBCDIC with fixed-format columns, so conversion choices change what the model sees. Columns 1 to 6 hold sequence numbers, column 7 is the indicator area for comments and continuation, Areas A and B carry code, and columns 73 to 80 often hold change tags. Packed-decimal (COMP-3) and binary fields in test files are not text at all.

Ask for UTF-8 conversion with the source code page recorded (for example IBM-037 or IBM-1047) and the original bytes kept for verification. Decide whether sequence numbers and change tags are stripped or retained; they leak change history and can make copied members look distinct to deduplication, which our page on deduplicating code training data covers. Never let a conversion silently mangle national characters, brackets or the cent sign.

Ownership traps in mainframe estates

Mainframe estates routinely mix code the company owns with code it only licenses. Vendor packages for banking cores, insurance policy administration or payroll are often customized in place, and their source belongs to the vendor. Code generated by 4GL or CASE tools may carry generator runtime libraries or license terms the operator cannot pass on.

These components need exclusion or a separate rights path. Copyright headers, vendor prefixes on program names and generator comment banners help identify them, but a person who knows the estate must confirm the split. Our guide to code ownership due diligence lists the documents to request, and retired systems raise their own questions, covered in licensing data from a legacy system you are retiring.

Evaluating models when unit tests do not exist

Many COBOL estates have no unit tests, so evaluation relies on compile success, SME-judged documentation and output comparison on test files. Ask whether regression test datasets and expected outputs exist for batch jobs; they let you compare a converted program's output byte for byte. Public benchmarks such as COBOLEval [4] are useful but become exposed to contamination once they circulate, a risk explained in code benchmark contamination.

Reserve part of each licensed estate as a held-out set at the application level, not the file level, so shared copybooks cannot leak between train and test. Mainframe SMEs should rate documentation for correctness of business rules, not prose quality. Credentials hard-coded in JCL or PROCs need scanning before delivery, as described in secrets removal for code datasets.

Request template for a COBOL estate

Write the request as a closure and rights specification, not a language filter. The template below adapts the general code dataset request specification to mainframes.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to specifyWhy it matters
ProgramsCOBOL batch and online programs, with dialect and compiler version if knownDialect and compiler options change valid syntax
Copy librariesAll copy members referenced, plus a COPY-to-member resolution manifestWithout them programs do not parse
Job controlJCL, PROCs, scheduler definitionsDefines batch flow and file mappings
Online layerCICS BMS mapsets, transaction and program definitionsNeeded for screen-flow explanation
Data definitionsDB2 DDL, DBRMs or bind data; IMS DBDs and PSBs; VSAM record layoutsLinks code to data
PreprocessingCompile listings or translated output where availableResolves EXEC SQL and EXEC CICS context
EncodingSource code page, UTF-8 conversion rule, original bytes keptVerifiable conversion
ExclusionsVendor packages, 4GL-generated code, third-party utilitiesOwnership and rights
Test assetsRegression inputs and expected outputs, maskedOutput-comparison evaluation
DocumentationRun books, operator notes, change historySupervision for explanation tasks
PrivacyPersonal data removed from comments, test files and literalsTest decks often contain customer records

If you need before-and-after migration code rather than COBOL alone, see legacy code translation pairs. For the wider buyer's map, start at the code and software engineering data hub, and check sampling advice in evaluating a code dataset sample.

How SourceX handles COBOL estate requests

SourceX sources operational datasets, including engineering records and documents, from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data, not the businesses, and SourceX looks for US companies that hold it, with every release approved by the supplying company. The general offer for source code is described on proprietary code datasets, and retirement-driven projects on legacy software retirements.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, account numbers and phone numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement; describe your COBOL requirements to SourceX to start the Find and Assess steps.

Sourcing COBOL training data for your modernization model

SourceX serves AI teams wherever they are based and manages the commercial process from Find through Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees. Prices are not published; terms are agreed per deal. Request COBOL and mainframe code data.

Sources

  1. arXiv, "COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation" (2026). https://arxiv.org/pdf/2604.03986
  2. arXiv, "XMainframe: A Large Language Model for Mainframe Modernization" (2024). https://arxiv.org/pdf/2408.04660
  3. Open Mainframe Project (Linux Foundation), "Open Mainframe Project post on the Zorse COBOL dataset effort". https://openmainframeproject.org/?p=14001
  4. ecosyste.ms (GitHub project listing), "zorse-project/COBOLEval". https://awesome.ecosyste.ms/projects/github.com%2Fzorse-project%2Fcoboleval
  5. arXiv, "Spec2COBOLRot: An Agentic-AI Degradation Loop for Realistic COBOL Corpus Generation" (2026). https://arxiv.org/pdf/2609.26835

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data