Skip to content

Code and software engineering data

Legacy Code Translation Pairs: COBOL-to-Java and Other Migration Data for AI

Quick answer

A useful COBOL-to-Java translation dataset is aligned before-and-after code from a completed enterprise migration, paired with the evidence that the new code behaves like the old: regression suites, parallel-run comparisons and reconciliation reports. As of October 2026, public corpora are small and mostly synthetic, so teams training or evaluating translation models usually need licensed pairs from companies that finished a migration. Specify alignment granularity, both conversion stages, equivalence evidence and a system-level eval split before you request anything.

By SourceX Editorial · Updated

Why public COBOL-to-Java corpora fall short

Public parallel data for COBOL is too small and too clean to represent production mainframe estates. The COBOL-Coder authors note that no benchmark had been curated specifically for COBOL translation and built COBOL-JavaTrans, 143 pairs derived from HumanEval-X Java solutions [1]. A patent on low-resource code translation describes generating synthetic parallel COBOL from Java because only about 2,500 COBOL documents exist in public repositories [2]. IBM Research's code translation work prioritizes COBOL-to-Java and commonly draws on CodeNet-style problem sets [3].

Those sources are useful for smoke tests, but they share three gaps. Algorithmic exercises contain no copybooks, no VSAM or DB2 access, no CICS transactions, no JCL job steps and no packed-decimal (COMP-3) arithmetic. Synthetic pairs translated from Java inherit Java idioms in the COBOL side, so a model learns to translate code that no 1987 programmer wrote. Real migration pairs fix both problems, which is why buyers look to companies with completed modernization programs; for single-language estates without a modern counterpart, see COBOL and mainframe code datasets.

Alignment granularity decides what you can train

Migration pairs are rarely one-to-one, so your request must state the alignment level you need and who produces the mapping. A COBOL program often becomes several Java classes, a paragraph or section becomes a method, a copybook becomes a DTO or entity class, and a batch JCL stream becomes a Spring Batch job. Some migrations also re-architect: three CICS transactions collapse into one REST service, and the trace from old to new exists only in a spreadsheet the integrator kept.

Three levels are common in practice:

  • Program-level: one COBOL program (plus its copybooks) mapped to the Java package that replaced it. Easiest to obtain, best for repository-level and agentic translation, weakest for SFT on short contexts.
  • Module or paragraph-level: PERFORM targets mapped to methods. Best for supervised fine-tuning, but the mapping is often reconstructed after the fact and can be wrong.
  • Function or statement-level: fine-grained mappings, usually present only when an automated converter emitted trace comments or a mapping file.

Ask who produced the mapping (converter, integrator team or reconstructed by the supplier's engineers) and whether it was reviewed. A mapping hypothesis is fine if it is labeled as one. Mixing reviewed and inferred alignments without a flag corrupts both training and evaluation. For context-window requirements across files, see repository-level code data.

Ask for both conversion stages, not just the final Java

Most large migrations run an automated converter first and then hand-refactor, so a complete pair has three snapshots: legacy source, tool output and the refactored code that shipped. Tool output alone teaches the tool's style, such as COBOL-shaped Java with fields named WS-CUST-REC, GOTO emulation and calls into the vendor's runtime library for PIC handling and file I/O. Research on evaluating the quality of automated COBOL-to-Java transformation output shows that this quality needs its own measurement, separate from whether the code runs [4].

The refactored stage is where the value is, because it shows how engineers turned literal translation into idiomatic code with BigDecimal arithmetic, JPA repositories and proper exceptions. The diff between tool output and final code is also a strong preference or reward signal. Record which converter and version produced stage two, because generated code may depend on, or embed, proprietary runtime classes that you are not licensing.

Equivalence evidence is the label

A translation pair without evidence that both sides behave the same is an unverified guess, usable as weak SFT data and unusable for evaluation. The strongest evidence comes from how the migration was accepted: regression suites run against both systems, parallel runs where the mainframe and the Java system processed the same production day, and reconciliation reports comparing output files, database tables and report totals record by record. Recent research on LLM-based COBOL translation uses symbolic execution and delta debugging to find and fix behavioral divergences [6], and AI-driven modernization studies likewise need checks beyond textual similarity [5].

Ask for the evidence at the same granularity as the pairs. A green parallel run for a whole nightly batch does not prove that a specific paragraph-level pair is equivalent. Useful evidence fields include test case IDs, input datasets (de-identified), expected and actual outputs, tolerance rules for rounding and date handling, and known accepted differences such as EBCDIC-to-UTF-8 collation changes that alter sort order. Those accepted differences are themselves valuable: they are exactly the edge cases a translation model gets wrong. For building executable checks, see unit test generation data and regression suites for production LLM applications.

Illustrative pair record

A well-formed pair record carries the code, the alignment, the stage and the evidence in one row.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "app07-pgm-ACCTUPD-para-2100",
  "system_id": "app07-billing",
  "source_lang": "COBOL (Enterprise COBOL)",
  "target_lang": "Java 17",
  "alignment_level": "paragraph_to_method",
  "alignment_source": "integrator_mapping_reviewed",
  "legacy_path": "src/cobol/ACCTUPD.cbl#2100-APPLY-PAYMENT",
  "copybooks": ["ACCTREC.cpy", "PAYREC.cpy"],
  "tool_output_path": "gen/com/app07/AcctUpd.java#para2100ApplyPayment",
  "final_path": "src/main/java/com/app07/billing/PaymentService.java#applyPayment",
  "converter": {"name": "redacted", "version": "x.y"},
  "runtime_dependencies": ["vendor-runtime-numeric"],
  "equivalence": {
    "regression_tests": ["RT-4412", "RT-4413"],
    "parallel_run": {"cycle": "2025-03-14", "status": "matched", "tolerance": "0.00"},
    "accepted_differences": ["sort order after EBCDIC to UTF-8"]
  },
  "deidentification": "test data synthetic; account numbers replaced",
  "split": "eval"
}

Split by system, not by file

Train and eval splits for migration pairs must be cut by application or system, never by file or pair. Programs in one estate share copybooks, utility subprograms and naming conventions, so a file-level split leaks the eval answer into training through shared record layouts. Hold out entire applications, and check that no copybook or common module in the eval set appears in training under another name.

Also check the Java side for public contamination: migrated services sometimes reuse open-source frameworks or snippets, which raises both leakage and license questions covered in code benchmark contamination and copyleft contamination in licensed code.

Rights: the client usually owns both versions

The company whose system was migrated is typically the party that must approve release, even when an integrator did the work. Integrators often deliver under client contracts that assign the work product to the client, so an integrator's copy of before-and-after code may not be theirs to license; see who owns data in a client project. Conversion-tool vendors may also assert rights in generated runtime code, and some license terms restrict reuse of their output.

Practical checks: confirm the supplier owns the legacy source and the final Java, identify third-party runtime libraries and exclude or separately clear them, scan both sides and the full history for credentials (see secrets in code datasets), and confirm test data is synthetic or de-identified. Deeper guidance sits in code ownership due diligence and, for decommissioned systems, data licensing for legacy software retirements.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request checklist for migration pair datasets

A precise request makes supplier assessment and sample review meaningful.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to specifyWhy it matters
Language pairsCOBOL to Java, PL/I to Java, RPG to Java, Natural to C#Dialects and target idioms differ sharply
Legacy contextCopybooks, JCL, CICS maps, DB2 DDL, VSAM layoutsPrograms do not compile or make sense without them
Alignment levelProgram, paragraph or statement; reviewed or inferredDetermines SFT versus eval use
StagesLegacy, converter output, final refactored codeAvoids learning one tool's style
Equivalence evidenceRegression IDs, parallel-run results, reconciliation reportsBecomes the label and the RL reward signal
Test dataSynthetic or de-identified inputs and outputsLets you re-execute pairs
Split unitWhole application or systemPrevents copybook leakage
ExclusionsVendor runtime, third-party and open-source codeKeeps rights clean
DeliveryGit repositories with history plus JSONL pair indexPreserves refactoring sequence

For general request structure, see how to specify a code dataset request; database-side conversions of stored procedures and embedded SQL belong in SQL dialect migration data.

Using pairs for SFT, evaluation and RL

The same pair set serves three uses if the evidence is attached. SFT can use all reviewed pairs, preferring final refactored targets over converter output. Evaluation should use only held-out systems with executable equivalence checks, scoring compilation, test pass rate and output reconciliation rather than BLEU or CodeBLEU alone.

For reinforcement learning, equivalence checks become the reward: compile, run the regression cases, compare outputs within the recorded tolerances. That only works if the supplier can deliver runnable harnesses or enough input and expected-output data to rebuild them. Before licensing, run your own pass over a sample with this code sample evaluation guide, and browse the wider code and software engineering data map for adjacent datasets such as dependency upgrade and API migration pairs.

How SourceX handles migration pair requests

SourceX sources operational datasets, including engineering records, from US companies and manages the licensing and ongoing purchases. Data is sourced on request rather than held in stock, so you describe the pairs and evidence you need, not the businesses, and a request does not guarantee a match. Every release is approved by the supplying company, rights-reviewed for ownership and consents, and delivered under a license that defines records, uses, term and delivery. You can describe your migration pair requirements to SourceX; related context is on proprietary code datasets with full Git history and software development agencies as data holders.

Sourcing COBOL-to-Java translation pairs with SourceX

SourceX looks for US companies that hold the migration records you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees. Start a buyer request.

Frequently asked questions

Can converter output alone train a good translation model?

It can teach literal translation, but it reproduces the converter's style and runtime dependencies. Pair it with the hand-refactored final code and flag which stage each target came from.

What if a migration has no paragraph-level mapping?

Program-level pairs are still useful for repository-level and agentic translation. Ask the supplier to label any reconstructed mapping as inferred so you can exclude it from evaluation.

Do I need the mainframe test data?

You need enough inputs and expected outputs to re-execute equivalence checks. Synthetic or de-identified test files and reconciliation outputs usually suffice; production records are rarely necessary.

Sources

  1. arXiv, "COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation" (2026). https://arxiv.org/pdf/2604.03986
  2. USPTO, "Large language models for creating a multi-lingual, low-resource code translation dataset". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/12579050
  3. IBM Research, "Code translation". https://research.ibm.com/projects/code-translation
  4. arXiv, "Quality Evaluation of COBOL to Java Code Transformation" (2025). https://arxiv.org/pdf/2507.23356
  5. arXiv, "Code Reborn: AI-Driven Legacy Systems Modernization from COBOL to Java" (2025). https://arxiv.org/pdf/2504.11335
  6. arXiv, "SEDCoT: Enhancing LLM-Based COBOL Code Translation via Symbolic Execution and Delta Debugging" (2026). https://arxiv.org/pdf/2607.04092

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data