Skip to content

Code and software engineering data

ABAP and SAP Custom Code Data for AI Coding Models

Quick answer

ABAP for LLM training is scarce because almost all of it lives inside private SAP systems, not public repositories [1]. Useful datasets therefore come from licensing a company's own custom code: customer-namespace programs, classes, enhancements and data dictionary (DDIC) definitions, exported with version and transport history, ideally paired with the ABAP Test Cockpit findings that drove S/4HANA remediation. Scope the namespace boundary, the export path and the remediation pairs before you discuss volume.

By SourceX Editorial · Updated

Why public ABAP data will not carry a migration assistant

Public ABAP is enough to measure a model, not to train one. A January 2026 benchmark built 180 ABAP tasks, 164 adapted from HumanEval and 16 SAP-specific scenarios, and scored them with the SAP compiler and unit tests; its authors describe ABAP as low-resource precisely because code is proprietary and SAP-embedded [1]. The companion public dataset on Hugging Face is MIT-licensed and under 1,000 rows, which is evaluation scale [2].

Broad surveys of LLMs for software engineering show datasets concentrated on mainstream languages such as Python and Java [3]. Scraped ABAP fragments from forums and blogs add little: they are short, often pre-7.40 syntax, rarely compile in isolation and carry unclear rights. For continued pre-training or supervised fine-tuning on real enterprise patterns (BAdI implementations, user exits, ALV reports, BAPI wrappers, CDS views, RAP behavior definitions), you need code licensed from the companies that wrote it. Practitioner research on training data quality also treats properties such as correctness and relevance as quality criteria, not volume alone [4].

The ownership boundary: customer namespace versus SAP-delivered code

The licensable unit is code the company wrote, not code SAP shipped. Customer objects conventionally sit in the Z and Y namespaces or in a registered customer namespace such as /ACME/, and the company that commissioned them can usually license them. SAP standard code, and SAP objects a customer modified in place, remain subject to the customer's SAP agreement; treat them as out of scope unless counsel confirms otherwise after reading that agreement.

Three edge cases cause most disputes:

  • Modifications and implicit enhancements. A modification to an SAP include changes SAP code; an implicit enhancement section inserted into it is customer code anchored in SAP code. Export only the enhancement body and a reference to the anchor point, not the surrounding SAP source.
  • Integrator-written code. Code built by a systems integrator may belong to the client or the integrator depending on the services contract. Ask for the IP assignment clause, as covered in code ownership due diligence.
  • Partner add-ons. Objects in a third-party vendor namespace belong to that vendor even when installed in the customer's system.

Licensing matters here because the U.S. Copyright Office's Part 3 report, still a pre-publication version as of October 2026, concludes that many acts in AI training implicate the reproduction right [5]. A clean namespace filter plus an ownership attestation is cheaper than an unwinding exercise later.

Exporting ABAP history: transports and version management, not git

ABAP history lives in the SAP system's version management and transport requests, not in a git log. Each object has versions stored in the repository, and changes move through the landscape in transport requests with a short description, an owner and an object list (E071 entries). Some teams also mirror code to git with abapGit, but that mirror usually begins at adoption and loses older history.

Ask the supplier how versions were extracted, from which system (development, quality or production), and whether transport descriptions and linked change tickets are included as change rationale. A transport text such as "Fix rounding in Z_FI_ACCRUAL for S/4 decimal shift" is a weak but real commit message. Expect gaps: deleted objects, versions purged by housekeeping, and transports merged during client copies. For general repository delivery conventions, compare with delivering code repositories with full history.

Remediation pairs: the most valuable unit for S/4HANA assistants

For migration assistants, the strongest training and evaluation unit is a remediation pair: custom code before the S/4HANA check, the finding that flagged it, and the code after the fix. In SAP conversion projects, ABAP Test Cockpit (ATC) readiness checks compare custom code against a simplification database of affected or deprecated objects, and the resulting findings list drives effort estimates and the adaptation worklist. Teams commonly run these checks from a central ATC system that scans the legacy landscape remotely, so the findings exist as structured records rather than only in developer memory.

That makes a well-run conversion project a natural source. Typical findings include direct reads of tables replaced in S/4HANA (for example index tables in finance and logistics), field length extensions such as the 40-character material number, SELECT statements that relied on implicit sort order on the old database, and obsolete function modules. A pair is only useful if the finding ID, check variant and simplification item are kept with it, and if the after version actually passed a re-run of the check.

Hold back a slice of pairs from systems the model never sees for evaluation; see held-out coding agent evaluation sets. The pattern is close to legacy code translation pairs, but here the language stays ABAP while the platform underneath changes.

DDIC structure yes, production table contents no

Deliver data dictionary definitions, never production table data, unless that data goes through a separate review. Table, structure, data element, domain and CDS view definitions are code-like metadata that a model needs to resolve types and joins. The rows inside a Z table, a customizing table or a change-document log (CDHDR/CDPOS) are business records that can contain customer names, vendor bank details, employee numbers and pricing.

Also scan for secrets: hardcoded RFC destination passwords, SM59 credentials pasted into comments, API keys in HTTP client code and personal user IDs in AUTHORITY-CHECK workarounds. The method in secrets removal for code datasets applies to ABAP source and transport texts alike.

ABAP dataset request specification

A short, explicit specification gets better supplier answers than "ABAP code." Adapt the template below, and see how to specify a code dataset request for the general fields.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample requirement
Object scopeCustomer namespace only (Z*, Y*, registered /NAMESPACE/); exclude SAP standard, modifications and third-party namespaces
Object typesPROG, CLAS, INTF, FUGR, ENHO, DDLS, BDEF, TABL/DTEL/DOMA definitions
Release coverageSource release (e.g., ECC 6.0 EHP7) and target S/4HANA release stated per object
HistoryAll repository versions with timestamps; transport ID, description and owner role per change
Remediation pairsBefore code, ATC finding (check ID, message, simplification item), after code, re-check result
TestsABAP Unit classes and last run result where present
ExcludedTable contents, change documents, credentials, personal user names (pseudonymized)
FormatOne JSON Lines record per object version, UTF-8, plus a manifest with object counts by type
Rights evidenceOwnership attestation per namespace; integrator IP terms where code was outsourced

Before signing, run the checks in evaluating a code dataset sample: does a random sample of objects compile against a stated release, are includes complete, and do the remediation pairs reproduce their findings?

How SourceX approaches ABAP requests

SourceX sources operational datasets from US companies, including engineering records, and manages the commercial process through licensing and ongoing purchases. Data is sourced on request rather than held in stock, so an ABAP request does not guarantee a match; buyers describe the data they need, and SourceX looks for US businesses that hold it. Every dataset is rights-reviewed for ownership and consents, every release is approved by the supplying company, and delivery runs through private, access-controlled workflows only after an executed agreement. If you are scoping an SAP migration model, you can describe the ABAP data you need to SourceX. Broader options for licensed repositories are on the proprietary codebases page and the software buyers page, and the code data buyer's map and AI data hub cover adjacent sources such as COBOL and mainframe code.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request licensed ABAP custom code for your model

SourceX works with AI teams wherever they are based and follows a Find, Assess, Agree, Transact and Manage process; nothing is contracted until a supplier agrees. Personal details are removed or replaced before delivery, the method is recorded, and a sample is checked. Tell SourceX what ABAP data you need.

Sources

  1. alphaXiv / arXiv, "ABAP code generation benchmark for large language models (arXiv 2601.15188)" (2026). https://www.alphaxiv.org/abs/2601.15188
  2. Hugging Face (timkoehne), "LLM-ABAP-Code-Generation-Benchmark dataset". https://huggingface.co/datasets/timkoehne/LLM-ABAP-Code-Generation-Benchmark/blob/main/dataset.jsonl
  3. arXiv, "Large Language Models for Software Engineering: A Systematic Literature Review" (2023). https://arxiv.org/pdf/2308.10620
  4. ASE 2024 (researchr), "What Makes a High-Quality Training Dataset for Large Language Models: A Practitioners' Perspective" (2024). https://conf.researchr.org/details/ase-2024/ase-2024-research/53/What-Makes-a-High-Quality-Training-Dataset-for-Large-Language-Models-A-Practitioners
  5. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data