Text and language data
Terminology Databases and Termbases as Training Data for LLM Translation
Quick answer
A licensed terminology database gives an LLM translation system concept-level ground truth: one concept ID, approved terms per language, a definition, and a status that marks terms as preferred, admitted or forbidden. That is different from sentence pairs. Buyers use termbases to build terminology-constrained SFT data, to supply reward signals for terminology-aware RL, to ground prompts at inference time, and to score term accuracy. What matters is entry-level validation, an explicit status field, and clear ownership of client-specific glossaries.
By SourceX Editorial · Updated
Why termbases still add value when LLMs translate fluently
Term accuracy is the weak point of otherwise fluent LLM translation, which is why a curated termbase can measurably change output quality. A 2025 multi-domain evaluation reports terminology accuracy of roughly 52–53% for Chinese to English and 37–39% for German to English on the WMT23 terminology task [2]. A 2026 study of LLMs for specialised terminology concludes they can help specialised translators but cannot yet replace specialised terminology resources and corpora [3].
The practical reading is that fluency and terminology are separate problems. A model can produce a natural sentence that still uses an admitted synonym where the client mandates the preferred term, or a forbidden legacy product name. Only a termbase with status data can tell you which error happened, so it doubles as an evaluation asset.
How termbase data differs from translation memories and parallel corpora
A termbase is concept-oriented: each record describes one concept and the terms that designate it in each language, not an aligned sentence. Translation memories (TMX) and parallel corpora carry segment pairs with context but no authoritative statement of which term is correct. Our comparison of multilingual training data types sets out where each fits, and the translation memory licensing page covers segment-level data.
Termbases are small by row count but dense in decisions. A 20,000-concept engineering termbase may encode years of terminologist review, including forbidden terms that never appear in any corpus because they were edited out. That negative evidence is hard to reconstruct from parallel corpora or from post-editing data, which records corrections but rarely the rule behind them.
What a usable termbase entry contains
A termbase is training-grade when every concept carries an ID, language-tagged terms, a status per term, a domain, a source and a review date. TBX, standardized as ISO 30042:2019, defines a concept-entry, language-section and term-section metamodel with data categories such as administrative status, part of speech and definition, serialized in DCA or DCT XML styles. Many enterprise termbases live in tools such as Trados MultiTerm (formerly SDL MultiTerm), memoQ or Excel exports rather than clean TBX, so ask what the native format is and what is lost on export.
The illustrative record below shows the fields worth requesting, flattened to JSON Lines for pipeline use.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"concept_id": "C-004817",
"domain": "industrial hydraulics",
"subject_field_code": "HYD-VALVE",
"definition_en": "Valve that limits system pressure by diverting flow when a set pressure is reached.",
"definition_source": "internal engineering standard, rev. 4",
"terms": [
{"lang": "en-US", "term": "pressure relief valve", "status": "preferred", "pos": "noun"},
{"lang": "en-US", "term": "blow-off valve", "status": "forbidden", "note": "legacy marketing term"},
{"lang": "de-DE", "term": "Druckbegrenzungsventil", "status": "preferred", "gender": "n"},
{"lang": "de-DE", "term": "Überdruckventil", "status": "admitted"},
{"lang": "ja-JP", "term": "リリーフ弁", "status": "preferred"}],
"usage_note": "Do not translate as safety valve unless the part is certified to the safety class.",
"created_by_role": "terminologist", "last_reviewed": "2025-11-03",
"owner": "supplier", "client_specific": false}
Common failure modes to screen for: status fields left at a tool default so everything reads as preferred, definitions copied from public dictionaries with unclear rights, duplicate concepts split across languages, and stale entries with no review date.
TBX versus simpler serializations in prompts and training
Keep TBX as the archival and exchange format, but expect to flatten it for LLM use. Lackner and colleagues tested TBX and alternative structured formats as prompt input and found TBX verbose, with formatting strategy modestly affecting term accuracy; terminology-augmented generation with capable LLMs performed competitively in their setup [1]. Verbose XML also burns context tokens when you inject dozens of entries per segment.
A workable pipeline keeps the TBX or tool export untouched for provenance, converts it to JSONL or a compact "source term → target term (status)" list, and records the conversion script version. Ask suppliers for the original export, not only a flattened CSV, because CSV exports often drop status, notes and per-language attributes.
Training uses: SFT, terminology-aware RL, grounding and evaluation
Termbases support four distinct uses, and the license should name each one you plan. For SFT, buyers pair termbase entries with source segments containing the concept to build terminology-constrained translation examples. For RL, termbase status supplies a checkable reward: TAT-R1 trains terminology-aware translation with reinforcement learning, using word alignment to reward correct term translation [4].
For grounding, entries are retrieved per segment and placed in the prompt, which keeps terminology current without retraining; if you do this, plan how entries leave your index when a license ends (see removing licensed content from vector indexes). For evaluation, hold out concepts, not just segments, so the model is scored on terms it never saw in training.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use | Fields you need | Main risk if missing |
|---|---|---|
| Terminology-constrained SFT | preferred term per language, domain, example context | Model learns admitted synonyms as targets |
| Terminology-aware RL reward | status incl. forbidden, inflection or POS data | Reward misfires on inflected forms |
| Prompt grounding / RAG | compact serialization, concept ID, review date | Stale terms injected at inference |
| Held-out term evaluation | concept-level split, definitions | Train/test leakage across languages |
Ownership and licensing checks specific to termbases
Confirm who owns each termbase before anything else, because client-specific glossaries built by a language service provider may belong to the end client rather than the provider. Under US copyright law, copyright vests initially in the author, the employer or other person for whom a work made for hire was prepared is considered the author, and ownership can be transferred [5]. Localization contracts often assign deliverables, including glossaries, to the client.
Ask for a per-termbase ownership statement (supplier-built, client-assigned, or derived from a public standard), the client contract clause where relevant, and the source of definitions. If a termbase is EU-hosted or EU-built, also review the EU database right, since a curated termbase may be protected as a database independent of any copyright in individual entries. Personal data is usually minimal, but creator and approver name fields should be removed or replaced.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How SourceX handles termbase requests
SourceX sources operational datasets from US companies on request; it does not hold termbases in stock, and a request does not guarantee a match. Buyers describe the data, such as languages, domains, entry fields and status coverage, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. You can start a termbase request as a buyer.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names and emails are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. For context on what localization data supports, see AI use cases for translation and localization agency data and the text data hub.
Sourcing a terminology database for LLM training
If your translation model needs validated, multilingual term entries with status and definitions, describe the languages, domains and fields you need, and SourceX will look for US companies that hold them and manage the licensing. Pricing and allowed uses are agreed per deal in a license. Describe your termbase data needs to SourceX.
Sources
- Padova University Press, "Lackner et al. 2025 (JDTL)" (2025). https://jdtl.padovauniversitypress.it/system/files/papers/jdtl_lackner_et_al_2025.pdf
- arXiv, "How Well Do Large Reasoning Models Translate? A Comprehensive Evaluation for Multi-Domain Machine Translation" (2025). https://arxiv.org/pdf/2505.19987
- arXiv, "On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?" (2026). https://arxiv.org/pdf/2607.24784
- arXiv, "TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment" (2025). https://arxiv.org/pdf/2505.21172
- U.S. Government Publishing Office (govinfo), "17 U.S.C. 201 - Ownership of copyright" (2024). https://www.govinfo.gov/content/pkg/USCODE-2024-title17/html/USCODE-2024-title17-chap2-sec201.htm
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.