Skip to content

Text and language data

Translation Post-Editing and Quality-Annotation Data for MT and LLM Translation

Quick answer

Machine translation post-editing data is the record of a human translator correcting an MT draft: the source segment, the raw MT output (with engine and version), the post-edited segment, and ideally effort signals and MQM-style error labels. Unlike a translation memory, which keeps only the approved target, a triplet preserves what the machine got wrong and how a professional fixed it. That delta is the training signal for automatic post-editing (APE), quality estimation (QE), translation reward models and evaluation sets.

By SourceX Editorial · Updated

Why post-edit triplets carry a signal translation memories lack

Post-edit triplets are valuable because the correction itself is the label. A TMX export from a localization tool keeps the final approved target, which is fine for supervised translation but tells a model nothing about where an MT engine fails (see licensing translation memories for that adjacent asset). APE research benchmarks use the same shape: (source, MT, post-edit) triplets, with systems scored by HTER (human-targeted translation edit rate) against the human post-edit and compared with a do-nothing baseline that leaves the MT unchanged. Buying from production localization workflows gives you that shape in your own domains and language pairs rather than a benchmark's.

Strong engines still fail in predictable places, which is why fresh corrections stay useful. A 2025 multi-domain evaluation of large reasoning models found terminology and domain-specific errors persist even in these systems [1]. Localization post-edits concentrate on exactly these errors: client glossary violations, domain senses, untranslatable product names, number and unit formats, and register. For how post-edits compare with native text, parallel corpora and termbases, see multilingual training data types compared.

Which model tasks each record type supports

Each downstream task needs a different subset of fields, so specify the task before you specify the data. QE needs labels on MT output without a reference; APE needs the full triplet; reward and preference training needs a ranked pair; evaluation needs error spans with severities from qualified raters.

Downstream taskMinimum recordLabel or targetCommon failure in purchased data
Automatic post-editingsource, MT, post-editpost-edit as target; HTER for rankingMT column regenerated later, not the draft the editor saw
Sentence-level QEsource, MT, effort signalHTER, edit time, keystrokesTime logs include breaks or review rounds
Word-level QEsource, MT, post-editOK/BAD tags from TER alignmentTokenization differs from the alignment used to derive tags
Preference / reward trainingMT draft vs post-edit, or two MT draftschosen vs rejected pair; DPO-style training [3]Light post-editing leaves "chosen" barely better than "rejected"
Terminology-aware RLsource, MT, post-edit, term listalignment-based term reward [2]Glossary version not recorded per segment
Human evaluation setsource, MT, error spansMQM category and severityAnnotators lacked document context

Effort labels are noisier than they look. Sentence-level QE is commonly trained on HTER, post-editing time and keystrokes, and word-level OK/BAD tags are usually derived by aligning MT output to its post-edit with TER, so tokenization and how shifts are handled change the labels. Ask the supplier which tool computed each effort field and with what settings. If a supplier's CAT tool only exports final segments, you can compute HTER yourself, but you cannot recover time or keystrokes after the fact.

MQM error annotations versus edit distance

MQM annotations tell you why a segment was wrong, while edit distance only tells you how much changed. Multidimensional Quality Metrics (MQM) has professional raters mark error spans with a category and a severity, which separates a meaning-changing mistranslation from a comma. For buyers this matters in two ways: MQM spans (accuracy/mistranslation, terminology, fluency, style, locale convention) with minor/major/critical severity make far better eval and reward data than raw HTER, and MQM collected without document context will understate discourse and consistency errors.

Edit distance also hides stylistic preference edits. A reviewer who rewrites a correct sentence to match house style inflates HTER without the MT having made an error. Ask whether the supplier's workflow separates "correction" from "preferential" edits, or whether a review step (LQA) logged MQM categories on a sample. Use the annotation quality audit method to check agreement on a double-annotated subset before you accept a delivery.

Record schema to request from localization suppliers

Ask for segment-level records with workflow metadata, not a flattened bilingual file. The schema below is a request template you can adapt; fields marked optional raise value but are often missing from CAT-tool exports such as XLIFF 2.x or TMX.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "segment_id": "proj-0412/doc-07/seg-0193",
  "doc_id": "doc-07",
  "segment_index": 193,
  "source_lang": "en-US",
  "target_lang": "de-DE",
  "domain": "medical-device IFU",
  "source": "Replace the filter every 30 days.",
  "mt_output": "Ersetzen Sie den Filter alle 30 Tage.",
  "mt_engine": "vendor-NMT",
  "mt_engine_version": "2025-11 custom model",
  "mt_settings": {"glossary_applied": true, "glossary_version": "v14"},
  "post_edit": "Tauschen Sie den Filter alle 30 Tage aus.",
  "editor_id_hash": "a91f...",
  "editor_tier": "professional, domain-qualified",
  "edit_seconds": 14.2,
  "keystrokes": 31,
  "hter": 0.25,
  "edit_type": "terminology",
  "mqm_errors": [{"span": [0, 8], "category": "terminology", "severity": "minor"}],
  "review_round": "post-edit + LQA sample",
  "ai_assistance_disclosed": "none beyond MT draft",
  "client_consent_basis": "client authorized AI-training use",
  "deid_method": "names, emails, phones, account numbers replaced with typed placeholders"
}

Keep doc_id and segment_index so you can rebuild document context for MQM and for splits. Split train and test by document or by client project, never by segment, or near-duplicate segments from the same manual will leak across folds.

Localization content almost always belongs to the agency's end client, so the agency alone usually cannot license it for AI training. Ask who owns the source text, whether the client authorized training use, and whether the MT engine's terms of service restrict reuse of its outputs, because the MT column is a third-party system's output. Large text-dataset audits have found license information frequently missing or wrong on hosting sites [4], so require segment- or project-level provenance rather than a single dataset-wide statement; human annotation provenance covers annotator agreements and AI-assistance disclosure.

De-identification has to handle both columns. Customer names, emails, phone numbers and account numbers appear in the source and are often copied into the MT and post-edit, so masking one column but not the others leaves the data exposed. Tools such as Presidio help detect PII, but the project itself warns detection is incomplete [5]; see how to de-identify translation memories for a field-level approach. For parallel-data licensing pitfalls more broadly, read parallel corpora for commercial translation models.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Acceptance checks for a post-editing data sample

Accept a delivery only after checking that the MT column is the original draft and the effort labels mean what they claim. Commercial human-translation data is an established product [6], but triplets with intact drafts and effort logs are less standardized, so expect variation.

  • Draft integrity: confirm mt_output was stored at edit time; if HTER is near zero on most segments, the column may be the post-edit copied over.
  • Engine traceability: each segment names engine and version; mixed engines without labels make QE models learn engine identity.
  • Edit mix: sample 200 segments and classify edits as correction versus preference.
  • Effort plausibility: cap or flag edit_seconds outliers; check that keystrokes exist for segments with non-zero HTER.
  • Terminology coverage: match edits against the client glossary and record glossary version; compare with your termbase data needs.
  • Context: confirm document order is preserved for MQM and discourse evaluation.
  • PII: run detection across all three text columns, then spot-check by hand.

How SourceX helps buyers source post-editing data

SourceX sources operational datasets, including documents and workflow records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced with the method recorded and a sample checked, and every release is approved by the supplying company. Describe the language pairs, domains and fields you need on the SourceX buyer page, or start from the text and language data hub.

Request MT post-editing and MQM data

Tell SourceX the triplet fields, effort signals and error annotations your APE, QE or reward model needs. SourceX assesses data and licensing permissions with suppliers and agrees allowed uses in a license before anything is delivered. Describe your translation data request.

Sources

  1. arXiv, "How Well Do Large Reasoning Models Translate? A Comprehensive Evaluation for Multi-Domain Machine Translation" (2025). https://arxiv.org/pdf/2505.19987
  2. arXiv, "TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment" (2025). https://arxiv.org/pdf/2505.21172
  3. Rafailov et al., Stanford, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  4. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  5. Microsoft (microsoft/presidio), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  6. TAUS, "Data for AI". https://www.taus.net/data-solutions/data-for-ai

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data