Skip to content

Text and language data

Parallel Corpora for Commercial Translation Models: Licensing Pitfalls

Quick answer

A parallel corpus is only usable for a commercial translation model when its license explicitly grants commercial training and says who may use the resulting model. Many well-known bitext releases are research-only, and some extend that restriction to derived data and to any translator trained on them [1]. Before you buy bilingual sentence pairs, confirm the grant in a signed contract rather than a dataset card, check upstream sources such as patents or web crawls, and price against the licensee tier that actually applies to you.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why research-only bitext reaches into your model

Research-only bitext terms frequently restrict the model, not just the files. NTT's JParaCrawl, one of the largest public English-Japanese corpora, is released for research, and its terms exclude commercial use of derived data and of translators trained on the corpus; commercial use requires a separate arrangement with NTT [1]. A team that fine-tunes an NLLB-style or LLM translator on such data and then ships it inside a product has a problem that deleting the corpus afterward does not fix.

The same pattern shows up when bitext travels through aggregators. Aggregators such as OPUS redistribute many corpora with citation requests, but the originating provider's own terms still govern what you can do with the data. Treat an aggregator download as a pointer to the original license, never as the license itself. For the general escalation route, see getting commercial rights for research-only datasets.

Where parallel corpora come from, and the rights each origin carries

Each bitext origin carries a different upstream rights question, so classify the corpus by origin before you read the price. The categories below are the ones procurement teams meet most often; the comparison of multilingual data types covers how they differ for training.

  • Web-mined bitext. ParaCrawl-style pipelines crawl multilingual websites and align pages with tools such as Bitextor. The compiler may release its alignments under permissive terms, but the underlying sentences remain the website owners' text, so the copyright and opt-out posture is yours to assess.
  • Patent-derived bitext. Corpora built from parallel patent publications, such as Chinese-English patent sentence collections, carry distribution terms from the compiler and sometimes the patent office feed [2][8]. Check both layers; the patent full-text corpus guide explains the source data.
  • Institutional and legislative text. Parliamentary and treaty translations are often reusable, but reuse terms vary by institution and some exclude bulk or commercial reuse.
  • Translation memories and post-edits. Enterprise TMX exports and post-edited MT output usually belong to the client whose content was translated, not only the LSP that holds the files. Translation memories have their own owner page: license translation memories for AI training. Quality-annotated pairs are covered in post-editing and quality data.
  • Vendor-commissioned bitext. Commercial vendors produce or curate pairs for sale; the commercial grant is usually clear, but verify what the vendor itself is entitled to sublicense.

Dataset cards are not licenses

A license field on a dataset card is metadata, and it can be wrong or contradict another version of the same data. On the Hugging Face Hub, the card is the repository README, and its YAML block records license, language and size to drive filtering [6]. One vendor's English-Russian bitext card lists a non-commercial CC BY-NC-SA license while another version of the same dataset is tagged commercial [4]. A Japanese-English listing tagged with a commercial license turns out to be a preview sample of a paid dataset sold separately [5].

This is not an edge case. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular hosting sites [7]. Put the governing terms, version and record scope in the executed agreement, and attach the card only as a description.

Derived-model clauses to settle before signature

Derived-model language decides whether the money you spend on bitext produces a deployable translator. Ask counsel to resolve each of these in the contract, not in an email thread.

Illustrative example: invented to show structure; it does not describe an available dataset.

ClauseWhat to pin downFailure mode if silent
Permitted useTraining, fine-tuning, evaluation, synthetic back-translation, distillation into a smaller modelLicensor later reads "research and development" as excluding production
Derived dataStatus of filtered subsets, back-translations, synthetic pairs generated by your modelSynthetic data inherits the original restriction, as in JParaCrawl-style terms [1]
Model ownershipWho owns weights trained on the bitext; whether outputs are restrictedProduct team cannot ship or open-weight the model
Hosting and customersWhether the model can be served via API or delivered to customers on-premisesDeal covers internal use only
Language-pair and domain scopeExact pairs and directions (en-ja vs ja-en), domains, record countsReverse-direction training is disputed
Term and terminationEffect of expiry on trained models, not just on filesRetraining obligation after term ends
Upstream warrantyLicensor's statement on source rights, including web or patent originsYou inherit claims from upstream rights holders
DisclosurePermission to name the dataset in training-data summariesConflict with confidentiality when you publish documentation

Disclosure deserves attention in 2026. EU AI Act Article 53 has required general-purpose model providers to keep a copyright compliance policy, including honoring Article 4(3) reservations of rights, since 2 August 2025, with AI Office enforcement for new models from 2 August 2026 [10]. California AB 2013 required developers of generative AI systems offered to Californians to post documentation of their training data, with disclosures due 1 January 2026 [11]. Your license must let you describe the bitext at the level those regimes require.

Jurisdiction does not rescue a research license. In the UK, the CDPA s29A text and data analysis exception remains limited to non-commercial research as of October 2026 [12].

How bitext is priced, and what "price per sentence pair" hides

Bitext prices vary mostly by licensee type, volume unit and rights scope, so a quoted per-pair or per-word rate is meaningless until those three are fixed. The Linguistic Data Consortium lists its Chinese-English patent sentence corpus at US$25 for non-profits and US$5,000 for for-profit organizations, or US$4,000 for members [2]. LDC's for-profit membership agreement is the gateway to commercial use, and non-members generally cannot use LDC data in commercial products [3]. Commercial vendors tend to price by word volume and run time-limited promotions, so a historical rate card is a poor benchmark [13].

When comparing quotes, normalize them to cost per deduplicated, filtered pair in the exact language direction you need. Raw web-mined bitext can lose a large share of pairs to language-ID, length-ratio and alignment-score filters, which changes the effective price. Ask whether the quote covers updates, additional pairs or new language pairs, and whether derived-model rights cost extra.

Provenance and alignment metadata to require with delivery

Alignment metadata is what lets you filter, audit and later prove where each pair came from. Request the fields below as delivery conditions, and register the resource and version with a persistent identifier such as an ISLRN so licenses map to exact releases [9]. The broader field list is in metadata fields to require with licensed text corpora.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "enja-000184223",
  "src_lang": "en", "tgt_lang": "ja",
  "src_text": "Replace the filter cartridge every 500 hours.",
  "tgt_text": "フィルターカートリッジは500時間ごとに交換してください。",
  "source_doc_id_src": "manual-7731-en-r4",
  "source_doc_id_tgt": "manual-7731-ja-r4",
  "origin": "supplier_owned_documentation",
  "translation_direction": "en->ja",
  "tgt_is_machine_translated": false,
  "post_edited": true,
  "alignment_method": "sentence_aligner_v2 + human_review",
  "alignment_score": 0.91,
  "license_ref": "agreement-2026-014, schedule B",
  "resource_pid": "ISLRN placeholder",
  "release_version": "1.2.0"
}

Two fields matter most. tgt_is_machine_translated flags translationese and MT output that can teach a model the errors of an earlier system, and translation_direction tells you which side is the original, which affects quality for the reverse direction. Treat the presence of these fields in a vendor sample as a hypothesis to test, not a given.

A pre-purchase checklist for bitext buyers

Run these checks on every parallel corpus before procurement signs. They take hours, not weeks, and they catch most of the problems described above.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Identify the origin of each side: web crawl, patent, institutional, TM, post-edit, vendor-produced.
  2. Read the provider's own terms, not the aggregator or card; record the URL and date.
  3. Confirm the grant covers commercial training, derived data and deployed models in writing.
  4. Confirm the licensor can sublicense upstream content, especially for web and patent text.
  5. Pull a sample of 500 to 1,000 pairs and measure language-ID accuracy, duplicate rate, length ratio and share of MT output.
  6. Check for personal data in support, legal or HR-derived bitext and agree on how it is removed.
  7. Normalize price to filtered pairs per direction; list which rights are priced separately.
  8. Register the release with a persistent identifier and store the license with the dataset in your training corpus provenance audit.

Where company-held bitext fits

Company-held multilingual documents, such as localized manuals, contracts and support content, can supply domain bitext that public releases rarely contain, because they are not on the open web. SourceX sources operational datasets, including documents, from US companies on request and manages the licensing process; it does not source scraped web content and does not hold bitext in stock, so a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, released only with the supplying company's approval, and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Buyers can describe the bitext they need on the SourceX buyers page.

If you hold translation memories rather than buy them, the guide for companies selling translation memories covers that side. For pre-training text beyond bitext, see licensing non-English text corpora and the text and language data hub, or browse all categories from the AI data hub.

License parallel corpora with explicit commercial and model rights

SourceX looks for US businesses that hold the bilingual or multilingual data you describe, assesses the data and its licensing permissions, and agrees pricing and allowed uses in a license; nothing is contracted until a supplier agrees. Prices are not published and terms are agreed per deal. Describe the bitext you need to license.

Sources

  1. NTT Communication Science Laboratories, "JParaCrawl". https://www.rd.ntt/cs/team_project/icl/lirg/jparacrawl/
  2. Linguistic Data Consortium, "Chinese-English Parallel Sentences Extracted from Patents (LDC2016T22)" (2016). https://catalog.ldc.upenn.edu/LDC2016T22
  3. Linguistic Data Consortium, "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
  4. Nexdata on Hugging Face, "English-Russian Parallel Corpus Data (dataset card README)". https://huggingface.co/datasets/Nexdata/English-Russian_Parallel_Corpus_Data/blob/main/README.md
  5. Nexdata on Hugging Face, "Japanese-English Parallel Corpus Data". https://huggingface.co/datasets/Nexdata/Japanese-English_Parallel_Corpus_Data
  6. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  7. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. ACL Anthology, "ParaPat: The Multi-Million Sentences Parallel Corpus of Patents Abstracts" (2020). https://aclanthology.org/2020.lrec-1.465/
  9. International Standard Language Resource Number (ISLRN), "ISLRN resource record 471-919-856-164-1". https://www.islrn.org/resources/471-919-856-164-1/
  10. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  11. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  12. UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 (consolidated), section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  13. TAUS, "TAUS data sale to boost multilingual LLMs". https://www.taus.net/resources/blog/taus-data-sale-to-boost-multilingual-llms

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data