Skip to content

For AI teams

License translation memories for AI training

Short answer

SourceX sources translation memories from U.S. businesses that own them. Each dataset is rights-reviewed, prepared with client names, confidential source text and personal data in segments removed, and approved by the supplying company before delivery under a use-limited license. Teams use it for machine translation, localization and multilingual models.

What's included#

Typically aligned source and target segments with edits.

Why it matters for models#

These records are valuable because professionally translated pairs are high-quality language data.

Common uses#

  • Machine translation
  • Localization
  • Multilingual models

Removed before delivery#

  • Client names
  • Confidential source text
  • Personal data in segments

How a request works#

  • Describe the records, volume and uses you need.
  • SourceX matches suitable U.S. suppliers and runs a rights review.
  • A prepared sample is shared under NDA after supplier approval.
  • The license sets allowed uses, term and deletion terms; resale and re-identification are barred.

Frequently asked questions

Who supplies the data?

U.S. businesses that own the records. Every release is approved by the supplying company.

Can we see a sample?

Yes, once an NDA is in place and the supplier approves a prepared sample.

Is pricing published?

No. Terms depend on volume, history, exclusivity and allowed uses, and are agreed per deal.

See if you qualify