Skip to content

Question

What is model memorization?

By SourceX Editorial · Updated

Short answer

Memorization is when a model stores specific training examples closely enough to reproduce them. It is more likely for text that appears many times or is unusual. It is why sensitive details should be removed before data is used for training.

Part of: How licensing works

Explanation#

Labs use deduplication and filtering to reduce it.

Contracts can prohibit attempts to extract training data.

Rights and privacy#

  • Confirm customer, vendor and employee agreements allow the use
  • Remove or de-identify personal details before anything leaves the company
  • Approve the final dataset and permitted uses in writing

How licensing works#

A business describes its data, buyer demand is assessed, scope and permitted use are proposed, the data is prepared and de-identified, an agreement is signed, and the approved dataset is delivered. See data licensing for detail.

Related questions

Sources

  1. Carlini et al., Extracting Training Data from Large Language Models (2021)

General information, not legal advice. Editorial policy.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify