Skip to content

Question

How do AI companies value datasets?

By SourceX Editorial · Updated

Short answer

AI companies value a dataset by how much it is expected to improve a specific model or product, relative to what it costs to obtain and prepare. Key factors are relevance to the target task, scarcity, quality and labelling, volume and recency, clarity of rights, exclusivity, and the privacy and cleaning work required. There is no public price list; valuations are negotiated deal by deal.

Part of: Data Value Index

Explanation

Buyers often test a sample first, measuring whether it improves performance on their evaluations.

A smaller dataset that covers a capability gap can be worth more than a much larger generic one.

Examples

  • Specialist medical coding decisions: scarce, high task relevance
  • Generic marketing copy: abundant, low incremental value

What determines value

  • Task relevance and measured model impact
  • Scarcity and uniqueness
  • Label quality and outcomes
  • Rights clarity and exclusivity
  • Preparation and privacy cost (reduces net value)

Rights and privacy

  • Buyers discount data with unclear provenance
  • Exclusivity usually commands a premium

How licensing works

A business describes its data, buyer demand is assessed, scope and permitted use are proposed, the data is prepared and de-identified, an agreement is signed, and the approved dataset is delivered. See data licensing for detail.

Related questions

Sources

  1. Ouyang et al., Training language models to follow instructions with human feedback (2022)

General information, not legal advice. Editorial policy.

Related resources

Explore a data partnership

Tell us what data your company holds. No data is shared during the initial assessment.

Estimate my data's value