Skip to content

Privacy and preparation

Dataset watermarking and canary records: tracing a leak to its source

By SourceX Editorial · Updated

Short answer

Canary records and dataset watermarks trace a leak by giving each recipient a uniquely marked copy: planted fictional entries, distinct placeholders or contact points that only that copy contains. If a copy surfaces, the markers show which delivery it came from. They are evidence for enforcing a license, not protection against misuse, and only work with a documented registry.

Key takeaways

  • Give each recipient a different set of markers, or the markers cannot tell copies apart.
  • Combining canary records with per-recipient pseudonym keys makes a stronger fingerprint.
  • Filtering, deduplication or rewriting can strip markers, so contract terms carry the enforcement.
  • Canary records must use identities and contact points the supplier controls, never real people.
  • A marker registry ties each delivery to its markers, files, keys and recipient.

How do canary records and dataset watermarks work?#

Canary records work by planting a small number of fictional entries in a dataset, each set unique to one recipient. A canary support ticket might describe a made-up customer at an email domain the supplier controls, with a phone number that rings a monitored line. If that ticket turns up in a leaked file, or that address starts receiving mail, the supplier knows whose copy escaped.

Dataset watermarking is the broader idea: make each delivered copy differ in ways that do not change its usefulness, and record the differences. Canary records are one method. Others change how existing data is encoded, such as which placeholder each person receives, rather than adding records.

Marker methods compared#

No single method survives every kind of handling, which is why careful suppliers combine two or three. The comparison shows what each one withstands and what defeats it.

Marker methods compared
MethodHow it marks a copySurvivesBreaks when
Canary recordsUnique fictional entries per recipientCopying and sampling that keep themUnusual records are filtered or deduplicated out
Honeytoken contact pointsEmails, phone numbers or URLs that alert when usedUse of the data outside the datasetNo one ever contacts them
Per-recipient pseudonym keysThe same person gets different placeholders in each copyMost edits that keep placeholdersPlaceholders are re-mapped
Record selection and orderEach copy holds a slightly different subset or orderBulk copyingData is shuffled or resampled
File hashesA fingerprint of each delivered fileExact copiesAny edit at all

Designing canary records that do not harm the data#

A good canary record is indistinguishable from real records to a casual reader and harmless to a model that learns from it. The rules below keep both properties.

Disclosure matters as much as design. Buyers care about training quality, and synthetic records hidden in a dataset can look like a breach of trust if found without warning. A clause stating that deliveries may contain recipient-specific markers keeps the practice transparent without revealing which records they are. Check the license's data warranties with counsel too: if the supplier promises that every record is a genuine business record, the marker clause has to carve canaries out of that promise.

  • Match the record type: a canary nonconformance report should look like a real one, with plausible parts, causes and outcomes.
  • Use only identities and contact points you control: a domain you own, phone numbers you hold, and names checked against your customer and employee tables.
  • Keep canaries few relative to the dataset so they do not distort what a model learns.
  • Make each recipient's set unique and record it before delivery.
  • Avoid anything harmful if a model reproduced it: no false safety instructions, no defamatory statements, no real addresses.
  • State in the license that markers may be present, without identifying them.

When are markers worth the effort?#

Markers are worth the effort when the same records go to more than one place. Non-exclusive licenses to several developers, recipients that rely on contractors or cloud partners, and long license terms all raise the odds that a copy turns up where it should not, and all make it harder to say which copy it was without markers.

They add less when there is one exclusive recipient with strong audit rights, or when the delivery is aggregated statistics with no record-level detail. In those cases a file hash and a clean chain of custody record may be enough.

Plan monitoring before you plant anything. Honeytoken inboxes and phone lines need an owner who checks them and knows what an alert means, or the most sensitive part of the method goes unwatched.

What markers cannot do#

Markers cannot prevent a leak or misuse. They help after the fact, by showing which copy left its permitted environment. A determined recipient can strip many markers by filtering unusual records, deduplicating, regenerating identifiers or paraphrasing text.

Markers also say little about model training. A model trained on a dataset may never reproduce a specific canary, so the absence of a canary in model output proves nothing. For questions about training use, rely on contract terms, audit rights and deletion certificates rather than detection.

A marker match is evidence, not a verdict. Two recipients may share records if your process reused them, and a contractor working for the recipient may be the actual source. The registry has to be precise enough to support the conversation that follows.

How to document markers: the registry#

The marker registry is what turns a suspicious file into an answer. Without it, even well-designed canaries are just odd records nobody can match to a delivery.

Store the registry with the supplier, apart from the data, and limit access to a small group. It belongs in the same controlled place as any pseudonym keys, and it should feed the chain of custody record for each delivery.

The registry can sit alongside openly verifiable integrity data. C2PA's AI/ML guidance describes a Training Data Set Content Credential, in which a collection data hash assertion can describe each folder of a training dataset. A signed record like that shows what was delivered and that it is intact; the private registry shows whose copy it was.

How to document markers: the registry
Registry fieldWhy it is needed
Recipient and license referenceTies markers to the obligations that apply
Delivery ID and dateSeparates repeat deliveries to the same recipient
File list with hashesProves exactly which files were sent
Canary record IDs and contentsLets you search a leaked file quickly
Pseudonym key IDLinks a placeholder pattern to one copy
Honeytoken contact points and monitoring ownerMakes sure an alert reaches a person
Registry access listKeeps the markers themselves secret

Illustrative: a manufacturer licensing quality records twice#

Illustrative: a fictional precision parts manufacturer licensed de-identified nonconformance reports and corrective actions from its QMS to two AI developers on non-exclusive terms. Its CTO wanted to be able to tell the two copies apart.

Each delivery used its own pseudonym key for customer and supplier placeholders and held a small set of canary nonconformance reports describing fictional parts, with contact details on a domain the company owned. The registry recorded files, hashes, keys and canaries per delivery, and both licenses stated that recipient-specific markers might be present. When a sample of quality records later appeared in a third party's product demo, the placeholders matched one delivery's key, and the company raised the matter under that license's audit clause.

Recipient markers in SourceX deliveries#

Recipient markers are an optional Preparation choice in SourceX deliveries, made by the supplier rather than required. Where they are used, the marker method and the license disclosure are settled before the supplier's approval, so the buyer signs terms that already mention them.

The SourceX Evidence Packet records the marker method and the registry location alongside provenance, permitted use and release authorization, so each delivery can be traced back to its terms. The registry contents stay with the supplier.

Frequently asked questions

Do canary records lower the value of a dataset?

Not when they are few, realistic and disclosed. A small set of extra records rarely changes what a model learns. They lower value when they are numerous, implausible or hidden, because a buyer who finds them may start to question the rest of the data.

Should each delivery to the same buyer get new markers?

Yes. New markers per delivery show which shipment leaked and keep a buyer's old and new copies distinguishable. Record each set separately in the registry, with its own delivery ID.

Should evaluation samples carry markers too?

Yes, if samples go to more than one prospective buyer. Samples are shared early, often with less formal terms than a full license, so a light set of canaries and a per-recipient key make it possible to tell whose sample was passed along. Record them in the same registry.

Can a buyer remove the canaries?

A buyer that finds them can filter them, and a license cannot stop that technically. That is why the contract matters: work with counsel on terms that prohibit removal of markers and require notice of any unauthorized disclosure, so tampering itself becomes a breach.

Are recipient markers the same as C2PA content credentials?

No. A C2PA content credential is a signed, cryptographically bound record of where an asset came from, built to be read and checked openly. Recipient markers are meant to go unnoticed and to identify one copy. The two can coexist in a well-documented delivery.

Do markers help if data leaks from our own systems?

Only for delivered copies. A leak from your own storage would carry no recipient markers, which is itself informative. Internal access logs and security controls are what protect the source.

Sources

  • The C2PA Explainer defines a Content Credential, also called a C2PA Manifest, as a cryptographically bound structure that records an asset's provenance. Source
  • C2PA's AI/ML guidance describes a Training Data Set Content Credential for AI-ML training datasets and notes that a collection data hash assertion can be used to describe each folder of a training data set. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify