Provenance, rights and permitted use
Buying Web-Derived Datasets: Provenance Records to Demand from Crawl Vendors
Quick answer
Web-scraped dataset provenance means proof of how, when and under what signals each page was collected. Before buying a crawl-derived corpus, require per-URL fetch timestamps, the exact user-agent string, robots.txt and terms-of-service snapshots captured at fetch time, the opt-out and TDM-reservation checks applied, and a takedown process that reaches copies already delivered. Without those records you cannot show that collection respected the signals in force when it happened, and site preferences shifted quickly in 2023–2024 [1].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why crawl-time evidence matters more than today's robots.txt
A domain's current robots.txt tells you almost nothing about whether a crawl made eighteen months ago was permitted. The Data Provenance Initiative's audit of about 14,000 web domains found that restrictions on AI crawling rose sharply within a single year across 2023–2024 [1], and mainstream coverage described the web as becoming markedly more hostile to crawlers [2]. A vendor that re-checks robots.txt on the day you ask is showing you the wrong moment.
The fix is temporal: every record in the corpus should resolve to the robots.txt body, HTTP status and fetch time that governed it. That is the only way to answer "was this allowed when it was taken?" for a specific URL. It also lets you decide what to do with content whose owner later opted out.
The same audit found that terms of service and robots.txt frequently disagree [1]. A site can allow a crawler in robots.txt while its terms prohibit scraping or AI training, or the reverse. Treat the two as separate signals with separate evidence.
The core crawl-time records, field by field
The minimum viable provenance record for web-derived data is per fetch, not per dataset. Dataset-level statements such as "we respect robots.txt" are marketing until backed by per-URL logs; vendors in this market publish provenance explainers as a selling point, and those claims are self-reported until you sample them [3]. For a broader framing of when row-level evidence is necessary, see record-level provenance.
Ask for these fields on every delivered document or WARC record:
- Canonical URL and final URL after redirects, plus the HTTP status code. Redirect chains often cross domains, and the governing robots.txt belongs to the final host.
- Fetch timestamp (UTC) from the WARC-Date header or equivalent, not the date the vendor packaged the shard.
- User-agent string and crawler identity, including whether the crawler identified itself with a token that site owners could target (for example CCBot or a vendor-named bot) or used a generic browser string. A generic string makes AI-specific disallow rules unenforceable and is a red flag.
- Robots.txt snapshot reference: a hash or archived copy of the robots.txt body, its HTTP status (a 404, 5xx or timeout each implies a different crawl decision), and the fetch time of that file.
- Robots.txt evaluation result: which group matched (named token or
*), and the allow/disallow decision for this path. - Terms-of-service snapshot reference for the domain, with the clause classification the vendor applied (no restriction, scraping prohibited, AI or TDM prohibited, licensing required).
- Page-level signals checked: meta robots and X-Robots-Tag values (
noindex,noai-style directives where present), TDMRep headers or/.well-known/tdmrep.json, and any machine-readable license file. - Access conditions: confirmation that the fetch did not pass a login, paywall or CAPTCHA. Common Crawl, for comparison, says it follows robots.txt, avoids paywalled and login-protected content and rate-limits its crawler [7].
- Processing lineage: extraction tool and version, language ID, dedup cluster ID, quality and toxicity filter scores, PII handling, and which filter removed or kept the record.
For a structured way to carry these fields, the MLCommons Croissant-RAI vocabulary provides machine-readable slots for responsible-AI and data life cycle metadata [11]. The Data Provenance Standards explainer covers industry metadata standards built on the same ideas.
Opt-out and rights-reservation handling to verify
Robots.txt was designed for crawl access, not for expressing training permissions, and standards bodies have convened specifically on its limits for AI use [8]. A buyer therefore needs to know which additional opt-out mechanisms the vendor checked and how it resolved conflicts. Ask for the written decision logic, not a summary.
EU text-and-data-mining reservations. Under Article 53(1)(c) of the EU AI Act, general-purpose AI model providers must maintain a copyright policy that identifies and complies with rights reservations expressed under Article 4(3) of Directive (EU) 2019/790, including through state-of-the-art technologies [4]. Those duties have applied since 2 August 2025, and AI Office enforcement for new models began 2 August 2026 (status as of October 2026). The Copyright chapter of the GPAI Code of Practice gives signatories a route to show compliance, including commitments about lawful access and honoring machine-readable reservations [5]. If your model will be placed on the EU market, the vendor's opt-out evidence becomes part of your own copyright policy file.
Machine-readable protocols beyond robots.txt. The W3C TDM Reservation Protocol lets a rightsholder declare a TDM reservation via HTTP header, HTML meta tag or a well-known JSON file, optionally pointing to a TDM policy [6]. Really Simple Licensing, launched in September 2025, lets publishers state crawl and licensing terms in machine-readable form, but it depends on crawler operators choosing to honor it [9]. Ask which of these the vendor parsed, from which date, and what it did on a match: exclude, quarantine pending license, or ignore.
Retroactive opt-outs. Decide in the contract what happens when a domain adds a restriction after the crawl. Some buyers accept historical snapshots that were permitted at fetch time; others require exclusion from future deliveries and refreshes. Either way, the vendor needs the timestamped evidence above to apply the rule consistently.
Takedowns, refreshes and propagation into delivered snapshots
A takedown process is only useful if it reaches every copy you already hold. Require the vendor to maintain a removal log keyed by URL, domain and content hash, with the request date, requester type and reason. Each refresh should ship with a delta manifest of removed record IDs so your pipeline can purge them from training shards and RAG indexes.
Failure modes to probe:
- Hash drift. The same article re-fetched with a changed ad slot produces a different hash, so removal by hash alone misses near-duplicates. Ask whether removals propagate across MinHash or similar dedup clusters.
- Mirror and syndication copies. Content removed at the origin often survives on aggregator or syndication domains. Ask how the vendor handles requests that name the work rather than the URL.
- Frozen derivatives. Tokenized shards, embeddings and filtered subsets built before the removal still contain the content. Your own provenance audit should trace these.
Benchmark contamination is a related exclusion problem. BIG-bench embeds a canary GUID in its task files precisely so builders of web-scraped corpora can filter benchmark data out [12]. Ask whether the vendor scans for known canary strings and evaluation sets and logs what it removed.
A provenance request template for crawl vendors
The most efficient way to compare web-data vendors is to send the same structured request and score the answers. The template below asks for evidence at three levels: crawler, domain and record.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Level | Record to request | Acceptable evidence | Red flag |
|---|---|---|---|
| Crawler | User-agent string(s) and published bot documentation | Named token, documented IP ranges, public opt-out instructions | Generic browser UA; rotating residential proxies |
| Crawler | Crawl window per shard | Start and end dates per WARC segment | Only a "collected 2023–2025" range |
| Domain | Robots.txt snapshot at fetch time | Archived body plus status code and timestamp, hash-linked to records | "We check robots.txt" with no archive |
| Domain | Terms-of-service snapshot and classification | Archived ToS with clause tags and reviewer notes | ToS not reviewed because robots.txt allowed |
| Domain | TDM and licensing signals parsed | TDMRep, RSL, meta and header checks with first-parse date | Signals checked only for new crawls |
| Record | Fetch metadata | URL, final URL, status, WARC-Date, UA, access path | Missing timestamps on extracted text |
| Record | Filter and dedup lineage | Tool versions, filter scores, cluster IDs | No way to map text back to source URL |
| Process | Takedown log and delta manifests | Per-refresh removal list with hashes and cluster IDs | Removals applied only to future crawls |
| Process | Sampling access | Right to pull 500+ random records with full lineage | Curated samples only |
An illustrative record that passes this spec might look like this:
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "shard-0412/rec-00918231",
"url": "https://example-news.test/2024/05/story",
"final_url": "https://www.example-news.test/2024/05/story",
"http_status": 200,
"warc_date": "2024-06-03T14:22:09Z",
"user_agent": "ExampleVendorBot/2.1 (+https://vendor.test/bot)",
"robots_snapshot": {"sha256": "9f2c...", "fetched": "2024-06-03T14:20:51Z", "status": 200, "matched_group": "*", "decision": "allow"},
"tos_snapshot": {"sha256": "a71e...", "classification": "no_ai_restriction", "reviewed": "2024-05-28"},
"tdm_signals": {"tdmrep": "absent", "meta_robots": "index,follow", "x_robots_tag": null},
"access_path": "public_no_login",
"dedup_cluster": "mh-55120943",
"removed_on": null
}
Then test it: pull a random sample, re-resolve each robots_snapshot hash against the archive, and compare the classification against the archived terms. The sample-testing guide covers sample sizes and pass thresholds.
Disclosure duties that depend on vendor records
Several regimes now ask model developers to describe their training data, which means your vendor's records become your disclosure inputs. California's AB 2013 (Civil Code Section 3111) requires developers of generative AI systems made available to Californians to post documentation about training datasets [10]; those postings were due 1 January 2026 (status as of October 2026). In the EU, GPAI providers must publish a training-content summary and keep the copyright policy described above [4][5].
Map each disclosure field to a vendor record before you sign. Collection period comes from WARC dates, source description from domain lists, and opt-out compliance from the robots, ToS and TDM evidence. If a field has no backing record, either negotiate for it or plan to exclude the data.
Where licensed operational data fits instead
Some buyers decide that crawl provenance cannot be made strong enough for their risk posture and shift part of the corpus to directly licensed sources. The trade-offs between scraped, synthetic and licensed data are covered in Licensed vs Synthetic vs Scraped AI Training Data, and the distinction between public and proprietary data in Public data vs proprietary data. For overlap risk, see whether licensed data is already in public web crawls.
SourceX does not source scraped web content. It sources operational datasets from US companies, such as support and sales histories, engineering records, documents, and finance and legal workflows, on request, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. If your web corpus has gaps that proprietary business records could fill, you can describe the data you need to SourceX. For the wider framework, start at the provenance hub or the AI data hub, and check common warning signs in provenance red flags.
Sourcing operational data alongside web-derived corpora
SourceX looks for US businesses that hold the data you describe and manages licensing through Find, Assess, Agree, Transact and Manage; nothing is contracted until a supplier agrees, and a request does not guarantee a match. Data is sourced on request rather than held in stock, and every release is approved by the supplying company. Tell SourceX what operational data your team needs.
Sources
- Longpre et al. (Data Provenance Initiative), arXiv / NeurIPS 2024, "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
- The Register, "AI Training Data Is Shrinking as Websites Block Crawlers (coverage of Consent in Crisis)" (2024). https://www.theregister.com/2024/07/22/ai_training_data_shrinks/
- Zyte, "What Is AI Data Provenance?". https://dev.zyte.com/learn/what-is-ai-data-provenance/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- W3C TDM Reservation Protocol Community Group, "TDM Reservation Protocol (TDMRep), Final Community Group Report" (2022). https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20220216/
- Common Crawl Foundation, "Common Crawl's Submission to the UK's Copyright and AI Consultation". https://commoncrawl.org/uk-copyright-and-ai-consultation
- IETF / Internet Architecture Board, "IAB Workshop on AI-CONTROL (aicontrolws) materials" (2024). https://datatracker.ietf.org/group/aicontrolws/materials/
- Wikipedia, "Really Simple Licensing". https://en.wikipedia.org/wiki/Really_Simple_Licensing
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Jain et al., MLCommons Croissant RAI task force (arXiv), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI (Croissant-RAI)" (2024). https://arxiv.org/pdf/2407.16883
- Srivastava et al. (BIG-bench), arXiv, "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models" (2022). https://arxiv.org/pdf/2206.04615
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.