Skip to content

Retrieval, RAG and grounding data

Product search relevance data: queries, catalogs and judgments

Quick answer

A usable product search relevance dataset has three linked parts: real shopper or buyer queries, the catalog as it looked when each query ran, and relevance evidence for query-product pairs, either graded human labels (Exact, Substitute, Complement, Irrelevant) or logged impressions and clicks with positions. Public benchmarks such as Amazon's Shopping Queries Dataset cover evaluation research, but production click logs are rarely released, so teams training or evaluating a domain ranker usually need to license them from a company that runs the search.

By SourceX Editorial · Updated

What a product search relevance dataset should contain

A complete dataset joins query logs, catalog snapshots and judgments on stable keys, so every label can be traced to the exact product record the searcher saw. Without that join, you cannot tell whether a "bad" result was a ranking failure or a product that was out of stock, repriced or renamed after the fact.

The minimum tables are:

  • Queries: query_id, raw query text, normalized text, timestamp, locale, channel (web, app, punchout, ERP-integrated), session_id, and for B2B, an account segment rather than an account name.
  • Impressions: query_id, product_id, rank position, page number, filters and sort applied, and the ranker or experiment variant that produced the list.
  • Interactions: clicks, add-to-cart, quote requests, purchases, and dwell or return-to-results events, each with timestamp and position.
  • Catalog snapshot: product_id, title, attributes, category path, brand, price, availability and variant relationships as of the query timestamp.
  • Judgments: query_id, product_id, grade, assessor type (human, LLM, rule), guideline version and adjudication status.

For the catalog-text side alone, such as descriptions for pretraining or attribute extraction, the owner page is licensing product catalogs and descriptions for AI training. This page covers the search layer on top of it.

Public product search datasets and where they stop

Public data is good for benchmarking methods but rarely matches your catalog, query distribution or license needs. The best-known public resource is Amazon's Shopping Queries Dataset (often called ESCI), whose query-product pairs carry four labels: Exact, Substitute, Complement and Irrelevant. Check the license file in its official repository, and the scale and language coverage in its paper, before any commercial use.

Click data is scarcer. The ORCAS authors note that click logs are rarely published because they are commercially sensitive, and ORCAS itself covers web documents from the TREC Deep Learning corpus, not products [1]. Its distribution terms are restrictive, so confirm whether they allow commercial model training before relying on it [2]. For a wider survey, see which public retrieval datasets allow commercial use.

The practical gap: public sets lack your vertical's vocabulary (part numbers, SKU fragments, unit-of-measure queries like "3/4 npt brass elbow"), lack query-time prices and stock, and lack the B2B signals that matter, such as contract pricing and reorder behavior.

Why the catalog snapshot at query time matters

Clicks only mean something relative to what the searcher saw, so the dataset must pin the catalog state at query time. A product that was out of stock, priced 40% above a substitute, or missing an image will collect fewer clicks regardless of relevance, and a model trained on that signal learns availability and price, not relevance.

Ask suppliers for one of three snapshot designs: full daily catalog dumps keyed by date, a slowly changing dimension table (valid_from and valid_to per product version), or impression-level copies of the displayed fields. The second is usually the best balance of size and fidelity. Pinning also makes offline evaluation reproducible; see pinned corpus snapshots for reproducible evaluation.

Click logs versus graded labels for training and evaluation

Use logged clicks for training at scale and graded human judgments for evaluation, because clicks are biased by position and presentation while labels are expensive but interpretable. Joachims and colleagues showed that clicks are skewed by rank position and proposed propensity-weighted, counterfactual learning-to-rank to correct for it [3]. Propensities can also be estimated from logs without randomized result swaps, provided the logs record the ranking actually shown and come from more than one ranker [4].

That makes the ranker variant field and full impression lists (not only clicked items) non-negotiable. A log of clicks without impressions cannot be debiased.

For labels, ESCI-style four-grade scales fit commerce well because "Substitute" and "Complement" capture commercial value that a binary scale loses. Store judgments in a qrels-compatible layout (query, product, grade) so standard evaluation tools read them [5]. For assessor choices and scale design, see buying relevance judgments (qrels), LLM labels versus human assessors and graded scales for domain experts.

Request template for a product search data license

A precise request describes the data and its use, not the companies you hope hold it. The template below is a starting point for suppliers or intermediaries.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample specification
DomainB2B industrial supplies catalog, US buyers
Query volume12 months of search queries, all sessions, with timestamps
ImpressionsTop 48 results per query with rank, page, filters, sort, ranker variant
InteractionsClick, add-to-cart, quote request, order line, with positions
CatalogSlowly changing dimension table: title, attributes, category, price band, stock flag
Labels5,000 queries x 20 products graded E/S/C/I by trained assessors, guideline attached
IdentifiersAccount and user IDs replaced with stable pseudonyms; free text scanned for PII
FormatParquet, partitioned by date, with a data dictionary and data card
UseTrain and evaluate a product ranker; offline evaluation; no redistribution
DeliveryAccess-controlled share or bucket; no email attachments

Ask for a data card covering sources, collection, annotation method and known gaps [8]. Delta Sharing or a cloud bucket share works well for large partitioned logs [7].

Privacy and commercial sensitivity in query logs

Query logs carry two separate risks: personal data typed into the search box and commercially sensitive signals about the supplier's customers and pricing. Shoppers paste order numbers, emails, phone numbers and names into search boxes, and B2B users type customer names and PO numbers. Automated scanners such as Presidio help, but their own authors caution that no tool finds everything, so require a recorded method plus a manual sample review [6].

Commercial sensitivity is often the bigger blocker. Suppliers will typically bucket prices into bands, drop margin fields and replace account identifiers with pseudonyms; agree on which fields survive before scoping labels. The guide to de-identifying search query logs covers techniques and residual risk.

Diligence checklist before you sign

Check that the supplier can show the data was collected by its own search system, that its terms allow the use you need, and that joins hold up on a sample.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Sample of 1,000 queries joins cleanly to impressions, interactions and the catalog version valid at query time.
  • Impression lists include non-clicked results and the ranker variant.
  • Bot and internal-test traffic flagged or removed, with the rule documented.
  • Label guideline version, assessor agreement figures and adjudication process provided.
  • De-identification method recorded and a reviewed sample available.
  • License defines records, permitted uses (training, evaluation, or both), term and delivery.
  • Refresh cadence agreed if you need ongoing logs for drift monitoring.

For clause-level detail, see RAG content license terms; for click logs used as general retriever training data, see search query and click logs for retriever training.

How SourceX approaches product search data requests

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. You describe the data, such as queries, impressions, catalog snapshots and labels; SourceX looks for US businesses that hold it, and every release is approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. SourceX does not train models and does not publish prices; terms are agreed per deal. You can describe your product search data needs to SourceX, and the retrieval data hub covers related corpora.

License product search relevance data

SourceX finds US companies that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before any delivery. Nothing is contracted until a supplier agrees. Start a product search data request.

Frequently asked questions

Can I train a commercial ranker on ESCI?

Possibly, but check the license in the official repository before you do. Even where terms allow it, ESCI reflects one retailer's catalog and query mix, so most teams use it for method comparison and license domain data for production.

Do I need human labels if I have clicks?

Yes for evaluation. Debiased clicks can train a ranker [3], but a graded, guideline-based label set gives a stable test that does not move with the ranker you are testing.

How many labeled queries does an evaluation set need?

It depends on the metric variance you can tolerate and how many query segments you report. Stratify by head, torso and tail queries and by category so each reported segment has enough queries to compare rankers.

Sources

  1. arXiv (Craswell et al., Microsoft), "ORCAS: 18 Million Clicked Query-Document Pairs for Analyzing Search" (2020). https://arxiv.org/pdf/2006.05324
  2. Microsoft (MS MARCO), "ORCAS". https://microsoft.github.io/msmarco/ORCAS.html
  3. arXiv (Joachims, Swaminathan, Schnabel), "Unbiased Learning-to-Rank with Biased Feedback" (2016). https://arxiv.org/pdf/1608.04468
  4. arXiv (Agarwal et al.), "Estimating Position Bias without Intrusive Interventions" (2018). https://arxiv.org/pdf/1812.05161
  5. NIST TREC, "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
  6. Microsoft (microsoft/presidio), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  7. Delta Lake project, "Read Delta Sharing Tables". https://docs.delta.io/delta-sharing/
  8. arXiv (Pushkarna, Zaldivar, Kjartansson, Google Research), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data