Skip to content

Retrieval, RAG and grounding data

De-identifying search query logs before licensing

Quick answer

To anonymize search query logs for licensing, a supplier should scan raw query text for direct identifiers, drop or coarsen user and session keys, aggregate to query-level or query-document counts, and suppress any query issued by fewer than k distinct users. Aggregation plus a k-anonymity filter is the pattern the public ORCAS click release used [1]. Each step costs retrieval value: thresholds cut the long tail and session cuts remove reformulation signal, so buyers should negotiate the release spec, not just accept "anonymized."

By SourceX Editorial · Updated

This page is for the privacy engineer or procurement lead who has to sign off on a query-log dataset for retriever, reranker or query-rewriting work. For what makes query and click logs useful in the first place, see search query and click logs for retriever training; for the general, supplier-side method across all company data, the owner page is de-identifying company data for AI training.

Why query logs are harder to de-identify than ordinary tables

Query logs are hard because the identifying content sits inside free text that users typed themselves, and the sequence of queries from one person is itself a fingerprint. A row like user_id, timestamp, query, clicked_url, rank looks tabular, but the query column can contain a name, a street address, an order number or a medical condition, and no column-level rule catches that.

The canonical failure is the 2006 AOL research release, in which user IDs were replaced by numbers but every query from one pseudonym still sat together, so the histories could be linked back to people. Pseudonymizing the user key does nothing to stop that kind of linking. NIST's survey of re-identification documents the same pattern across domains: removing direct identifiers while leaving rich records intact has repeatedly been reversed [5].

Query logs also carry commercial sensitivity on top of personal risk. The ORCAS authors note that click logs reveal both personal and commercially sensitive information, which is why their release was filtered and aggregated rather than published raw [1]. For an enterprise supplier, raw internal search logs can expose product roadmaps, customer names, deal codes and unreleased SKUs, so expect the supplier's legal team to want the same filters for its own reasons.

The four controls a query-log release should apply

A defensible release combines free-text scrubbing, key removal, aggregation and a minimum-user threshold; any one alone is insufficient. The order matters, because thresholds computed before scrubbing will miscount near-duplicate queries that differ only by an embedded account number.

  1. Free-text PII and secret scanning on the query string. Run an entity detector (Microsoft Presidio is a common open-source choice) plus regexes for emails, phone numbers, card and account numbers, order IDs, tracking numbers, IPs and internal ticket keys. Presidio's own documentation cautions that ML-based detection gives no guarantee of finding everything [6], so pair it with domain patterns and a manual sample. Internal search boxes also attract pasted API keys and passwords; see credentials and secrets in ticket, chat and log datasets.
  2. Remove or coarsen linkage keys. Drop user_id, cookie, device and IP fields; truncate timestamps (hour or day); and decide explicitly whether session IDs survive. Research on event logs shows that sequences plus timestamps make individual cases unique even without names [8].
  3. Aggregate. Publish (query, document, click_count) or (query, count) tuples rather than per-event rows. ORCAS released query-document click pairs, not user histories [1].
  4. Apply a k-user threshold. Keep a query (or query-document pair) only if at least k distinct users issued it. This is k-anonymity applied to the query string as the quasi-identifier [3]; read the k-anonymity glossary entry for the general definition.

Treat k-anonymity as a floor, not a proof. Sweeney's original paper already notes that attacks can succeed on k-anonymous releases unless accompanying policies are in place [3], and NIST SP 800-188 contrasts these traditional techniques with formal methods such as differential privacy [4].

What each control costs in retrieval value

Every control trades training signal for privacy, and the biggest losses fall on exactly the queries a domain retriever most needs. Head queries survive thresholds easily; rare, specific, long-tail queries (exact part numbers, error strings, policy clauses) are disproportionately removed, and those are often where an off-the-shelf embedding model fails.

Illustrative example: invented to show structure; it does not describe an available dataset.

ControlWhat it removesRetrieval costMitigation to negotiate
k-user threshold (k = 5 to 50)Queries issued by fewer than k usersLong-tail and zero-result queries; rare entity lookupsRelease a canonicalized template of rare queries (error code <CODE> on <MODEL>) instead of dropping them
PII span maskingNames, emails, account and order numbers inside queriesEntity-matching signal; navigational queriesTyped placeholders (<ORDER_ID>) rather than deletion, so query shape is preserved
Session ID removalReformulation chainsQuery-rewrite and conversational retrieval training pairsRelease adjacent-pair reformulations (q1 -> q2) that each pass the threshold, without session keys
Timestamp coarseningFine-grained timeFreshness and seasonality modelingDay or week buckets
Click aggregationPer-impression rank, dwell, skipsPosition-bias correction for click modelsAggregate click-through by rank bucket per query
Dropping the query text entirelyEverythingMost of the valueRarely acceptable; consider synthetic queries only as a supplement

Measure the cost rather than guess it. Ask the supplier for distribution statistics before and after each filter: share of unique queries retained, share of total query volume retained, and the query-length histogram. A release that keeps most volume but a small share of unique queries is a head-query dataset, which matters when you pilot-test a source for retrieval lift.

Thresholds versus formal privacy for query logs

A frequency threshold is the established practical pattern, while differential privacy is the option when the release must withstand a motivated adversary. NIST SP 800-188 sets out this distinction between traditional de-identification and formal privacy methods, and cautions that traditional techniques have inherent limits [4].

In practice, DP for query logs means a randomized threshold plus Laplace or Gaussian noise on counts, with a per-user contribution cap. The buyer consequences are concrete: counts are noisy (so click-through ratios on mid-frequency queries become unreliable), the per-user cap discards heavy users' tails, and the epsilon value becomes a license-relevant parameter you should see documented. Use the same NIST guidance as the vocabulary for describing the chosen model and its governance in the spec.

Free-text risks a threshold does not catch

Even a query issued by many users can still be sensitive, and some identifiers repeat across users. Examples include an employee's name searched by a whole team, a customer company name in an enterprise search log, or a viral phone number. Thresholds count users, not meaning.

Three residual risks deserve an explicit check:

  • Quasi-identifier combinations. A query mentioning a rare condition plus a small town plus an employer can point to one person even with no direct identifier. The Text Anonymization Benchmark distinguishes direct identifiers from quasi-identifiers and confidential attributes for exactly this reason [9].
  • Inference by your own models. LLMs can infer personal attributes such as location or occupation from ordinary text, not only from memorized strings [7]. A rewriter trained on scrubbed queries can still learn and surface those correlations.
  • Linkage to the document side. If clicked URLs point to user-specific pages (profile pages, order-status URLs with tokens in the query string), the document column re-identifies the user. URL parameters need the same scanning as queries.

The re-identification glossary entry covers the general attack types; for corpus-side scrubbing that preserves retrievability, see de-identifying a RAG corpus without breaking retrieval.

A release spec to request from a query-log supplier

Ask for a written release spec before you evaluate samples, so that the de-identification is reviewable rather than a claim. The checklist below is what a privacy reviewer typically needs to approve the data.

Illustrative example: invented to show structure; it does not describe an available dataset.

query_log_release_spec:
  source_system: "internal site search (Elasticsearch query logs)"
  period: "2025-01-01 to 2025-12-31"
  unit_of_release: "query-document pair with click_count"
  removed_fields: [user_id, cookie_id, ip_address, user_agent, session_id]
  timestamp_granularity: "day"
  free_text_scrub:
    detectors: ["Presidio NER", "regex: email, phone, card, order_id, account_no"]
    replacement: "typed placeholders, e.g. <ORDER_ID>"
    url_parameter_scrub: true
    secret_scan: true
  threshold:
    rule: "keep query only if issued by >= k distinct users"
    k: 20
    applied_after_scrub: true
  formal_privacy: "none (or: DP with stated epsilon, delta, contribution cap)"
  retention_stats:
    unique_queries_retained_pct: "<reported>"
    query_volume_retained_pct: "<reported>"
  manual_review: "random sample of retained queries reviewed; residual findings logged"
  known_limits: "detector recall is not 100 percent; quasi-identifier combinations not exhaustively tested"

Before signing, confirm three things against the spec. First, that thresholds were computed on scrubbed, normalized strings (lowercased, whitespace-collapsed), not raw text. Second, that the sample you evaluate went through the same pipeline as the full delivery. Third, that the license names the allowed uses (retriever training, evaluation, query rewriting) and whether derived embeddings or rewrites can be retained; RAG content license terms covers the clause side.

How ORCAS frames the trade-off for commercial buyers

ORCAS is the reference point for what a filtered public click release looks like, and also a reminder that public options rarely solve the commercial problem. It provides about 18 million query-document connections over roughly 10 million distinct queries and 1.4 million TREC Deep Learning documents, and it is released for non-commercial use [2].

So a lab building a production retriever usually needs licensed logs from an operator whose queries match its domain: e-commerce, support portals, technical documentation or product search relevance data. Those logs will be smaller and more domain-specific than web-scale releases, which raises the threshold cost: with fewer users, more queries fall below k. Budget for that in your volume expectations, and use the retrieval cluster hub to compare query logs with other relevance-signal sources such as qrels and ticket-to-article links.

If you would rather describe the query data you need and have the commercial process handled, SourceX sources operational datasets from US companies on request; you can describe your query-log requirements to SourceX.

Licensing de-identified search query logs through SourceX

SourceX looks for US businesses that hold the query data you describe, and every release is approved by the supplying company. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Start by describing the query logs you need.

Frequently asked questions

Is a k-anonymity threshold on query strings enough on its own?

No. It prevents publishing queries unique to one user, but it does not catch identifiers shared by many users, quasi-identifier combinations, or linkage through clicked URLs [3][5]. Pair it with free-text scrubbing and key removal.

What value of k should we expect?

There is no standard value; it is a risk decision for the supplier and should be stated in the spec. Ask for retention statistics at two or three candidate k values so you can see the long-tail cost before agreeing.

Should session data ever be released?

Only if your use case needs reformulation chains, and then preferably as thresholded adjacent query pairs without session identifiers. Sequences of events are highly unique and raise re-identification risk sharply [8].

Sources

  1. arXiv (Craswell et al., Microsoft), "ORCAS: 18 Million Clicked Query-Document Pairs for Analyzing Search (arXiv:2006.05324)" (2020). https://arxiv.org/pdf/2006.05324
  2. Microsoft (MS MARCO), "ORCAS: Click data for TREC Deep Learning". https://microsoft.github.io/msmarco/ORCAS.html
  3. Latanya Sweeney, Data Privacy Lab, "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/people/sweeney/kanonymity.html
  4. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  5. National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  6. Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  7. arXiv (Staab, Vero, Balunovic, Vechev; ICLR 2024), "Beyond Memorization: Violating Privacy Via Inference with Large Language Models (arXiv:2310.07298)" (2023). https://arxiv.org/pdf/2310.07298
  8. arXiv (Nunez von Voigt et al.; CAiSE 2020), "Quantifying the Re-identification Risk of Event Logs for Process Mining (arXiv:2003.10707)" (2020). https://arxiv.org/pdf/2003.10707
  9. arXiv (Pilan et al.), "The Text Anonymization Benchmark (TAB) (arXiv:2202.00443)" (2022). https://arxiv.org/pdf/2202.00443

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data