Retrieval, RAG and grounding data
Multi-source enterprise search corpora: wikis, chat, tickets and drives
Quick answer
An enterprise search benchmark dataset worth buying is a pinned snapshot of several systems from one real company: a wiki, chat channels, a ticketing tool and a shared drive, with the cross-links, permissions, duplicates and stale pages left intact. Most public enterprise benchmarks are generated or fictional, so they miss that mess. Buyers should specify the systems, linkage keys, access-control metadata, judgment format and license scope (evaluation or training) before approaching a supplier.
By SourceX Editorial · Updated
Why public enterprise RAG benchmarks fall short
Public enterprise benchmarks are useful baselines, but the strongest ones simulate a company rather than sample one. Onyx's EnterpriseRAG-Bench spans nine enterprise source types, yet it is built from a generated company with injected distractors and released under the MIT license [1]. Its accompanying paper argues that company internal knowledge is a distinct retrieval target from open-web QA [2], which is exactly why a synthetic stand-in leaves a gap.
Citation-grounded corpora are smaller still. RAG-Multi-Corpus contains 236 documents across five fictional organizations, with 786 query-answer pairs and ground-truth citations across mixed file formats [3]. WixQA is closer to real operations because it is anchored in one company's support content [4], but it covers a single knowledge domain rather than the linked application estate a copilot actually indexes. BEIR remains the reference for zero-shot heterogeneity [5], and none of its datasets is an internal company corpus.
What generated corpora cannot reproduce:
- Organic cross-references. A Jira key pasted into a Slack thread, a Confluence page linked from a Zendesk macro, a Google Drive URL in a ticket comment.
- Version sprawl. "Pricing_v3_FINAL(2).docx" next to the superseded v2, and a wiki page last edited four years ago that still ranks.
- Vocabulary drift. Internal codenames, product renames and acronyms that change between teams and years.
- Permission structure. Private channels, restricted spaces and drive folders shared with three people, which decide what a user is allowed to see.
For a broader map of what public collections permit commercially, see which public retrieval datasets allow commercial use.
Which systems to include and what each contributes
A useful cross-application corpus combines at least three systems with different document shapes, because retrievers fail differently on long pages, short messages and structured records. The table below shows what each source class adds to evaluation and where it typically breaks a pipeline.
| System class | Typical tools | Retrieval value | Common failure it exposes |
|---|---|---|---|
| Wiki / knowledge base | Confluence, Notion, SharePoint pages | Canonical answers, long-form procedures | Stale pages outranking current ones |
| Chat | Slack, Microsoft Teams | Tacit knowledge, recent decisions | Short context; answers split across thread replies |
| Tickets | Zendesk, Jira Service Management, ServiceNow | Real questions paired with resolutions | Internal notes leaking into customer-facing answers |
| Engineering tracker | Jira, Linear, GitHub Issues | Bug history, root causes, codenames | Identifier-only queries ("PLAT-4412") |
| Shared drives | Google Drive, OneDrive, Box | Specs, decks, contracts, spreadsheets | Near-duplicates, scanned PDFs, embedded tables |
| Exchange, Gmail | Decisions and external context | Quoted-reply duplication, signatures |
Ticket data deserves particular care because the audit trail records each comment and field change, including whether a comment was public or internal [8]. That flag matters: an assistant evaluated on internal notes may learn to surface them to customers. Single-system collections are covered on the owner pages for enterprise document archives and workplace email and chat; this page is about linking them.
Cross-links and ACLs make the corpus realistic
The value of a multi-source corpus sits in the joins between systems and in the permission metadata, so both must survive export and de-identification. Ask for an explicit link table rather than relying on URLs left in text, because URL rewriting during redaction often breaks them silently.
Permissions are the field most often dropped. Enterprise search products must enforce document-level access, so a corpus without principal lists cannot test permission-aware retrieval or leakage across groups. Request pseudonymized group and user identifiers that stay consistent across systems, so "eng-platform" in Slack and in Confluence map to the same token.
Ticket-to-article links are the cheapest source of relevance labels; support tickets linked to knowledge articles explains how to turn them into training pairs. For process timelines across ERP, CRM and ITSM, see cross-system workflow records, which is a different use than search.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "conf:SPACE-OPS:184220",
"source_system": "confluence",
"doc_type": "wiki_page",
"title": "Failover runbook: billing-api",
"created_at": "2022-03-14T09:12:00Z",
"last_modified_at": "2024-11-02T16:40:00Z",
"version": 17,
"superseded_by": null,
"author_pid": "u_7f3a",
"acl": {"groups": ["g_eng_platform", "g_sre"], "users": [], "public_to_company": false},
"links_out": [
{"target": "jira:PLAT-4412", "relation": "mentions"},
{"target": "gdrive:1aXc...redacted", "relation": "attachment"}
],
"links_in": [
{"source": "slack:C02OPS:1730561230.004100", "relation": "shared_in_thread"},
{"source": "zendesk:ticket:88213", "relation": "macro_reference"}
],
"deid": {"method": "entity_replacement", "entities_replaced": ["PERSON", "EMAIL", "PHONE"]},
"snapshot_id": "snap-2026-09-30"
}
Turning the corpus into a test collection
A corpus becomes an enterprise search benchmark only once it has topics, judgments and a frozen snapshot, and you can usually build those yourself if the supplier provides the raw material. The TREC 2025 RAG track distributes its corpus, topics, relevance judgments and nuggets as separate files [6], and classic TREC qrels give a standard format for graded judgments [7]. Mirroring that split keeps your evaluation portable across vendors and internal tools.
For generation-side metrics, frameworks such as Ragas expect each record to hold a question, retrieved contexts, an answer and a ground truth [9]. Sources of real questions inside a company corpus include ticket subjects, Slack messages ending in a question mark that received a reply, and search logs if the supplier can license them. Method detail is on building a private domain retrieval test collection and pinned corpus snapshots for reproducible RAG evaluation.
Keep the duplicates. Superseded drafts and near-copies are what make ranking hard in production, and deduplicating them before evaluation inflates scores; superseded versions, drafts and near-duplicates covers how to label them instead of deleting them.
Rights review differs by system
Each system in the corpus carries a different rights profile, so review needs to run per source rather than once for the whole export. Treat the following as a buyer's diligence list, not a legal standard.
- Employee-authored content (wiki, chat, internal email): confirm the company's employment and acceptable-use policies support licensing it, and check works-council or union constraints for non-US staff.
- Customer-authored content (ticket bodies, shared files from customers): check what the supplier's customer terms and privacy notice said when the data was collected. The FTC has warned that quietly adopting more permissive practices, such as using data for AI training or sharing it with third parties, may be unfair or deceptive [10].
- Third-party content (vendor PDFs in drives, partner contracts, pasted articles): often not the supplier's to license; ask for exclusion rules by file origin or folder.
- SaaS platform terms: confirm that exporting and relicensing workspace data is allowed under the supplier's agreements with the tools themselves.
- Personal data: names, emails, handles and phone numbers appear everywhere in chat and tickets. See PII redaction for LLM training data for measuring what redaction misses, and the chain of title documents to request.
Evaluation-only versus training licenses
Decide early whether you need the corpus to measure retrieval or to train retrievers and rerankers, because suppliers price and approve those uses differently. An evaluation-only license may restrict copies, require a held-out environment and forbid fine-tuning on judgments; a training license must address embeddings, model weights and what happens to derived indexes when the term ends. Grounding license vs training license compares the two.
Spell out three specifics in any request: whether embeddings computed from the corpus may persist after the license ends, whether benchmark scores on the corpus may be published, and whether the supplier's name may be disclosed. Many suppliers will allow internal evaluation long before they allow public leaderboards.
Request template for a cross-application corpus
A precise request shortens supplier assessment, because the holder of the data can check it against what their systems actually contain.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example specification |
|---|---|
| Systems | Wiki, chat, ticketing, engineering tracker, shared drive (minimum three) |
| Organization profile | US B2B software company, 200 to 2,000 employees, at least three years of history |
| Snapshot | One pinned export date; optional quarterly refresh |
| Linkage | Cross-system link table; consistent pseudonymous user and group IDs |
| Permissions | Per-document ACL principals; private versus public channel flag |
| Versions | Revision history or superseded-by pointers kept |
| Formats | JSONL metadata plus original files (DOCX, PDF, PPTX, XLSX, HTML) |
| Labels | Ticket-to-article links; optional query set with qrels |
| De-identification | Entity replacement, method documented, sample check |
| Use | Evaluation only, or training of retrievers and rerankers |
How SourceX approaches multi-source corpora
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and documents; nothing is held in stock, and a request does not guarantee a match. Buyers describe the data, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees.
Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For the evaluation use case specifically, see knowledge retrieval evaluation and RAG evaluation datasets from real company documents, or start a request on the SourceX buyer page.
Source an enterprise search corpus from real company systems
SourceX serves AI teams wherever they are based and sources operational datasets from US companies, managing licensing and ongoing purchases. Describe the systems, linkage and license scope you need, and SourceX will look for a supplier willing to approve the release. Describe your corpus requirements to SourceX.
Frequently asked questions
Can a cross-application corpus come from more than one company?
It can, but each company's systems should stay a separate tenant with its own ACLs and snapshot, since cross-company retrieval is not a realistic enterprise scenario. If you need breadth across suppliers, single-source or multi-source supplier strategy covers the trade-offs.
Does de-identification break cross-links?
It can, when user handles, emails or URLs are redacted inconsistently across systems. Ask for deterministic replacement so the same person or link maps to the same token everywhere, and test a sample of link-table rows after redaction.
What size is realistic?
Size depends on the supplier's estate and what they approve for release, so specify minimum coverage per system and history length rather than a document count. The retrieval and RAG data hub lists related corpus types.
Sources
- Onyx, "EnterpriseRAG-Bench". https://onyx.app/enterpriserag-bench
- arXiv, "EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge" (2026). https://arxiv.org/pdf/2605.05253
- arXiv, "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation" (2026). https://arxiv.org/pdf/2604.04936
- arXiv, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
- arXiv, "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- NIST TREC, "2025 RAG Search Track" (2025). https://trec.nist.gov/data/rag2025.html
- NIST TREC, "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
- Zendesk Developer Docs, "Ticket Audit events reference". https://developer.zendesk.com/documentation/ticketing/reference-guides/ticket-audit-events-reference/
- Ragas documentation, "Prepare your test dataset". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.