Tables, time series and transactional data
Private Text-to-SQL Evaluation Sets on Enterprise Databases
Quick answer
An enterprise text-to-SQL benchmark you can trust is a private, held-out set of natural-language questions paired with execution-verified gold SQL, run against a snapshot of a real business database plus its documentation. Top scores on public suites such as Spider 1.0 and BIRD can reflect training exposure and label noise as much as skill. A licensed private set, built from real query logs, refreshed on a schedule and never published, gives a better signal of how models will perform on a customer's warehouse.
By SourceX Editorial · Updated
Why Spider and BIRD scores stop predicting enterprise performance
Public text-to-SQL benchmarks overstate production readiness because their schemas are small, their values are clean, and their test items have circulated for years. The Spider 2.0 authors ran the same o1-preview-based agent across suites: it solved 91.2% of Spider 1.0 and 73.0% of BIRD, but only 17.0% of Spider 2.0's 632 enterprise workflow tasks in the arXiv version [1]. The ICLR 2025 version reports 21.3% for that model, so cite both figures and the version when you quote them [2].
The gap comes from what real warehouses look like. Spider 2.0 databases on BigQuery and Snowflake often exceed 1,000 columns, and tasks require reading project documentation and codebases, not just emitting one SELECT [2]. BIRD moved closer to reality with 12,751 question/SQL pairs, dirty cell values and "evidence" strings carrying external knowledge, but its databases are still public and its items are known to model trainers [3].
Contamination is the second failure mode. Test items that leak into pretraining corpora inflate scores until a benchmark stops discriminating [7]. OpenAI stopped reporting SWE-bench Verified for exactly this reason in the coding domain [8], and the same pressure applies to any SQL suite that sits on GitHub and Hugging Face.
Gold labels are the weakest link in public SQL benchmarks
Never treat a public gold query as ground truth without auditing it, because benchmark labels carry measurable error. Northcutt and colleagues estimated an average label error rate of at least 3.3% across ten widely used test sets [6]. For BIRD specifically, a secondhand report describes an audit that found errors in a large share of financial-domain items [5]; treat that as a warning rather than a measured rate, and run your own audit.
Typical SQL label defects include:
- Ambiguous questions where two semantically different queries are both reasonable ("active customers" by last login vs. by open contract).
- Gold SQL that returns the right rows for the wrong reason, such as a missing
DISTINCTthat happens not to matter on the snapshot. - Evidence strings that encode the answer, which reward copying instead of schema reasoning.
- Execution-match false positives when the result set is empty or a single
NULL.
A private set lets you fix these at authoring time: every item gets a written disambiguation note, a second reviewer, and a result set checked against a non-trivial expected cardinality.
What a buyable enterprise text-to-SQL evaluation set contains
A usable private set ships as a bundle: tasks, gold SQL, a frozen database snapshot, documentation and a scoring harness specification. Missing any of these turns the set into a pile of strings you cannot execute. The checklist below is what an evaluation lead should require in the statement of work.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Component | What to require | Why it matters |
|---|---|---|
| Question | Natural-language text as a real analyst would write it, with an ambiguity note | Measures intent resolution, not template matching |
| Gold SQL | Dialect-tagged (Snowflake, BigQuery, PostgreSQL, T-SQL), executed against the snapshot | Execution accuracy needs a runnable reference |
| Expected result | Stored result set with row count, column types and ordering flag | Enables order-insensitive comparison and catches empty-result false passes |
| Database snapshot | Frozen DDL plus de-identified data at a stated as-of date | Gold results drift if the data changes |
| Documentation | Data dictionary, metric definitions, dbt or semantic-layer files | Spider 2.0-style tasks depend on documentation lookups [2] |
| Difficulty tags | Join depth, window functions, CTEs, nested aggregation, date logic | Lets you slice regressions by capability |
| Provenance | Origin (query log, analyst ticket, authored), reviewer IDs, review date | Supports label audits and refresh decisions |
| Efficiency baseline | Gold query runtime or bytes scanned | Supports a BIRD-style Valid Efficiency Score [3] |
For a data-dictionary layer that grounds these tasks, see data dictionaries and schema documentation as AI grounding data.
How private SQL benchmarks are curated from real query logs
The most credible source of enterprise tasks is the supplier's own SQL history, back-translated into questions and reviewed by people who know the schema. BenchPress describes this workflow: an LLM drafts natural-language descriptions of logged enterprise SQL, and domain experts review and correct them, which cuts annotation time while keeping a human on every item [4]. The advantage is distributional realism; the questions reflect what analysts actually ask of that warehouse.
A practical curation pipeline looks like this:
- Sample query logs (for example, Snowflake
QUERY_HISTORYor BigQueryINFORMATION_SCHEMA.JOBS) and drop service-account, DDL and failed queries. - Deduplicate by normalized query fingerprint so dashboard refreshes do not dominate.
- Draft questions with an LLM, then have a schema owner rewrite them in analyst voice and note ambiguity.
- Re-execute gold SQL on the frozen snapshot and store the result set and runtime.
- Second-review a stratified sample and record disagreement rates per difficulty bucket.
- Add multi-step tasks that require cleaning, transformation or reading documentation, not only single queries [1].
Logs mean the questions inherit the supplier's business logic, which is the point. They also mean query text can contain literals such as customer names or account numbers, so de-identification has to cover SQL strings, not just table rows.
Contamination controls for a held-out SQL set
A private set stays useful only as long as its items stay out of training corpora, so contamination control is a contract and operations question, not just a technical one. The working pattern for evaluation data is:
- Never publish items, including in papers, blog posts or leaderboard appendices.
- License the set for evaluation use only, with an explicit prohibition on training, fine-tuning or distillation.
- Log access to the snapshot and task files, and keep the scoring harness separate from model-serving infrastructure that might cache prompts.
- Refresh periodically with newly authored questions, the approach LiveBench uses for general LLM evaluation [7].
- Keep a canary string or hashed item list so you can check whether items appear in a model's outputs or a crawl.
Third-party curation carries its own risk. Researchers have pointed out that private evaluation by data curators can introduce bias and conflicts of interest when the curator also sells training data to the vendors being evaluated [9]. Ask any supplier how evaluation material is segregated from training material. SourceX's guide to contamination checks for licensed eval data covers the checks in more depth, and the broader trade-off is laid out in private evaluation sets vs. public benchmarks.
Scoring data agents, not just single queries
Data-agent evaluation needs execution-based scoring, cost tracking and multi-turn traces, because agents explore schemas, run intermediate queries and recover from errors. Score at least four things:
- Execution accuracy: does the final result set match gold, order-insensitive unless the question demands ordering.
- Efficiency: runtime or bytes scanned relative to gold; BIRD's Valid Efficiency Score rewards correct queries that are also efficient [3].
- Exploration cost: number of tool calls, intermediate queries and tokens before the final answer.
- Safety: whether the agent attempted writes,
SELECT *on wide tables, or cross-schema access it should not have.
Run agents against a read-only copy. Warehouse-native sharing helps here: Snowflake Secure Data Sharing exposes objects read-only to the consumer without copying data between accounts [10], which suits a supplier that will not export a snapshot. For multi-turn follow-ups ("now break that down by region"), pair this set with conversational text-to-SQL data, and for broader environment design see agent evaluation task suites.
Privacy and rights questions specific to SQL evaluation data
Enterprise SQL evaluation data exposes personal and commercially sensitive information in three places: table rows, query literals and documentation. Masking columns alone leaves WHERE customer_name = '...' in the gold SQL and in the question text. Consistent pseudonymization matters too: if ACME-1042 becomes a token in the table, it must become the same token in every question and gold query, or execution accuracy breaks.
Traditional de-identification has known limits compared with formal privacy methods, and NIST SP 800-188 says so directly [11]. For evaluation, the practical controls are pseudonymization that preserves join keys, removal of free-text columns not needed by any task, and a reviewed sample of questions and SQL strings. Confirm with the supplier which tables are in scope, who owns them, and whether the license permits sharing scores publicly, since a published per-item breakdown can leak the set.
Training sets and evaluation sets should be bought separately
Buy text-to-SQL training pairs and evaluation sets under separate agreements, because they need different license scope, refresh cadence and contamination controls. A training license expects the data to enter weights; an evaluation license forbids it. Mixing them under one delivery makes it hard to prove later that eval items never reached training. For the training side, see text-to-SQL training data from real enterprise schemas and schema linking data for wide warehouses; the hub for this cluster is tabular, time-series and transactional data for AI.
How SourceX sources enterprise SQL evaluation data
SourceX sources operational datasets from US companies on request, including engineering records, documents and finance workflows, and manages licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. You can describe the evaluation set you need to SourceX; more on evaluation use cases is on AI evaluation data and evaluation datasets built from real business work.
Request a held-out enterprise text-to-SQL evaluation set
Describe the warehouse type, dialect, domain and task mix you need, and SourceX looks for US businesses that hold matching data. Pricing and allowed uses are agreed in a license per deal, and nothing is contracted until a supplier agrees. Start a buyer request at SourceX.
Frequently asked questions
Is Spider 2.0 enough as an enterprise text-to-SQL benchmark?
Spider 2.0 is one of the most realistic public suites, with 632 tasks, many on large cloud warehouses [1], but it is public. Use it for comparability and a licensed private set for decisions, since public items can enter training data over time [7].
How many items does a private text-to-SQL eval set need?
Size depends on how many capability slices you track. A set stratified by dialect, join depth and task type needs enough items per slice to detect the regressions you care about; plan the slices first, then the count.
Can we report scores on a licensed private set publicly?
Only if the license allows it. Aggregate scores are usually less risky than per-item results, which can reveal questions or schema details.
Sources
- arXiv (Lei et al.), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows (arXiv:2411.07763)" (2024). https://www.arxiv.org/pdf/2411.07763
- ML Anthology / ICLR 2025, "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows (ICLR 2025)" (2025). https://mlanthology.org/iclr/2025/lei2025iclr-spider
- arXiv (Li et al.), "Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs (BIRD)" (2023). https://arxiv.org/abs/2305.03111v2
- arXiv, "BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation" (2025). https://arxiv.org/pdf/2510.13853
- Beancount.io Bean Labs, "BIRD benchmark: the text-to-SQL real database gap" (2026). https://beancount.io/bean-labs/research-logs/2026/06/06/bird-benchmark-text-to-sql-real-database-gap
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (White et al.), "LiveBench: A Challenging, Contamination-Free LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
- Snowflake Documentation, "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
- NIST, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.