Tables, time series and transactional data
Text-to-SQL Training Data from Real Enterprise Schemas
Quick answer
Text-to-SQL training data is a set of records that pair a database schema and its values with a natural-language question and a gold SQL query that has been checked by execution. For enterprise models, the useful version comes from real warehouses: wide tables, cryptic column names, coded values and a named dialect such as Snowflake, BigQuery, T-SQL or PostgreSQL. Buyers should specify the schema snapshot, the question source, the gold-SQL verification method, the dialect, the de-identification approach and the license scope before anything is collected.
By SourceX Editorial · Updated
Why real enterprise schemas beat academic and fully synthetic ones
Real schemas teach the failure modes that academic schemas hide: hundreds of columns, near-duplicate tables, abbreviations like CUST_ACCT_NBR, and values that only make sense with a code list. Synthetic pipelines know this. NVIDIA's NeMo Data Designer write-up for Nemotron injects distractor tables and columns, runs separate validators per dialect (MySQL, PostgreSQL, SQLite) and adds LLM critics, all to imitate the noise of real systems [1]. Researchers building triples from real schemas also report that existing real-schema resources remain small for training large models [2].
Values matter as much as structure. BIRD was built because Spider and WikiSQL focused on schema with few rows of content; it pairs 12,751 questions with SQL over 95 databases totaling 33.4 GB and adds external-knowledge "evidence" to resolve dirty values [3]. Spider 2.0 moved further toward production, with 632 workflow problems on databases from real applications, including BigQuery and Snowflake warehouses [4]. Both are evaluation-oriented; your training set should cover the same difficulty without overlapping them. Keep held-out sets separate, as described in our guide to private text-to-SQL evaluation sets on enterprise databases.
What one deliverable record should contain
A usable record bundles the schema, the values the query touches, the question, the gold SQL and the expected result, so you can re-execute it. Public enterprise-style sets show the shape: one Hugging Face set offers 500 question/gold-SQL pairs on an Odoo 17 supply-chain schema, execution-verified against a demo database, plus 50 multi-turn items [6]. A commercial delivery should be stricter, because you will train on it at scale.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "t2s-000184",
"db_id": "wholesale_erp_snapshot_2026q2",
"dialect": "snowflake",
"schema_ref": "ddl/wholesale_erp_snapshot_2026q2.sql",
"snapshot_ref": "parquet/wholesale_erp_snapshot_2026q2/",
"question": "Which three distributors had the highest return rate last quarter?",
"evidence": "Return rate = RMA lines / shipped lines; STATUS_CD 'RT' means returned.",
"gold_sql": "SELECT d.DIST_NAME, ... QUALIFY ROW_NUMBER() OVER (ORDER BY rate DESC) <= 3",
"expected_result_ref": "results/t2s-000184.parquet",
"result_row_count": 3,
"tables_used": ["DIST_MASTER", "SHIP_LINE", "RMA_LINE"],
"difficulty": "hard",
"question_source": "query_log_backtranslated",
"reviewer_status": "expert_verified",
"deid_method": "consistent_pseudonymization_v2"
}
The evidence field connects to data dictionaries and metric definitions used as grounding data. The tables_used field doubles as supervision for schema linking and schema retrieval. Ask for dataset-level metadata in a machine-readable format such as Croissant, which describes resources and record structure in JSON-LD [9].
Sourcing question/SQL pairs from analyst query history
Real query history is the strongest seed for question/SQL pairs, because it reflects what analysts actually ask of a warehouse. The practical route is back-translation: an LLM drafts the natural-language intent behind an existing query, then a domain expert corrects the question, and an execution check confirms the SQL still runs. BenchPress documents this human-in-the-loop pattern for curating text-to-SQL items from logs [5].
Query logs need hygiene before they become training data:
- Literals: scrub or tokenize literal values in
WHEREandINclauses consistently, so the same customer ID maps to the same token across queries and the snapshot. - Schema drift: keep the schema snapshot as of query time; a query written before a column rename will fail against today's DDL.
- Templated dashboards: deduplicate queries generated by BI tools that differ only in date filters, or they will dominate the distribution.
- Broken or abandoned queries: keep only queries that executed successfully, unless you want error-correction examples labeled as such.
- Access context: record which role ran the query; row-level security can make the same SQL return different results.
Multi-turn follow-ups ("now break that down by region") belong in a different shape, covered in multi-turn conversational text-to-SQL data.
Naming the dialect and the hard constructs
Specify the target dialect in the request, because gold SQL that parses in one engine often fails in another. Per-dialect validation is standard practice even in synthetic pipelines [1], and enterprise benchmarks run on cloud warehouses such as BigQuery and Snowflake [4]. The constructs that trip models are predictable: date arithmetic (DATEADD vs DATE_ADD vs INTERVAL), semi-structured access (Snowflake VARIANT with : paths and LATERAL FLATTEN, BigQuery STRUCT/ARRAY with UNNEST, PostgreSQL jsonb operators), QUALIFY, TOP vs LIMIT in T-SQL, and identifier quoting with backticks, double quotes or square brackets.
Ask suppliers to report the dialect mix as counts per engine and version, and to state whether any record was transpiled from another dialect. Transpiled queries are useful but should be flagged, since SQL-to-SQL conversion is a separate task from natural-language translation.
Buyer specification checklist
A short written specification prevents most disputes about what a text-to-SQL delivery should contain.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Item | What to specify | Failure it prevents |
|---|---|---|
| Schema scope | Domains (ERP, CRM, billing), table count range, column width | Toy schemas that do not transfer |
| Value snapshot | Format (Parquet, SQL dump), row sampling, as-of date | Gold SQL that returns empty sets |
| Question source | Query-log back-translation, analyst-written, or synthetic, with shares | Unrealistic phrasing |
| Gold verification | Execution against the delivered snapshot; stored result set | Silent wrong answers |
| Dialect | Engine and version per record | Parse failures at inference |
| Difficulty tags | Join count, nesting, window functions, evidence needed | Skewed easy-only mix |
| De-identification | Method per column, consistency across tables | Broken joins or leaked values |
| Overlap check | Hash comparison against BIRD and Spider 2.0 | Benchmark contamination |
| License scope | Training use, schema and values both covered, term | Unclear rights in vendor schemas |
Privacy and rights in schemas and cell values
Cell values in a snapshot can contain personal data, so de-identification must be consistent across tables for joins and filters to still return meaningful results. Replace names, emails and account numbers with stable pseudonyms rather than nulls, and check that aggregates still behave plausibly. NIST SP 800-188 is a useful reference for choosing techniques and is explicit about the limits of traditional de-identification [8]. For health data, HIPAA de-identification follows Safe Harbor or Expert Determination [7].
Rights questions differ from those in a source-code purchase (see licensing source code for AI training). Packaged-application schemas can be vendor intellectual property, so the license should cover the schema, the values and the derived questions. Provenance matters: an audit of 1,800+ public datasets found license omission above 70% and error rates above 50% on popular hosting sites [10]. For background on how table data differs from text, see structured vs unstructured AI training data.
How SourceX approaches text-to-SQL data requests
SourceX sources operational datasets, including engineering records and finance and legal workflows, from US companies, and manages the licensing process. Data is sourced on request rather than held in stock, so a request does not guarantee a match. Buyers describe the data, not the businesses; SourceX looks for US companies that hold it, and every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Pricing is agreed per deal. You can submit a text-to-SQL data request as a buyer, or browse the wider tables, time series and transactional data guide and the AI data hub.
License text-to-SQL training data on real schemas
If your model needs schemas, values and gold SQL from real enterprise systems, describe the domains, dialects and record format you need. SourceX follows a Find, Assess, Agree, Transact and Manage process, and nothing is contracted until a supplier agrees. Describe the text-to-SQL data you need.
Sources
- NVIDIA NeMo Data Designer documentation, "Text-to-SQL for Nemotron Super". https://docs.nvidia.com/nemo/datadesigner/latest/dev-notes/text-to-sql-for-nemotron-super
- arXiv, "SQaLe: schema, question and query triples from real database schemas" (2026). https://arxiv.org/html/2602.22223v1
- arXiv, "Can LLM Already Serve as a Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs (BIRD)" (2023). https://arxiv.org/abs/2305.03111v2
- arXiv, "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
- arXiv, "BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation" (2025). https://arxiv.org/pdf/2510.13853
- Hugging Face (AniruddhaAI), "scm-sql". https://huggingface.co/datasets/AniruddhaAI/scm-sql
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.