Code and software engineering data
Jupyter Notebook Datasets for Data Science Agents
Quick answer
A Jupyter notebook dataset for agent training is a licensed set of .ipynb files in which code cells, their saved outputs and the markdown reasoning around them stay together, ideally paired with the report or decision the analysis fed. Public corpora drawn from GitHub and Kaggle are useful for pretraining, but they rarely show how company data teams actually explore messy warehouse tables. Enterprise notebooks fill that gap, provided outputs are scrubbed item by item, execution order is preserved and duplicate templates are removed.
By SourceX Editorial · Updated
Why notebooks need their own sourcing spec, separate from repository code
Notebooks interleave three kinds of content (code, outputs and narrative) in one JSON document, so the preparation rules for repository code do not transfer cleanly. The nbformat specification, published by Project Jupyter, stores each notebook as JSON with a list of cells; code cells carry source, metadata, an execution_count and an outputs array, and every output has an output_type of stream, display_data, execute_result or error. Rich outputs are keyed by MIME type, so a single cell can hold text/html tables, image/png charts and text/plain fallbacks at once.
That structure is what makes notebooks valuable for data science agents. The model sees the hypothesis in markdown, the pandas or SQL call, the resulting dataframe preview and the analyst's next move. Repository code shows the finished pipeline; a notebook shows the search that led to it. For finished, versioned code with full history, see the owner page on proprietary code datasets with full Git history, and for production pipelines see dbt, Airflow and Spark pipeline code.
What public notebook corpora already cover
Public research corpora give scale and reproducible baselines, but they skew toward tutorials, competitions and student work. JetBrains Research published a corpus of 847,881 properly licensed notebooks as an 18.4 GB PostgreSQL dump [1]. KGTorrent pairs Kaggle Python notebooks with a MySQL database of metadata and user activity [2], and DistilKaggle reports more than 12 million code cells and metrics for roughly 542,000 Kaggle notebooks from 2015 to 2023 [3]. Research on computational notebooks also distinguishes exploration notebooks from explanation notebooks, a useful lens when you decide which kind your agent needs [4].
What these corpora lack is the enterprise context: proprietary schemas, Snowflake or BigQuery connectors, internal helper libraries, stakeholder questions and the memo that closed the loop. Kaggle notebooks start from a clean CSV and a known target; a company analyst starts from a vague request and twelve joined tables. As of October 2026, benchmarks such as Spider 2.0 explicitly target that gap with 632 problems built on databases from real applications [6].
Output scrubbing: the main privacy and confidentiality risk
Saved outputs are where notebooks leak data, so treat every output as a data record, not as code. A df.head() renders real rows into text/html; a print(customer) lands in a stream output; a matplotlib chart embeds a base64 PNG that may show account names on an axis; an error output's traceback can expose file paths, hostnames or connection strings.
Agree a per-output policy before delivery rather than stripping everything. Tools like nbstripout remove all outputs, which is safe but discards the observation step agents most need to learn from. A better pattern is to classify each output: keep aggregate statistics and charts of non-sensitive metrics, replace row-level previews with schema-faithful synthetic rows, and drop images that cannot be inspected. Run secret scanning across cell sources and outputs as well; the approach on secrets removal in code datasets applies to notebooks, since connection strings and API keys can sit in a setup cell.
Execution order, reruns and evaluation without the data
Execution counts are the cheapest signal of whether a notebook tells a coherent story, so require them to be retained. If cells show counts 3, 1, 7, 2, the analyst ran them out of order and the file will likely fail a top-to-bottom rerun. Keep the counts and compute an ordering flag per notebook, so you can filter strictly linear notebooks for SFT and keep messy ones for robustness evaluation.
Re-execution usually requires warehouse access the buyer will never receive. Plan evaluation around three options:
- Static grading: score the agent's next cell or narrative against the saved outputs and the analyst's actual next step.
- Synthetic stand-in data: the supplier provides tables with the same schema and plausible distributions, so cells can run in a sandbox.
- Report-level grading: compare the agent's conclusion to the final memo, using a rubric built with the domain experts who review evaluation rubrics.
Pairing notebooks with the reports they fed
The notebook-to-report link is what turns analysis code into supervision for judgment, so ask for it explicitly where the company allows. A pair might be a churn exploration notebook plus the slide or decision memo it produced, which lets you train an agent to go from evidence to recommendation and grade it on whether it reached the same decision. Spreadsheet models are often the other half of that chain; the owner page on spreadsheet and financial model datasets covers that format.
Reports frequently carry more sensitive content than notebooks, including revenue figures and named customers. Expect suppliers to release them only with redaction, and specify the link fields (notebook path, report ID, date) so pairs survive de-identification.
Deduplication and template detection
Shared team folders are full of copied notebooks, so near-duplicate removal matters more here than in most code. Teams clone a "standard EDA" template, change one table name and save it forty times. Research on language-model training data shows that near-duplicates are common and that deduplication reduces verbatim memorized output [5]. For notebooks, hash normalized cell sources (strip outputs, whitespace and execution counts), then cluster at the notebook level with MinHash over cell sequences. Keep one exemplar per template cluster, and check against public corpora too, using the methods in code benchmark contamination.
Notebook dataset request template
A precise request is what separates a usable notebook corpus from a folder dump. Adapt the fields below; the broader code dataset request specification covers languages, history and build requirements.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
| Notebook types | Exploratory analysis, model prototyping, ad hoc reporting | Exploration and explanation notebooks train different skills [4] |
| Kernels and libraries | Python 3 with pandas, scikit-learn, SQL magics; R optional | Matches your agent's tool surface |
| Data sources referenced | Warehouse SQL (Snowflake, BigQuery, Redshift), parquet, CSV | Tests realistic connector and join behavior |
| Output policy | Aggregates and charts kept; row-level previews replaced; tracebacks path-scrubbed | Main confidentiality control |
| Execution metadata | execution_count retained; ordering flag computed | Detects non-linear notebooks |
| Paired artifacts | Final memo or slide per notebook where allowed | Enables report-level grading |
| Deduplication | Template clusters collapsed to one exemplar | Reduces memorization [5] |
| Delivery format | Original .ipynb plus one JSON Lines record per cell | Raw fidelity plus training-ready rows [8] |
| Documentation | Data card covering source teams, time span, prep steps | Supports review and reuse [7] |
A per-cell JSON Lines record (one UTF-8 JSON value per line [8]) might carry notebook_id, cell_index, cell_type, execution_count, source, output_types, output_policy_applied and report_id. Keeping the original .ipynb alongside lets you rebuild context windows differently later. See dataset delivery formats and schemas for transfer options.
Sample checks before you license
A short structured review of a sample catches most problems before contract. Check that nbformat versions validate, that outputs follow the agreed policy in every MIME type (including images), that execution counts survived, that markdown cells contain real reasoning rather than empty headers, and how many notebooks collapse into template clusters. Ask how analyst names in metadata and comments were handled. The general method is in evaluating a code dataset sample, and the code and software engineering data hub maps adjacent formats such as developer session recordings for coding agents.
How SourceX approaches notebook data requests
SourceX sources operational datasets, including engineering records and documents, from US companies on request; it does not hold notebooks in stock, and a request does not guarantee a match. You describe the data you need, not the businesses, and SourceX looks for US companies that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect. You can start a request on the SourceX buyer page.
Request enterprise Jupyter notebook datasets
If your data science agent needs real analyst notebooks with outputs and paired reports, describe the notebook types, output policy and delivery format you need. SourceX runs Find, Assess, Agree, Transact and Manage, nothing is contracted until a supplier agrees, and delivery follows an executed license that defines records, uses, term and delivery. Describe your notebook data needs to SourceX.
Sources
- Zenodo, "Dataset of Jupyter notebooks (JetBrains Research, MSR'22)" (2022). https://zenodo.org/records/6383115
- arXiv, "KGTorrent: A Dataset of Python Jupyter Notebooks from Kaggle" (2021). https://arxiv.org/abs/2103.10558v1
- Zenodo, "DistilKaggle" (2023). https://zenodo.org/records/10317389
- UC San Diego Library Digital Collections, "Data from: Exploration and Explanation in Computational Notebooks". https://library.ucsd.edu/dc/collection/bb6931851t
- arXiv (Lee et al.; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- arXiv, "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
- arXiv (Pushkarna, Zaldivar, Kjartansson; FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.