Schemas, packaging and delivery
Zero-Copy Warehouse Data Sharing for Licensed Datasets
Quick answer
Zero-copy sharing lets a supplier expose live, read-only tables to your warehouse account without shipping files: Snowflake Secure Data Sharing, the open Delta Sharing protocol and BigQuery sharing listings all work this way. It suits licensed operational records you query, sample and refresh. It stops being zero-copy once a training job exports data to object storage, so the license must name those copies and revoking the share will not delete them.
By SourceX Editorial · Updated
What zero-copy actually means between two companies
Zero-copy means the consumer queries the provider's stored data through metadata grants instead of receiving a duplicate. Salesforce's market definition is sharing between data stores "without duplicating or moving" data [5], and the warehouse vendors implement that idea differently. In Snowflake, no data is copied or transferred between accounts when a provider shares a database; the consumer imports the share and every shared object is read-only [1]. In BigQuery, Google's sharing documentation describes subscribing to a listing as creating a read-only linked dataset in your project that points at the publisher's tables, again without a copy.
The phrase hides two caveats that matter for licensed data. First, cross-region and cross-cloud sharing usually does involve a copy, made by the platform on the provider's side; Snowflake's listing auto-fulfillment, for example, provisions Snowflake-managed share areas in remote regions and copies the product there, per Snowflake's provider documentation as of October 2026. Second, nothing stops your own pipeline from materializing the data, which is exactly what most training jobs do.
For bulk file delivery patterns, compare this page with cross-account bucket delivery and the general overview at how licensed data is delivered.
Snowflake, Delta Sharing and BigQuery compared
The three models differ mainly in who must run which platform, how cross-region access works and what controls the provider holds. Snowflake and BigQuery are closed-platform shares: both sides need an account on the same platform (Snowflake can provision reader accounts for consumers without their own). Delta Sharing is an open protocol: Databricks launched it in 2021 as an HTTP/REST API over Delta Lake and Apache Parquet files in S3, ADLS or GCS [2], so a recipient can read with open-source connectors, pandas or Spark rather than a Databricks workspace [4].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Dimension | Snowflake Secure Data Sharing | Delta Sharing (open protocol) | BigQuery sharing (Analytics Hub) |
|---|---|---|---|
| Consumer object | Database created from a share; read-only [1] | Table addressed as share.schema.table [3] | Linked dataset in subscriber project; read-only |
| Recipient requirement | Snowflake account or provider-managed reader account | Any client holding a profile file (endpoint plus bearer token) [3] | Google Cloud project with BigQuery |
| Cross-region / cross-cloud | Listings with auto-fulfillment copy to remote regions | Recipient reads files from provider storage; egress cost and latency apply | Regional; check the listing's region against your datasets |
| Provider export controls | Read-only objects; no native block on CTAS into consumer tables | Token scope and expiry [3] | Egress controls can block copy/export of data and query results |
| Provider usage visibility | Listing usage views (confirm current views in Snowflake docs) | Server-side request logs on the sharing server | INFORMATION_SCHEMA.SHARED_DATASET_USAGE, optional subscriber email logging |
| Typical fit | Both parties on Snowflake | Mixed stacks, Spark or Python training pipelines | Both parties on Google Cloud |
Treat the table as a starting map, not a spec. Shareable object types, region rules and logging views change with releases, so verify each row against the vendor's live documentation before you write it into a technical schedule.
How Delta Sharing works for a recipient outside the provider's platform
A Delta Sharing recipient gets a JSON profile file containing the sharing server endpoint and a bearer token, which may carry an expiration time [3]. The client calls list endpoints such as GET /shares, then asks for a table's metadata and receives short-lived pre-signed URLs to the underlying Parquet files, which it downloads directly from cloud storage [2][3]. That design is why Delta Sharing scales to large tables without a warehouse in the middle.
It also means the "zero-copy" boundary is thin. A Spark job reading the share via delta_sharing.load_as_spark() streams Parquet files to your cluster, and anything you write afterwards is your copy. Store the profile file like a credential (secrets manager, not a notebook), note the token expiry in your runbook, and expect the provider to rotate it at renewal.
Why training pipelines break the zero-copy model
Most training and evaluation stacks read from object storage, not from a SQL engine, so teams export shared tables to Parquet, JSONL or WebDataset shards before a run. Tokenization, deduplication, PII rescans and train/eval splits all produce derived artifacts that persist after the source share is gone. If the license only contemplates querying in place, those exports may fall outside permitted use.
Two practical consequences follow. On BigQuery, a publisher that enables egress controls can block table copy and export, and optionally export of query results, which can make a standard EXPORT DATA to GCS fail; ask before you design the pipeline. On Snowflake, a CREATE TABLE AS SELECT or COPY INTO @stage from the shared database into your own account is technically allowed, so the constraint lives in the license, not the platform.
For format choices on those exports, see Parquet vs JSONL for training data; for keeping record identity across exports, see stable record IDs and join keys.
Worked checklist: before the first export from a share
Illustrative example: invented to show structure; it does not describe an available dataset.
- License names each copy class it permits: query-in-place, exported snapshots, derived features, tokenized shards, model checkpoints.
- Export target is an access-controlled bucket with object versioning and a retention tag such as
license_id=LIC-2026-014. - Every export writes a manifest with share name, table, snapshot timestamp or Delta version, row count and SHA-256 per file.
- Platform egress controls (BigQuery) or token expiry (Delta Sharing) are documented in the runbook with their end dates.
- A lineage record ties each export to the training or eval runs that consumed it.
- Deletion procedure covers exports, derived tables and caches, with evidence suitable for a destruction certificate.
Revocation is an end-of-term control, not a deletion
Revoking a share cuts live access, but it does nothing to data you already exported. In Snowflake the provider removes your account from the share; in Delta Sharing the provider revokes or lets the recipient token expire [3]; in BigQuery the publisher can remove the subscription or listing. Each stops new queries against the provider's storage. None reaches into your buckets, feature stores or checkpoints.
That gap is why exit terms for shared data still need the same mechanics as file deliveries: an inventory of copies, a deletion procedure and, where the contract asks for it, a certificate of data destruction. Plan it before the first export, because reconstructing which tables went where after a year of reruns is slow and error-prone.
What the provider can see, and what that means for you
Providers on these platforms typically see more about consumption than a file sender does. BigQuery exposes INFORMATION_SCHEMA.SHARED_DATASET_USAGE to publishers, and with subscriber email logging enabled it records the principal that ran each job in job_principal_subject, according to Google's documentation as of October 2026. Delta Sharing servers can log each table query, because every read starts with an authenticated API call before the client fetches files through pre-signed URLs [3]. Snowflake offers listing usage views to providers; confirm the current view names and granularity in its documentation.
For a buyer this has three effects. Query patterns can reveal what you are building, so route research queries through a service account rather than individual analysts if the license allows. Usage logs can serve as evidence in a dispute about scope, in either direction. And a provider that sees heavy scans near term end may ask about exports, so keep your copy inventory current.
Freshness, versioning and schema drift on a live share
A live share means the provider's next write is your next read, which helps for recurring purchases and hurts reproducibility. Delta tables carry a version history, so pin training reads to a specific table version or timestamp and record it; on Snowflake and BigQuery, snapshot the shared tables into your own account at a known time and record the snapshot ID. Auto-fulfilled Snowflake listings also refresh on a provider-set schedule, which providers can inspect through LISTING_REFRESH_HISTORY.
Column renames, type changes and dropped fields reach you immediately with no file boundary to catch them. Add schema contract tests on the share before downstream jobs run, and see schema evolution across recurring deliveries and versioning licensed datasets for release-note conventions.
Traceability for model documentation
Shared tables still need dataset-to-model lineage, because regulatory summaries ask what data trained a model. The EU AI Office's template for the public summary of training content under Article 53(1)(d) applies to general-purpose AI model providers [6], and that summary is far easier to produce when every export and snapshot carries a license ID. The pattern is covered in tracing which models trained on which licensed dataset.
If you run a lakehouse, register exported snapshots in the same catalog as the share (Unity Catalog, Polaris or Glue) with tags for license and expiry; the data lakehouse glossary entry explains the storage model. Access policy for the copies belongs in your normal controls, described in access controls for licensed training data.
When to choose a share over a file delivery
Choose a share when both parties already run the same platform or can use Delta Sharing, the data is tabular or semi-structured, and you want ongoing refreshes. Choose a file delivery when you need fixed, checksummed snapshots, the dataset holds large binaries such as audio or documents, or the provider will not allow exports and your training stack cannot read in place. Many deals combine both: a share for exploration and incremental updates, and a manifest-verified snapshot for each training run, as described in manifest and checksum verification.
SourceX sources operational datasets such as support and sales histories, engineering records and finance workflows from US companies, and every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. You can describe the data you need on the buyer page, and the delivery cluster hub and AI data hub cover the rest of the delivery decisions.
Licensing operational records for warehouse and lakehouse teams
SourceX sources operational data from US companies on request, so categories are not inventory and a request does not guarantee a match. Each release is approved by the supplying company, and nothing is contracted until a supplier agrees on pricing and allowed uses in a license. Describe the records and fields you need at sourcex.si/buyers.
Sources
- Snowflake Inc., "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
- delta-io/delta-sharing (Mintlify rendering), "Delta Sharing protocol: REST APIs". https://www.mintlify.com/delta-io/delta-sharing/protocol/rest-apis
- Delta Lake project, "Read Delta Sharing Tables". https://docs.delta.io/delta-sharing/
- Salesforce, "Guide to Zero Copy". https://www.salesforce.com/blog/zero-copy/
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.