Skip to content

Schemas, packaging and delivery

Who Pays for Data Egress? Transfer Costs in Dataset Delivery

Quick answer

Egress is billed by the cloud that holds the bytes when they leave a region, a provider or the internet edge, so by default the supplier's account pays to send a dataset to you. That default can be moved: requester-pays buckets bill the downloader, in-place sharing avoids copies, and same-region delivery avoids most transfer charges. For a 100 TB corpus the difference is material, so write the payer, route and region into the delivery specification before anything moves.

By SourceX Editorial · Updated

How cloud egress is billed when a dataset changes hands

Cloud transfer charges attach to the account that owns the source storage, not to whoever benefits from the data. On Amazon S3 the bucket owner pays for storage and, by default, for data transfer out of the bucket [1]; Google Cloud Storage bills the project that contains the bucket unless a requester-pays request names another billing project, and Azure bills data transfer out of its regions and between regions, with rates that vary by source region, routing and negotiated agreement.

Three cost classes matter for a licensed delivery, and they are priced very differently:

  • Same region, same provider. Copies between accounts in one region usually carry only request (API operation) charges, not per-GB transfer. This is the cheapest route and the one to aim for.
  • Cross-region, same provider. Inter-region replication or copy is billed per GB by the provider, typically below internet rates but not free.
  • Cross-cloud or to on-premises. Data leaving the provider for the internet or another cloud is billed at internet egress rates, the most expensive class, with volume tiers.

Request charges are small per call but add up when a corpus is stored as millions of small objects. A text or document corpus delivered as 40 million individual JSON files costs far more in PUT, GET and LIST operations than the same bytes packed into a few thousand Parquet files or tar shards; see WebDataset tar shards for a packaging pattern that also helps throughput.

Requester-pays buckets: shifting download costs to the buyer

A requester-pays bucket makes the downloading account pay request and data transfer charges while the owner keeps paying for storage [1]. Google Cloud Storage offers an equivalent setting. It is the cleanest way to move egress to the buyer without invoicing transfer as a separate line, and it suits suppliers who hold data in their own cloud and do not want an open-ended network bill.

The prerequisites differ by provider and break pipelines when missed:

  • Amazon S3. Anonymous access is not allowed, every request must be authenticated, and SOAP requests are not supported [1]. Each request must carry the x-amz-request-payer: requester header (in the AWS CLI, --request-payer requester); tools that omit it receive 403 errors. Cross-account access still needs both a bucket policy from the owner and an IAM policy on the buyer's role [2].
  • Google Cloud Storage. Requests must name a billing project (--billing-project or the userProject parameter); without one they fail with a UserProjectMissing error. Storage charges stay with the bucket's project, and some managed services, such as Cloud SQL import and export, cannot read from requester-pays buckets. Check Google's current Requester Pays documentation [9] for the full list.
  • Billing visibility. Requester charges may not appear as a separate line item, so if the billing project also holds its own buckets, delivery costs can blend into normal download charges. Use a dedicated billing project or AWS cost allocation tags if finance needs a clean per-deal figure.

Requester pays only moves the bill. If your compute is in another region or another cloud, you still pay the full cross-region or internet rate, now on your own invoice. For bucket policy, role and KMS key patterns, see cross-account bucket delivery.

Why same-region delivery is the default to negotiate

Placing the delivered copy in the same provider and region as your training cluster removes per-GB transfer from both the delivery and every later read. The common, expensive failure is a single delivery that is cheap once and then read repeatedly across regions: a team pulls a 100 TB corpus from us-east-1 into GPU capacity in us-west-2 for each training run, and the cross-region bill recurs per epoch or per experiment.

Ask early where the supplier's data physically sits. If it is in a different region from your compute, the decision is who pays one copy into your region, not who pays every read. One cross-region copy into your own bucket, verified against a manifest, is almost always cheaper than streaming reads; dataset manifests and checksums covers verification.

Warehouse-native sharing changes the arithmetic. Snowflake's Cross-Cloud Auto-Fulfillment replicates a listing into the consumer's region and bills the provider for replication and transfer [3], with an Egress Cost Optimizer feature aimed at reducing that provider cost [4]. BigQuery sharing lets subscribers query linked datasets in place inside the provider's region [5], and Delta Sharing serves Delta Lake and Parquet tables from S3, ADLS or GCS over a REST protocol [6]. In each model, ask which side pays when the consumer sits in another region or cloud, because the docs place real transfer costs somewhere [3].

Cross-cloud transfer: what moving 100 TB actually involves

Moving 100 TB between clouds is billed as internet egress by the source provider, plus request charges at both ends and any compute you run to drive the copy. Destination ingress is generally not charged per GB, but the source-side bill is the dominant number. Published rates are tiered by monthly volume and differ by source region, so the same 100 TB costs a different amount leaving North America than leaving Asia or South America.

Read the worked example below as a budgeting method, not a quote. Check the current price sheet for your source region and any negotiated discounts before committing a number.

Illustrative example: invented to show structure; it does not describe an available dataset.

Line itemAssumptionFormulaEstimate
Internet egress from source cloud100 TB = 102,400 GB at an assumed blended $0.05/GB102,400 x 0.05about $5,120
Source GET requests2 million objects at an assumed $0.40 per million2 x 0.40under $1
Destination PUT requests2 million objects at an assumed $5.00 per million2 x 5.00about $10
Same delivery as 40 million small filesRequest charges scale 20x(40 x 0.40) + (40 x 5.00)about $216
Repeated cross-region reads (avoidable)10 training runs read the corpus across regions at an assumed $0.02/GB10 x 102,400 x 0.02about $20,480

The last row is the lesson: recurring cross-region reads can cost several times the one-time delivery. Tooling choice (rclone, gsutil/gcloud storage, AzCopy, Storage Transfer Service, DataSync) affects throughput but not the provider's per-GB rate; see network transfer tools and tuning.

Egress waivers and switching rules: check current terms

Provider egress waivers exist but are narrow, request-based and subject to change, so treat them as a negotiation input, not a budget assumption. As of October 2026, AWS, Google Cloud and Azure [10] have each announced programs that waive internet egress for customers moving their data off the platform. As publicly described, these are requested through support or an account team, reviewed case by case and often paid as credits rather than applied automatically.

These programs target customers migrating away from a provider, not routine transfers to a business partner. A supplier sending you a dataset while staying on its cloud is unlikely to qualify. Read the current eligibility text, since the terms have changed since launch and may change again.

In the EU, the Data Act (Regulation (EU) 2023/2854) includes cloud-switching provisions that phase out switching charges, including data egress charges, for customers changing data processing providers [8]. As of October 2026, those provisions govern provider switching, not commercial data licensing, so they do not by themselves make a supplier-to-buyer delivery free. Confirm their application to your contracts with counsel.

When shipping physical media beats the network

Shipping drives or an appliance wins when network time or egress cost exceeds courier, handling and import fees, which usually means hundreds of terabytes or limited uplink bandwidth. At a sustained 1 Gbps, 100 TB takes roughly nine days of continuous transfer; at 10 Gbps it is under a day, and egress, not time, becomes the deciding cost.

Plan around current availability. AWS Snowball Edge has been closed to new customers since November 7, 2025, and AWS no longer offers any Snow Family device to new customers [7]. Other options include provider appliances still on sale, encrypted commodity drives with a chain-of-custody record, and colocation cross-connects. See offline transfer appliances and encryption and key exchange for the security side.

Naming the payer in the delivery specification

Every transfer charge should have a named payer, a named route and a cap before the first byte moves. Disputes typically arise when a supplier's finance team discovers an unplanned egress line weeks after delivery, or when a buyer finds that "delivered to your bucket" meant a cross-region copy billed to the buyer's account.

Illustrative example: invented to show structure; it does not describe an available dataset.

transfer_costs:
  source_location: { provider: aws, region: us-east-1, account_owner: supplier }
  destination: { provider: aws, region: us-east-1, account_owner: buyer }
  route: same_region_cross_account_copy   # alternatives: requester_pays, share_in_place, cross_cloud, physical_media
  egress_payer: buyer                     # supplier | buyer | split
  request_charges_payer: buyer
  requester_pays_enabled: true
  billing_reference: "buyer cost allocation tag: delivery-2026-q4"
  estimated_volume_gb: 102400
  object_packaging: parquet_and_tar_shards_256MB_to_1GB
  cost_cap_usd: "to be agreed; stop and confirm if exceeded"
  incremental_deliveries: same_route_and_payer
  retransmission_on_checksum_failure: supplier_pays
  waiver_or_credit_programs: none_assumed

Two clauses deserve attention. Retransmission after a failed checksum should be charged to whoever caused the failure, and ongoing purchases should repeat the same route and payer so that incremental deliveries do not reopen the question each month. Fold the block into your delivery specification template and the wider delivery formats and transfer hub.

Keep transfer costs separate from the price of the data itself. The license fee reflects the dataset's value and rights; egress reflects where bytes sit. For the commercial side, see what drives the price of licensed enterprise data.

A pre-transfer checklist for ML platform leads

Before approving a large delivery, confirm these items in writing:

  • Source provider, region and account owner are known.
  • Destination region matches the training compute, or one landing copy is planned.
  • Route is chosen: same-region copy, requester pays, in-place share, cross-cloud or physical media.
  • Payer is named for egress, requests and retransmission.
  • Requester-pays prerequisites are tested with your tooling (headers, billing project, IAM) [1][2].
  • Object packaging avoids millions of tiny files.
  • Any waiver is confirmed in writing by the provider, not assumed.
  • A manifest with checksums travels with the data.
  • A cost cap and stop condition are agreed.

Buyers sourcing operational data from US companies can describe the dataset and delivery constraints to SourceX at the start, so region and transfer questions surface during scoping. More delivery guidance sits in the AI data hub.

Planning delivery of a licensed dataset with SourceX

SourceX sources operational datasets from US companies on request and manages the commercial process, from assessing data and licensing permissions to agreeing pricing and allowed uses in a license. Each dataset is delivered under a license that defines records, uses, term and delivery, through private, access-controlled workflows only after an executed agreement and supplier approval. Describe the data you need on the SourceX buyers page.

Sources

  1. Amazon Web Services (Amazon S3 User Guide), "Using Requester Pays general purpose buckets for storage transfers and usage". https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
  2. Amazon Web Services (Amazon S3 User Guide), "Walkthroughs that use policies to manage access to your Amazon S3 resources". https://docs.aws.amazon.com/AmazonS3/latest/userguide/example-walkthroughs-managing-access.html
  3. Snowflake (documentation), "Auto-fulfillment costs". https://docs.snowflake.com/en/collaboration/provider-understand-cost-auto-fulfillment
  4. Snowflake (documentation), "Optimizing data transfer costs with Egress Cost Optimizer". https://docs.snowflake.com/en/collaboration/provider-listings-auto-fulfillment-eco
  5. Google Cloud (BigQuery documentation), "Introduction to BigQuery sharing". https://docs.cloud.google.com/bigquery/docs/analytics-hub-introduction
  6. Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing". https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
  7. Amazon Web Services (Snowball Edge Developer Guide), "AWS Snowball Edge availability change". https://docs.aws.amazon.com/snowball/latest/developer-guide/snowball-edge-availability-change.html
  8. EUR-Lex (Official Journal of the European Union), "Regulation (EU) 2023/2854 (Data Act)". https://eur-lex.europa.eu/eli/reg/2023/2854/oj
  9. Google Cloud, "Requester Pays (Cloud Storage)". https://docs.cloud.google.com/storage/docs/requester-pays
  10. Amazon Web Services, "Free data transfer out to internet when moving out of AWS". https://aws.amazon.com/blogs/aws/free-data-transfer-out-to-internet-when-moving-out-of-aws/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data