Skip to content

Schemas, packaging and delivery

Offline Transfer Appliances for Petabyte-Scale Dataset Delivery

Quick answer

An offline transfer appliance is a provider-owned, encrypted storage device that the destination cloud ships to the data holder, who fills it and returns it for import into the buyer's storage account. As of October 2026, the major-cloud options for new customers are Azure Data Box and Google Cloud Transfer Appliance; AWS closed its Snow Family to new customers in November 2025 [1]. Ship rather than send when a sustained link would take weeks, but budget several weeks of logistics and verify every object against a manifest after import.

By SourceX Editorial · Updated

How an appliance import actually works

An appliance import is an order placed on the destination cloud account, filled at the source site and uploaded by the provider into a bucket or storage account you name. That ordering model is what separates it from a buyer-procured encrypted drive: the cloud provider owns the hardware, runs the upload facility and performs the post-upload wipe, while you own the destination and, ideally, the keys.

The sequence is broadly the same across providers:

  1. Order the device against the destination subscription or project, naming the target storage account (Azure) or Cloud Storage bucket (Google).
  2. Ship the device to the data holder's site, where it is racked or placed on a desk and connected to the local network.
  3. Copy data onto it over the supplier's LAN, typically through NFS or SMB shares exposed by the device; insist on encrypted SMB 3.x or a private VLAN for the copy.
  4. Return the device by courier to the provider's upload facility.
  5. Upload into the named destination, followed by a provider-run sanitization of the device, which should follow NIST SP 800-88 [3].
  6. Verify the imported objects against the source manifest before anyone treats the delivery as complete.

For a licensed dataset, step 3 happens on the supplier's network, not yours. That has practical consequences: the supplier needs rack space, a 10 GbE or faster port, and someone to run the copy, and the device sits on their premises for days. Agree on who does each of these before the order is placed.

Network transfer or appliance: the arithmetic that decides it

The decision rests on effective throughput: at a sustained 1 Gbps, 100 TB takes about nine days, and at 10 Gbps a petabyte takes the same nine days. The raw formula is bytes × 8 ÷ bits per second, so 100 × 10^12 × 8 ÷ 10^9 is 800,000 seconds, or 9.3 days, before protocol overhead, retries and the share of the link you are actually allowed to use.

Illustrative example: invented to show structure; it does not describe an available dataset.

Corpus size1 Gbps at 70% efficiency10 Gbps at 70% efficiencyAppliance (provider cycle)
50 TB~6.6 days~0.7 daysNetwork wins
250 TB~33 days~3.3 daysAppliance only if source uplink is ≤1 Gbps
1 PB~132 days~13 daysClose call at 10 Gbps; appliance wins below that
3 PB~397 days~40 daysAppliances, multiple devices in parallel

The fourth column is not a speed number. Plan on a multi-week appliance cycle end to end (confirm the current figure in the provider's documentation), which covers shipping both ways, local copy and facility upload. An appliance therefore beats the network only when the network estimate exceeds that cycle by a wide margin, or when the supplier's uplink, firewall policy or egress budget rules out a sustained transfer. The egress question is covered in who pays for data egress; tuned parallel transfers are covered in moving multi-terabyte datasets over the network.

Two failure modes skew the math. First, the copy to the appliance is itself a network transfer over the supplier's LAN, and millions of small files (JPEG frames, short WAV clips) can cut throughput far below the port speed. Second, appliances do not fix a slow source: if the corpus sits on a NAS that reads at 300 MB/s, a 300 TB fill takes about 11.5 days regardless of the device.

Current appliance options as of October 2026

Among the three largest clouds, only two appliance lines are orderable by new customers in October 2026, so the comparison is short, and capacities, models and cipher details should be checked on the live provider pages before ordering.

Question to answerAzure Data Box familyGoogle Cloud Transfer ApplianceAWS Snowball Edge
Can a new customer order it?Yes, per Microsoft documentationYes, per Google documentationNo; existing customers only since 7 Nov 2025 [1]
Where does data land?Azure Storage account named on the orderCloud Storage bucket named on the orderAmazon S3 (existing customers only)
How does the supplier copy?SMB or NFS shares on the deviceNFS mount, or SCP/SSHNot applicable for new orders
What protects data at rest?Device encryption with credentials held in the Azure order; check whether double encryption is offeredDevice encryption; check whether customer-managed Cloud KMS keys are supported for your modelNot applicable for new orders
What happens after upload?Provider wipe recorded against the order; confirm the NIST 800-88 revision attestedProvider wipe; ask for a wipe certificate and the NIST 800-88 revisionNot applicable for new orders
Which sizes exist?Several device sizes plus SSD disk kits; verify current usable capacityRackable and freestanding models; verify current usable capacityNot applicable for new orders

AWS documentation now points new customers to AWS DataSync, an online service, rather than a device [2]. If your training platform is on AWS and the corpus is too large for the network, the realistic paths are a network transfer over a provisioned link, an appliance import into another cloud followed by a cloud-to-cloud copy, or a buyer-procured encrypted drive. Each of those has its own custody and cost profile; do not assume a Snowball order is available because a runbook from 2024 mentions it.

Provider pages for these devices change often and have at times listed inconsistent capacities and cipher strengths across models, so read the current security page for the exact model you order rather than a third-party summary table. NIST finalized SP 800-88 Rev. 2 [3]; if your security team requires the newer revision on a wipe certificate, confirm which revision the provider currently attests to.

Keys, encryption and chain of custody in transit

The appliance protects data in transit only as well as its key handling, so decide who holds the device credentials and keys before the device ships. Where the provider supports customer-managed keys in its key service, the buyer can hold the key so the upload facility needs the buyer's grant to decrypt. Where device passwords live in the cloud portal under the order, restrict that order to a small named group in your account.

For a licensed dataset the cleanest pattern is: the buyer places the order and holds the keys, the supplier receives a time-limited device credential for the copy window only, and the credential is not reused for later deliveries. Avoid sending device passwords by email or chat; pass them through the same private channel used for other deal documents, and see encrypting dataset deliveries for key-exchange patterns.

Chain of custody should be a written record, not an assumption. Capture at minimum:

  • Order ID, device serial and the destination account or bucket.
  • Who opened the device for copying, when, and from which host.
  • Copy start and end times, file and byte counts, and the manifest hash.
  • Courier tracking numbers for both legs and tamper-evident seal numbers if used.
  • Upload completion time, provider copy logs and the erasure record or wipe certificate, referencing NIST SP 800-88 [3].

If the corpus contains recordings of people, such as meeting video or call audio, ask whether faces, voices or names were treated before the copy. An appliance does not change your privacy obligations, and biometric statutes can apply to voiceprints and face geometry; see biometric data in AI training datasets.

Packaging a media corpus so the appliance copy stays fast

Shard small files into large archives before the copy, because per-file overhead on NFS and SMB is the most common reason appliance fills run slower than planned. A video or image corpus of tens of millions of objects copies far faster as tar shards; the WebDataset convention, with shards often around 1 GB and full datasets reaching multiple terabytes, is widely supported by training loaders [4].

Practical rules for the supplier's copy job:

  • Keep one top-level directory per delivery version (for example release=2026-10-01/) so the import lands in a predictable prefix.
  • Write sidecar metadata (clip IDs, durations, codecs, consent or license flags) as Parquet or JSONL next to the shards, not embedded inside them; see packaging video datasets.
  • Generate the manifest from the source volume before the copy, not from the appliance afterward, so it reflects what the supplier intended to send.
  • Check object-name rules on the destination: cloud object stores handle path separators, maximum key length and special characters differently from POSIX file systems, and failures there show up only at upload time.

Verifying the import against the manifest

An appliance delivery is complete only when every object in the destination matches the supplier's manifest by path, size and checksum. Provider upload logs confirm that the facility copied what was on the device; they do not confirm that the device held what the supplier meant to deliver.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"path": "release=2026-10-01/video/shard-000412.tar", "bytes": 1073741312, "sha256": "9f2c...e41a", "records": 1840, "source_volume": "nas-02"}
{"path": "release=2026-10-01/meta/clips-000412.parquet", "bytes": 2219904, "sha256": "41bb...0c7d", "records": 1840, "source_volume": "nas-02"}

A verification pass that holds up in procurement review:

  1. Count objects under the destination prefix and compare with manifest line count.
  2. Compare byte sizes per object; mismatches usually indicate truncated copies.
  3. Recompute SHA-256 on the destination for every object, or on a statistically chosen sample if compute cost is prohibitive and the provider supplies per-file MD5 or CRC values for the rest.
  4. Reconcile record counts in sidecar metadata against shard contents.
  5. Record the result, sign it off and only then confirm acceptance to the supplier.

The full method, including how to handle partial failures and re-sends, is on dataset manifests and checksums. After acceptance, apply the same access controls you would for any licensed training data: lock the import bucket to the training service accounts that the license allows, as described in access controls after delivery.

When an appliance is the wrong tool

An appliance is a poor fit for ongoing purchases, small increments and data that must stay in the supplier's environment. Monthly refreshes of a few terabytes are better handled as incremental network deliveries (see incremental vs full refresh deliveries), and supplier-to-bucket copies are often simpler with cross-account cloud bucket delivery. Appliances are also the wrong choice when the supplier cannot accept a third-party device on site, which some regulated and secure facilities refuse outright.

Use this quick decision checklist before ordering:

  • Network estimate at realistic efficiency exceeds roughly twice the provider's appliance cycle.
  • The supplier has rack space or desk space, a 10 GbE or faster port and an operator for the copy.
  • The destination cloud is Azure or Google Cloud, or a second cloud hop is acceptable.
  • Keys and device credentials are held by the buyer, with a defined handover window.
  • A manifest with SHA-256 per object exists before the copy starts.
  • The license and the delivery specification name the delivery method; use the technical delivery specification template to record it.

Where SourceX fits in large deliveries

SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and manages the commercial process through licensing and ongoing purchases. Delivery runs through private, access-controlled workflows, never email attachments, and only after an executed agreement and supplier approval. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, so the delivery method sits inside that license rather than being improvised later. For the broader set of formats and transfer methods, start at the dataset delivery hub or read how licensed data is delivered; to describe a corpus you need, submit a buyer request.

Bringing a very large licensed corpus into your cloud

If you are planning a multimodal or speech corpus large enough to need an appliance, describe the data you need rather than the businesses that might hold it. SourceX looks for US businesses that hold the described data, assesses data and licensing permissions, and nothing is contracted until a supplier agrees; a request does not guarantee a match. Describe the large dataset you need.

Sources

  1. Amazon Web Services, "AWS Snowball Edge availability change" (2025). https://docs.aws.amazon.com/snowball/latest/developer-guide/snowball-edge-availability-change.html
  2. Amazon Web Services, "Using S3-compatible storage on Snowball Edge (availability notice)". https://docs.aws.amazon.com/snowball/latest/developer-guide/s3compatible-on-snow.html
  3. NIST Computer Security Resource Center, "NIST SP 800-88 Rev. 2, Guidelines for Media Sanitization" (2025). https://csrc.nist.gov/pubs/sp/800/88/r2/final
  4. Hugging Face Hub documentation, "WebDataset". https://huggingface.co/docs/hub/datasets-webdataset
  5. Google Cloud, "Specifications (Transfer Appliance)". https://docs.cloud.google.com/transfer-appliance/docs/4.0/specifications?authuser=5
  6. Microsoft, "Azure Data Box for data transfer". https://azure.microsoft.com/products/databox/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data