Schemas, packaging and delivery
Moving Multi-Terabyte Datasets Over the Network: Tools and Tuning
Quick answer
To transfer large datasets between clouds reliably, use a managed service such as Google Storage Transfer Service for bucket-to-bucket moves above roughly 1 TiB [5], rclone when you need one tool across many backends and fine control, and Globus or a UDP-based accelerator for high-latency research or on-premises links. Whichever you choose, tune parallelism to the object size mix, make the job resumable, and verify a checksum manifest at the destination instead of trusting exit codes.
By SourceX Editorial · Updated
This page is for the engineer who receives a licensed delivery and has to land it in the team's own storage. Channel setup lives in cross-account bucket delivery, and drive shipment in offline transfer appliances; this page covers the network move itself. The wider context is in the delivery formats and transfer hub.
Which transfer tool fits which move
The right tool depends on the source and destination pair, who operates the job, and the link between them. Managed services remove the need to run workers but only cover the endpoints they support. Sync tools run anywhere but make you responsible for workers, retries and memory. Research-network tools excel on long, lossy paths where TCP throughput drops.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Situation | Reasonable first choice | Why | Watch out for |
|---|---|---|---|
| S3 bucket to GCS bucket, more than 1 TiB | Storage Transfer Service, agentless | Managed, scheduled, retries built in; Google positions it for transfers above roughly 1 TiB [5] | AWS egress and request charges on the source side |
| S3-compatible store (MinIO, Ceph, Wasabi) or on-premises NAS to GCS | Storage Transfer Service, agent-based | Agents run in Docker in an agent pool you place on your network, so you choose the route | You run and patch the agents; sizing the pool is on you |
| Mixed backends (S3, GCS, Azure Blob, R2, SFTP) or one-off copies | rclone copy or sync | One binary, dozens of backends, server-side copy where the provider allows it | Memory per worker, list cost on huge prefixes, verification is a separate step |
| Many small files within one cloud | Provider-native batch tools (S3 Batch Operations, gcloud storage cp with parallelism, AzCopy) | Higher request concurrency, no data leaving the provider | Request-rate limits per prefix |
| University or HPC endpoint, long round-trip time | Globus | Managed endpoints, automatic retry and checksum-based integrity checking after transfer | Both sides need endpoints (collections) configured |
| Intercontinental link with packet loss, commercial setting | UDP-based acceleration (IBM Aspera FASP and similar) | Throughput less sensitive to latency and loss than single TCP streams | Licensing cost, firewall rules for UDP ports |
| More bytes than your link can move in the window | Offline appliance or encrypted drives | Physics beats tuning | AWS Snowball Edge is closed to new customers as of November 2025 [1]; AWS points new customers to DataSync [3] |
Ask the supplier which endpoints they can expose before you pick. A supplier who can only grant read access to an S3 prefix with Requester Pays enabled changes the math, since you then pay request and download charges and anonymous access is refused [2]. Cost allocation is covered in who pays for data egress.
Managed cloud transfer: Storage Transfer Service from S3 to GCS
Storage Transfer Service is the low-effort path when both ends are object stores it supports. For an Amazon S3 source, you grant read access on the bucket, and two Google identities need permissions: the user creating the job and the Google-managed service agent, project-PROJECT_NUMBER@storage-transfer-service.iam.gserviceaccount.com. For a licensed delivery, that means the supplier's bucket policy or your assumed role must allow s3:ListBucket and s3:GetObject on exactly the delivery prefix.
Google documents three egress routes for S3 sources: the default agentless path where AWS charges egress, routing through a CloudFront distribution, and a Google-managed private network where you pay Google a per-GiB rate instead of S3 egress but may still incur AWS operation charges for LIST and GET calls, according to Google Cloud documentation as of October 2026 [5]. Pick the route before the job runs, because egress is usually the largest line item on a multi-terabyte move. Use include prefixes rather than copying the whole bucket so you do not pull files outside the licensed scope.
Practical settings worth knowing:
- Overwrite policy. Set the job to skip objects that already exist with matching metadata; this is how reruns resume after a failure.
- Delete options. Leave "delete from source" off; you rarely own the supplier's bucket, and "delete objects in destination not in source" can wipe earlier versions you meant to keep.
- Logging. Enable transfer logs to Cloud Logging so you have a per-object record of successes and failures for the delivery file.
- Schedules. For incremental deliveries, a recurring job with a
last modified sincefilter avoids relisting what you already hold.
On AWS, AWS DataSync plays a similar role for NFS, SMB, HDFS and object stores, and Azure offers AzCopy and Azure Storage Mover. The same questions apply: which identity reads the source, which network path carries the bytes, and what record the service keeps of each object.
rclone vs managed transfer services
rclone wins on reach and control; managed services win on operations. With rclone you can copy from an SFTP server to Azure Blob to R2 using one syntax, tune every buffer, and run the job on a VM sitting in the destination region so the expensive hop is cloud-to-cloud. The cost is that you own the worker: its memory, its network, its restarts and its logs.
The settings that matter on large object-store moves:
--transferssets how many files move at once. Raise it for many medium files.--s3-upload-cutoffsets the size at which uploads switch to multipart, and--s3-chunk-sizesets the part size used for multipart uploads.--s3-upload-concurrencysets how many parts of one file upload in parallel. Raise it for a few huge files.--checkerscontrols how many listing and comparison operations run in parallel; on prefixes with millions of keys, listing can dominate wall time.--fast-listtrades memory for fewer list calls on backends that support recursive listing.
Memory scales roughly as transfers times upload concurrency times chunk size, because each in-flight part is buffered. Eight transfers with eight concurrent 128 MiB parts can buffer several gigabytes, which is a common reason rclone workers are killed by the OOM killer mid-job. Check the current flag defaults on rclone.org [1] before relying on them, since they change between releases.
One misunderstanding causes most silent failures: rclone's --checksum flag changes how rclone decides which files to copy (hash comparison instead of size and modification time); it does not by itself prove the bytes landed intact. Run rclone check after the copy, or generate a hash manifest with rclone hashsum at the source and check it at the destination.
Parallelism, multipart sizing and request-rate limits
Throughput on object stores comes from parallel streams, not a single fast connection. Each multipart upload is split into parts that upload independently, and S3 allows at most 10,000 parts per upload, so part size must grow with object size. A 1 TiB object split into 64 MiB parts would need over 16,000 parts and fail; 128 MiB parts fit.
Use this worksheet before you start a large job.
Illustrative example: invented to show structure; it does not describe an available dataset.
Delivery profile
total bytes: 38 TiB
object count: 1.9 million
size mix: 92% < 8 MiB (JSONL shards, sidecars), 8% 2-40 GiB (tar shards, video)
Link and worker
worker: VM in destination region, 25 Gbit/s NIC, 64 GiB RAM
round-trip to source: ~70 ms
Tuning decisions
small files: --transfers 64, --checkers 64 (request-bound, not byte-bound)
large files: part size 256 MiB (40 GiB / 256 MiB = 160 parts, far below 10,000)
--s3-upload-concurrency 8
memory check: 64 transfers x 8 parts x 256 MiB is far too much for 64 GiB
-> split into two jobs: small-file pass, then large-file pass
prefix layout: keys spread across many prefixes to stay under per-prefix request rates
Expected bottleneck: object count (listing and PUT rate), not bandwidth
Request-rate limits bite before bandwidth when objects are small. S3 scales request capacity per prefix, so a delivery with millions of objects under one prefix sees throttling (HTTP 503 Slow Down) that looks like a network problem. Ask suppliers to shard by a hashed or dated prefix, or repackage small files into larger tar or Parquet shards; see Parquet vs JSONL for training data. Exponential backoff in the client is required, not optional.
Verify checksums after transfer, not just exit codes
A zero exit code means the tool believes it finished; it does not mean every byte matches the supplier's copy. Truncated parts, a worker killed between batches, a sync that skipped files with matching size, or an encoding layer that rewrote content can all end with success status. The only defensible acceptance test is a manifest of per-file hashes produced at the source and checked at the destination.
Know what each store's native checksum means. Amazon S3 can compute and store additional checksums such as CRC32C or SHA-256 on upload, including for multipart uploads, and returns them on request [4]. A multipart object's ETag is not the MD5 of the whole file, so comparing ETags across clouds or against an md5sum fails for large objects. Google Cloud Storage stores CRC32C on objects but omits MD5 for composite objects, so CRC32C is usually the common ground between S3 and GCS.
A verification runbook that holds up in an audit:
- Receive the supplier's manifest (path, size, SHA-256 or CRC32C) before the transfer, signed or delivered over a separate channel. The format is covered in dataset manifests and checksums.
- Transfer with the tool's built-in integrity checking enabled (rclone hash comparison where backends share a hash type; Globus integrity checking; S3 additional checksums on upload) [4].
- Recompute hashes at the destination, or read stored checksums where both sides support the same algorithm, and diff against the manifest.
- Treat missing, extra, size-mismatched and hash-mismatched files as four separate failure classes and log each.
- Record the manifest hash, the tool version and the job ID in your delivery record so the dataset can be traced to later model runs; see dataset-to-model traceability.
If the delivery is encrypted at rest with PGP or age, verify hashes of the ciphertext against the supplier's manifest, then of the plaintext after decryption. Key exchange is covered in encrypting dataset deliveries.
Resume and incremental sync for interrupted transfers
Every multi-terabyte transfer will be interrupted at least once, so design for idempotent reruns. The safe pattern is a copy that skips objects already present with matching size and hash, rerun until the skip count equals the manifest count. Managed services do this with an overwrite policy; rclone does it by default with copy, and more strictly with --checksum where both backends expose a compatible hash.
Three failure modes to plan for:
- Orphaned multipart uploads. An interrupted multipart upload leaves parts that are billed but invisible as objects. Set a lifecycle rule to abort incomplete multipart uploads after a few days.
syncdeleting data.rclone syncmakes the destination match the source, including deletions. If the supplier removes or renames a folder mid-delivery, your copy follows. Usecopyfor deliveries, and--backup-dirif you must usesync.- Moving targets. If the supplier is still writing files during your transfer, you copy a mixed state. Agree on a frozen snapshot or a version ID per release, as described in versioning licensed datasets.
Access and logging during the move
The transfer path is part of your access control story, not just plumbing. Use short-lived credentials scoped to the delivery prefix, run workers in a network segment that only reaches the source and destination, and keep transfer logs with the delivery record. On Google Cloud, Data Access audit logs for Cloud Storage are reportedly off by default, so enable them on the landing bucket if you need a read trail. After landing, apply the controls in access controls for licensed training data.
When you source data through SourceX, delivery runs through private, access-controlled workflows rather than email attachments, and only after an executed agreement and supplier approval. Each dataset is delivered under a license that defines records, uses, term and delivery, so agree the endpoint, the manifest format and the transfer path while the license is being drafted. More on how handover works is on how licensed data is delivered, and you can describe the data you need on the SourceX buyer page.
Get operational data delivered to a channel you can verify
SourceX sources operational datasets from US companies on request and manages the commercial process, from assessing data and licensing permissions to agreeing allowed uses in a license. Personal details are removed or replaced before delivery, and delivery runs through private, access-controlled workflows after supplier approval. Describe the dataset you need delivered.
Sources
- Amazon Web Services (AWS Snowball Edge Developer Guide), "AWS Snowball Edge availability change notice" (2025). https://docs.aws.amazon.com/snowball/latest/developer-guide/snowball-edge-availability-change.html
- Amazon Web Services (Amazon S3 User Guide), "Using Requester Pays general purpose buckets for storage transfers and usage". https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
- Amazon Web Services (AWS Snowball Edge Developer Guide), "Amazon S3 compatible storage on Snow Family devices (availability notice)". https://docs.aws.amazon.com/snowball/latest/developer-guide/s3compatible-on-snow.html
- Amazon Web Services (Amazon S3 User Guide), "Checking object integrity in Amazon S3". https://docs.aws.amazon.com/hi_in/AmazonS3/latest/userguide/checking-object-integrity.md
- Google Cloud, "Transfer from Amazon S3 to Cloud Storage". https://docs.cloud.google.com/storage-transfer/docs/create-transfers/agentless/s3
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.