Use rsync over SSH for anything large. It resumes after a dropped connection, skips files that already arrived, and runs the same way in both directions. scp is fine for a single file, sftp for browsing and resuming by hand, rclone for object storage, and the Hugging Face Hub for models and datasets. On QuantaCloud the copy off the instance is the step that cannot be skipped: stopping an instance terminates it and deletes its disk, and there are no volumes or snapshots to fall back on.
Every example below uses the quanta-gpu alias from the SSH and VS Code guide, which logs in as ubuntu on the instance's public IP. QuantaCloud does not charge for data transfer in either direction. What a slow transfer does cost is GPU time, because the instance bills from launch to stop, including the hours spent waiting on a copy.
Which tool for which job#
The rule I follow is to pick the tool by where the data lives, not by habit:
| Job | Tool | Resumes after a drop |
|---|---|---|
| A folder from your laptop to the instance, or results back | rsync -avP | Yes: rerun the same command |
| One file, quickly | scp | No |
| Browsing the instance's files, resuming by hand | sftp with reget and reput | Yes |
| Tens of thousands of small files | tar piped over ssh | No, but it is one stream |
| Data kept in an S3-compatible bucket | rclone copy | Yes, file by file |
| Public models and datasets, and your own finished models | hf download and hf upload | Yes |
| Moving to a new instance | rsync from the old instance to the new one | Yes |
Check the disk first#
The disk is fixed per configuration and included in the hourly price. In the catalog on 2026-09-27, a single RTX A6000 came with 256 GB, an RTX PRO 6000 Blackwell with 725 GB, an H200 NVL with 750 GB and an H100 PCIe with 1,250 GB, and 8-GPU configurations went up to 5,000 GB. Before a big upload, check what is free:
ssh quanta-gpu 'df -h'
The operating system, Docker images and model downloads share that disk with your data, so a 256 GB instance holds less than it sounds. The Hugging Face download guide works through the arithmetic for a 70B model.
rsync: resumable copies in both directions#
rsync is the tool I reach for first. Ubuntu 22.04 ships rsync 3.2.7, and it has to be installed on both ends: if rsync --version fails on either side, install it with sudo apt-get install -y rsync. To upload a dataset and to bring results back:
rsync -avP ./dataset/ quanta-gpu:~/dataset/
rsync -avP quanta-gpu:~/outputs/ ./outputs/
-a copies folders with their permissions and timestamps, -v lists each file, and -P keeps partly transferred files and shows progress. The trailing slash on the source matters: ./dataset/ copies the folder's contents into ~/dataset, while ./dataset without the slash would create ~/dataset/dataset.
Resuming is the reason to use rsync. If the connection drops, run the same command again. rsync skips every file whose size and modification time already match, and -P kept the partial file, so the next run finishes it much faster. For one very large file you are resuming, --append-verify appends to the partial copy and then checks the whole file. The rsync manual calls it dangerous unless every file in the transfer is a shared, growing file, so use it only to finish that one file.
A few options change the picture for big transfers:
| Option | What it does | When I use it |
|---|---|---|
--info=progress2 | One progress line for the whole transfer instead of one per file, in rsync 3.1.0 and newer | Folders with many files |
-z | Compresses data in transit, which the rsync manual calls useful over a slow connection | A slow home connection |
--exclude "*.tmp" | Skips matching files | Caches and scratch files you can rebuild |
-n, --dry-run | Shows what would be copied without copying | Before any run that uses --delete |
On a Mac, the rsync that ships with the system does not accept --info=progress2: older macOS releases ship rsync 2.6.9, and macOS 15.4 and newer ship openrsync. -avP works with both.
scp and sftp for quick copies#
scp is the shortest way to move one file or a small folder. Copying a single file into a folder that does not exist yet fails, so create the folder first:
ssh quanta-gpu 'mkdir -p ~/data'
scp ./train.jsonl quanta-gpu:~/data/
scp -r quanta-gpu:~/outputs ./outputs-$(date +%F)
It cannot resume, so a dropped connection means copying that file again from the start, which makes it a poor fit for anything that takes more than a few minutes. -r copies folders and follows symbolic links, -C compresses and -p keeps modification times.
sftp opens an interactive session on the instance, which helps when you do not remember where a file is:
sftp quanta-gpu
sftp> ls outputs
sftp> get -R outputs
sftp> reget checkpoints/final.safetensors
sftp> reput dataset.tar
reget and reput resume an interrupted download or upload. They assume the partial copy matches the original, so use them only to finish a transfer that sftp itself started.
Many small files: stream a tar archive#
A folder of tens of thousands of small files, such as an image dataset, can go over as a single tar stream through SSH. tar writes the archive to standard output, and a second tar on the other side unpacks it, so nothing is stored in between:
tar -C ./dataset -cf - . | ssh quanta-gpu 'mkdir -p ~/dataset && tar -C ~/dataset -xf -'
mkdir -p ./outputs && ssh quanta-gpu 'tar -C ~/outputs -cf - .' | tar -C ./outputs -xf -
rsync and scp handle each file separately, while tar sends one continuous stream. The catch is that a tar stream cannot resume: if it breaks, finish the job with rsync, which copies only the files that are missing or incomplete.
Object storage with rclone#
rclone is the tool for data that lives in a bucket. It speaks the S3 API, and its documentation lists dozens of S3-compatible stores alongside its other storage backends. Install it on the instance with the script from the rclone docs:
sudo -v ; curl https://rclone.org/install.sh | sudo bash
rclone can take its whole configuration from environment variables, which suits an instance that starts clean every time. Read the keys without leaving them in your shell history:
export RCLONE_CONFIG_STORE_TYPE=s3
export RCLONE_CONFIG_STORE_PROVIDER=Other
export RCLONE_CONFIG_STORE_ENDPOINT=https://your-object-store-endpoint
read -rs -p "Access key ID: " RCLONE_CONFIG_STORE_ACCESS_KEY_ID; echo
read -rs -p "Secret access key: " RCLONE_CONFIG_STORE_SECRET_ACCESS_KEY; echo
export RCLONE_CONFIG_STORE_ACCESS_KEY_ID RCLONE_CONFIG_STORE_SECRET_ACCESS_KEY
rclone lsd store:
Other is rclone's setting for any S3-compatible store, and the endpoint comes from your object store's settings. Then copy in either direction:
rclone copy store:my-bucket/datasets/train ~/data/train --progress
rclone copy ~/outputs store:my-bucket/runs/run-01 --progress
rclone copy skips files that are already identical at the destination, so rerunning the same command after an interruption picks up where it stopped, one file at a time, and it never deletes anything at the destination. It moves 4 files at once by default, and --transfers 16 helps when you have many files and a fast connection.
Models and datasets through the Hugging Face Hub#
The Hub is the direct route for public models and datasets: the instance downloads them itself, without going through your laptop. If hf is not installed yet, the Hugging Face download guide installs it with uv (uv tool install hf). Datasets then download with the same command as models:
hf download HuggingFaceH4/ultrachat_200k --repo-type dataset --local-dir ~/data/ultrachat
The Hub also works in the other direction for anything you want to keep, such as a fine-tuned model or a finished LoRA adapter. hf upload creates the repo if it does not exist, and --private keeps it private:
hf upload your-name/my-run ./outputs --private
It needs a token with write access. Large uploads resume on their own: run the same command again and files that already arrived are skipped. For a long training run, hf upload your-name/my-run ./checkpoints --every 10 --private pushes new files every 10 minutes while the job runs, so an interrupted run still leaves its most recent uploaded checkpoint on the Hub. Leave it running in tmux next to the training job. The Hugging Face download guide covers tokens, filters and the cache layout.
Move data from one instance to another#
Moving to a bigger GPU means moving the data too, and the shortest path is directly between the two instances, which keeps your own connection out of it. Launch the new instance with the same SSH key, load that key into your laptop's agent, and connect to the old instance with agent forwarding:
ssh-add ~/.ssh/quantacloud_ed25519
ssh -A ubuntu@OLD_INSTANCE_IP
rsync -avP ~/outputs/ ubuntu@NEW_INSTANCE_IP:~/outputs/
The OpenSSH manual advises enabling agent forwarding with caution, because anyone with root on the old instance could use your agent while you are connected, so use it only on instances you control and log out when the copy finishes. Without forwarding, scp -r ubuntu@OLD_INSTANCE_IP:~/outputs ubuntu@NEW_INSTANCE_IP:~/ also works: scp routes a copy between two remote hosts through your laptop by default, which is simpler and as slow as your own connection.
Before you stop: a checklist#
Stopping an instance terminates it and deletes its disk, and nothing brings it back. Run through these steps every time:
-
See what is on the disk:
ssh quanta-gpu 'du -sh ~/outputs ~/checkpoints' -
Copy it off with
rsync -avP,rclone copyorhf upload. -
Check that the copy is complete. A dry run with checksums lists every file that still differs, and prints no file names when the two sides match:
rsync -avnc quanta-gpu:~/outputs/ ./outputs/ -
Stop the instance in the console.
A stop is not the only way an instance ends. If your balance cannot cover the next hour, QuantaCloud terminates the instance and deletes its disk. A low-balance email goes out below $2, but termination does not wait for it: an instance whose next hour costs more than the remaining balance ends even while the balance is above $2. Optional auto top-up on the Billing page refills the balance automatically (billing in the docs). For runs that last longer than you can watch, copy results off as they appear, with hf upload --every or an rclone copy in a loop. The idle GPU guide shows how to stop an instance automatically once the work is done, and the guides index has the rest of the setup guides.
File transfer questions#
Does QuantaCloud charge for uploads or downloads?
No. There are no egress or ingress charges. You pay for the time the instance runs, from launch to stop, including the time a transfer takes, so a slow upload on a large GPU costs GPU time (pricing).
How do I resume an interrupted transfer?
Run the same rsync -avP command again. It skips the files that arrived and continues the partial one. With sftp, use reget or reput. scp has no resume option, which is why I do not use it for large files.
Can I keep the disk and come back to it later?
No. Stopping terminates the instance and deletes its disk, and QuantaCloud has no volumes or snapshots. Keep code in Git, data in a bucket or on the Hub, and copy results off before every stop. If a project needs a machine that keeps its disk between jobs, that is a dedicated GPU server, which we build to order.
Should I use rsync or scp for large files?
rsync. Both run over SSH, and only rsync resumes an interrupted copy. For many small files, compare rsync with a tar stream over SSH.
My rule for any instance is to decide where the results go before the job starts: rsync to my laptop for small outputs, a bucket or the Hub for anything large, and hf upload --every for long training runs. Then copy, verify with a checksum dry run, and only then stop. Start on an RTX A6000 with the Bare Metal template, and see running Docker with a GPU if your job runs in a container.