GPU guide

Download Hugging Face models fast on a GPU server

Use the Hugging Face CLI's hf download with Xet, a token for gated models and file filters, then script it for every fresh GPU instance.

Faiz Ahmed9 min read

The fastest way to download a Hugging Face model on a GPU server is the Hugging Face CLI's hf download with a file filter, so you fetch only the weights you will load. If you came looking for huggingface-cli download, that command is gone: huggingface_hub 2.0.0, released on 2026-09-24, removed huggingface-cli and kept hf. The hf_transfer speed-up is gone as well, because every transfer from the Hub now runs through Xet.

On QuantaCloud the download is part of every session. Stopping an instance deletes its disk, so each launch starts with an empty cache, and the minutes spent downloading are billed like any other GPU minutes (pricing). Filter hard, and put the download in a script.

Launch an RTX A6000 to follow along

Install the hf CLI with uv#

The rule I follow on a fresh instance is to install hf with uv, which does not depend on the image's Python packages:

Terminal
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
uv tool install hf
hf version

The hf package on PyPI is the CLI from huggingface_hub 2.0.0, and on x86_64 machines it pulls in hf_xet, the Xet client that does the actual transfers. Inside an existing Python environment, pip install -U huggingface_hub gives you the same hf command. hf env prints the installed versions, the cache path and the token path.

Get access to gated models before you launch#

The one thing I always do first is request access on huggingface.co, before the GPU is running. Gated repos need an access request made in the browser while you are logged in, and the model's authors choose between automatic and manual approval. On 2026-09-28 the Hub API listed black-forest-labs/FLUX.1-dev as automatic and meta-llama/Llama-3.1-8B-Instruct as manual, and a manual request waits until the authors act on it.

Then create a token under Settings, Access Tokens on huggingface.co: a read token, or a fine-grained token that can read only the repos you need. On the instance, log in and confirm the account:

Terminal
hf auth login
hf auth whoami

hf auth login asks how you want to log in. Log in with your browser prints a URL and a code that you open on your laptop, and Paste an access token reads the token as hidden input. For gated models I paste a token I created myself, so I know exactly what it can read. Either way the token ends up in ~/.cache/huggingface/token. If the request is still pending, or the token belongs to an account without access, the download stops with:

Output
Error: Access denied. Model 'meta-llama/Llama-3.1-8B-Instruct' requires approval.

For ComfyUI, the repackaged files from Comfy-Org and other mirrors are not gated, but the upstream license still applies: FLUX.1 [dev] stays non-commercial for the model itself whichever repo you pull it from.

Download only the files you load#

The biggest speed-up is skipping files you will never load. Many repos carry the same weights in two or three formats, and the Hub API showed these totals on 2026-09-28:

RepoWhole repoWith the filterFilter
openai/gpt-oss-120b195.8 GB65.3 GB--exclude "original/*" --exclude "metal/*"
openai/gpt-oss-20b41.3 GB13.8 GBThe same two excludes
meta-llama/Llama-3.3-70B-Instruct282.3 GB141.1 GB--exclude "original/*"
meta-llama/Llama-3.1-8B-Instruct32.1 GB16.1 GB--exclude "original/*"
black-forest-labs/FLUX.1-dev, for ComfyUI57.9 GB24.1 GBName flux1-dev.safetensors and ae.safetensors
Qwen/Qwen3-8B16.4 GB16.4 GBNone needed

The original/ folders hold checkpoints for the vendors' own reference code, such as a 16.1 GB consolidated.00.pth in Llama-3.1-8B-Instruct, and gpt-oss adds a third copy in metal/. Transformers and vLLM load the files in the repo root, so exclude the rest:

Terminal
hf download openai/gpt-oss-120b --exclude "original/*" --exclude "metal/*"
hf download meta-llama/Llama-3.1-8B-Instruct --exclude "original/*"

Repeat --include or --exclude once per pattern. Unlike shell globs, the patterns match across folders and are case-sensitive, so --include "*.safetensors" on gpt-oss-120b also takes the safetensors in original/ and downloads 130.5 GB. Add --dry-run to see the file list and the total before anything downloads:

Terminal
hf download openai/gpt-oss-120b --exclude "original/*" --exclude "metal/*" --dry-run

The same filters work from Python:

Python
from huggingface_hub import snapshot_download

path = snapshot_download("openai/gpt-oss-120b", ignore_patterns=["original/*", "metal/*"])
print(path)

The gpt-oss GPU requirements guide covers which GPU each gpt-oss size needs once it is on disk.

Where the files land#

By default everything lands in ~/.cache/huggingface. Models go to hub/models--<org>--<name>/snapshots/<commit>/ as symlinks into a blobs/ folder, Xet keeps its own cache in xet/, and the token file sits next to them. hf download prints the snapshot path when it finishes, --quiet prints only that path for scripts, and hf cache ls lists what is cached and how big it is.

My rule is to keep the default cache for anything that loads models by repo ID, such as Transformers, vLLM and diffusers, and to use --local-dir only for tools that want files at a path. ComfyUI is the usual case. With ComfyUI installed by comfy-cli in ~/comfy/ComfyUI on the Bare Metal template:

Terminal
hf download black-forest-labs/FLUX.1-dev flux1-dev.safetensors --local-dir ~/comfy/ComfyUI/models/diffusion_models
hf download black-forest-labs/FLUX.1-dev ae.safetensors --local-dir ~/comfy/ComfyUI/models/vae

--local-dir writes the files straight into that folder, next to a small .cache/huggingface/ metadata folder, so they take the space once. A file that is already in the main cache gets copied instead, which takes the space twice. Running FLUX in ComfyUI picks up from here, and running ComfyUI on a cloud GPU covers the template. Getting models into ComfyUI on a fresh instance scripts a whole model set.

Containers see the cache when you mount it, which is what the vLLM Docker example does with -v ~/.cache/huggingface:/root/.cache/huggingface. Download first, then start the server with HF_HUB_OFFLINE=1 in its environment, so it reads only local files and makes no calls to the Hub. Deploying vLLM with Docker uses this layout.

Check the disk before you download#

Check free space before the first download. My rule of thumb is to budget twice the model's size when I plan to convert or quantize it. Start with:

Terminal
df -h

If the big disk in df -h is mounted somewhere other than the filesystem that holds your home directory, move the cache there before downloading with export HF_HOME=/path/on/that/disk/huggingface. HF_HOME moves the token file along with the cache, and it has to be set before Python imports huggingface_hub.

Single-GPU configurationDisk includedRAM
RTX A6000 1x256 GB24 to 64 GB
RTX 6000 Ada 1x350 GB72 GB
L40 1x, L40S 1x625 GB72 GB
A100 SXM4 80GB 1x625 to 1,000 GB100 to 120 GB
RTX PRO 6000 Blackwell 1x725 GB144 GB
H200 NVL 1x750 GB180 GB
H100 PCIe 1x1,250 GB128 GB

Those are the configurations in the catalog on 2026-09-27, and they change, so each GPU page lists the live ones. Our calculation for the smallest disk: the whole Llama-3.3-70B-Instruct repo is 282.3 GB, more than an RTX A6000's 256 GB, while the filtered 141.1 GB leaves 256 minus 141.1, about 115 GB, before the operating system takes its share. Docker images and your outputs live on the same disk. How much VRAM you need is the other half of sizing.

Make the download faster#

The honest answer is that Xet already does most of the work. hf_xet downloads large files as parallel chunk ranges and tunes its concurrency to the network, starting at 1 stream and scaling up to 64, and Hugging Face says the defaults saturate the available bandwidth on most machines, data-center ones included. The options compare like this:

MethodWhat it doesUse it when
hf download with defaultsXet transfers with adaptive concurrencyAlmost always
HF_XET_HIGH_PERFORMANCE=1 hf downloadRaises the concurrency limits and memory buffersHigh bandwidth and at least 64 GB of RAM
snapshot_download() in PythonThe same engine as hf downloadThe download lives inside your own code
curl or wget on a file URLOne HTTP stream per file, without Xet's parallel chunksOne small file and nothing else installed
git clone with git-xetA Git checkout of every file in the repoYou want Git history or plan to push changes
HF_HUB_ENABLE_HF_TRANSFER=1Nothing now: hf_transfer belonged to the old storage backendNever

The one switch worth testing is high-performance mode:

Terminal
HF_XET_HIGH_PERFORMANCE=1 hf download openai/gpt-oss-120b --exclude "original/*" --exclude "metal/*"

Hugging Face recommends HF_XET_HIGH_PERFORMANCE=1 only for machines with high bandwidth and at least 64 GB of RAM, and warns that it can slow smaller machines down. The RTX A6000 1x configurations had 24 to 64 GB of RAM in the catalog on 2026-09-27, at best the bare minimum, so I leave it off there, while the H200 NVL 1x had 180 GB. --max-workers, 8 by default, sets how many files download at once. To see the baseline an instance gets from Hugging Face's CDN, run hf extensions install julien-c/hf-speedtest and then hf speedtest.

Script it for every launch#

Put the download in a script, because every launch starts with an empty disk. Keep the script next to your code and run it the moment a new instance reaches running:

Terminal
#!/usr/bin/env bash
# fetch-models.sh: run on a fresh instance as ubuntu
set -euo pipefail
export PATH="$HOME/.local/bin:$PATH"

if ! command -v hf >/dev/null 2>&1; then
  curl -LsSf https://astral.sh/uv/install.sh | sh
  uv tool install hf
fi

# Full commit hashes, so every launch gets the same files
hf download Qwen/Qwen3-8B \
  --revision b968826d9c46dd6066d109eabc6255188de91218
hf download meta-llama/Llama-3.1-8B-Instruct \
  --revision 0e9e39f249a16976918f6564b8830bc894c89659 \
  --exclude "original/*"

hf cache ls

From your laptop, with the Host quanta-gpu entry from the SSH and VS Code guide, copy your token over and run the script:

Terminal
ssh quanta-gpu 'mkdir -p ~/.cache/huggingface && umask 077 && cat > ~/.cache/huggingface/token' < ~/.config/hf-token
ssh quanta-gpu 'bash -s' < fetch-models.sh

The first line streams the token from a local file that only you can read into the file hf reads by default, so it never appears in a command line or a shell history. The second runs the script on the instance. --revision needs the full 40-character commit hash, not the short form. For a download that takes more than a few minutes, start the script inside tmux on the instance, so a dropped connection cannot kill it.


My default for a new instance: hf download with an exclude filter, pinned to a commit, run from a script as soon as the instance is up, with HF_XET_HIGH_PERFORMANCE=1 only on configurations with more than 64 GB of RAM. Whatever you create is deleted with the disk, so before you stop, push it to a private repo with hf upload your-name/my-model ./outputs --private, which needs a token with write access, or copy it to your laptop. Then serve the model with vLLM straight from the same cache.

Launch an H200 NVL for large downloads

Keep building

Choose your next step.