GPU guide

Run llama.cpp with full GPU offload on a cloud GPU

Run llama.cpp on an NVIDIA cloud GPU: the CUDA server image or a source build, full offload with -ngl, GGUF quant choice and an OpenAI-compatible API.

Faiz Ahmed11 min read

The quickest way to run llama.cpp on an NVIDIA GPU is its official CUDA server image, started with --gpus all, a GGUF model and -ngl all. Current llama.cpp already tries to keep the whole model on the GPU by default. When the model and its context do not fit, it quietly lowers the context (if you did not set one) and then moves layers to the CPU, which is the usual reason llama.cpp feels slow. -ngl all turns that into an error you can see. On a QuantaCloud Bare Metal VM, Ubuntu 22.04 with the NVIDIA driver and Docker, the Docker route needs no build step, and a native build needs the CUDA toolkit first.

Pick a GGUF file that fits#

The rule I follow is to pick the largest quantization whose file plus KV cache fits the GPU with a gigabyte to spare, starting from Q4_K_M. llama.cpp's quantize tool lists how much each type raises perplexity on Llama-3-8B, where lower means closer to the 16-bit model, and the file sizes below are the real ones on Hugging Face:

TypeBits per weightQwen3-32B fileLlama 3.3 70B filePerplexity increase, Llama-3-8B
Q2_K3.16Not published26.38 GB+3.5199
Q3_K_M4.00Not published34.27 GB+0.6569
Q4_K_M4.8919.76 GB42.52 GB+0.1754
Q5_K_M5.7023.21 GB49.95 GB+0.0569
Q6_K6.5626.88 GB57.89 GB+0.0217
Q8_08.5034.82 GB74.98 GB+0.0026
16-bit16.0065.52 GB (BF16 weights)141.12 GBReference

Q8_0 is almost indistinguishable from 16-bit on that measure, Q4_K_M gives up a little, and Q2_K gives up about 20 times as much as Q4_K_M (our calculation from the tool's figures). The bits per weight are llama.cpp's own figures for Llama 3.1 8B. A Q4_K_M file works out closer to 4.9 bits than 4, because each block of weights carries its own scales and the quantizer keeps some tensors at a higher-precision type. The Qwen3-32B files come from Qwen's own GGUF repository, which has no Q2 or Q3 files, and the Llama files from a community repository. gpt-oss needs no choice, since it ships in MXFP4: ggml-org's GGUFs are 12.11 GB for gpt-oss-20b and 63.39 GB for gpt-oss-120b.

How much GPU memory llama.cpp needs#

llama.cpp needs the model file, the KV cache and a compute buffer, and by default it keeps 1,024 MiB of each GPU free on top. The KV cache is 16-bit by default: 2 x layers x KV heads x head size x 2 bytes for every token of context. When you leave -c unset, llama-server sizes the cache for the model's full trained context and shares it among its 4 default slots. These totals are our calculation, before the compute buffer:

Model and fileContextWeightsKV cacheQuantaCloud GPU that holds it
Qwen3-8B Q8_032,7688.1 GiB4.5 GiBAny 48 GB card
gpt-oss-20b MXFP4131,07211.3 GiB3.0 GiBAny 48 GB card
Qwen3-32B Q4_K_M32,76818.4 GiB8.0 GiBAny 48 GB card
Qwen3-32B Q8_032,76832.4 GiB8.0 GiBA 48 GB card, with little to spare
Llama 3.3 70B Q4_K_M8,19239.6 GiB2.5 GiBA 48 GB card, short context only
Llama 3.3 70B Q4_K_M131,072, q8_0 cache39.6 GiB21.3 GiB80 GB: A100 or H100 PCIe
gpt-oss-120b MXFP4131,07259.0 GiB4.5 GiB80 GB
Llama 3.3 70B Q8_032,76869.8 GiB10.0 GiB96 GB RTX PRO 6000 or 141 GB H200 NVL

-ctk q8_0 -ctv q8_0 stores the cache at 34 bytes per 32 values instead of 64, about 53% of the 16-bit size, which is how the 70B model gets its full 131,072 tokens on an 80 GB card. The 8-bit V cache needs Flash Attention, and llama.cpp's default -fa auto switches it on for that. The KV cache guide explains the formula, and the Ollama GPU requirements apply the same arithmetic to Ollama's tags, which are GGUF files too.

Launch an Ubuntu GPU VM#

Start with the Bare Metal template on a 48 GB RTX A6000, which holds everything up to Qwen3-32B at Q8_0. The GPU VPS page describes the VM itself.

Launch an Ubuntu GPU VM

Most single-GPU VMs are running in about 3 minutes (median). Connect as ubuntu (connecting over SSH) and check the GPU and the driver:

Terminal
ssh ubuntu@YOUR_INSTANCE_IP
nvidia-smi --query-gpu=name,driver_version,memory.total,compute_cap --format=csv

The server-cuda image is a CUDA 12.8 build, which runs on an R570 or newer driver. Its server-cuda13 sibling is a CUDA 13 build for R580 or newer. Checking your GPU, driver and CUDA version explains the numbers.

Run llama-server with Docker#

The image to use is ghcr.io/ggml-org/llama.cpp:server-cuda, pinned to a build number. llama.cpp publishes several builds a day: b11223 came out on 2026-09-27, about four hours after b11222. First check that a container can see the GPU:

Terminal
docker run --rm --gpus all ghcr.io/ggml-org/llama.cpp:server-cuda-b11223 --list-devices

It should list your GPU as a CUDA device. If Docker answers with permission denied, add ubuntu to the docker group, and if it complains that it cannot select a GPU device driver, install the NVIDIA Container Toolkit. Running Docker with NVIDIA GPUs shows both fixes. Then create an API key and start the server:

Terminal
export LLAMA_API_KEY=$(openssl rand -hex 32)
echo "$LLAMA_API_KEY"

docker run -d --name llama --gpus all \
  -p 127.0.0.1:8080:8080 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -e LLAMA_API_KEY \
  ghcr.io/ggml-org/llama.cpp:server-cuda-b11223 \
  -hf Qwen/Qwen3-32B-GGUF:Q4_K_M --alias qwen3-32b \
  -ngl all -c 32768

docker logs -f llama

The image's entrypoint is llama-server, so the arguments after the image name are server flags. -hf downloads Qwen3-32B-Q4_K_M.gguf, 19.76 GB, into the Hugging Face cache that the volume keeps on the VM's disk. -ngl all asks for every layer on the GPU. -c 32768 sets the context that the 4 slots share. The server reads the key from LLAMA_API_KEY, and --alias gives the model a short name for API calls. The image sets LLAMA_ARG_HOST=0.0.0.0 inside the container, so the 127.0.0.1: in the port mapping is what keeps it off the public interface. The pull is 2.59 GB compressed.

Check that every layer is on the GPU#

The log says how many layers went to the GPU:

Terminal
docker logs llama 2>&1 | grep -E "offloaded|buffer size"
nvidia-smi --query-gpu=memory.used,memory.total --format=csv

For Qwen3-32B you want load_tensors: offloaded 65/65 layers to GPU: 64 transformer layers plus the output layer. The CUDA0 model buffer size, KV buffer size and compute buffer size lines show where the memory went. Any lower first number means part of the model runs on the CPU.

That is what the defaults do when memory runs short. -ngl defaults to auto and --fit to on: llama.cpp estimates the memory, lowers the context towards 4,096 tokens when you left -c unset, then keeps as many layers on the GPU as fit, and for mixture-of-experts models it moves expert weights to the CPU first. It only changes settings you left at their defaults, so an explicit -ngl all makes the load fail with an out-of-memory error instead. On a rented GPU I want that failure, because a quiet CPU fallback costs the same per hour and runs slower.

If the log says warning: no usable GPU found, --gpu-layers option will be ignored, the binary has no GPU backend or cannot see the GPU. The usual causes are a CPU-only image such as the plain server tag, or a container started without --gpus all.

Call the OpenAI-compatible API#

llama-server speaks the OpenAI API under /v1, with the key as a Bearer token:

Terminal
curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $LLAMA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-32b", "messages": [{"role": "user", "content": "Explain GGUF quantization in two sentences."}], "max_tokens": 1000}'

Qwen3 thinks before it answers. llama-server returns the thinking in reasoning_content, separate from the answer, and it counts against max_tokens, so leave room. Every response also carries a timings object, where predicted_per_second is the generation speed and prompt_per_second the prompt processing speed. /health answers without a key, which suits monitoring and nothing else.

From your laptop, open a tunnel and point any OpenAI client at it:

Terminal
ssh -N -L 8080:127.0.0.1:8080 ubuntu@YOUR_INSTANCE_IP
Python
import os
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key=os.environ["LLAMA_API_KEY"])
reply = client.chat.completions.create(
    model="qwen3-32b",
    messages=[{"role": "user", "content": "Explain GGUF quantization in two sentences."}],
    max_tokens=1000,
)
print(reply.choices[0].message.content)

The same tunnel serves llama-server's built-in chat page at http://127.0.0.1:8080. The SSH and VS Code guide covers keys and port forwarding.

Build llama.cpp from source instead#

Build it yourself when you want native binaries or a commit newer than the image. The Ubuntu CUDA archives on the release page do not run on this VM. Their libraries need glibc 2.38 and a newer libstdc++, while Ubuntu 22.04 ships glibc 2.35 and the libstdc++ of GCC 12, so llama-server from the b11223 archive stops at once:

Output
./llama-server: /lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.38' not found (required by ./libllama-server-impl.so)

The build needs NVIDIA's CUDA toolkit, which you install from NVIDIA's repository. Version 12.8 suits an R570 or newer driver:

Terminal
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
sudo apt-get install -y cuda-toolkit-12-8 build-essential cmake git libssl-dev
export PATH=/usr/local/cuda-12.8/bin:$PATH

git clone --depth 1 --branch b11223 https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86
cmake --build build --config Release -j "$(nproc)" -t llama-server
./build/bin/llama-server --list-devices

Set CMAKE_CUDA_ARCHITECTURES to your GPU's compute capability without the dot: 80 for the A100, 86 for the RTX A6000, 89 for the L40, L40S and RTX 6000 Ada, 90 for the H100 and H200, and 120 for the RTX PRO 6000. It matters here because llama.cpp detects the installed GPU only with CMake 3.24 or newer, and Ubuntu 22.04 ships 3.22.1. Without the flag, the build compiles the default list the Docker image is built with: native code for compute capability 8.6, 8.9 and 12.0, and PTX for the rest, which the driver compiles when an A100, H100 or H200 first loads it. The binary takes the same flags as the container. It listens on 127.0.0.1:8080 by default, and the key goes in --api-key or LLAMA_API_KEY:

Terminal
./build/bin/llama-server -hf Qwen/Qwen3-32B-GGUF:Q4_K_M --alias qwen3-32b -ngl all -c 32768

More than one GPU, or a model larger than the GPU#

With several GPUs, llama.cpp splits the model by layers by default (-sm layer), which puts consecutive layers and their KV cache on each GPU in turn. -ts 3,1 changes the proportions, -sm row splits each weight matrix instead, and -sm tensor is marked experimental. For serving many users across GPUs, vLLM's tensor and pipeline parallelism is the better tool, as the vLLM multi-GPU guide explains.

A mixture-of-experts model larger than the GPU can still run with its experts in system RAM: --cpu-moe keeps all of them there, and --n-cpu-moe N keeps those of the first N layers. That trades speed for size, and it needs RAM. The RTX A6000 1x offers came with 24 to 64 GB of RAM in the 2026-09-27 catalog, against 180 GB on an H200 NVL 1x, so check the offer card before you plan on it.

Before you stop the instance#

Stopping a QuantaCloud instance deletes its disk, including the downloaded GGUF files and the pulled image. The next launch downloads the 2.59 GB image and every model again, so keep the docker run line in a script on your laptop. The Hugging Face download guide shows how to fetch large files faster. The first hour is charged at launch, and the unused seconds are refunded when you stop.

llama.cpp GPU FAQ#

Why is llama.cpp not using my GPU?

Three causes cover most cases. A build or image without CUDA logs no usable GPU found. A container started without --gpus all cannot see the GPU at all. A model that does not fit gets some of its layers moved to the CPU by the default --fit. The offloaded N/M layers line tells you which one you have.

What does -ngl 99 mean?

It asks for up to 99 layers on the GPU, which older guides used to mean every layer. Current builds accept -ngl all for that, and their default, auto, fits as many layers as free memory allows.

Should I use llama.cpp or vLLM?

llama.cpp for GGUF files, a single user or a small team, and models that have to spill into system RAM. vLLM for many concurrent users, where its batching and multi-GPU support pay off. The vLLM Docker guide covers that side, and vLLM vs Ollama compares the two approaches: Ollama 0.34.4 runs llama.cpp's llama-server underneath. Self-hosted LLM inference compares the serving stacks side by side.

Which GGUF quantization should I download?

Q4_K_M first, because it is the default of -hf and gives up little quality. Move to Q5_K_M, Q6_K or Q8_0 when the GPU has room after the KV cache, and go below Q4 only when nothing else fits.


My rule for llama.cpp on a rented GPU: run the pinned server-cuda image, pass -ngl all and a fixed -c, and read the offloaded line before anything else. Pick the largest GGUF that leaves room for the context you need, which on a 48 GB RTX A6000 means Q8_0 up to 32B models and Q4_K_M for 70B only at short context.

Launch an Ubuntu GPU VM for llama.cpp

Keep building

Choose your next step.