Self-hosted LLM inference on QuantaCloud means renting an NVIDIA GPU VM, running an open-weight model on it with vLLM, SGLang, Ollama or llama.cpp, and calling it through an OpenAI-compatible API that only you control. There is no managed endpoint and no serverless API. QuantaCloud provides the GPU instance, with Ubuntu 22.04, the NVIDIA driver and Docker on the Bare Metal template. The serving stack, the model and their security are yours to run. In return you pick the model, the precision and the context length, and no model provider sees your prompts.
A 48 GB RTX A6000 runs from $0.48/GPU-hr, and the 141 GB H200 NVL from the console price. You pay from launch to stop: the first hour is charged at launch, and the unused seconds of the current hour are refunded when you stop. Prices checked 6 Oct 2026, 01:40 UTC
Which GPU fits which model#
The GPU to rent is the smallest one that holds the weights plus the KV cache for your context. vLLM claims 92% of the card by default, loads the weights, and turns what is left into KV cache, the memory that stores every token of every active request. The table works this out for common open models, using the real size of each checkpoint on Hugging Face rather than parameters x bytes.
| Model and precision | Checkpoint | KV cache per token | Room for KV cache at vLLM's default | Where I would run it |
|---|---|---|---|---|
| Qwen3-8B, BF16 | 16.4 GB | 144 KiB | About 190,000 tokens on 48 GB | Any 48 GB card |
| gpt-oss-20b, MXFP4 | 13.8 GB | 24 KiB | About 1.2 million tokens on 48 GB | Any 48 GB card, after a test run |
| Qwen3-32B, per-channel FP8 | 34.3 GB | 256 KiB | About 38,000 tokens on 48 GB, 170,000 on 80 GB | L40S, L40 or RTX 6000 Ada for short contexts, 80 GB for long ones |
| Llama 3.3 70B, 4-bit AWQ | 39.8 GB | 320 KiB | About 14,000 tokens on 48 GB, 119,000 on 80 GB | An 80 GB card |
| gpt-oss-120b, MXFP4 | 65.2 GB | 36 KiB | About 360,000 tokens on 80 GB, 2 million on 141 GB | 80 GB for a few users, H200 NVL for many |
| Qwen3-32B, BF16 | 65.5 GB | 256 KiB | About 50,000 tokens on 80 GB, 110,000 on 96 GB | RTX PRO 6000 |
| Llama 3.3 70B, FP8 | 72.7 GB | 320 KiB | About 66,000 tokens on 96 GB, 200,000 on 141 GB | H200 NVL |
The room column is our calculation: 92% of the memory each card reports in nvidia-smi (41.4, 73.3, 87.9 or 129.2 GiB, taking ECC on for the 48 GB cards and the H100 PCIe for 80 GB) minus the checkpoint, divided by the KV cache per token at BF16, before activations. It is a pool that every concurrent request shares, so 190,000 tokens is room for nearly six full 32k conversations at once, or dozens of short ones. The 48 GB cards are the RTX A6000, RTX 6000 Ada, L40 and L40S, the 80 GB cards the A100 and H100 PCIe, and the 96 GB card the RTX PRO 6000 Blackwell. The KV cache guide has the formula for any model, and how much VRAM you need covers the rest.
Fitting is not the same as being supported. vLLM's gpt-oss recipe says the model runs on the H100, H200 and B200, gives the A100 its own instructions, and still lists Ampere and Ada Lovelace support as work in progress, so treat gpt-oss on a 48 GB card as something to test before you build on it. The gpt-oss guide has the detail.
Precision decides speed as well as fit. FP8 math needs compute capability 8.9 or newer: the Ada cards (L40S, L40, RTX 6000 Ada), Hopper (H100, H200) and Blackwell (RTX PRO 6000). On Ada, vLLM 0.30.0 uses FP8 math for per-channel FP8 checkpoints such as RedHatAI/Qwen3-32B-FP8-dynamic, while block-scaled ones, including Qwen's own FP8 releases, run weight-only below Hopper. The Ampere cards, the RTX A6000 and A100, run every FP8 checkpoint weight-only, so you get the memory saving without the FP8 speed. gpt-oss ships in MXFP4, which needs compute capability 8.0, and NVFP4 runs natively only on Blackwell. FP8 vs FP16 vs BF16 and vLLM quantization go further.
Launch an RTX PRO 6000 (96 GB) Launch an H200 NVL (141 GB)Models bigger than one GPU
A model that does not fit one card can be split across the GPUs of one VM. On 2026-09-27 the largest configurations were 8x A100 SXM4 with 640 GB, 4x RTX PRO 6000 and 8x L40S or RTX 6000 Ada with 384 GB each, and 2x H200 NVL with 282 GB. vLLM splits every layer across the GPUs with --tensor-parallel-size, and on cards without NVLink, such as the L40S, its docs recommend --pipeline-parallel-size instead. The catalog flags NVLink on the A100 SXM4 and H200 NVL offers, so check what your VM actually has with nvidia-smi topo -m before you pick one. vLLM across 2 to 8 GPUs and DeepSeek GPU requirements work through real cases.
Some open models barely fit any single on-demand VM. Kimi-K2.6 is 595.2 GB on disk, 554.3 GiB, which leaves only about 34.5 GiB of the 588.8 GiB vLLM would claim on 8x A100 SXM4 at its default (our calculation: 8 x 81,920 MiB / 1,024 x 0.92), and the deployment guide in its Hugging Face repo serves it on one 8-GPU H200 server. That is a job for a dedicated 8-GPU server, built to order.
Pick a serving stack#
The rule I follow: vLLM when more than one person or program calls the model, Ollama for a private chat, and llama.cpp when the build of the model that fits is a GGUF file.
| vLLM | SGLang | Ollama | llama.cpp | |
|---|---|---|---|---|
| Latest release on 2026-09-28 | 0.30.0 | 0.5.20 | 0.34.4 | 0.5.0 |
| What it loads | Hugging Face checkpoints: BF16, FP8, AWQ, GPTQ, MXFP4, NVFP4 | Hugging Face checkpoints | Models pulled from the Ollama library | GGUF files, 1.5-bit to 8-bit |
| Listens on by default | All interfaces, port 8000 | 127.0.0.1, port 30000 | 127.0.0.1, port 11434 | 127.0.0.1, port 8080 |
| API key | --api-key, which guards the /v1 routes only | --api-key | None on the local API | --api-key, off by default |
| Best fit | An API for many users or programs | The same job as vLLM, so benchmark both on your model | A private chat for one person | GGUF builds, and models larger than GPU memory |
| On QuantaCloud | Docker on Bare Metal | Bare Metal, with an R580 or newer driver | The Open WebUI + Ollama template, or Bare Metal | Bare Metal |
| Guide | vLLM with Docker | vLLM vs SGLang | Ollama API | llama.cpp on a GPU |
vLLM is the default I would pick for an API. It batches concurrent requests on the GPU and hands out KV cache in blocks as each request grows, the PagedAttention design from the paper that introduced it, so little memory is wasted and more requests fit in one batch. The official vllm/vllm-openai Docker image is the quickest route on the Bare Metal template, and installing vLLM with pip or uv is the other. SGLang serves the same kind of OpenAI-compatible API, and its 0.5.20 release ships only CUDA 13 builds, which need an R580 or newer driver.
Ollama is the quickest way to chat with a model you pulled a minute ago. QuantaCloud's Open WebUI + Ollama template starts it next to a chat interface, and on that template Ollama's port stays inside the app container, so scripts call Open WebUI's API instead. Ollama processes one request at a time per model unless you raise OLLAMA_NUM_PARALLEL, which is why I keep it for one person rather than an API that several programs share. vLLM vs Ollama and Ollama GPU requirements put numbers on the trade-off.
llama.cpp is the tool for GGUF quantizations from 1.5-bit to 8-bit. It can also run a model larger than GPU memory by keeping part of it on the CPU, which its README calls partial acceleration: only the layers on the GPU run at GPU speed. Its server speaks the OpenAI API too, with continuous batching for several users.
Launch vLLM in four steps#
Four steps take you from an empty VM to an API on your laptop. The vLLM Docker guide walks through each one, including the fixes for when Docker cannot see the GPU or asks for sudo.
-
Launch the Bare Metal template on the GPU the sizing table points to (deploying a GPU) and connect as
ubuntu(connecting over SSH). Most single-GPU VMs are running in about 3 minutes (median). -
Pick the image tag from the driver version. R580 or newer runs
vllm/vllm-openai:v0.30.0, and older drivers needv0.30.0-cu129, so let the driver decide:DRIVER_MAJOR=$(nvidia-smi --query-gpu=driver_version --format=csv,noheader | head -n1 | cut -d. -f1) if [ "$DRIVER_MAJOR" -ge 580 ]; then VLLM_TAG=v0.30.0; else VLLM_TAG=v0.30.0-cu129; fi -
Start the server with an API key, published on 127.0.0.1 only, and copy the key it prints:
export VLLM_API_KEY=$(openssl rand -hex 32) echo "$VLLM_API_KEY" docker run -d --name vllm --gpus all --ipc=host \ -p 127.0.0.1:8000:8000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ "vllm/vllm-openai:$VLLM_TAG" \ Qwen/Qwen3-8B --max-model-len 32768 --api-key "$VLLM_API_KEY" -
From your laptop, open a tunnel with
ssh -N -L 8000:127.0.0.1:8000 ubuntu@YOUR_INSTANCE_IPand point any OpenAI SDK athttp://127.0.0.1:8000/v1with the key.
Lock it down#
The safest default is a server that listens on 127.0.0.1 and an SSH tunnel to reach it, because then the API is reachable only through SSH on port 22.
| Server | Listens on by default | Authentication | What I do |
|---|---|---|---|
| vLLM | All interfaces, port 8000 | --api-key guards /v1, /v2, /inference and /cohere only | Publish on 127.0.0.1 and tunnel, or put a TLS proxy in front that passes /v1 only |
| SGLang | 127.0.0.1, port 30000 | --api-key | Keep the default host, set a key and tunnel |
| Ollama | 127.0.0.1, port 11434 | None on the local API | Keep the default host and tunnel |
| llama.cpp server | 127.0.0.1, port 8080 | --api-key, off by default | Set a key, keep the default host and tunnel |
An API key alone is not enough to open vLLM's port to the internet. /invocations runs the same inference with no key, and /pause and /abort_requests let anyone stop generation. Docker makes it worse: a port published with -p 8000:8000 bypasses ufw, and Docker releases older than 28.0.0 let hosts on the same network segment reach even ports published on 127.0.0.1. When an app on another machine needs the API, put a reverse proxy with TLS in front, let only /v1 through and keep the key on. The vLLM Docker guide has a Caddy config for exactly that, and the Ollama API guide does the same for Ollama.
What it costs#
You pay for the GPU by the hour from launch to stop, and nothing per token. The first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. An endpoint that runs around the clock uses 730 hours a month (8,760 hours a year divided by 12). At the 2026-09-27 prices that is $350.40 on one RTX A6000, $1,744.70 on one RTX PRO 6000 and $2,503.90 on one H200 NVL (our calculation: 730 x $0.48, $2.39 and $3.43).
To compare that with a per-token API, turn throughput into cost per million output tokens: the hourly price divided by (output tokens per second x 3,600), times 1,000,000. Today's prices:
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| H100 PCIe | 80 GB | $2.59/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| H200 NVL | - | Not listed | No |
Two rules protect a running endpoint. If your balance cannot cover the next hour, QuantaCloud terminates the instance and deletes its disk, so turn on auto top-up for anything that serves traffic. A low-balance email goes out when the balance drops below $2. Stopping deletes the disk too, with the image and the weights on it, so every launch downloads them again: 65.2 GB for gpt-oss-120b. Keep the setup in a script, and see downloading Hugging Face models fast to shorten the wait. The pricing page has the full billing rules.
When to reserve capacity instead#
Reserve when the endpoint has to run around the clock for months, or when the model needs more memory than one VM holds. Reserved capacity is built to order: we order and build the hardware to your spec, from one dedicated GPU server to a GPU cluster, and the configuration, lead time and terms come back in writing before you commit. For inference that usually means one 8-GPU server, such as an HGX H200 board with 8 x 141 GB = 1,128 GB, which holds a model like Kimi-K2.6 with room for its cache.
QuantaCloud does not publish an SLA. If the service you build needs service terms, list them in the brief, together with the model, the precision, the context length and the traffic you expect. Send a capacity brief
Questions before you self-host#
Is there a managed endpoint or serverless API?
No. QuantaCloud sells GPU instances, not tokens: there is no hosted model API, no autoscaling and no scale to zero. You run the server on your VM, and it bills from launch to stop whether or not it receives requests.
Where do my prompts go?
From your client to your VM, where the model processes them. The weights download from Hugging Face or the Ollama library. Ollama's cloud models, the tags ending in cloud, run on Ollama's own servers rather than your GPU, so skip them for private work. QuantaCloud runs in US regions only: us-east-1 in Virginia, and us-midwest-1, us-midwest-2 and us-midwest-4 in the Midwest.
Does the VM keep the model between sessions?
No. Stopping terminates the instance and deletes its disk, and there are no volumes or snapshots. Every launch starts clean, so script the setup and download the weights at the start of each session.
Can my team share one endpoint?
Yes, through the API. QuantaCloud accounts are single-user, but a vLLM server behind a TLS proxy with an API key serves any client that has the key, and it batches their requests together. Open WebUI can sit in front of it as the chat interface for people who do not write code.
Can I serve a model I fine-tuned?
Yes. vLLM loads a merged model folder like any other checkpoint, or serves the base model with your LoRA adapter on top through --enable-lora and --lora-modules. Fine-tuning on QuantaCloud sizes the GPU for the training run.
My rule for self-hosted inference: start on a 48 GB RTX A6000 with an 8B model or gpt-oss-20b, and read the KV cache line vLLM prints at startup. Move to the RTX PRO 6000 or the H200 NVL only when that line says your context or your concurrency does not fit, and send a capacity brief once the endpoint has to run around the clock. The vLLM Docker guide is the place to start.
Launch a GPU for vLLM