GPU guides

Build something with your GPU.

Step-by-step guides for running ComfyUI, vLLM, Ollama and fine-tuning on cloud GPUs, plus sizing guides for GPU memory and prices.

56 guides6 topics

Find what you need

56 guides
LLM inference11 min read

Batch inference with vLLM on a cloud GPU

Run offline batch inference with vLLM's LLM class: a runnable Python script, how continuous batching and prefix caching work, and the knobs for throughput.

Read the guide
LLM inference11 min read

Chat with your documents in Open WebUI (RAG) on your own GPU

Chat with your documents in Open WebUI on a GPU VM: knowledge bases, the embedding model and its settings, GPU memory, and saving the knowledge base.

Read the guide
Setup & operations11 min read

Check your GPU, NVIDIA driver and CUDA version

nvidia-smi shows the newest CUDA your driver supports, not what is installed. Check driver, toolkit and PyTorch versions, then fix the usual mismatches.

Read the guide
Memory & sizing10 min read

ComfyUI and Stable Diffusion GPU requirements: VRAM by model

The GPU for Stable Diffusion and ComfyUI depends on the model: the VRAM SD 1.5, SDXL, FLUX, Qwen-Image and Wan 2.2 need, and what offloading changes.

Read the guide
Setup & operations9 min read

Connect to a cloud GPU with SSH, VS Code and port forwarding

Connect to a cloud GPU as ubuntu with one ~/.ssh/config entry, open it in VS Code Remote-SSH and tunnel Jupyter, ComfyUI, vLLM and Ollama.

Read the guide
Memory & sizing11 min read

DeepSeek GPU requirements: what runs on 48 to 141 GB GPUs

What DeepSeek needs to run: real file sizes for the R1 distills and the 671B models, what fits one 48 to 141 GB GPU, what needs more, and the licences.

Read the guide
Training & fine-tuning13 min read

DeepSpeed ZeRO on one multi-GPU VM: stages, memory and configs

DeepSpeed ZeRO stages and offload, memory per GPU from the ZeRO paper and DeepSpeed's estimator, configs, launching on 2 to 8 GPUs, and when to use FSDP.

Read the guide
LLM inference11 min read

Deploy vLLM with Docker on a GPU server

Serve an open model behind an OpenAI-compatible API with vLLM's Docker image: driver check, pinned tag, Compose file, API key and which GPU fits.

Read the guide
Setup & operations9 min read

Download Hugging Face models fast on a GPU server

Use the Hugging Face CLI's hf download with Xet, a token for gated models and file filters, then script it for every fresh GPU instance.

Read the guide
Training & fine-tuning12 min read

Fine-tune an LLM with Unsloth on a cloud GPU (QLoRA, step by step)

Unsloth QLoRA fine-tuning on one RTX A6000: install, train Gemma 4 E4B, push the LoRA adapter to Hugging Face before the VM stops, and size bigger runs.

Read the guide
Training & fine-tuning11 min read

Fine-tune gpt-oss-20b on a single cloud GPU

LoRA fine-tuning of gpt-oss-20b with TRL and PEFT on one 80 GB GPU: memory by recipe, the harmony format, the script, and serving the result with vLLM.

Read the guide
Setup & operations11 min read

Fix "CUDA out of memory" (and know when you need a bigger GPU)

Read the CUDA out of memory error, find what holds GPU memory, apply fixes in order of cost from batch size to a bigger GPU, and fix vLLM's startup errors.

Read the guide
LLM inference12 min read

Fix vLLM out-of-memory errors

Fix vLLM out-of-memory errors: the five startup errors in vLLM 0.30.0, what each memory flag trades, and worked maths for 48 and 80 GB GPUs.

Read the guide
Setup & operations11 min read

FP8 vs FP16 vs BF16: choosing precision for training and inference

Bits, range and precision of FP32, TF32, FP16, BF16, FP8 and FP4, which NVIDIA GPUs compute each in hardware, and which to use for training and inference.

Read the guide
Images & video12 min read

Generate video with Wan 2.2 in ComfyUI on a cloud GPU

Run Wan 2.2 video in ComfyUI: the 5B and A14B files and sizes, VRAM with and without offloading, resolution and length settings, and which GPU to rent.

Read the guide
Images & video8 min read

Get models into ComfyUI on a fresh cloud instance

Where each file goes in ComfyUI's models folder, and a script and model list that fetch your set from Hugging Face and Civitai on every new instance.

Read the guide
Pricing & planning8 min read

Google Colab pricing vs renting a GPU by the hour

Colab Pro is $9.99 a month for 100 compute units and Pro+ $49.99 for 600. When renting a 48 GB GPU by the hour costs less, with the arithmetic.

Read the guide
Memory & sizing11 min read

gpt-oss-120b and gpt-oss-20b: GPU requirements

How much GPU memory gpt-oss-120b and gpt-oss-20b need, which GPUs fit them, and how to serve them on H100, H200 NVL, RTX PRO 6000 and RTX A6000.

Read the guide
Training & fine-tuning10 min read

How much does it cost to fine-tune an LLM on rented GPUs? Worked examples

Estimate a fine-tune's cost: count the tokens, get tokens per second from a published or measured run, turn that into hours, and price them at live rates.

Read the guide
Memory & sizing11 min read

How much VRAM do you need for AI? LLMs, images, video and fine-tuning

How much VRAM you need for LLM inference, image and video generation, and fine-tuning: the formulas, and which workloads fit 48, 80, 96 and 141 GB GPUs.

Read the guide
Training & fine-tuning16 min read

How much VRAM do you need to fine-tune an LLM?

How much GPU memory QLoRA, LoRA and full fine-tuning need for 7B to 70B models, why published tables disagree, and which GPU setups fit.

Read the guide
Pricing & planning10 min read

How to rent a GPU for AI

Rent a cloud GPU for AI step by step: choose GPU memory, pick a template, add credit, launch, connect over SSH, save your results and stop.

Read the guide
Images & video11 min read

How to run ComfyUI on a cloud GPU

Run ComfyUI on a rented NVIDIA GPU: pick a GPU for your models, launch the template or install it yourself, load models, and save outputs before you stop.

Read the guide
Setup & operations10 min read

InfiniBand vs NVLink: what connects what in a GPU cluster

NVLink links GPUs inside a server at 900 GB/s to 1.8 TB/s per GPU. InfiniBand links servers at 400 or 800 Gb/s per port. What that means for your cluster.

Read the guide
Images & video9 min read

Install ComfyUI-Manager and custom nodes on a remote GPU

Install ComfyUI-Manager on a remote GPU, what its security levels allow, how to pin custom node versions, and how to reinstall your nodes on each launch.

Read the guide
LLM inference8 min read

Install vLLM and serve a model with vllm serve

Install vLLM 0.30.0 with uv or pip, match the wheel to your NVIDIA driver (CUDA 13 or cu129), then serve an OpenAI-compatible API with vllm serve.

Read the guide
LLM inference10 min read

KV cache explained: how much GPU memory LLM serving really needs

What the KV cache is, the formula to size it, worked numbers for Llama, Qwen, gpt-oss and Kimi, and how it decides between 48, 80, 96 and 141 GB GPUs.

Read the guide
Setup & operations11 min read

Move datasets and checkpoints to and from a cloud GPU

Copy data to and from a GPU VM with rsync, scp, sftp, rclone and the Hugging Face Hub, resume broken transfers and save results before the disk is deleted.

Read the guide
Training & fine-tuning11 min read

Multi-GPU fine-tuning with Axolotl on one 8-GPU VM

Install Axolotl, write a LoRA config, spread it over 8 GPUs with FSDP2 or DeepSpeed, size jobs with Axolotl's table, and push the adapter off the VM.

Read the guide
Training & fine-tuning17 min read

Multi-GPU training on one node with PyTorch DDP and FSDP

DDP or FSDP2 on one 2 to 8 GPU VM: memory per GPU, a runnable torchrun script, communication without NVLink, and checkpoints that survive the VM.

Read the guide
Pricing & planning9 min read

NVIDIA A100 price: buying vs renting per hour

What an NVIDIA A100 80GB costs to buy now that it is past end of sale, against renting one by the hour, with the break-even arithmetic shown.

Read the guide
Pricing & planning9 min read

NVIDIA H100 price: buying vs renting per hour

What an NVIDIA H100 costs to buy, from dated sources, against renting an H100 PCIe by the hour, with the break-even arithmetic shown.

Read the guide
Pricing & planning8 min read

NVIDIA H200 price: buying vs renting per hour

What an NVIDIA H200 costs to buy, from dated sources, against renting an H200 NVL by the hour, with the break-even arithmetic shown.

Read the guide
Memory & sizing11 min read

Ollama GPU requirements: which models fit 48, 80, 96 and 141 GB

How much GPU memory Ollama models need: library sizes by quantization, the KV cache of Ollama's default context, CPU offload and which GPU fits.

Read the guide
Images & video10 min read

Open-source video generation models and the GPUs they need

Wan 2.2, HunyuanVideo 1.5, LTX-2.3, LTX-2.5, Kandinsky 5 and MiniMax H3 compared: parameters, licences, file sizes, VRAM and ComfyUI support.

Read the guide
Images & video10 min read

Run ComfyUI in Docker on an Ubuntu GPU server

Build your own ComfyUI Docker image from the official install steps, give it the GPU, keep models on the VM's disk and reach it through an SSH tunnel.

Read the guide
Images & video7 min read

Run ComfyUI workflows from Python with the ComfyUI API

Export an API-format workflow, queue it with POST /prompt from Python, follow progress over the websocket and download the images through an SSH tunnel.

Read the guide
Setup & operations10 min read

Run Docker containers with NVIDIA GPUs

Install the NVIDIA Container Toolkit, pick GPUs with --gpus, choose CUDA images your driver runs, use GPUs in Compose, keep ports on 127.0.0.1, fix errors.

Read the guide
Images & video9 min read

Run FLUX in ComfyUI on a cloud GPU

Which FLUX model fits a 48, 80 or 96 GB GPU, where each file goes in ComfyUI, the settings that work, and what the FLUX dev and schnell licenses allow.

Read the guide
LLM inference11 min read

Run llama.cpp with full GPU offload on a cloud GPU

Run llama.cpp on an NVIDIA cloud GPU: the CUDA server image or a source build, full offload with -ngl, GGUF quant choice and an OpenAI-compatible API.

Read the guide
Images & video9 min read

Run Qwen-Image in ComfyUI on a cloud GPU

Qwen-Image and Qwen-Image-Edit in ComfyUI: file sizes, fp8 or bf16, the Qwen2.5-VL text encoder, VRAM on each GPU, workflow settings and the licence.

Read the guide
Images & video8 min read

Run Stable Diffusion WebUI Forge on a cloud GPU with Docker

AUTOMATIC1111 and the original Forge have stalled, and Forge Neo carries on. Run it in Docker on an Ubuntu GPU VM, on localhost, through an SSH tunnel.

Read the guide
LLM inference9 min read

Set up Open WebUI with Ollama on a cloud GPU

Set up Open WebUI with Ollama on an NVIDIA cloud GPU two ways: the ready template, or a pinned Docker Compose file on Ubuntu. Then save your chats.

Read the guide
Setup & operations10 min read

Stop paying for idle GPUs: auto-stop a QuantaCloud instance

A watchdog script that reads nvidia-smi and stops an idle QuantaCloud GPU instance through the REST API, run by systemd or cron. Copy results off first.

Read the guide
Setup & operations12 min read

SXM vs PCIe GPUs and NVLink: what it means for multi-GPU AI

PCIe cards, SXM modules, NVLink bridges and NVSwitch with NVIDIA's bandwidth figures, when the link between GPUs matters, and how to check it on a GPU VM.

Read the guide
Training & fine-tuning10 min read

Train a FLUX LoRA on a cloud GPU (AI Toolkit or FluxGym)

Train a FLUX.1 [dev] LoRA on a 48 GB cloud GPU with AI Toolkit, and why not FluxGym: dataset, config, memory, copying the LoRA off, and the licence.

Read the guide
Images & video10 min read

Use Ollama inside ComfyUI for prompt generation on one GPU

Run Ollama next to ComfyUI on one cloud GPU, add the comfyui-ollama node, split GPU memory between them and build a workflow that writes its own prompts.

Read the guide
LLM inference9 min read

Use the Ollama API on a remote GPU server

Call the Ollama API on a remote NVIDIA GPU over an SSH tunnel: curl, Python and the OpenAI-compatible endpoint, with port 11434 never exposed.

Read the guide
LLM inference14 min read

vLLM on multiple GPUs: tensor and pipeline parallelism

Split a model across 2, 4 or 8 GPUs with vLLM: tensor or pipeline parallelism, the flags, how to check the GPU links, and which VMs to rent.

Read the guide
LLM inference12 min read

vLLM quantization: which format to run on which GPU

The quantization formats vLLM 0.30 loads, what AWQ, GPTQ, FP8, NVFP4 and MXFP4 need from the GPU, real checkpoint sizes, and what fits on 48 to 141 GB.

Read the guide
LLM inference13 min read

vLLM vs Ollama: which one to run on a GPU server

vLLM vs Ollama on one GPU: design, GGUF vs safetensors, batching and concurrency, memory use and OpenAI compatibility, and when to use each.

Read the guide
LLM inference11 min read

vLLM vs SGLang: how the two serving engines differ

vLLM vs SGLang from both projects' docs and source: prefix caching, scheduling, structured output, memory flags, security defaults and which to run.

Read the guide
Pricing & planning7 min read

What is a neocloud? How GPU clouds differ from hyperscalers

A neocloud is a cloud provider built around renting GPU compute for AI. How neoclouds differ from hyperscalers, and what to check before you choose one.

Read the guide
Pricing & planning8 min read

What is GPU as a service (GPUaaS)?

GPU as a service means renting GPU compute by the hour or for a term instead of buying hardware. The four models, how pricing works and what to compare.

Read the guide
Setup & operations14 min read

What is NVFP4? 4-bit inference on NVIDIA Blackwell

NVFP4 is NVIDIA's 4-bit float for Blackwell, with an FP8 scale per 16 values. How it compares with MXFP4 and FP8, real checkpoint sizes and where it runs.

Read the guide
Setup & operations11 min read

Why GPU utilization is low during training, and how to fix it

Why GPU utilization is low in training: measure it with nvidia-smi, dmon and the PyTorch profiler, then fix data loading, small batches and sync points.

Read the guide