GPU guides
Build something with your GPU.
Step-by-step guides for running ComfyUI, vLLM, Ollama and fine-tuning on cloud GPUs, plus sizing guides for GPU memory and prices.
Batch inference with vLLM on a cloud GPU
Run offline batch inference with vLLM's LLM class: a runnable Python script, how continuous batching and prefix caching work, and the knobs for throughput.
Chat with your documents in Open WebUI (RAG) on your own GPU
Chat with your documents in Open WebUI on a GPU VM: knowledge bases, the embedding model and its settings, GPU memory, and saving the knowledge base.
Check your GPU, NVIDIA driver and CUDA version
nvidia-smi shows the newest CUDA your driver supports, not what is installed. Check driver, toolkit and PyTorch versions, then fix the usual mismatches.
ComfyUI and Stable Diffusion GPU requirements: VRAM by model
The GPU for Stable Diffusion and ComfyUI depends on the model: the VRAM SD 1.5, SDXL, FLUX, Qwen-Image and Wan 2.2 need, and what offloading changes.
Connect to a cloud GPU with SSH, VS Code and port forwarding
Connect to a cloud GPU as ubuntu with one ~/.ssh/config entry, open it in VS Code Remote-SSH and tunnel Jupyter, ComfyUI, vLLM and Ollama.
DeepSeek GPU requirements: what runs on 48 to 141 GB GPUs
What DeepSeek needs to run: real file sizes for the R1 distills and the 671B models, what fits one 48 to 141 GB GPU, what needs more, and the licences.
DeepSpeed ZeRO on one multi-GPU VM: stages, memory and configs
DeepSpeed ZeRO stages and offload, memory per GPU from the ZeRO paper and DeepSpeed's estimator, configs, launching on 2 to 8 GPUs, and when to use FSDP.
Deploy vLLM with Docker on a GPU server
Serve an open model behind an OpenAI-compatible API with vLLM's Docker image: driver check, pinned tag, Compose file, API key and which GPU fits.
Download Hugging Face models fast on a GPU server
Use the Hugging Face CLI's hf download with Xet, a token for gated models and file filters, then script it for every fresh GPU instance.
Fine-tune an LLM with Unsloth on a cloud GPU (QLoRA, step by step)
Unsloth QLoRA fine-tuning on one RTX A6000: install, train Gemma 4 E4B, push the LoRA adapter to Hugging Face before the VM stops, and size bigger runs.
Fine-tune gpt-oss-20b on a single cloud GPU
LoRA fine-tuning of gpt-oss-20b with TRL and PEFT on one 80 GB GPU: memory by recipe, the harmony format, the script, and serving the result with vLLM.
Fix "CUDA out of memory" (and know when you need a bigger GPU)
Read the CUDA out of memory error, find what holds GPU memory, apply fixes in order of cost from batch size to a bigger GPU, and fix vLLM's startup errors.
Fix vLLM out-of-memory errors
Fix vLLM out-of-memory errors: the five startup errors in vLLM 0.30.0, what each memory flag trades, and worked maths for 48 and 80 GB GPUs.
FP8 vs FP16 vs BF16: choosing precision for training and inference
Bits, range and precision of FP32, TF32, FP16, BF16, FP8 and FP4, which NVIDIA GPUs compute each in hardware, and which to use for training and inference.
Generate video with Wan 2.2 in ComfyUI on a cloud GPU
Run Wan 2.2 video in ComfyUI: the 5B and A14B files and sizes, VRAM with and without offloading, resolution and length settings, and which GPU to rent.
Get models into ComfyUI on a fresh cloud instance
Where each file goes in ComfyUI's models folder, and a script and model list that fetch your set from Hugging Face and Civitai on every new instance.
Google Colab pricing vs renting a GPU by the hour
Colab Pro is $9.99 a month for 100 compute units and Pro+ $49.99 for 600. When renting a 48 GB GPU by the hour costs less, with the arithmetic.
gpt-oss-120b and gpt-oss-20b: GPU requirements
How much GPU memory gpt-oss-120b and gpt-oss-20b need, which GPUs fit them, and how to serve them on H100, H200 NVL, RTX PRO 6000 and RTX A6000.
How much does it cost to fine-tune an LLM on rented GPUs? Worked examples
Estimate a fine-tune's cost: count the tokens, get tokens per second from a published or measured run, turn that into hours, and price them at live rates.
How much VRAM do you need for AI? LLMs, images, video and fine-tuning
How much VRAM you need for LLM inference, image and video generation, and fine-tuning: the formulas, and which workloads fit 48, 80, 96 and 141 GB GPUs.
How much VRAM do you need to fine-tune an LLM?
How much GPU memory QLoRA, LoRA and full fine-tuning need for 7B to 70B models, why published tables disagree, and which GPU setups fit.
How to rent a GPU for AI
Rent a cloud GPU for AI step by step: choose GPU memory, pick a template, add credit, launch, connect over SSH, save your results and stop.
How to run ComfyUI on a cloud GPU
Run ComfyUI on a rented NVIDIA GPU: pick a GPU for your models, launch the template or install it yourself, load models, and save outputs before you stop.
InfiniBand vs NVLink: what connects what in a GPU cluster
NVLink links GPUs inside a server at 900 GB/s to 1.8 TB/s per GPU. InfiniBand links servers at 400 or 800 Gb/s per port. What that means for your cluster.
Install ComfyUI-Manager and custom nodes on a remote GPU
Install ComfyUI-Manager on a remote GPU, what its security levels allow, how to pin custom node versions, and how to reinstall your nodes on each launch.
Install vLLM and serve a model with vllm serve
Install vLLM 0.30.0 with uv or pip, match the wheel to your NVIDIA driver (CUDA 13 or cu129), then serve an OpenAI-compatible API with vllm serve.
KV cache explained: how much GPU memory LLM serving really needs
What the KV cache is, the formula to size it, worked numbers for Llama, Qwen, gpt-oss and Kimi, and how it decides between 48, 80, 96 and 141 GB GPUs.
Move datasets and checkpoints to and from a cloud GPU
Copy data to and from a GPU VM with rsync, scp, sftp, rclone and the Hugging Face Hub, resume broken transfers and save results before the disk is deleted.
Multi-GPU fine-tuning with Axolotl on one 8-GPU VM
Install Axolotl, write a LoRA config, spread it over 8 GPUs with FSDP2 or DeepSpeed, size jobs with Axolotl's table, and push the adapter off the VM.
Multi-GPU training on one node with PyTorch DDP and FSDP
DDP or FSDP2 on one 2 to 8 GPU VM: memory per GPU, a runnable torchrun script, communication without NVLink, and checkpoints that survive the VM.
NVIDIA A100 price: buying vs renting per hour
What an NVIDIA A100 80GB costs to buy now that it is past end of sale, against renting one by the hour, with the break-even arithmetic shown.
NVIDIA H100 price: buying vs renting per hour
What an NVIDIA H100 costs to buy, from dated sources, against renting an H100 PCIe by the hour, with the break-even arithmetic shown.
NVIDIA H200 price: buying vs renting per hour
What an NVIDIA H200 costs to buy, from dated sources, against renting an H200 NVL by the hour, with the break-even arithmetic shown.
Ollama GPU requirements: which models fit 48, 80, 96 and 141 GB
How much GPU memory Ollama models need: library sizes by quantization, the KV cache of Ollama's default context, CPU offload and which GPU fits.
Open-source video generation models and the GPUs they need
Wan 2.2, HunyuanVideo 1.5, LTX-2.3, LTX-2.5, Kandinsky 5 and MiniMax H3 compared: parameters, licences, file sizes, VRAM and ComfyUI support.
Run ComfyUI in Docker on an Ubuntu GPU server
Build your own ComfyUI Docker image from the official install steps, give it the GPU, keep models on the VM's disk and reach it through an SSH tunnel.
Run ComfyUI workflows from Python with the ComfyUI API
Export an API-format workflow, queue it with POST /prompt from Python, follow progress over the websocket and download the images through an SSH tunnel.
Run Docker containers with NVIDIA GPUs
Install the NVIDIA Container Toolkit, pick GPUs with --gpus, choose CUDA images your driver runs, use GPUs in Compose, keep ports on 127.0.0.1, fix errors.
Run FLUX in ComfyUI on a cloud GPU
Which FLUX model fits a 48, 80 or 96 GB GPU, where each file goes in ComfyUI, the settings that work, and what the FLUX dev and schnell licenses allow.
Run llama.cpp with full GPU offload on a cloud GPU
Run llama.cpp on an NVIDIA cloud GPU: the CUDA server image or a source build, full offload with -ngl, GGUF quant choice and an OpenAI-compatible API.
Run Qwen-Image in ComfyUI on a cloud GPU
Qwen-Image and Qwen-Image-Edit in ComfyUI: file sizes, fp8 or bf16, the Qwen2.5-VL text encoder, VRAM on each GPU, workflow settings and the licence.
Run Stable Diffusion WebUI Forge on a cloud GPU with Docker
AUTOMATIC1111 and the original Forge have stalled, and Forge Neo carries on. Run it in Docker on an Ubuntu GPU VM, on localhost, through an SSH tunnel.
Set up Open WebUI with Ollama on a cloud GPU
Set up Open WebUI with Ollama on an NVIDIA cloud GPU two ways: the ready template, or a pinned Docker Compose file on Ubuntu. Then save your chats.
Stop paying for idle GPUs: auto-stop a QuantaCloud instance
A watchdog script that reads nvidia-smi and stops an idle QuantaCloud GPU instance through the REST API, run by systemd or cron. Copy results off first.
SXM vs PCIe GPUs and NVLink: what it means for multi-GPU AI
PCIe cards, SXM modules, NVLink bridges and NVSwitch with NVIDIA's bandwidth figures, when the link between GPUs matters, and how to check it on a GPU VM.
Train a FLUX LoRA on a cloud GPU (AI Toolkit or FluxGym)
Train a FLUX.1 [dev] LoRA on a 48 GB cloud GPU with AI Toolkit, and why not FluxGym: dataset, config, memory, copying the LoRA off, and the licence.
Use Ollama inside ComfyUI for prompt generation on one GPU
Run Ollama next to ComfyUI on one cloud GPU, add the comfyui-ollama node, split GPU memory between them and build a workflow that writes its own prompts.
Use the Ollama API on a remote GPU server
Call the Ollama API on a remote NVIDIA GPU over an SSH tunnel: curl, Python and the OpenAI-compatible endpoint, with port 11434 never exposed.
vLLM on multiple GPUs: tensor and pipeline parallelism
Split a model across 2, 4 or 8 GPUs with vLLM: tensor or pipeline parallelism, the flags, how to check the GPU links, and which VMs to rent.
vLLM quantization: which format to run on which GPU
The quantization formats vLLM 0.30 loads, what AWQ, GPTQ, FP8, NVFP4 and MXFP4 need from the GPU, real checkpoint sizes, and what fits on 48 to 141 GB.
vLLM vs Ollama: which one to run on a GPU server
vLLM vs Ollama on one GPU: design, GGUF vs safetensors, batching and concurrency, memory use and OpenAI compatibility, and when to use each.
vLLM vs SGLang: how the two serving engines differ
vLLM vs SGLang from both projects' docs and source: prefix caching, scheduling, structured output, memory flags, security defaults and which to run.
What is a neocloud? How GPU clouds differ from hyperscalers
A neocloud is a cloud provider built around renting GPU compute for AI. How neoclouds differ from hyperscalers, and what to check before you choose one.
What is GPU as a service (GPUaaS)?
GPU as a service means renting GPU compute by the hour or for a term instead of buying hardware. The four models, how pricing works and what to compare.
What is NVFP4? 4-bit inference on NVIDIA Blackwell
NVFP4 is NVIDIA's 4-bit float for Blackwell, with an FP8 scale per 16 values. How it compares with MXFP4 and FP8, real checkpoint sizes and where it runs.
Why GPU utilization is low during training, and how to fix it
Why GPU utilization is low in training: measure it with nvidia-smi, dmon and the PyTorch profiler, then fix data loading, small batches and sync points.