GPU guide

How to rent a GPU for AI

Rent a cloud GPU for AI step by step: choose GPU memory, pick a template, add credit, launch, connect over SSH, save your results and stop.

Faiz Ahmed10 min read

The short version is five moves: pick the GPU memory your job needs, add prepaid credit, launch a VM from a template, connect, and copy your results off before you stop. That last move is the one to learn before you spend anything. Stopping a QuantaCloud instance terminates it and deletes its disk, and there are no volumes or snapshots to fall back on.

Before you start: four decisions#

The four decisions are memory, template, GPU count and budget, and memory comes first because nothing else matters if the model does not fit.

How much GPU memory

Size the GPU to the model, not the other way round. Published figures for common jobs:

JobMemory it takesWhere I would start
Images with SDXLStability says 8 GB consumer GPUs are enoughRTX A6000, 48 GB
Images with FLUX.1 [dev] at full 16-bit34.2 GB of files: model, text encoders and VAE (our calculation: 23.80 + 9.79 + 0.25 + 0.34 GB)RTX A6000
Chat with an 8B model in BF1616.1 GB of weights for Llama 3.1 8BRTX A6000
QLoRA fine-tune of a 70B model40 to 48 GB, by Axolotl's estimate at short contextA100 80GB, for headroom
Serve gpt-oss-120bFits one 80 GB GPU, per OpenAIH100 PCIe, 80 GB
Serve Llama 3.3 70B in FP8 with a full 128k contextAbout 115.6 GB (our calculation: a 72.7 GB FP8 checkpoint plus 42.9 GB of KV cache)H200 NVL, 141 GB

For anything else, how much VRAM you need has the arithmetic, and LoRA, QLoRA and full fine-tuning VRAM covers training.

Template or plain VM

Pick a template when the app is the job, and the plain VM when you bring your own code. QuantaCloud has four: ComfyUI for image and video workflows, Open WebUI + Ollama for chatting with open models, PyTorch + Jupyter for notebooks, and Bare Metal, a plain Ubuntu 22.04 VM with the NVIDIA driver and Docker, for everything else. The template does not change the price, but app templates take longer to start because the app container has to come up.

One GPU or several

Start with one GPU unless the model does not fit on one. Self-serve VMs come with 1, 2, 4 or 8 GPUs in a single machine, and more GPUs buy memory for a model you split across cards, or throughput for data-parallel training. Training across several machines is not self-serve: it is a reserved GPU cluster built to order.

How long you will run it

Budget the hours you expect plus one, because the first hour is charged at launch and your balance must cover each next hour as it comes. On 2026-09-27 one RTX A6000 cost $0.48 an hour (today: $0.48/GPU-hr), so a 6-hour session cost $2.88 and a $10 deposit covered about 20 hours (our calculations: 6 x $0.48, and $10 / $0.48 = 20.8). The pricing page has the billing rules and three worked examples.

Step 1: create an account#

Sign up in the QuantaCloud console with GitHub, Google, an emailed magic link, or an email and password. A password sign-up sends a verification link, and the console stays locked until you click it.

QuantaCloud also creates a managed SSH key pair for your account, and its private key can be downloaded only once. Save that .pem file the moment the console offers it, and run chmod 600 on it. If you prefer your own key, add its public half under SSH Keys before you launch: Ed25519 or RSA, because ECDSA keys are not installed on the VM (SSH keys docs).

If you clicked a launch link before signing up and the console opens on the dashboard after you verify your email, go to the marketplace and pick the offer again.

Step 2: add credit#

Add credit on the Billing page before you choose a GPU. The minimum deposit is $5 and the quick amounts run from $10 to $500. A new card goes through a checkout page, and a saved card is charged at once. Your balance must cover one hour of the configuration you launch, and each next hour as it comes.

Auto top-up is optional: when your balance falls below a threshold you set, between $2 and $1,000, it charges your saved card for an amount you choose, between $10 and $25,000. I would turn it on for any run you will not be watching, because a balance that cannot pay the next hour terminates the instance and deletes its disk (adding credits).

Step 3: pick a GPU offer#

Pick by memory first, then price, then region. The marketplace filters by GPU model and GPU count, and each offer shows the GPU and its memory, the vCPUs, RAM and included disk, the region and the hourly price. The same GPU can appear in several offers: on 2026-09-27, single RTX A6000 offers came with 6 to 12 vCPUs and 24 to 64 GB of RAM at different prices, so compare the rows rather than taking the first. The regions are us-east-1 in Virginia, and us-midwest-1, us-midwest-2 and us-midwest-4 in the Midwest. The GPU catalog shows every GPU with its live from-price.

Step 4: choose a template and launch#

The deploy page is a single screen: the offer, the template, the SSH key (your managed key is preselected) and a cost summary with the hourly rate (deploying a GPU). The disk is fixed by the offer, so there is nothing to size. Click Deploy and the first hour is charged. Stop the instance before it reaches running and that hour comes back in full.

If your balance is short, the button reads Fund & Deploy instead. It takes you to add credit, and when you return you still have to click Deploy. Check the template at that point, because the console can reset it to Bare Metal after funding.

Step 5: connect#

Connect over SSH as ubuntu, not root, on port 22, with the key you chose at launch. The Deployments page shows the IP address.

Terminal
chmod 600 ~/.ssh/quantacloud.pem
ssh -i ~/.ssh/quantacloud.pem ubuntu@<instance-ip>
nvidia-smi

The last command should list the GPU you rented.

App templates add an Open Application button, which opens the app at a private address under app.quantacloud.net after a QuantaCloud login, and only your account can open it. Open WebUI has its own sign-in as well.

For editing code on the VM, VS Code's Remote - SSH extension works with a host entry like this in ~/.ssh/config:

Output
Host quanta-gpu
  HostName <instance-ip>
  User ubuntu
  IdentityFile ~/.ssh/quantacloud.pem

Then run Remote-SSH: Connect to Host and pick quanta-gpu. The SSH and VS Code guide adds port forwarding for Jupyter and ComfyUI, and the docs cover connecting over SSH.

Step 6: do the work#

Start from the guide for your job rather than a blank terminal:

JobTemplateGuide
Image workflowsComfyUIRun ComfyUI on a cloud GPU
FLUX imagesComfyUIFLUX in ComfyUI
A private chat assistantOpen WebUI + OllamaOpen WebUI with Ollama
An OpenAI-compatible APIBare MetalDeploy vLLM with Docker
gpt-oss-20b or gpt-oss-120bBare Metalgpt-oss GPU requirements
Fine-tuning with QLoRABare Metal or PyTorch + JupyterFine-tune an LLM with Unsloth

Whatever the job, download model weights on the VM instead of uploading them from your laptop.

Step 7: save your results, then stop#

Copy everything you want to keep off the VM before you press Stop, because Stop terminates the instance and deletes its disk. From your own machine:

Terminal
rsync -avz -e "ssh -i ~/.ssh/quantacloud.pem" ubuntu@<instance-ip>:~/outputs/ ./outputs/
scp -i ~/.ssh/quantacloud.pem ubuntu@<instance-ip>:~/checkpoints/final.tar.gz .

App templates run the app in a Docker container on the VM, so either download outputs through the app in your browser, or copy them out of the container with sudo docker cp and then over SSH. Then press Stop on the Deployments page, and the unused seconds of the current hour go back to your balance.

The clock starts when you click Deploy, because the minutes spent provisioning count once the instance reaches running. At the 2026-09-27 price of $0.48 an hour (today: $0.48/GPU-hr), the billing rules predict $0.12 for 15 minutes from Deploy to Stop (our calculation: 0.25 h x $0.48). The billing docs show where the charge and refund lines appear.

Prices checked 6 Oct 2026, 00:15 UTC

Common mistakes#

These are the mistakes I would warn any first-time renter about, and most of them come back to the stop rule and the balance:

MistakeWhat happensHow to avoid it
Not saving the managed SSH private keyIt can be downloaded only onceSave the file at sign-up, or add your own key
Picking too little GPU memoryThe job fails with CUDA out of memorySize the model first, and read fixing CUDA out of memory
Logging in as rootThe login failsLog in as ubuntu
Leaving an instance runningIt charges every hour, busy or idleStop it as soon as the job ends
Pressing Stop before copying outputsThe disk is deleted with the instanceCopy with rsync or scp first
Letting the balance run out mid-jobThe instance is terminated at the next unpaid hour, disk includedFund the whole run plus an hour, or turn on auto top-up

When renting is not the right answer#

Renting by the hour is the wrong tool at both ends. For a short notebook you watch for an hour, Colab's free tier costs nothing, and Colab pricing vs renting a GPU shows where the line sits. For work that keeps the same GPUs busy for months, or training across several machines, reserved capacity built to order fits better: we order and build the hardware to your spec and confirm the configuration, lead time and terms in writing. Send a capacity brief.

Frequently asked questions#

Do I need a credit card?

Yes. Credit is bought by card, $5 at minimum.

Is there a free trial?

No. New accounts start with a $0 balance, and there is no sign-up credit.

Can I pause an instance and resume it later?

No. Stop terminates the instance and deletes its disk. Keep code in Git and outputs on your own machine, and launch a fresh instance when you come back.

Where are the GPUs?

In the US: us-east-1 in Virginia, and us-midwest-1, us-midwest-2 and us-midwest-4 in the Midwest.

Can I use my own Docker image?

Yes, on the Bare Metal template, which comes with Docker: pull your image and run it with --gpus all. There is no custom-image template in the console, and the GPU VPS page lists the checks to run on a new VM first.


My rule: rent by the hour when the job has a start and an end, and you can copy its results off before you stop.

Create an account and launch

Keep building

Choose your next step.