QuantaCloud

Choose your starting point.

QuantaCloud has four GPU templates: Bare Metal (Ubuntu 22.04, NVIDIA driver, Docker), PyTorch + Jupyter, Open WebUI + Ollama and ComfyUI. What each runs.

Start with the environment

Four ways to start building.

Ubuntu + Docker

SSH, the NVIDIA driver, and Docker. Start with your own containers.

Explore the template Launch on a GPU

PyTorch + Jupyter

A GPU notebook environment for experiments and training.

Explore the template Launch on a GPU

Open WebUI + Ollama

A private chat interface for the open models you choose.

Explore the template Launch on a GPU

ComfyUI

A node-based workspace for your image and video workflows.

Explore the template Launch on a GPU

A QuantaCloud template decides what is running on your GPU instance when it boots, and there are four. Bare Metal is a plain Ubuntu 22.04 VM with the NVIDIA driver and Docker. PyTorch + Jupyter, Open WebUI + Ollama and ComfyUI start their app in a container on the same kind of VM and give it a private address that only your QuantaCloud account can open. Every template runs on every GPU in the catalog, from the 48 GB RTX A6000 to the 141 GB H200 NVL, and the template does not change the price.

A single RTX A6000 runs from $0.48/GPU-hr, whichever template you pick. Prices checked 6 Oct 2026, 01:40 UTC

The four templates side by side#

The main difference is what starts on boot, and everything else follows from it.

Bare MetalPyTorch + JupyterOpen WebUI + OllamaComfyUI
What startsUbuntu 22.04 with the NVIDIA driver and Docker, and nothing elseJupyterLab with PyTorch and GPU support, in a containerOpen WebUI with Ollama next to it, in a containerComfyUI, in a container
How you open itSSH as ubuntu on port 22Open Application, or SSHOpen Application, or SSHOpen Application, or SSH
Sign-inYour SSH keyQuantaCloud login, no Jupyter tokenQuantaCloud login, then Open WebUI's own accountQuantaCloud login
App port behind the private addressNone888880808188
ModelsNoneNoneNone until you pull one in the appAdd your own after launch
Best forvLLM, llama.cpp, your own Docker images and training scriptsNotebooks, experiments and single-GPU fine-tuningA private chat with open modelsImage and video workflows
MoreGPU VPSJupyter on a GPUOpen WebUIComfyUI

The templates docs describe the same four from the console's side.

How the private address works#

Open Application on the Deployments page opens an app template at a private address under app.quantacloud.net. QuantaCloud checks that you are logged in and that the deployment is yours before it passes a request on, so nobody else can open the app through that address, and that includes teammates: QuantaCloud accounts are single-user today. Jupyter and ComfyUI rely on that login alone, while Open WebUI adds its own account screen, where the first account created becomes the administrator.

Bare Metal has no app address. You connect over SSH as ubuntu with your key, and forward a port through SSH for anything with a web interface, as the SSH and VS Code guide shows.

Bare Metal: Ubuntu, the driver and Docker#

Bare Metal is the name of the plain template, not the hardware: every on-demand instance is a VM. It is the template for anything QuantaCloud does not start for you, such as a vLLM, SGLang or llama.cpp server, a training script, your own Docker image or a VS Code remote session. There is no custom-image template, so your own container runs here with Docker, and deployments created through the REST API always use Bare Metal. The GPU VPS page lists the checks I run on a fresh VM, and Docker with NVIDIA GPUs, vLLM with Docker and checking the driver and CUDA version take it from there.

Launch Bare Metal on an RTX A6000

PyTorch + Jupyter#

PyTorch + Jupyter opens JupyterLab in your browser with PyTorch and GPU support ready, behind your QuantaCloud login, so there is no Jupyter token to copy. It suits notebooks, quick experiments and single-GPU fine-tuning, and it has no idle timeout: it runs, and bills, until you stop it. Install packages from a cell with %pip install, and open a terminal from the Launcher for everything else. Jupyter on a GPU compares it with hosted notebooks, and the Unsloth tutorial runs a full QLoRA fine-tune on it.

Launch PyTorch + Jupyter on an RTX A6000

Open WebUI + Ollama#

Open WebUI + Ollama starts a chat interface with Ollama next to it, and no model: you pull one from the Ollama library inside the app, such as gpt-oss:20b. After the QuantaCloud login, Open WebUI asks you to create its own account, and the first account on a fresh install becomes the administrator. Ollama's port stays inside the app container on this template, so scripts reach the model through Open WebUI's API over an SSH tunnel to port 8080.

Open WebUI on your own GPU covers which models fit which GPU, and the setup guide, the Ollama API guide and chatting with your documents go further. For an API that several programs share, self-hosted LLM inference compares the serving stacks.

Launch Open WebUI + Ollama on an RTX A6000

ComfyUI#

ComfyUI opens the node editor in your browser behind your QuantaCloud login. Add your models after launch, by uploading them or downloading them on the instance, and copy your outputs off before you stop. The template's ComfyUI version is not pinned, so when a workflow needs an exact release or your own set of custom nodes, install ComfyUI yourself on Bare Metal instead.

ComfyUI on QuantaCloud has the GPU picker, and running ComfyUI on a cloud GPU covers the template and the do-it-yourself path. The guides to loading models on a new instance, FLUX in ComfyUI, ComfyUI Manager and custom nodes, the ComfyUI API from Python and ComfyUI in Docker go deeper.

Launch ComfyUI on an RTX A6000

Versions and boot times#

The app templates run container images whose versions are not pinned yet, so what starts can change from one launch to the next. Check the version inside the app at the start of a session rather than relying on an old number.

Most single-GPU VMs are running in about 3 minutes (median). App templates take longer: the deployment stays in Connecting while the app's container image is pulled and until the app answers its health check.

Launch a template in five steps#

Five steps take you from the catalog to a running app.

  1. Pick a GPU offer in the console, or use a button on this page, which opens the deploy page with the offer and the template already selected (deploying a GPU).
  2. Check the template on the deploy page before you click Deploy. If you add credit on the way, the page can come back set to Bare Metal, so select your template again.
  3. Wait for Running on the Deployments page.
  4. Click Open Application for an app template, or connect with ssh ubuntu@<instance-ip> for Bare Metal (connecting over SSH).
  5. Copy your work off before you press Stop, because stopping terminates the instance and deletes its disk.

What a template costs#

Every template bills the same way: the GPU's hourly price from launch to stop, with the first hour charged at launch and the unused seconds of the current hour refunded when you stop. The template adds nothing. Whichever template you pick, stopping deletes the disk, so notebooks, chats, pulled models and generated images go with it, and there are no volumes or snapshots to restore from. The pricing page has the full billing rules, and the GPU catalog has every configuration.

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
H200 NVL-Not listedNo

Template questions#

Is there a vLLM or custom Docker template?

No. The four templates on this page are the whole list. For vLLM, SGLang, llama.cpp or an image of your own, launch Bare Metal and run the container with Docker, as the vLLM Docker guide shows.

Can I switch templates on a running instance?

The template is chosen when you deploy. To use another one, copy what you need off the instance, launch a new one with the other template, and stop the old one. On any template you can still install software over SSH.

Can the API launch a template?

Not today. Deployments created through the REST API use Bare Metal, and templates are chosen in the console.

Which template should I use for fine-tuning?

PyTorch + Jupyter for single-GPU runs you watch from a notebook, and Bare Metal for long runs over SSH and multi-GPU trainers. Fine-tuning on QuantaCloud sizes the GPU for each method.


My rule: take the app template when the app is the whole job, a notebook, a chat or a ComfyUI canvas, and take Bare Metal as soon as you need an exact version, a server of your own or a Docker image. Whichever you pick, copy your work off before you press Stop.

Launch an RTX A6000

Keep building

Choose your next step.