GPU guide

Connect to a cloud GPU with SSH, VS Code and port forwarding

Connect to a cloud GPU as ubuntu with one ~/.ssh/config entry, open it in VS Code Remote-SSH and tunnel Jupyter, ComfyUI, vLLM and Ollama.

Faiz Ahmed9 min read

The fastest way to work on a cloud GPU is one Host entry in ~/.ssh/config that your terminal and VS Code Remote-SSH both use. On QuantaCloud you log in as ubuntu to the instance's public IP on port 22, with the key you chose at launch. A Jupyter, ComfyUI, vLLM or Ollama server you start yourself then reaches your laptop through an SSH tunnel, so nothing except SSH has to listen on the internet.

Every QuantaCloud instance is an Ubuntu 22.04 VM with the NVIDIA driver and Docker installed, so these steps are the same on an RTX A6000 as on an H200 NVL. One part is specific to QuantaCloud and comes last: stopping an instance deletes it, so expect a new IP address and a new host key every time you launch. If you have not launched one yet, how to rent a GPU covers that first.

Launch an RTX A6000 to follow along

Create a key just for GPU instances#

The rule I follow is one Ed25519 key for GPU instances and nothing else. If it ever leaks, you delete one key in one place. Create it on your laptop:

Terminal
ssh-keygen -t ed25519 -f ~/.ssh/quantacloud_ed25519 -C "gpu-instances"

Paste the contents of ~/.ssh/quantacloud_ed25519.pub into the console under SSH Keys, Add Existing Key (SSH keys in the docs). QuantaCloud takes Ed25519 and RSA public keys and skips ECDSA. The deploy page preselects your managed key, so pick this one in the SSH key dropdown every time you launch: each instance is set up with the key chosen at launch.

The managed key works too. QuantaCloud creates it for your account, and its private key downloads only once as a .pem file, so store it right away and lock its permissions:

Terminal
mv ~/Downloads/your-key.pem ~/.ssh/quantacloud.pem
chmod 600 ~/.ssh/quantacloud.pem

Without the chmod, ssh prints WARNING: UNPROTECTED PRIVATE KEY FILE and This private key will be ignored, skips the key and fails to log in. A lost managed key cannot be downloaded again: generate a new key or upload your own, and launch with that one.

Save the instance in ~/.ssh/config#

The first thing I set up is a Host alias, because every later command, and VS Code, uses it. Add this to ~/.ssh/config on your laptop, with the IP address shown on the Deployments page (connecting over SSH):

Output
Host quanta-gpu
    HostName 203.0.113.10
    User ubuntu
    Port 22
    IdentityFile ~/.ssh/quantacloud_ed25519
    IdentitiesOnly yes
    ServerAliveInterval 60
    StrictHostKeyChecking accept-new
    UserKnownHostsFile ~/.ssh/known_hosts_quantacloud

Replace 203.0.113.10 with your instance's IP, and point IdentityFile at ~/.ssh/quantacloud.pem if you use the managed key. OpenSSH takes the first value it finds for each setting, so keep this block above any Host * defaults.

LineWhat it does here
User ubuntuQuantaCloud instances log in as ubuntu, not root
IdentitiesOnly yesssh offers only this key, even when your agent holds others
ServerAliveInterval 60After 60 seconds of silence ssh checks the connection, and it closes a dead one after 3 unanswered checks
StrictHostKeyChecking accept-newThe first host key an instance shows is saved without a prompt, and a changed key is still refused
UserKnownHostsFileHost keys from short-lived instances stay out of your main known_hosts

Now connect and look at the GPU:

Terminal
ssh quanta-gpu
nvidia-smi

Checking the driver and CUDA version explains what nvidia-smi prints. The first connection also prints a line that starts with Warning: Permanently added and names the host key type.

Two errors can show up right after launch. Permission denied (publickey) means the instance was launched with a different key from the one in IdentityFile. A refused or timed-out connection in the first minutes usually means the instance is still booting.

Open the instance in VS Code Remote-SSH#

VS Code Remote-SSH reads the same ~/.ssh/config, so the quanta-gpu alias is all it needs:

  1. Install the Remote - SSH extension (ms-vscode-remote.remote-ssh) in VS Code on your laptop.
  2. Open the Command Palette with F1, run Remote-SSH: Connect to Host and pick quanta-gpu.
  3. If VS Code asks for the remote platform, choose Linux. It saves the answer in the remote.SSH.remotePlatform setting.
  4. Wait while VS Code installs its server in ~/.vscode-server on the instance.
  5. Choose File, Open Folder and open /home/ubuntu.

Extensions such as Python and Jupyter run on the instance, not on your laptop, and a new instance starts with none of them. List them once in your VS Code settings.json, and VS Code installs them on every SSH host you open:

json
"remote.SSH.defaultExtensions": [
    "ms-python.python",
    "ms-toolsai.jupyter"
]

With the Jupyter extension on the instance you can skip a Jupyter server altogether: open a .ipynb file in the remote window and choose a Python environment on the instance as the kernel. VS Code offers to install ipykernel into that environment if it is missing.

Tunnel Jupyter, ComfyUI, vLLM and Ollama#

The honest answer for web UIs on a GPU server is to forward a port, not to open one. This is for apps you start yourself on the Bare Metal template: the app templates open from the console instead, as the end of this section explains. -L 8888:127.0.0.1:8888 tells ssh to listen on port 8888 on your laptop and deliver the traffic to 127.0.0.1:8888 as the instance sees it, and -N skips the remote shell:

App you start yourselfDefault portListens on by defaultTunnel command on your laptop
Jupyter8888localhostssh -N -L 8888:127.0.0.1:8888 quanta-gpu
ComfyUI8188127.0.0.1ssh -N -L 8188:127.0.0.1:8188 quanta-gpu
vLLM (vllm serve)8000Every interface, so start it with --host 127.0.0.1ssh -N -L 8000:127.0.0.1:8000 quanta-gpu
Ollama11434127.0.0.1ssh -N -L 11434:127.0.0.1:11434 quanta-gpu

Leave the tunnel running and open http://127.0.0.1:8188, or the matching port, in your laptop's browser. A Jupyter server you start yourself prints a URL with a token, so open the same path and token on your side. If port 8888 is taken on the instance, Jupyter moves to the next free port, so forward the one it prints. For the APIs, test from your laptop with curl http://127.0.0.1:8000/v1/models for vLLM or curl http://127.0.0.1:11434/api/tags for Ollama.

One ssh command can carry several tunnels, as in ssh -N -L 8888:127.0.0.1:8888 -L 8188:127.0.0.1:8188 quanta-gpu. When a port is already taken on your laptop, change only the first number: -L 18888:127.0.0.1:8888 puts the instance's Jupyter on local port 18888. For ports you want every time, add a line such as LocalForward 8188 127.0.0.1:8188 to the Host quanta-gpu block, and the tunnel opens with every connection, VS Code included. For a one-off port inside VS Code, run Forward a Port from the Command Palette or use Add Port in the Ports view.

When the app runs in Docker, publish it on loopback only, as in -p 127.0.0.1:8188:8188 for ComfyUI in Docker. A plain -p 8188:8188 listens on every interface, and Docker's rules bypass ufw.

The app templates need none of this. With PyTorch + Jupyter, ComfyUI or Open WebUI + Ollama, Open Application on the Deployments page opens the app at a private URL behind your QuantaCloud login. The PyTorch + Jupyter template has no token to copy, because the QuantaCloud login takes its place, and Open WebUI asks for its own account as well. Each template's container publishes only its own app port on 127.0.0.1, so a tunnel to port 11434 does not reach the Ollama inside the Open WebUI + Ollama template: the Ollama row above is for an Ollama you install yourself. The Jupyter, ComfyUI and Open WebUI pages describe each template. Tunnels are for what you run yourself on the Bare Metal template, such as vLLM in Docker.

Keep long jobs alive with tmux#

Long jobs belong inside tmux, not in a bare SSH session. A process started in a plain session stops when that session drops, and a sleeping laptop or a VS Code reload is enough to drop it:

Terminal
tmux new -s train        # start a session called train
python3 train.py         # run the job inside it
# detach: press Ctrl-b, then d
tmux ls                  # later, from any SSH or VS Code terminal
tmux attach -t train     # reattach and watch the output

If tmux is missing, sudo apt-get update && sudo apt-get install -y tmux adds it. A tmux session protects a job from a dropped connection. Nothing protects it from a stop, which terminates the instance and deletes its disk.

Keep every port except SSH closed#

My rule is that only SSH listens on the instance's public IP. The defaults show why. The Ollama API has no authentication, and neither does ComfyUI. vllm serve binds every interface, and its --api-key guards the /v1 routes while /health, /tokenize and others stay open. A Jupyter token is all that stands between the internet and a terminal on your GPU.

When someone else needs the app, put a reverse proxy with TLS and authentication in front of it and allow only their addresses in a firewall. The template's private URL cannot be shared that way, because it opens only for the account that owns the deployment. Using the Ollama API on a remote GPU walks through a locked-down Ollama setup.

Move to a new instance#

Every launch is a new machine, because stopping a QuantaCloud instance terminates it and deletes its disk. There is no restart, so moving over takes five steps:

  1. Copy what you need off the old instance before you stop it:

    Terminal
    scp -r quanta-gpu:~/outputs ./outputs-$(date +%F)
    
  2. Stop the old instance, launch the new one and pick your key in the SSH key dropdown again.

  3. Put the new IP from the Deployments page in the HostName line of the Host quanta-gpu block.

  4. Remove the old instance's host key:

    Terminal
    ssh-keygen -R 203.0.113.10 -f ~/.ssh/known_hosts_quantacloud
    
  5. Run ssh quanta-gpu. VS Code reconnects to the same alias and installs its server and your default extensions again.

If ssh refuses the connection with the warning below, the IP address now belongs to a machine with a different host key:

Output
@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
@    WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED!     @
@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
IT IS POSSIBLE THAT SOMEONE IS DOING SOMETHING NASTY!

If you just launched a new instance on that IP, this is expected: run ssh-keygen -R for the IP and connect again. Do not silence it with StrictHostKeyChecking no, because that also hides the day the warning is real. Your models and packages are gone too, so script the Hugging Face downloads and rerun them after each launch.


The decision rule is short: if the app has a QuantaCloud template, open it with Open Application, and for everything you run yourself, tunnel it over SSH and keep the port closed. Launch an RTX A6000 on the Bare Metal template, paste its IP into the Host quanta-gpu block and connect VS Code, then check the driver and CUDA version before you install anything.

Launch an RTX A6000 with the Bare Metal template

Keep building

Choose your next step.