The safest way to use the Ollama API on a remote GPU server is an SSH tunnel: Ollama keeps listening on 127.0.0.1:11434 on the server, and your laptop reaches it as localhost:11434. Ollama's local API has no authentication, so anyone who can reach port 11434 can run your models, pull new ones and delete the ones you have. Keep it on localhost and let SSH decide who gets in.
The setup is two install lines on the server and one SSH command on your laptop. After that, curl, Ollama's Python library and any OpenAI client talk to localhost, and the requests run on the remote GPU.
Run Ollama on the GPU server#
For an API server I would pick the Bare Metal template: a plain Ubuntu 22.04 VM with the NVIDIA driver, where Ollama runs as a system service and nothing else competes for the GPU.
Launch an RTX A6000 with Bare Metal-
SSH in as
ubuntu, with the key you chose at deploy time. The SSH docs show where the address is.ssh -i ~/.ssh/your_key ubuntu@<instance-ip> -
Install a pinned Ollama version with the official script. Ollama's Linux packages are zstd archives, and the script stops if the
zstdtool is missing, so install it first:sudo apt-get update && sudo apt-get install -y zstd curl -fsSL https://ollama.com/install.sh | OLLAMA_VERSION=0.34.4 shThe script finds the NVIDIA driver already installed, sets Ollama up as a systemd service and ends with
The Ollama API is now available at 127.0.0.1:11434.That loopback address is the default, and it is the one you want. Ollama needs driver 550 or newer. -
Pull a model and check the server:
ollama pull gpt-oss:20b curl http://127.0.0.1:11434/api/version
gpt-oss:20b is a 14 GB download and fits a 48 GB RTX A6000 with room for a long context. The Open WebUI page has a model-by-GPU table for larger models, and Ollama GPU requirements covers which models fit on 48, 96 and 141 GB.
If you prefer Docker, run the official image with a pinned tag, published on the loopback address only:
sudo docker run -d --gpus=all --restart unless-stopped \
-v ollama:/root/.ollama -p 127.0.0.1:11434:11434 \
--name ollama ollama/ollama:0.34.4
Ollama's own Docker example uses -p 11434:11434, which publishes the port on every interface, and Docker's published ports bypass ufw. Keep the 127.0.0.1: prefix. Docker releases older than 28.0.0 also let hosts on the same network segment reach ports published to 127.0.0.1, so check docker version first.
With Docker, the CLI runs inside the container: sudo docker exec ollama ollama pull gpt-oss:20b.
On the Open WebUI + Ollama template#
On the template, Ollama runs inside the app container next to Open WebUI, and the VM publishes only Open WebUI's port, at 127.0.0.1:8080. Port 11434 stays inside the container, so your code goes through Open WebUI: it passes Ollama's API through under /ollama and serves an OpenAI-compatible endpoint at /api/chat/completions, both behind an Open WebUI API key.
-
In Open WebUI, open Settings > Admin > Authentication, switch on API Keys and click Save.
-
Open Settings > Account, click Show next to Secrets in the API keys section, then Create new secret key, and copy the key that starts with
sk-. -
From your laptop, forward Open WebUI's port:
ssh -N -L 8080:127.0.0.1:8080 -i ~/.ssh/your_key ubuntu@<instance-ip>. -
Call it with the key:
export OPENWEBUI_KEY=sk-... curl http://localhost:8080/ollama/api/tags \ -H "Authorization: Bearer $OPENWEBUI_KEY" curl http://localhost:8080/api/chat/completions \ -H "Authorization: Bearer $OPENWEBUI_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "gpt-oss:20b", "messages": [{"role": "user", "content": "Say hello in one sentence."}]}'
The key acts as your Open WebUI user and does not expire. Each account holds one key, and creating a new one replaces the old. The Open WebUI setup guide covers the template itself.
Open the tunnel from your laptop#
One SSH command opens the tunnel:
ssh -N -L 11434:127.0.0.1:11434 -o ServerAliveInterval=60 \
-i ~/.ssh/your_key ubuntu@<instance-ip>
-L maps port 11434 on your laptop to 127.0.0.1:11434 on the server, -N skips the remote shell because you only want the forward, and ServerAliveInterval=60 sends a keepalive every 60 seconds when the connection is quiet. Leave it running in its own terminal. If Ollama also runs on your laptop, 11434 is already taken: use -L 11435:127.0.0.1:11434 and call port 11435 instead.
Check it from the laptop:
curl http://localhost:11434/api/version
It should print the server's version, {"version":"0.34.4"} with the install above. For a reusable host entry in ~/.ssh/config and VS Code on the same machine, see connecting to a GPU server over SSH.
Call the API with curl#
Every Ollama endpoint lives under /api. List the models on the server:
curl http://localhost:11434/api/tags
Send a chat and get one JSON object back:
curl http://localhost:11434/api/chat -d '{
"model": "gpt-oss:20b",
"messages": [{"role": "user", "content": "Why is the sky blue?"}],
"stream": false
}'
Leave out "stream": false and the reply arrives as a stream of JSON objects, one per chunk, which is what a chat interface wants. The final object carries timings in nanoseconds: eval_count tokens were generated in eval_duration, so tokens per second is eval_count divided by eval_duration, times 1,000,000,000.
Two request fields are worth setting on most calls: "keep_alive": "30m" keeps the model in GPU memory for 30 minutes after the request instead of the default 5, which saves a reload between calls, and "options": {"num_ctx": 16384} sets the context length for that request.
Call it from Python#
Ollama's Python library talks to any host you give it:
pip install ollama
from ollama import Client
client = Client(host="http://localhost:11434")
response = client.chat(
model="gpt-oss:20b",
messages=[{"role": "user", "content": "Why is the sky blue?"}],
)
print(response.message.content)
print(f"{response.eval_count / response.eval_duration * 1e9:.1f} tokens/s")
for chunk in client.chat(
model="gpt-oss:20b",
messages=[{"role": "user", "content": "Write a haiku about GPUs."}],
stream=True,
):
print(chunk.message.content, end="", flush=True)
With stream=True the chunks arrive as the model writes. gpt-oss is a reasoning model, and Ollama returns its reasoning separately in message.thinking, so message.content holds only the answer.
Use the OpenAI-compatible endpoint#
Ollama also speaks a subset of the OpenAI API under /v1, so code written for the OpenAI SDK works by changing the base URL:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1/",
api_key="ollama", # required by the client, ignored by Ollama
)
completion = client.chat.completions.create(
model="gpt-oss:20b",
messages=[{"role": "user", "content": "Say this is a test"}],
)
print(completion.choices[0].message.content)
The same base URL serves /v1/chat/completions, /v1/completions, /v1/models, /v1/embeddings and a stateless /v1/responses. One gap: the OpenAI API has no field for context size, so set it on the server with OLLAMA_CONTEXT_LENGTH, or create a model variant with PARAMETER num_ctx in a Modelfile. The same call with curl:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "gpt-oss:20b", "messages": [{"role": "user", "content": "Say this is a test"}]}'
Tune the server for API traffic#
Ollama reads its settings from environment variables, and on the systemd install you set them with an override:
sudo systemctl edit ollama
[Service]
Environment="OLLAMA_KEEP_ALIVE=1h"
Environment="OLLAMA_NUM_PARALLEL=4"
Environment="OLLAMA_CONTEXT_LENGTH=32768"
sudo systemctl daemon-reload
sudo systemctl restart ollama
The defaults are 5 minutes of keep-alive and 1 request at a time per model, with up to 512 requests queued before Ollama answers 503. Parallel slots cost memory: it scales with OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH, so four slots of 32k need the KV cache of a 131,072-token context (our calculation: 4 x 32,768). The KV cache guide shows how many gigabytes that is for a given model. With Docker, pass the same variables with -e. vLLM vs Ollama compares the two servers under concurrent requests.
Serve an app without a tunnel#
The rule I follow: if anything other than my own laptop calls Ollama, it goes through a reverse proxy that checks a token over HTTPS, and Ollama itself stays on 127.0.0.1. Setting OLLAMA_HOST=0.0.0.0, or publishing 11434 on a public interface, puts an API with no login on the internet.
Caddy does the job in a few lines and fetches its own certificate. You need a domain name pointed at the VM's public IP, and inbound ports 80 and 443 open to the VM.
-
Make a token with
openssl rand -hex 32. -
Write
~/caddy/conf/Caddyfilewith your domain and the token:ollama.example.com { @unauthorized not header Authorization "Bearer PASTE_THE_TOKEN_HERE" respond @unauthorized 401 reverse_proxy 127.0.0.1:11434 { header_up Host localhost:11434 } } -
Run Caddy on the host network, so it can reach Ollama on 127.0.0.1:
sudo docker run -d --name caddy --network host --restart unless-stopped \ -v ~/caddy/conf:/etc/caddy -v caddy_data:/data caddy:2.11.4
Requests without the right Authorization header get a 401, and the rest reach Ollama with the Host header it expects. The OpenAI SDK already sends its key as a bearer token, so pass the token as api_key and https://ollama.example.com/v1/ as base_url. With curl or Ollama's library, add the header yourself:
from ollama import Client
client = Client(
host="https://ollama.example.com",
headers={"Authorization": "Bearer PASTE_THE_TOKEN_HERE"},
)
Before you stop the instance#
Stopping the instance deletes every model you pulled. QuantaCloud's Stop terminates the VM and deletes its disk, including any model you built with ollama create, and there are no volumes or snapshots to bring it back. Keep your Modelfiles in version control on your laptop, and keep a pull list so the next instance is ready with one command:
for m in gpt-oss:20b qwen3.6:27b; do ollama pull "$m"; done
Billing stops when you stop, and the unused seconds of the current hour are refunded. The pricing page has the details.
Troubleshooting#
Each failure below has one likely cause and one fix.
| Symptom | Likely cause | Fix |
|---|---|---|
connection refused on localhost:11434 | The tunnel is not running, or Ollama is not | Restart the ssh -N -L command. On the server, run systemctl status ollama |
Address already in use when you open the tunnel | Something on your laptop already uses 11434, often a local Ollama | Forward 11435 instead |
| 404 and an error that the model does not exist | The model was never pulled on this instance | Pull it first. Every new instance starts empty |
| 503 from the API | The request queue is full | Raise OLLAMA_NUM_PARALLEL if memory allows, or slow the client down |
| The first reply after a pause is slow | Ollama unloaded the model after 5 idle minutes | Set keep_alive per request, or OLLAMA_KEEP_ALIVE on the server |
My rule for the Ollama API: an SSH tunnel for anything you run yourself, a token-checking HTTPS proxy for anything that runs without you, and port 11434 never on a public interface. Start with a Bare Metal RTX A6000, install Ollama, open the tunnel and point your existing OpenAI client at localhost. If you want a chat interface on the same GPU as well, set up Open WebUI with Ollama. To write ComfyUI prompts with it, see Ollama inside ComfyUI.