GPU guide

Stop paying for idle GPUs: auto-stop a QuantaCloud instance

A watchdog script that reads nvidia-smi and stops an idle QuantaCloud GPU instance through the REST API, run by systemd or cron. Copy results off first.

Faiz Ahmed10 min read

QuantaCloud has no built-in idle shutdown, so the way to stop paying for an idle GPU is a small watchdog on the VM itself. It reads GPU utilization from nvidia-smi, and once every GPU has been idle for a set time, it stops the deployment through QuantaCloud's REST API. The script below is about 60 lines of bash, runs under systemd or cron, and starts in a dry-run mode, so you can watch it work before it can stop anything.

Two facts decide how you use it. Billing runs from launch to stop: the first hour is charged at launch, each further hour when the previous one is used up, and only the unused seconds of the current hour come back when you stop, so an idle GPU costs exactly what a busy one does. And stopping terminates the instance and deletes its disk. There are no volumes or snapshots, so anything you want to keep has to be off the disk before the watchdog fires.

What an idle GPU costs#

The cost of forgetting is the hourly price times the idle hours. These are 1x and 2x VMs at the prices on 2026-09-27 (our calculation).

VMPrice per hour on 2026-09-2712 idle hours overnight60 idle hours over a weekend
1x RTX A6000$0.48$5.76$28.80
1x H100 PCIe$2.59$31.08$155.40
1x H200 NVL$3.43$41.16$205.80
2x H200 NVL$6.86$82.32$411.60

Today's prices per GPU-hour:

GPUMemoryFromAvailable now
RTX A600048 GB$0.48/GPU-hrYes
H100 PCIe80 GB$2.59/GPU-hrYes
H200 NVL-Not listedNo

Prices checked 6 Oct 2026, 01:40 UTC

An idle instance also drains the balance a running job needs. When the balance cannot cover the next hour, QuantaCloud terminates the instance and deletes its disk, with no grace period. The pricing page has the full billing rules.

How the watchdog decides a GPU is idle#

The watchdog reads one number per GPU: utilization.gpu from nvidia-smi. NVIDIA defines it as the percentage of time over the last sample period, between 1/6 of a second and 1 second depending on the product, in which one or more kernels ran on the GPU. The script takes a reading every 30 seconds, looks at the busiest GPU, and treats 5 percent or less as idle. Any busier reading resets the idle clock. When the clock reaches 30 minutes, the script asks the API whether the deployment is still active, and only then stops it.

What the VM is doingWhat utilization.gpu readsWhat the watchdog does
Training or batch inferenceRises whenever kernels runResets the idle clock
Downloading weights or a dataset0, since no kernels runCounts idle minutes
A notebook paused between cells0Counts idle minutes
A model server with no requests0, even with the model loadedCounts idle minutes
nvidia-smi fails, hangs or gives no number for one GPUNo readingCounts the check as busy

The rule I follow is that the idle window must be longer than the longest quiet stretch of real work, such as a large download or a long pause in a notebook. I would start at 30 minutes and pause the watchdog during big downloads. If the VM serves a model to users, the watchdog stops it after a quiet half hour, which is either exactly what you want or a reason not to install it there.

Move results off the disk first#

The watchdog is only as safe as your data habits. When it stops the instance, everything on the disk goes: checkpoints, outputs, downloaded models and the watchdog itself. Two patterns keep that from hurting.

The first is to write results off the VM while the job runs. hf upload with --every pushes a folder to a Hugging Face repository at a fixed interval in minutes, and an rsync or object-storage sync from your training loop does the same for any destination. A stop then costs you nothing but the rest of the idle time.

The second is the watchdog's pre-stop command. Put a command in BEFORE_STOP, such as an rsync to a machine you control, and the watchdog stops the instance only when that command succeeds. If it fails, the instance keeps running and billing, and the watchdog tries again at the next check. I would rather pay for an idle hour than delete the only copy of a checkpoint. The file transfer guide covers rsync and object storage, and the data persistence docs list everything that lives on the disk.

Store an API key on the VM#

The watchdog authenticates with a QuantaCloud API key. Create one in the console under API Keys, name it after this VM, and copy it at once, because the console shows it only once. Keys start with gpu_live_. Treat the key like a password: it is not limited to one instance, and whoever holds it can list, launch and stop deployments on your account and spend your credit. So it lives only on this VM, in a file only the ubuntu user can read, and you revoke it in the console once the instance is gone, because deleting the disk does not revoke it. The API keys docs show where the keys live in the console.

  1. SSH in as ubuntu (connecting over SSH covers keys), then create the settings file with permissions for your user only and type the key into it. read -s keeps the key off the screen and out of your shell history:

    Terminal
    mkdir -p ~/.config ~/bin
    install -m 600 /dev/null ~/.config/gpu-idle-stop.env
    read -rsp "QuantaCloud API key: " QC_API_KEY; echo
    echo "QC_API_KEY=$QC_API_KEY" >> ~/.config/gpu-idle-stop.env
    
  2. Find this VM's deployment ID. The list endpoint returns each deployment with its public IP, so pick the active one whose IP you connected to:

    Terminal
    curl -fsS -H "Authorization: Bearer $QC_API_KEY" \
      "https://core.quantacloud.net/api/v1/deployments?size=100" | python3 -c '
    import json, sys
    for d in json.load(sys.stdin)["data"]:
        if d["status"] == "active":
            print(d["id"], d["gpu"]["count"], "x", d["gpu"]["type"], (d["connection"] or {}).get("public_ip"))
    '
    
  3. Add the settings, starting in dry-run mode, then open the file and replace the placeholder with the ID from step 2:

    Terminal
    cat >> ~/.config/gpu-idle-stop.env <<'EOF'
    DEPLOYMENT_ID=paste-the-id-here
    IDLE_MINUTES=30
    BUSY_ABOVE=5
    CHECK_SECONDS=30
    DRY_RUN=1
    BEFORE_STOP=''
    EOF
    nano ~/.config/gpu-idle-stop.env
    unset QC_API_KEY
    

For a pre-stop copy, the BEFORE_STOP line takes one shell command, for example BEFORE_STOP='rsync -a /home/ubuntu/outputs/ backup@203.0.113.20:gpu-outputs/', which needs an SSH key on this VM that the backup machine accepts.

Install the watchdog script#

The script reads the settings file, checks the GPUs every CHECK_SECONDS, and logs one line per idle check. Any DRY_RUN value other than 0 is a dry run: the script does everything except the stop. It confirms through the API that the deployment is active, runs BEFORE_STOP, and logs the stop it would make. Only DRY_RUN=0 lets it call the stop endpoint, and it exits once that request succeeds.

Terminal
cat > ~/bin/gpu-idle-stop <<'EOF'
#!/usr/bin/env bash
# Stop this QuantaCloud deployment once every GPU has been idle for IDLE_MINUTES.
set -u
CONFIG="${CONFIG:-/home/ubuntu/.config/gpu-idle-stop.env}"
. "$CONFIG"
: "${QC_API_KEY:?missing in $CONFIG}" "${DEPLOYMENT_ID:?missing in $CONFIG}"
IDLE_MINUTES="${IDLE_MINUTES:-30}"
BUSY_ABOVE="${BUSY_ABOVE:-5}"
CHECK_SECONDS="${CHECK_SECONDS:-30}"
DRY_RUN="${DRY_RUN:-1}"
BEFORE_STOP="${BEFORE_STOP:-}"
API=https://core.quantacloud.net/api/v1

log() { echo "$(date -u +%FT%TZ) $*"; }

for n in "$IDLE_MINUTES" "$BUSY_ABOVE" "$CHECK_SECONDS"; do
  [[ "$n" =~ ^[0-9]+$ ]] || { log "IDLE_MINUTES, BUSY_ABOVE and CHECK_SECONDS must be whole numbers"; exit 1; }
done

# curl reads the key from stdin, so it never appears in the process list.
api() { curl -fsS -m 30 -H @- "$@" <<< "Authorization: Bearer $QC_API_KEY"; }

# Prints the busiest GPU's utilization, or fails if any GPU gives no number.
busiest_gpu() {
  local out v max=-1
  out=$(timeout 60 nvidia-smi --query-gpu=utilization.gpu --format=csv,noheader,nounits) || return 1
  while read -r v; do
    [[ "$v" =~ ^[0-9]+$ ]] || return 1
    (( v > max )) && max=$v
  done <<< "$out"
  (( max >= 0 )) && echo "$max"
}

deployment_status() {
  api "$API/deployments/$DEPLOYMENT_ID" |
    python3 -c 'import json, sys; print(json.load(sys.stdin)["status"])' 2> /dev/null
}

idle_since=""
while true; do
  now=$(date +%s)
  if ! util=$(busiest_gpu); then
    log "no reading from nvidia-smi, counting this check as busy"
    idle_since=""
  elif (( util > BUSY_ABOVE )); then
    idle_since=""
  else
    idle_since="${idle_since:-$now}"
    idle_min=$(( (now - idle_since) / 60 ))
    log "busiest GPU at ${util}%, idle for ${idle_min} min"
    if (( idle_min >= IDLE_MINUTES )); then
      status=$(deployment_status)
      if [[ "$status" =~ ^(stopping|terminated|interrupted|failed)$ ]]; then
        log "deployment is already $status, so there is nothing to stop"
        exit 0
      elif [[ "$status" != active ]]; then
        log "could not confirm that the deployment is active; trying again next check"
      elif [[ -n "$BEFORE_STOP" ]] && ! bash -c "$BEFORE_STOP"; then
        log "BEFORE_STOP failed, so the instance stays up; trying again next check"
      elif [[ "$DRY_RUN" != 0 ]]; then
        log "DRY_RUN: deployment $DEPLOYMENT_ID is active and would be stopped now"
        idle_since=""
      elif api -X POST "$API/deployments/$DEPLOYMENT_ID/stop" > /dev/null; then
        log "stop requested for deployment $DEPLOYMENT_ID"
        exit 0
      else
        log "stop request failed; trying again next check"
      fi
    fi
  fi
  sleep "$CHECK_SECONDS"
done
EOF
chmod 700 ~/bin/gpu-idle-stop

Four choices in the script are deliberate. A failed reading counts as busy, whether nvidia-smi errors out, hangs for a minute or gives no number for one GPU, so a driver problem never stops the instance by accident. The busiest GPU decides, so one working GPU keeps a multi-GPU VM alive. Before every stop request the script asks the API for the deployment's status, and it stops only an active deployment. A stop request that fails, for example on a network error, is retried at the next check only if the deployment is still active, so the watchdog never stops the same deployment twice. And the key reaches curl on its standard input, not on its command line, where other processes on the VM could read it.

Run it with systemd or cron#

Systemd keeps the watchdog running and restarts it if it crashes. The unit runs the script as ubuntu, and Restart=on-failure brings it back after an error but not after the clean exit that follows a successful stop:

Terminal
sudo tee /etc/systemd/system/gpu-idle-stop.service > /dev/null <<'EOF'
[Unit]
Description=Stop this QuantaCloud deployment when its GPUs sit idle
After=network-online.target
Wants=network-online.target

[Service]
User=ubuntu
ExecStart=/home/ubuntu/bin/gpu-idle-stop
Restart=on-failure
RestartSec=30

[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now gpu-idle-stop
sudo journalctl -u gpu-idle-stop -f

Cron is the option without sudo. This line tries to start the script every minute as ubuntu, and flock -n makes each attempt exit at once while a copy is already running:

Terminal
(crontab -l 2>/dev/null; echo '* * * * * flock -n /tmp/gpu-idle-stop.lock /home/ubuntu/bin/gpu-idle-stop >> /home/ubuntu/gpu-idle-stop.log 2>&1') | crontab -
tail -f ~/gpu-idle-stop.log

Use one of the two, not both.

Test the dry run, then arm it#

The dry run proves the clock, the key and the deployment ID before the watchdog can delete anything.

  1. In ~/.config/gpu-idle-stop.env, set IDLE_MINUTES=2. Restart the watchdog with sudo systemctl restart gpu-idle-stop, or with cron run pkill -f bin/gpu-idle-stop and wait up to a minute.
  2. Leave the GPU idle and watch the log. After about 2 minutes it prints DRY_RUN: deployment with the ID and is active and would be stopped now. If it prints could not confirm that the deployment is active instead, the curl error just above that line gives the HTTP status: 401 means a wrong key, and 404 a wrong deployment ID.
  3. Start your GPU job and check that the idle lines stop while it runs.
  4. Copy off anything on the disk that is not stored somewhere else, because the first real stop deletes it and ends the billing for this instance. Then set IDLE_MINUTES=30 and DRY_RUN=0, and restart the watchdog the same way. From now on it stops the instance.

To pause it for a large download or a planned break, run sudo systemctl stop gpu-idle-stop, and sudo systemctl start gpu-idle-stop to resume. With cron, comment out the line in crontab -e and run pkill -f bin/gpu-idle-stop.

What happens when it fires#

The watchdog reads the deployment's status with GET /api/v1/deployments/{id}, then calls QuantaCloud's documented stop endpoint, POST /api/v1/deployments/{id}/stop. QuantaCloud terminates the instance, deletes its disk and refunds the unused seconds of the current hour to your balance, and the API keeps the deployment's record. From your laptop, GET /api/v1/deployments/{id} shows the status moving through stopping to terminated, and the deployments API reference lists every field.

Auto-stop questions#

Does QuantaCloud stop idle instances by itself?

No. An instance runs, and bills, until you stop it or until your balance cannot cover the next hour, when QuantaCloud terminates it and deletes its disk.

Do I pay for the rest of the hour after the watchdog stops the instance?

No. The unused seconds of the current hour are refunded when the instance stops, so there is no reason to wait for the end of an hour before stopping.

Can I pause an instance instead of stopping it?

No. QuantaCloud has no stop and start: stopping always terminates the instance and deletes its disk. Keep a setup script that rebuilds the next instance, model downloads included, as the Hugging Face download guide shows.

Is it safe to keep the API key on the VM?

It is as safe as the VM. The file is readable only by ubuntu and root, which on this VM means anyone who can SSH in. Use a key made for this VM, and revoke it in the console when the instance is gone.


My rule: install the watchdog on any VM you might forget, give it an idle window longer than your longest download, and move your results somewhere else before you set DRY_RUN=0. If the watchdog fires in the middle of training, the GPU is waiting on something else, and low GPU utilization covers the usual causes. More how-tos are in the guides.

Launch an RTX A6000 with Bare Metal

Keep building

Choose your next step.