For offline work, such as a dataset to label, summarize or translate, skip the server and pass every prompt to one call of vLLM's LLM class. vLLM's own docs say it plainly: the class batches the prompts automatically within the GPU's memory, and for the best performance you put all of them into a single list. Under the hood it runs the same continuous batching as vllm serve, so one GPU works through thousands of prompts with hundreds of them in flight at once. Use vllm run-batch when your requests are already in OpenAI's batch file format, and vllm serve when another program sends requests as they arrive.
How continuous batching works#
Continuous batching schedules one step at a time, not one batch at a time. A static batch waits for its longest answer to finish before new prompts can start, so short answers leave the GPU idle. The Orca paper replaced that with iteration-level scheduling, and vLLM does the same: at every step it generates one token for each running request, fills the rest of the step's budget with prompt tokens from waiting requests, and starts a new request the moment one finishes. Long prompts are split into chunks so they do not stall the requests that are generating.
PagedAttention is what makes that possible in memory. Each request's KV cache lives in blocks of 16 tokens that vLLM allocates as the request grows, instead of reserving the maximum length up front. The vLLM paper measured earlier serving systems using 20.4% to 38.2% of their KV cache memory for actual tokens, against 96.3% in vLLM, and reports 2 to 4 times the throughput of FasterTransformer and Orca at the same latency. More requests fit in the same memory, and more requests per step is where batch throughput comes from.
Four settings bound the batch:
| Setting | What it limits | Default in the LLM class |
|---|---|---|
max_num_seqs | Requests running in the same step | 256 on 48 GB cards and the A100, 1,024 on GPUs with 70 GiB or more |
max_num_batched_tokens | Tokens processed per step | 8,192 on 48 GB cards and the A100, 16,384 on GPUs with 70 GiB or more |
gpu_memory_utilization | Share of the GPU vLLM claims | 0.92 |
max_model_len | Longest prompt plus answer | The model's own maximum |
When the running requests fill the KV cache, vLLM preempts some and recomputes them later rather than failing them. The KV cache guide shows how many tokens a given model and GPU hold.
Launch a GPU and install vLLM#
A 48 GB RTX A6000 runs an 8B model in BF16 with room for hundreds of short requests at once. Launch it with the Bare Metal template, a plain Ubuntu 22.04 VM with the NVIDIA driver and Docker.
Launch an RTX A6000 (Bare Metal)Connect as ubuntu over SSH and install vLLM into a fresh environment. Installing vLLM explains the driver check that decides the wheel. On an R580 or newer driver:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
mkdir -p ~/batch && cd ~/batch
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
uv pip install vllm==0.30.0 --torch-backend=auto
A batch job can run for hours, so start it inside tmux, where it survives a dropped SSH connection. Install tmux with sudo apt-get update && sudo apt-get install -y tmux, open a session with tmux new -s batch, and activate the environment in that session with source ~/batch/.venv/bin/activate.
A batch job in one script#
The script reads prompts from a JSON Lines file, sends all of them to llm.chat in one call, and writes one answer per line with the original id. First, a sample input to test with, 2,000 support tickets, saved as make_input.py:
import json
topics = ["billing", "login", "latency", "data export", "permissions"]
with open("input.jsonl", "w") as f:
for i in range(2000):
topic = topics[i % len(topics)]
prompt = f"A customer reports a problem with {topic}. What is the most likely cause? (ticket {i})"
f.write(json.dumps({"id": i, "prompt": prompt}) + "\n")
Then the job itself, saved as batch.py:
import json
import sys
import time
from vllm import LLM, SamplingParams
SYSTEM = "You are a support analyst for a GPU cloud. Answer in at most three sentences."
def main(in_path: str, out_path: str) -> None:
with open(in_path) as f:
rows = [json.loads(line) for line in f]
conversations = [
[{"role": "system", "content": SYSTEM}, {"role": "user", "content": row["prompt"]}]
for row in rows
]
llm = LLM(model="Qwen/Qwen3-8B", max_model_len=8192)
params = SamplingParams(temperature=0.7, top_p=0.8, top_k=20, max_tokens=512)
start = time.perf_counter()
outputs = llm.chat(conversations, params, chat_template_kwargs={"enable_thinking": False})
elapsed = time.perf_counter() - start
generated = 0
with open(out_path, "w") as f:
for row, out in zip(rows, outputs):
completion = out.outputs[0]
generated += len(completion.token_ids)
f.write(json.dumps({
"id": row["id"],
"answer": completion.text,
"finish_reason": completion.finish_reason,
"cached_prompt_tokens": out.num_cached_tokens,
}) + "\n")
print(f"{len(rows)} prompts, {generated} output tokens in {elapsed:.0f} s, "
f"{generated / elapsed:.0f} output tokens/s")
if __name__ == "__main__":
main(sys.argv[1], sys.argv[2])
Run it:
python make_input.py
python batch.py input.jsonl output.jsonl
The first run downloads 16.4 GB of weights. While it works, a progress bar labelled Processed prompts shows the estimated input and output tokens per second, and the outputs come back in the same order as the prompts.
Three details in that script matter more than they look. max_tokens defaults to 16 in SamplingParams, so leave it out and every answer stops after 16 tokens. The sampling values are the ones Qwen recommends for Qwen3 without thinking, and enable_thinking goes to the chat template to switch the thinking off. And the if __name__ == "__main__": guard is what vLLM's troubleshooting docs ask for, because its worker processes can start with Python's spawn method, which imports your script again.
Prefix caching for shared prompts#
Prefix caching is on by default, and it rewards prompts that start the same way. vLLM hashes each full 16-token block of a prompt together with everything before it, so every request that begins with the same system prompt, instructions or few-shot examples reuses the cached blocks instead of computing them again. Put the shared text first and the per-row data last: a record ID at the top of each prompt breaks the match for everything after it.
The cached_prompt_tokens field in the output shows how much of each prompt came from the cache. Requests that start after the shared prefix has been computed report it as cached: with a one-line system prompt like the one above that is a block or two, and with long instructions or few-shot examples it is most of the prompt.
Tune the batch for throughput#
The first rule is to hand vLLM all the work at once. A loop that calls llm.chat once per prompt runs one request at a time and throws continuous batching away. Once the whole list goes in one call, these settings move throughput, in the order I would try them:
- Set
max_model_lento what the job needs. The script uses 8,192, which covers a short prompt plus 512 output tokens with plenty of room, and leaves the startup check nothing to fail on. - Read the startup log line
GPU KV cache size: N tokens, Maximum concurrency for 8,192 tokens per request. If the pool is small next to your number of prompts times their length, the batch is memory-bound, and smaller weights or a bigger GPU help more than any flag. If vLLM refuses to start, fixing vLLM out-of-memory errors goes through each error. - Raise
max_num_batched_tokens. vLLM's tuning guide recommends more than 8,192 for throughput, especially for smaller models on large GPUs, for exampleLLM(..., max_num_batched_tokens=16384). - Raise
max_num_seqswhen prompts are short and the KV cache has room, or lower it if requests keep getting preempted.LLM(..., disable_log_stats=False)turns on the periodic stats line, which showsPreemptions:and the KV cache usage. - Load FP8 or 4-bit weights to leave more of the GPU for the KV cache. The vLLM quantization guide matches formats to GPUs.
Use every GPU on a multi-GPU VM#
When the model fits on one GPU, run one copy per GPU on its own share of the input. That is data parallelism, and it avoids the communication that tensor parallelism adds between GPUs. On a 2x VM:
split -n l/2 -d input.jsonl part_
CUDA_VISIBLE_DEVICES=0 python batch.py part_00 out_00.jsonl &
CUDA_VISIBLE_DEVICES=1 python batch.py part_01 out_01.jsonl &
wait
cat out_00.jsonl out_01.jsonl > output.jsonl
split -n l/2 cuts the file into two parts without breaking a line, and CUDA_VISIBLE_DEVICES gives each process one GPU. Use tensor parallelism (LLM(..., tensor_parallel_size=2)) only when the model does not fit on one GPU, and vLLM across several GPUs covers that case.
OpenAI batch files with vllm run-batch#
If your requests already exist as an OpenAI batch file, vllm run-batch processes it without any code. Each line is one request:
{"custom_id": "ticket-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "Qwen/Qwen3-8B", "messages": [{"role": "user", "content": "A customer reports a problem with login. What is the most likely cause? (ticket 1)"}], "max_completion_tokens": 512, "chat_template_kwargs": {"enable_thinking": false}}}
vllm run-batch -i requests.jsonl -o results.jsonl --model Qwen/Qwen3-8B --max-model-len 8192
Each line of results.jsonl carries the custom_id and a full chat completion response. It takes the same engine flags as vllm serve and supports /v1/chat/completions, /v1/embeddings and a few other endpoints, which makes it the easy path for moving a batch job off a hosted API.
When to use the server instead#
Use vllm serve when the requests do not exist yet: an application calling the model, several clients at once, streaming answers, or code in another language. The server runs the same continuous batching, so throughput under load is the same engine at work, and deploying vLLM with Docker covers it. The LLM class fits jobs where the input is a file you have now and the output is a file you want at the end. For a dataset too big for one machine, vLLM's docs point to Ray Data, which runs vLLM across a cluster.
Throughput and cost per token#
The number that decides the GPU for a batch job is output tokens per second, because cost per token follows from it. vLLM's offline benchmark measures it on your instance: it runs the LLM engine on random prompts of fixed length and prints a Throughput: line with requests, total tokens and output tokens per second.
vllm bench throughput --model Qwen/Qwen3-8B --max-model-len 32768 \
--dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 1000
Cost per million output tokens is then the hourly price divided by (output tokens per second x 3,600), times 1,000,000.
Before you stop the instance#
Copy the results off before you stop. Stopping a QuantaCloud instance terminates it and deletes its disk, including output.jsonl, the virtual environment and the 16.4 GB of cached weights, so the next launch installs and downloads again. From your laptop:
scp ubuntu@YOUR_INSTANCE_IP:~/batch/output.jsonl .
Billing runs from launch to stop: the first hour is charged at launch and the unused seconds of the last hour are refunded when you stop, so a batch job costs what it runs (how billing works).
vLLM batch inference FAQ#
Why are my outputs only 16 tokens long?
Because SamplingParams defaults to max_tokens=16. Set it to the longest answer you want, as the script does with 512.
Does the LLM class use continuous batching?
Yes. It runs the same engine and scheduler as vllm serve, adding requests to each step as others finish, as long as you pass the prompts in one list.
How many prompts should I pass at once?
All of them, if they fit in host memory. vLLM keeps only as many running as the KV cache and max_num_seqs allow and queues the rest. For inputs too large for RAM, process the file in chunks of a few thousand rows per llm.chat call.
Can I get JSON output for every row?
Yes. Put StructuredOutputsParams(json=schema) from vllm.sampling_params into SamplingParams(structured_outputs=...), and every answer follows the JSON schema.
Which GPU gives the lowest cost per token?
The one that generates the most output tokens per dollar for your model, and only a run on your model tells you which. Run the vllm bench throughput command above on each candidate and divide its hourly price by its output tokens per hour. Unused seconds of the current hour are refunded when you stop, so a short comparison costs only the minutes it runs.
My rule for batch jobs: one llm.chat call with every prompt, the shared text at the front of each prompt, and max_model_len set to what the job needs. Start on an RTX A6000 at $0.48/GPU-hr, read the KV cache line and the tokens per second, and move to a bigger card from the GPU catalog only when the cost per million tokens says it pays. Self-hosted LLM inference covers the serving side of the same stack.