GPU guide

vLLM vs SGLang: how the two serving engines differ

vLLM vs SGLang from both projects' docs and source: prefix caching, scheduling, structured output, memory flags, security defaults and which to run.

Faiz Ahmed11 min read

The short answer is: start with vLLM on a single QuantaCloud GPU, and try SGLang when most of your traffic shares long prompt prefixes. The two engines now share most of their design: continuous batching, a paged KV cache, chunked prefill, automatic prefix caching, speculative decoding, the same parallelism options and an OpenAI-compatible API. SGLang's README even lists vLLM among the projects it learned the design from and reused code from. What separates them in a deployment is narrower: how each one caches and schedules shared prefixes, which grammars it accepts for structured output, how it budgets GPU memory, and what it exposes by default.

My reasons for vLLM as the default are practical. vLLM 0.30.0 ships a CUDA 12.9 build for older drivers, while SGLang 0.5.20 needs CUDA 13. Its image is an 8.7 GB download against 15.1 GB for SGLang, and a fresh QuantaCloud instance downloads the image again on every launch. SGLang's case is a scheduler that can reorder waiting requests to reuse cached prefixes, which vLLM does not offer, and safer defaults for its address and API key.

vLLM and SGLang side by side#

The table compares the current releases, vLLM 0.30.0 (2026-09-22) and SGLang 0.5.20 (2026-09-18), from their docs and source.

vLLM 0.30.0SGLang 0.5.20
Installuv pip install vllm, a CUDA 13.0 build, or the cu129 wheel for older driversuv pip install --prerelease=allow sglang, CUDA 13 only
Docker image, amd64 downloadvllm/vllm-openai:v0.30.0, 8.7 GBlmsysorg/sglang:v0.5.20, 15.1 GB, or 13.6 GB for -runtime
Start a servervllm serve <model>python3 -m sglang.launch_server --model-path <model>
Default addressEvery interface, port 8000127.0.0.1, port 30000
What --api-key covers/v1, /v2, /inference and /cohere routesEvery route except /health, /ready and /metrics
Memory setting--gpu-memory-utilization, 0.92 of the GPU for weights, activations and KV cache--mem-fraction-static, weights plus KV pool, set by a heuristic
Prefix cacheHash per full 16-token block, on by defaultRadixAttention radix tree, on by default, LRU eviction
Scheduling policiesfcfs (default), priorityfcfs (default), lpm, dfs-weight, lof, random, priority, routing-key, hrrn
CPU and GPU overlapAsync scheduling, on by defaultOverlap scheduler, on by default
Structured output backendsauto, xgrammar, guidance, outlines, lm-format-enforcerxgrammar (default), outlines, llguidance
ConstraintsChoice, regex, JSON schema, grammar, structural tagJSON schema, regex, EBNF, structural tag
APIsOpenAI (/v1 chat, completions, responses, embeddings), Anthropic-style /v1/messagesOpenAI, native /generate, Anthropic /v1/messages, Ollama-compatible
Offline batchLLM class, vllm run-batchsgl.Engine
Minimum NVIDIA GPUCompute capability 7.5sm80 per the quickstart

What the two engines share#

Most of the feature list is the same on both. Each batches continuously, adding and removing requests between steps, and keeps the KV cache in pages so memory is not reserved for tokens that may never be generated. Each splits long prompts into chunks so they do not stall requests that are generating. Each supports tensor, pipeline, data and expert parallelism, LoRA adapters served next to the base model, FP8 and FP4 checkpoints as well as AWQ and GPTQ, and prefill-decode disaggregation for large deployments.

Speculative decoding is on both lists too. vLLM documents EAGLE, MTP, n-gram, draft-model, MLP speculator and suffix decoding. SGLang documents EAGLE-2 and EAGLE-3, MTP, n-gram, draft-model, UNO and DFLASH. The old headline difference, SGLang's zero-overhead scheduler that prepares the next batch on the CPU while the GPU runs the current one, has a counterpart in vLLM's async scheduling, which vLLM 0.30.0 turns on unless another option rules it out.

Prefix caching: hashed blocks vs a radix tree#

Both reuse the KV cache of a prefix they have already computed, with different bookkeeping. vLLM hashes each full block of 16 tokens together with everything before it, so two requests that share a system prompt share every full block of it, and a partial block at the end is recomputed. SGLang's RadixAttention keeps cached prefixes in a radix tree, a structure built for finding the longest match, and evicts the least recently used entries first by default.

The difference that matters in practice is scheduling, not storage. SGLang's --schedule-policy lpm (longest prefix match) reorders the waiting queue so that requests sharing a cached prefix run together, and its docs recommend it when the workload has many shared prefixes, at the cost of more scheduling overhead. vLLM serves requests first come, first served, or by a priority you assign. Agents that resend a long history, chat with multi-turn context and few-shot prompts are the workloads where that reordering can pay, so test them with shared prefixes, as in the benchmark section further down, before you pick an engine.

Scheduling and overload#

The two schedulers split each step differently. vLLM puts every generating request into the step first and spends the rest of the step's token budget on prompt chunks. The budget is --max-num-batched-tokens, 2,048 by default for the server on 48 GB cards and the A100 and 8,192 on GPUs with 70 GiB or more, and up to --max-num-seqs requests run together, 256 or 1,024 on the same split. When the KV cache runs out, vLLM preempts requests and recomputes them later.

SGLang runs prompt batches and generation batches as separate steps by default, a prompt batch first whenever one is ready, and mixing the two in one batch is an option (--enable-mixed-chunk). Its limits are --chunked-prefill-size and --max-running-requests, sized automatically unless you set them. It admits new requests according to --schedule-conservativeness (1.0 by default), and when the KV pool fills it retracts requests and logs a warning that starts KV cache pool is full. Its tuning guide suggests 0.3 when the pool is underused while requests wait, and 1.3 when retractions are frequent.

Structured output#

Both constrain output to a JSON schema or a regular expression, with different request fields for everything else. vLLM's backend setting is auto by default, which picks per request, and xgrammar, guidance, outlines or lm-format-enforcer can be pinned. It takes a structured_outputs object with choice, regex, json, grammar or structural_tag. SGLang uses XGrammar by default, with Outlines and Llguidance as options (--grammar-backend), and takes json_schema, regex or ebnf, one constraint per request. Here is the same classification request, first for vLLM and then for SGLang, with client an OpenAI client pointed at each engine's base URL as in the vLLM Docker guide:

Python
# vLLM
client.chat.completions.create(
    model="Qwen/Qwen3-8B",
    messages=[{"role": "user", "content": "Sentiment of: the GPU arrived broken."}],
    extra_body={"structured_outputs": {"choice": ["positive", "negative"]}},
)

# SGLang
client.chat.completions.create(
    model="Qwen/Qwen3-8B",
    messages=[{"role": "user", "content": "Sentiment of: the GPU arrived broken."}],
    extra_body={"regex": "(positive|negative)"},
)

The portable option is the OpenAI field that both accept, response_format={"type": "json_schema", ...}, so client code that sticks to it moves between the engines unchanged.

GPU memory: two different fractions#

The two memory settings are not the same number, so do not copy one into the other. vLLM's --gpu-memory-utilization is the share of the GPU for everything vLLM does, 0.92 by default: weights, activations, CUDA graphs and the KV cache, which gets whatever is left. SGLang's --mem-fraction-static covers only the weights and the KV pool, and activations and CUDA graph buffers live outside it. SGLang sets it with a heuristic, and its tuning guide says 5 to 8 GB left free for activations is usually enough.

Each engine prints its KV capacity at startup, and that line is the one to compare. vLLM logs GPU KV cache size: N tokens. SGLang logs max_total_num_tokens= together with available_gpu_mem=. The KV cache guide explains how many tokens a given model needs, and fixing vLLM out-of-memory errors covers what vLLM prints when they do not fit.

Deploying on a QuantaCloud VM#

Both run as a container on the Bare Metal template, a plain Ubuntu 22.04 VM with the NVIDIA driver and Docker, and neither has a QuantaCloud template. Check the driver first: vLLM's default image and SGLang 0.5.20 are CUDA 13 builds that need an R580 or newer driver. On an older driver, vLLM has v0.30.0-cu129, and SGLang's last CUDA 12 image is v0.5.19-cu129. Checking your driver and CUDA version covers the command.

Launch an RTX A6000 (Bare Metal)

vLLM's full setup, with the driver check, the API key and Docker Compose, is in deploying vLLM with Docker, and installing vLLM covers pip and uv. SGLang follows the same Docker pattern. Inside the container it has to listen on 0.0.0.0 so Docker can forward to it, and Docker publishes the port on the VM's 127.0.0.1 only:

Terminal
export SGLANG_API_KEY=$(openssl rand -hex 32)
echo "$SGLANG_API_KEY"

docker run -d --name sglang --gpus all --ipc=host --shm-size 32g \
  -p 127.0.0.1:30000:30000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  lmsysorg/sglang:v0.5.20 \
  python3 -m sglang.launch_server --model-path Qwen/Qwen3-8B \
  --host 0.0.0.0 --port 30000 --context-length 32768 \
  --reasoning-parser qwen3 --api-key "$SGLANG_API_KEY"

docker logs -f sglang

Wait for The server is fired up and ready to roll!, then call http://127.0.0.1:30000/v1/chat/completions with the key as a Bearer token, on the VM or through an SSH tunnel from your laptop. SGLang's key protects every route except health, readiness and metrics. vLLM's key protects only the /v1-style routes and leaves /invocations and /pause open, so vLLM needs 127.0.0.1 even more. Stopping a QuantaCloud instance terminates it and deletes its disk, so the next launch pulls the image and the weights again, about 6.3 GB more for SGLang's image than for vLLM's (our calculation).

For two or more GPUs in one VM, both take a tensor parallel size (--tensor-parallel-size in vLLM, --tp-size in SGLang). vLLM's docs prefer pipeline parallelism on GPUs without NVLink, such as the L40S, and SGLang's tuning guide favours data parallelism for throughput whenever a copy of the model fits on each GPU. vLLM across several GPUs covers the choice.

How to benchmark them on one GPU#

A comparison is only fair with the same model, the same context limit, the same prompts and the same GPU, one engine at a time. Serve Qwen/Qwen3-8B in BF16 with a 32,768-token limit on both, for example on an RTX A6000, and use vllm bench serve as the client for both engines, because it speaks the OpenAI completions API that both serve. The client runs from the vLLM image, so set VLLM_TAG with the driver check in the vLLM Docker guide first.

  1. Random prompts, the method in the vLLM Docker guide: 1,024 input and 256 output tokens with --ignore-eos, 32 prompts at concurrency 1 and 256 at concurrency 16, with the same seeds on both engines and a new seed for every run on the same server.
  2. Shared prefixes: 256 requests built from 8 prefixes of 4,096 tokens, each with its own 256-token suffix and 256 output tokens, at concurrency 16, on a freshly started server. Run SGLang twice, once with its default policy and once with --schedule-policy lpm.

The shared-prefix run against SGLang, from a client container on the host network:

Terminal
docker run --rm --network host --entrypoint vllm \
  -e OPENAI_API_KEY="$SGLANG_API_KEY" \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  "vllm/vllm-openai:$VLLM_TAG" bench serve \
  --backend openai --host 127.0.0.1 --port 30000 --model Qwen/Qwen3-8B \
  --dataset-name prefix_repetition --prefix-repetition-prefix-len 4096 \
  --prefix-repetition-suffix-len 256 --prefix-repetition-num-prefixes 8 \
  --prefix-repetition-output-len 256 --ignore-eos \
  --max-concurrency 16 --num-prompts 256 --seed 16

For vLLM, change the port to 8000 and the key to $VLLM_API_KEY. Stop one server before starting the other, since each claims most of the GPU.

vLLM vs SGLang FAQ#

Is SGLang faster than vLLM?

The docs give no general answer. Both projects describe the same core techniques, and both engines overlap CPU scheduling with GPU work by default. On paper, SGLang's lpm scheduling suits traffic with long shared prefixes, and that is the case to test first.

Can I switch between them without changing client code?

Mostly. Both serve /v1/chat/completions and /v1/completions, so an OpenAI client needs only a new base URL and key. Structured-output extras differ, so use response_format with a JSON schema if you want requests that work on both. For Qwen3, both engines take --reasoning-parser qwen3 to return the thinking separately from the answer.

Does SGLang run on every QuantaCloud GPU?

On paper, yes. SGLang's quickstart asks for sm80 or newer, and the oldest GPUs in the catalog, the A100 and RTX A6000, are compute capability 8.0 and 8.6. The catch is the driver, because SGLang 0.5.20 needs CUDA 13 and so an R580 or newer driver.

Can SGLang run offline batch jobs?

Yes. sgl.Engine(model_path=...) runs the engine inside your Python process with no HTTP server, the same idea as vLLM's LLM class, which batch inference with vLLM walks through.

Where does Ollama fit?

Ollama is built for a different job: chat for one person or a few, with 4-bit models that load on demand. vLLM vs Ollama covers that comparison.


My rule: run vLLM unless your traffic is dominated by long shared prefixes, and in that case run the shared-prefix benchmark above against both engines, with prefix lengths that match your traffic, before you switch. Either way, keep the server on 127.0.0.1, and budget memory with the startup KV line rather than the memory flag. Self-hosted LLM inference puts both next to the other serving options.

Launch a GPU for vLLM or SGLang

Keep building

Choose your next step.