The L40S price on QuantaCloud starts at $1.09/GPU-hr, for NVIDIA's Ada Lovelace data-center GPU with 48 GB of GDDR6 and FP8 Tensor Cores, and you can run one, two, four or eight in a single VM. It suits image and video generation, FP8 inference of models up to about 32B, and QLoRA fine-tuning up to 32B, all of which fit in 48 GB.
On 2026-09-27 the catalog listed 1x, 2x, 4x and 8x L40S VMs, all in us-midwest-1 (Midwest). The table above is the live list.
L40S specs#
The L40S carries 48 GB of GDDR6 with ECC, not GDDR6X, at 864 GB/s.
| Spec | NVIDIA L40S |
|---|---|
| Architecture | Ada Lovelace, compute capability 8.9 |
| GPU memory | 48 GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s |
| CUDA cores | 18,176 |
| Tensor Cores | 568, fourth generation |
| RT Cores | 142, third generation |
| FP32 | 91.6 TFLOPS |
| TF32 Tensor Core, dense / sparse | 183 / 366 TFLOPS |
| BF16 and FP16 Tensor Core, dense / sparse | 362 / 733 TFLOPS |
| FP8 Tensor Core, dense / sparse | 733 / 1,466 TFLOPS |
| INT8 Tensor Core, dense / sparse | 733 / 1,466 TOPS |
| FP4 | No |
| Host link | PCIe Gen4 x16, 64 GB/s bidirectional |
| NVLink | No |
| Multi-Instance GPU | No |
| Video engines | 3 NVENC and 3 NVDEC, with AV1 encode and decode |
| Max power | 350 W |
| Form factor | Dual-slot, 4.4 x 10.5 in, passive cooling |
Two rows decide most workloads. FP8 at 733 dense TFLOPS is the reason to pick the L40S over Ampere cards like the A100 and RTX A6000, which have no FP8 at all. In vLLM that speed applies to per-channel FP8 checkpoints, the format llm-compressor produces. Block-scaled ones, such as Qwen's own Qwen3-32B-FP8, run weight-only below Hopper. NVLink is absent, so multi-GPU jobs talk over PCIe Gen4. The video engines are a quieter advantage: NVIDIA's own H100 whitepaper notes that the H100 and A100 have no NVENC encoder, so a video pipeline on an L40S can generate and encode on the same GPU.
What fits in 48 GB#
The working budget for an LLM on one L40S is 41.4 GiB, because vLLM claims 92 percent of GPU memory by default and the card reports 46,068 MiB with ECC on, its default (our calculation: 46,068 / 1,024 x 0.92). ComfyUI does not reserve memory that way, and it can offload weights to system RAM when a model does not fit, at a cost in speed.
| Workload | Memory it needs | On one L40S | Basis |
|---|---|---|---|
| SDXL 1.0 | Stability AI: "consumer GPUs with 8GB VRAM" | Yes | Published |
| FLUX.1 [dev] or [schnell] | 12.33 GB for BFL's FP8 file, about 23 GB for the full-precision file, plus text encoders | Yes, at either precision | Published |
| Qwen-Image at fp8 | 20.4 GB model file plus a 9.38 GB fp8 text encoder | Yes | Published |
| Wan 2.2 TI2V-5B, 720p | 22.9 GB peak in Wan's own test | Yes | Published |
| Wan 2.2 A14B | 41.3 GB peak at 480p and 59.8 GB at 720p in Wan's single-GPU test | 480p yes. 720p only with ComfyUI's offloading, which is slower | Published |
| Qwen3-8B at BF16, 32k context | 16.4 GB (15.3 GiB) of weights, plus 4.5 GiB of KV cache per 32k sequence | Yes, nearly six full 32k sequences: (41.4 - 15.3) / 4.5 = 5.8 | Our calculation |
| Qwen3-32B at FP8, per-channel | 34.3 GB checkpoint | Yes, with about 38,000 tokens of BF16 KV cache, or 76,000 with FP8 KV: (41.4 - 32.0) GiB / 256 KiB. That is before activations, so one full 32k request, which needs 8 GiB of BF16 KV, is borderline | Published size, our calculation |
| Llama-3.3-70B at 4-bit (AWQ) | 39.8 GB checkpoint | Tight: about 14,000 tokens of BF16 KV cache, (41.4 - 37.0) GiB / 320 KiB | Published size, our calculation |
| gpt-oss-20b | "within 16GB of memory" | Yes | Published |
| gpt-oss-120b | 65.2 GB of weights | Not on 1x. Yes on 2x with pipeline parallelism | Our calculation |
| QLoRA fine-tune, 32B model | 26 GB (Unsloth), 24 to 32 GB (Axolotl) | Yes | Published |
| QLoRA fine-tune, 70B model | 41 GB (Unsloth), 40 to 48 GB (Axolotl) | Tight: Unsloth reached 12,106 tokens of context on 48 GB | Published |
Check the licence before you build a product on an image model. FLUX.1 [dev] lets you use its outputs commercially, but using the model itself for revenue-generating work needs a licence from Black Forest Labs, while FLUX.1 [schnell] and Qwen-Image are Apache 2.0. The FLUX in ComfyUI guide covers the files and settings, ComfyUI GPU requirements lists VRAM by image and video model, and how much VRAM you need covers other models.
L40S vs L40, RTX 6000 Ada and A100#
The L40S sits between the other 48 GB cards and the 80 to 96 GB ones, and each neighbour wins on one line of the spec sheet.
| GPU | Memory | Bandwidth | FP8 Tensor | Pick it over the L40S when |
|---|---|---|---|---|
| L40 | 48 GB GDDR6 | 864 GB/s | 362 dense, 724 sparse TFLOPS | The job is bound by bandwidth, like single-user LLM decoding, and the L40 costs less |
| RTX 6000 Ada | 48 GB GDDR6 | 960 GB/s | 1,457 sparse TOPS | It costs less when you check: same generation, same memory, 11 percent more bandwidth |
| RTX A6000 | 48 GB GDDR6 | 768 GB/s | None | FP8 does not matter to the job and the hourly price does |
| A100 80GB | 80 GB HBM2e | 1,935 GB/s (PCIe) or 2,039 GB/s (SXM4) | None | The model needs more than 48 GB, or decode speed matters more than FP8 |
| RTX PRO 6000 Blackwell | 96 GB GDDR7 | 1,597 or 1,792 GB/s, by edition | Yes, plus FP4 | You need up to 96 GB on one GPU, or FP4 |
| GPU | Memory | From | Available now |
|---|---|---|---|
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
The honest comparison is with the RTX 6000 Ada. On paper it matches the L40S on FP8 within 1 percent (1,457 sparse TOPS against 1,466) and has 11 percent more memory bandwidth (our calculation: 960 / 864 = 1.11). If it is cheaper when you check, I would take it for a single GPU. The L40 has the same memory and bandwidth as the L40S with half the FP8 throughput (our calculation: 733 / 362 = 2.0), so my rule is simple: bandwidth-bound work goes to whichever of the two is cheaper, and compute-heavy work, meaning diffusion, prefill and large batches, goes to the L40S. The L40 vs L40S and L40S vs RTX 6000 Ada comparisons cover the rest of the differences.
Against the A100 the question is memory. The L40S has FP8 and the A100 does not, but the A100 has 80 GB and more than twice the bandwidth (our calculation: 1,935 / 864 = 2.2). The L40S vs A100 comparison puts both side by side, and L40S vs H100 does the same for the H100.
Two to eight L40S in one VM#
More L40S cards add memory, not a faster link. There is no NVLink, so the GPUs exchange data over PCIe Gen4, and vLLM's documentation names the L40S when it says GPUs without NVLink should use pipeline parallelism instead of tensor parallelism. In practice that means --pipeline-parallel-size 2 to fit gpt-oss-120b across two cards (our calculation: 2 x 41.4 = 82.8 GiB of budget against 60.8 GiB of weights), and one model copy per GPU whenever the model fits on one card. The 8x VM gives you 384 GB in total. It also takes longer to come up: across QuantaCloud, the median 8-GPU VM is running in about 10 minutes, against about 3 for a single GPU.
Launch it with a template#
The template decides what is running when the VM comes up, and it does not change the price. The deploy docs walk through the console steps.
| Template | What I would run on an L40S | Launch |
|---|---|---|
| ComfyUI | FLUX.1, Qwen-Image at fp8 or Wan 2.2 5B video, following the ComfyUI guide | Launch ComfyUI |
| Open WebUI + Ollama | A private chat on Qwen3-32B or gpt-oss-20b | Launch Open WebUI |
| PyTorch + Jupyter | QLoRA fine-tuning up to 32B, as in the Unsloth guide | Launch PyTorch + Jupyter |
| Bare Metal: Ubuntu 22.04, NVIDIA driver, Docker | vLLM in Docker serving per-channel FP8 checkpoints, which use Ada's FP8 Tensor Cores | Launch Bare Metal |
SSH in as ubuntu on any template (connect over SSH). The app templates open at a private URL behind your QuantaCloud login. Download model weights and checkpoints after boot, and copy your outputs off before you stop. The templates docs list what each one contains, and the ComfyUI page has a GPU picker.
FAQ#
Is the L40S memory GDDR6 or GDDR6X?
GDDR6 with ECC: 48 GB at 864 GB/s, per NVIDIA's spec sheet. GDDR6X is the memory on GeForce cards such as the RTX 4090.
Should I rent the L40S or the L40?
Rent the L40S when the job leans on tensor math: diffusion models, prefill-heavy serving, big batches. It has twice the L40's FP8 throughput. For single-user LLM decoding, which is limited by the same 864 GB/s on both cards, I would take whichever is cheaper.
Should I rent the L40S or the A100?
The L40S when the model fits in 48 GB and you want FP8. The A100 when you need 80 GB on one GPU, or more than twice the memory bandwidth for decoding.
Does the L40S support NVLink or MIG?
Neither. It connects over PCIe Gen4 x16, and NVIDIA lists MIG as unsupported. For multi-GPU inference, use pipeline parallelism or one model per GPU.
Can I get eight L40S in one VM?
Yes. The 1x, 2x, 4x and 8x configurations were all listed on 2026-09-27, and the 8x VM has 384 GB of GPU memory.
Is it a VM or a bare-metal server?
A VM. Every QuantaCloud instance runs Ubuntu 22.04 with the NVIDIA driver and Docker, and you connect over SSH as ubuntu. Bare Metal is the name of the plain Ubuntu template, not the hardware. If you need physical L40S servers to yourself, dedicated GPU servers are built to order: send a capacity brief.
Is it available right now?
The table at the top is live. On-demand capacity is not held for you, and there is no SLA. QuantaCloud runs in US regions only, and on 2026-09-27 every L40S configuration was in us-midwest-1.
What happens when I stop the instance?
Stopping terminates the VM and deletes its disk. There are no volumes or snapshots, so download generated images and trained adapters first. The unused seconds of the current hour are refunded to your balance. Pricing and the billing docs have the full rules.
The rule I follow for the L40S: rent it when the job fits in 48 GB and uses FP8 or heavy tensor math, like FLUX, Wan 2.2 5B or a 32B model at FP8. For plain decoding, check the L40 and RTX 6000 Ada prices in the GPU catalog first. When the job needs more than 48 GB on one GPU, go to the A100 or the RTX PRO 6000. Start with the ComfyUI template on a 1x VM, time one image at your own settings, and scale out from there. Launch ComfyUI on L40S