The NVIDIA L40 price on QuantaCloud starts at $0.94/GPU-hr, for a 48 GB Ada data-center GPU in VMs with 1, 2 or 4 cards. The L40 suits work that memory bandwidth limits, such as one user generating text: it has the same 48 GB and 864 GB/s as the L40S, with half of the L40S's dense tensor throughput. It also has FP8 Tensor Cores, which the older RTX A6000 lacks.
Each configuration is a VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu, and its local disk is part of the hourly price. On 2026-09-27 the 1x came with 14 vCPUs, 72 GB of RAM and 625 GB of disk, and the 4x with 50 vCPUs, 288 GB of RAM and 2,500 GB. Offers were in us-midwest-1 and us-midwest-2 that day, with the 4x in us-midwest-2 only. The first hour is charged at launch and the unused seconds of the current hour are refunded when you stop. The pricing page has the rules, and the GPU catalog has every other card.
NVIDIA L40 specs#
The L40 is a passive, 300 W Ada data-center card, and NVIDIA still marks its datasheet figures as preliminary.
| Spec | NVIDIA L40 |
|---|---|
| Architecture | NVIDIA Ada Lovelace (AD102), compute capability 8.9 |
| GPU memory | 48 GB GDDR6 with ECC, enabled by default, 384-bit |
| Memory bandwidth | 864 GB/s |
| CUDA cores | 18,176 |
| Tensor Cores | 568, fourth generation, with FP8 |
| RT Cores | 142, third generation |
| FP32 | 90.5 TFLOPS |
| TF32 Tensor | 90.5 TFLOPS, 181 with sparsity |
| BF16 and FP16 Tensor | 181.05 TFLOPS, 362.1 with sparsity |
| FP8 Tensor | 362 TFLOPS, 724 with sparsity |
| INT8 and INT4 Tensor | 362 and 724 TOPS, doubled with sparsity |
| NVLink and MIG | Neither is supported |
| System interface | PCIe Gen4 x16, 64 GB/s bidirectional |
| Power and cooling | 300 W maximum, passive, dual slot |
| Video engines | 3 NVENC and 3 NVDEC, with AV1 encode and decode |
The L40 is neither the L40S nor the L4. The L40S keeps the L40's memory and doubles its dense tensor rate at 350 W, and the L4 is a 24 GB, 72 W inference card.
L40 vs L40S, in plain terms#
The extra S doubles the tensor math and leaves the memory alone. Both cards have 48 GB of GDDR6 at 864 GB/s and 18,176 CUDA cores. The L40S lists 733 dense FP8 TFLOPS and 362 dense FP16 TFLOPS against the L40's 362 and 181, and draws up to 350 W against 300 W.
What that means depends on the job. NVIDIA describes LLM token generation as memory-bound and prompt processing (prefill) as a matrix-matrix operation that saturates the GPU. So the two cards start from the same bandwidth when one user is generating text, and the L40S has twice the datasheet throughput for long prompts, batched serving and fine-tuning. I treat image and video generation as compute-bound too, so the same gap applies there.
| L40 | L40S | RTX 6000 Ada | |
|---|---|---|---|
| Memory | 48 GB GDDR6, ECC | 48 GB GDDR6, ECC | 48 GB GDDR6, ECC |
| Bandwidth | 864 GB/s | 864 GB/s | 960 GB/s |
| FP8 Tensor, dense | 362 TFLOPS | 733 TFLOPS | About 728 TFLOPS (our calculation: half of 1,457 with sparsity) |
| FP16 Tensor, dense | 181 TFLOPS | 362 TFLOPS | Not listed |
| FP32 | 90.5 TFLOPS | 91.6 TFLOPS | 91.1 TFLOPS |
| Max power | 300 W | 350 W | 300 W |
| Cooling | Passive | Passive | Active |
| GPU | Memory | From | Available now |
|---|---|---|---|
| L40 | 48 GB | $0.94/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
The L40 vs L40S comparison goes further, and the L40S page has its live configurations, which went up to 8 per VM on 2026-09-27. L40S vs RTX 6000 Ada compares the other two cards in the table.
What fits in 48 GB on an L40#
Budget 44 GB, not 48: vLLM claims 92% of GPU memory by default (our calculation: 48 x 0.92 = 44.2 GB). These are the jobs I would size to one L40. vLLM uses its FP8 Tensor Cores for per-channel FP8 checkpoints and for FP8 quantized at load, while block-scaled FP8 checkpoints, such as Qwen's official FP8 releases, run weight-only below Hopper.
| Workload | Memory it needs | On one L40 | Basis |
|---|---|---|---|
| Qwen3.8-27B, official FP8, full 262k context | 30.9 GB checkpoint + 17.2 GB of BF16 KV = 48.1 GB, or 39.5 GB with FP8 KV | Yes with FP8 KV. With BF16 KV, shorten the context. Qwen's FP8 files are block-scaled, so vLLM runs them weight-only on the L40 | Our calculation |
| Gemma 4 31B, quantized to FP8 at load, 32k context | 31.3 GB of weights + 1.3 to 2.7 GB of KV | Yes | Our calculation |
| Muse Glimmer 30B, 4-bit | Under 20 GB | Yes | Published (Meta) |
| Qwen3-30B-A3B, QLoRA fine-tuning | 17.5 GB | Yes | Published (Unsloth) |
| gpt-oss-20b, BF16 LoRA fine-tuning | 44 GB | Tight | Published (Unsloth) |
| LTX-2 audio and video, fp8 with the fp8 text encoder | 27.1 GB + 13.2 GB = 40.3 GB of files | Yes, tight once activations are added | Published file sizes, our sum |
| HiDream-I1 Full | More than 27 GB at full precision, more than 16 GB with fp8 files | Yes | Published (Comfy docs) |
| HunyuanVideo 1.5 | 14 GB minimum with model offloading | Yes | Published (Tencent) |
| SD3.5 Large, fp8 all-in-one file | 14.9 GB | Yes | Published file size |
Check the licences before you build a product on these. SD3.5 is free for commercial use only under $1M of annual revenue, LTX-2 needs a paid licence from $10M, HunyuanVideo's licence excludes the EU, the UK and South Korea, and HiDream's text encoder carries Meta's Llama 3.1 licence. In vLLM, NVFP4 checkpoints load on the L40 but run as weight-only 4-bit, because FP4 math is Blackwell-only. The VRAM guide and the KV cache explainer show how to size other models, and fine-tuning with Unsloth covers the QLoRA rows.
Where the L40 falls short#
The L40 falls short on tensor throughput, and on 2026-09-27 a faster card also cost less. The RTX 6000 Ada has the same core counts, 11% more bandwidth and twice the dense FP8 rate on NVIDIA's datasheets. That day, 100 hours on a 1x cost $79 on the RTX 6000 Ada against $94 on the L40 (our calculation at $0.79 and $0.94 per hour).
What the L40's configurations had that day was disk: 625 GB on the 1x against 350 GB on the RTX 6000 Ada, and the same ratio on the 2x and 4x. There are no volumes on QuantaCloud, so the local disk is all the room you get for checkpoints, datasets and outputs, and that can decide a video or multi-model ComfyUI setup. The L40 also has no NVLink, no MIG and no FP4.
| Step up to | Memory and bandwidth | Step up when |
|---|---|---|
| RTX 6000 Ada | 48 GB GDDR6, 960 GB/s | You want twice the datasheet FP8 rate at the same 48 GB, or 8 GPUs in one VM (offered on 2026-09-27) |
| L40S | 48 GB GDDR6, 864 GB/s | You want the data-center card with twice the tensor rate |
| RTX PRO 6000 Blackwell | 96 GB GDDR7, 1,597 or 1,792 GB/s by edition | The model or video workflow does not fit in 48 GB, or you want FP4 |
| A100 80GB | 80 GB HBM2e, up to 2,039 GB/s | The model needs up to 80 GB and FP8 does not matter |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | 96 GB | $2.39/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
Measured on QuantaCloud#
Across the catalog, most single-GPU VMs are running in about 3 minutes (median).
Launch the L40 with a template#
Pick the template by the first job you will run. The template does not change the price, and none of them comes with models preinstalled.
| Template | What you get | Launch |
|---|---|---|
| Bare Metal | Ubuntu 22.04, the NVIDIA driver and Docker over SSH. The start for vLLM in Docker and your own containers | Ubuntu + Docker on L40 |
| PyTorch + Jupyter | JupyterLab with PyTorch in the browser, behind your QuantaCloud login | Jupyter on L40 |
| Open WebUI + Ollama | A private chat UI. Pull models from inside the app. Open WebUI also has its own sign-in | Open WebUI on L40 |
| ComfyUI | Image and video generation in the browser, with the L40's larger disk for checkpoints. Upload or download your models after launch | ComfyUI on L40 |
The templates docs list what each template contains, and connecting over SSH covers the Bare Metal login.
FAQ#
How much does an NVIDIA L40 cost per hour?
From $0.94/GPU-hr, and the table at the top has each configuration's hourly price. You pay for the time you use: a job stopped after 2 hours and 45 minutes is charged three hours as it runs, then the unused 15 minutes are refunded. At the 1x price on 2026-09-27, $0.94 per hour, that job costs about $2.59 (our calculation: 2.75 hours x $0.94).
Is the L40 the same as the L40S?
No. They share 48 GB and 864 GB/s, but the L40S has twice the dense tensor throughput and draws up to 350 W. The comparison above has the figures.
Can I get 2 or 4 L40s in one VM?
Yes. On 2026-09-27 the catalog offered 1, 2 or 4 per VM, with the 4x in us-midwest-2 only. The L40 has no NVLink, so to split one model across the cards, vLLM's docs recommend pipeline parallelism over tensor parallelism.
Is ECC on?
NVIDIA ships the L40 with ECC enabled by default, and software can turn it off.
Is it a VM or bare metal?
A VM with Ubuntu 22.04, the NVIDIA driver and Docker. "Bare Metal" is the name of the plain Ubuntu template, not of the hardware, and every instance is provided by QuantaCloud in our US regions.
What happens when I stop?
The VM is terminated and its disk deleted, and there are no volumes or snapshots, so copy your outputs off first. The unused seconds of the current hour are refunded.
My rule for the L40: rent it when the job is bound by memory bandwidth and the L40's live price is the lowest of the Ada cards, or when you need the bigger disk its configurations carry. For compute-heavy work at a similar price, take the RTX 6000 Ada or the L40S instead. Launch an L40 , and how to rent a GPU covers the first login.