On-demand GPU

NVIDIA L40

48 GB. Room to build.

Rent NVIDIA L40 GPUs (48 GB, FP8) by the hour: 1, 2 or 4 per VM, live prices, full specs, and how the L40 differs from the L40S in plain terms.

Full SSH accessYour choice of template
Starting from$0.94/ GPU-hr
Launch L40
Available nowSee configurations
NVIDIA L40ADA LOVELACE
VRAMVRAMVRAMVRAMVRAMVRAMADA LOVELACE
Room for the work ahead.48 GB GDDR6
GPU memory
48 GB GDDR6
Memory bandwidth
864 GB/s
Architecture
Ada Lovelace
Deployment
On demand

Your GPU. Your configuration.

Choose where you start.

3 configurations
1 GPULowest price

1x L40

Midwest · us-midwest-2

vCPU
14
RAM
72 GB
Disk
625 GB
$0.94/ instance-hour
Available nowLaunch configuration
2 GPUs

2x L40

Midwest · us-midwest-2

vCPU
26
RAM
144 GB
Disk
1,250 GB
$1.88/ instance-hour
Available nowLaunch configuration
4 GPUs

4x L40

Midwest · us-midwest-2

vCPU
50
RAM
288 GB
Disk
2,500 GB
$3.75/ instance-hour
Available nowLaunch configuration

Prices checked

The NVIDIA L40 price on QuantaCloud starts at $0.94/GPU-hr, for a 48 GB Ada data-center GPU in VMs with 1, 2 or 4 cards. The L40 suits work that memory bandwidth limits, such as one user generating text: it has the same 48 GB and 864 GB/s as the L40S, with half of the L40S's dense tensor throughput. It also has FP8 Tensor Cores, which the older RTX A6000 lacks.

Each configuration is a VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu, and its local disk is part of the hourly price. On 2026-09-27 the 1x came with 14 vCPUs, 72 GB of RAM and 625 GB of disk, and the 4x with 50 vCPUs, 288 GB of RAM and 2,500 GB. Offers were in us-midwest-1 and us-midwest-2 that day, with the 4x in us-midwest-2 only. The first hour is charged at launch and the unused seconds of the current hour are refunded when you stop. The pricing page has the rules, and the GPU catalog has every other card.

NVIDIA L40 specs#

The L40 is a passive, 300 W Ada data-center card, and NVIDIA still marks its datasheet figures as preliminary.

SpecNVIDIA L40
ArchitectureNVIDIA Ada Lovelace (AD102), compute capability 8.9
GPU memory48 GB GDDR6 with ECC, enabled by default, 384-bit
Memory bandwidth864 GB/s
CUDA cores18,176
Tensor Cores568, fourth generation, with FP8
RT Cores142, third generation
FP3290.5 TFLOPS
TF32 Tensor90.5 TFLOPS, 181 with sparsity
BF16 and FP16 Tensor181.05 TFLOPS, 362.1 with sparsity
FP8 Tensor362 TFLOPS, 724 with sparsity
INT8 and INT4 Tensor362 and 724 TOPS, doubled with sparsity
NVLink and MIGNeither is supported
System interfacePCIe Gen4 x16, 64 GB/s bidirectional
Power and cooling300 W maximum, passive, dual slot
Video engines3 NVENC and 3 NVDEC, with AV1 encode and decode

The L40 is neither the L40S nor the L4. The L40S keeps the L40's memory and doubles its dense tensor rate at 350 W, and the L4 is a 24 GB, 72 W inference card.

L40 vs L40S, in plain terms#

The extra S doubles the tensor math and leaves the memory alone. Both cards have 48 GB of GDDR6 at 864 GB/s and 18,176 CUDA cores. The L40S lists 733 dense FP8 TFLOPS and 362 dense FP16 TFLOPS against the L40's 362 and 181, and draws up to 350 W against 300 W.

What that means depends on the job. NVIDIA describes LLM token generation as memory-bound and prompt processing (prefill) as a matrix-matrix operation that saturates the GPU. So the two cards start from the same bandwidth when one user is generating text, and the L40S has twice the datasheet throughput for long prompts, batched serving and fine-tuning. I treat image and video generation as compute-bound too, so the same gap applies there.

L40L40SRTX 6000 Ada
Memory48 GB GDDR6, ECC48 GB GDDR6, ECC48 GB GDDR6, ECC
Bandwidth864 GB/s864 GB/s960 GB/s
FP8 Tensor, dense362 TFLOPS733 TFLOPSAbout 728 TFLOPS (our calculation: half of 1,457 with sparsity)
FP16 Tensor, dense181 TFLOPS362 TFLOPSNot listed
FP3290.5 TFLOPS91.6 TFLOPS91.1 TFLOPS
Max power300 W350 W300 W
CoolingPassivePassiveActive
GPUMemoryFromAvailable now
L4048 GB$0.94/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
RTX A600048 GB$0.48/GPU-hrYes

The L40 vs L40S comparison goes further, and the L40S page has its live configurations, which went up to 8 per VM on 2026-09-27. L40S vs RTX 6000 Ada compares the other two cards in the table.

What fits in 48 GB on an L40#

Budget 44 GB, not 48: vLLM claims 92% of GPU memory by default (our calculation: 48 x 0.92 = 44.2 GB). These are the jobs I would size to one L40. vLLM uses its FP8 Tensor Cores for per-channel FP8 checkpoints and for FP8 quantized at load, while block-scaled FP8 checkpoints, such as Qwen's official FP8 releases, run weight-only below Hopper.

WorkloadMemory it needsOn one L40Basis
Qwen3.8-27B, official FP8, full 262k context30.9 GB checkpoint + 17.2 GB of BF16 KV = 48.1 GB, or 39.5 GB with FP8 KVYes with FP8 KV. With BF16 KV, shorten the context. Qwen's FP8 files are block-scaled, so vLLM runs them weight-only on the L40Our calculation
Gemma 4 31B, quantized to FP8 at load, 32k context31.3 GB of weights + 1.3 to 2.7 GB of KVYesOur calculation
Muse Glimmer 30B, 4-bitUnder 20 GBYesPublished (Meta)
Qwen3-30B-A3B, QLoRA fine-tuning17.5 GBYesPublished (Unsloth)
gpt-oss-20b, BF16 LoRA fine-tuning44 GBTightPublished (Unsloth)
LTX-2 audio and video, fp8 with the fp8 text encoder27.1 GB + 13.2 GB = 40.3 GB of filesYes, tight once activations are addedPublished file sizes, our sum
HiDream-I1 FullMore than 27 GB at full precision, more than 16 GB with fp8 filesYesPublished (Comfy docs)
HunyuanVideo 1.514 GB minimum with model offloadingYesPublished (Tencent)
SD3.5 Large, fp8 all-in-one file14.9 GBYesPublished file size

Check the licences before you build a product on these. SD3.5 is free for commercial use only under $1M of annual revenue, LTX-2 needs a paid licence from $10M, HunyuanVideo's licence excludes the EU, the UK and South Korea, and HiDream's text encoder carries Meta's Llama 3.1 licence. In vLLM, NVFP4 checkpoints load on the L40 but run as weight-only 4-bit, because FP4 math is Blackwell-only. The VRAM guide and the KV cache explainer show how to size other models, and fine-tuning with Unsloth covers the QLoRA rows.

Where the L40 falls short#

The L40 falls short on tensor throughput, and on 2026-09-27 a faster card also cost less. The RTX 6000 Ada has the same core counts, 11% more bandwidth and twice the dense FP8 rate on NVIDIA's datasheets. That day, 100 hours on a 1x cost $79 on the RTX 6000 Ada against $94 on the L40 (our calculation at $0.79 and $0.94 per hour).

What the L40's configurations had that day was disk: 625 GB on the 1x against 350 GB on the RTX 6000 Ada, and the same ratio on the 2x and 4x. There are no volumes on QuantaCloud, so the local disk is all the room you get for checkpoints, datasets and outputs, and that can decide a video or multi-model ComfyUI setup. The L40 also has no NVLink, no MIG and no FP4.

Step up toMemory and bandwidthStep up when
RTX 6000 Ada48 GB GDDR6, 960 GB/sYou want twice the datasheet FP8 rate at the same 48 GB, or 8 GPUs in one VM (offered on 2026-09-27)
L40S48 GB GDDR6, 864 GB/sYou want the data-center card with twice the tensor rate
RTX PRO 6000 Blackwell96 GB GDDR7, 1,597 or 1,792 GB/s by editionThe model or video workflow does not fit in 48 GB, or you want FP4
A100 80GB80 GB HBM2e, up to 2,039 GB/sThe model needs up to 80 GB and FP8 does not matter
GPUMemoryFromAvailable now
RTX 6000 Ada48 GB$0.78/GPU-hrYes
L40S48 GB$1.09/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes

Measured on QuantaCloud#

Across the catalog, most single-GPU VMs are running in about 3 minutes (median).

Launch the L40 with a template#

Pick the template by the first job you will run. The template does not change the price, and none of them comes with models preinstalled.

TemplateWhat you getLaunch
Bare MetalUbuntu 22.04, the NVIDIA driver and Docker over SSH. The start for vLLM in Docker and your own containersUbuntu + Docker on L40
PyTorch + JupyterJupyterLab with PyTorch in the browser, behind your QuantaCloud loginJupyter on L40
Open WebUI + OllamaA private chat UI. Pull models from inside the app. Open WebUI also has its own sign-inOpen WebUI on L40
ComfyUIImage and video generation in the browser, with the L40's larger disk for checkpoints. Upload or download your models after launchComfyUI on L40

The templates docs list what each template contains, and connecting over SSH covers the Bare Metal login.

FAQ#

How much does an NVIDIA L40 cost per hour?

From $0.94/GPU-hr, and the table at the top has each configuration's hourly price. You pay for the time you use: a job stopped after 2 hours and 45 minutes is charged three hours as it runs, then the unused 15 minutes are refunded. At the 1x price on 2026-09-27, $0.94 per hour, that job costs about $2.59 (our calculation: 2.75 hours x $0.94).

Is the L40 the same as the L40S?

No. They share 48 GB and 864 GB/s, but the L40S has twice the dense tensor throughput and draws up to 350 W. The comparison above has the figures.

Can I get 2 or 4 L40s in one VM?

Yes. On 2026-09-27 the catalog offered 1, 2 or 4 per VM, with the 4x in us-midwest-2 only. The L40 has no NVLink, so to split one model across the cards, vLLM's docs recommend pipeline parallelism over tensor parallelism.

Is ECC on?

NVIDIA ships the L40 with ECC enabled by default, and software can turn it off.

Is it a VM or bare metal?

A VM with Ubuntu 22.04, the NVIDIA driver and Docker. "Bare Metal" is the name of the plain Ubuntu template, not of the hardware, and every instance is provided by QuantaCloud in our US regions.

What happens when I stop?

The VM is terminated and its disk deleted, and there are no volumes or snapshots, so copy your outputs off first. The unused seconds of the current hour are refunded.


My rule for the L40: rent it when the job is bound by memory bandwidth and the L40's live price is the lowest of the Ada cards, or when you need the bigger disk its configurations carry. For compute-heavy work at a similar price, take the RTX 6000 Ada or the L40S instead. Launch an L40 , and how to rent a GPU covers the first login.

Keep building

Choose your next step.