On-demand GPU

NVIDIA L40S

48 GB. Room to build.

L40S price per hour, live: 48 GB GDDR6 with FP8, 1, 2, 4 or 8 per VM. Specs, what fits in 48 GB, and when the L40, RTX 6000 Ada or A100 is better.

Full SSH accessYour choice of template
Starting from$1.09/ GPU-hr
Launch L40S
Available nowSee configurations
NVIDIA L40SADA LOVELACE
VRAMVRAMVRAMVRAMVRAMVRAMADA LOVELACE
Room for the work ahead.48 GB GDDR6
GPU memory
48 GB GDDR6
Memory bandwidth
864 GB/s
Architecture
Ada Lovelace
Deployment
On demand

Your GPU. Your configuration.

Choose where you start.

6 configurations
1 GPULowest price

1x L40S

Midwest · us-midwest-1

vCPU
12
RAM
72 GB
Disk
625 GB
$1.09/ instance-hour
Available nowLaunch configuration
1 GPU

1x L40S

Midwest · us-midwest-2

vCPU
12
RAM
72 GB
Disk
625 GB
$1.09/ instance-hour
Available nowLaunch configuration
2 GPUs

2x L40S

Midwest · us-midwest-2

vCPU
24
RAM
144 GB
Disk
1,250 GB
$2.18/ instance-hour
Available nowLaunch configuration
2 GPUs

2x L40S

Midwest · us-midwest-1

vCPU
24
RAM
144 GB
Disk
1,250 GB
$2.18/ instance-hour
Available nowLaunch configuration
4 GPUs

4x L40S

Midwest · us-midwest-1

vCPU
46
RAM
288 GB
Disk
2,500 GB
$4.36/ instance-hour
Available nowLaunch configuration
8 GPUs

8x L40S

Midwest · us-midwest-1

vCPU
94
RAM
576 GB
Disk
5,000 GB
$8.72/ instance-hour
Available nowLaunch configuration

Prices checked

The L40S price on QuantaCloud starts at $1.09/GPU-hr, for NVIDIA's Ada Lovelace data-center GPU with 48 GB of GDDR6 and FP8 Tensor Cores, and you can run one, two, four or eight in a single VM. It suits image and video generation, FP8 inference of models up to about 32B, and QLoRA fine-tuning up to 32B, all of which fit in 48 GB.

On 2026-09-27 the catalog listed 1x, 2x, 4x and 8x L40S VMs, all in us-midwest-1 (Midwest). The table above is the live list.

L40S specs#

The L40S carries 48 GB of GDDR6 with ECC, not GDDR6X, at 864 GB/s.

SpecNVIDIA L40S
ArchitectureAda Lovelace, compute capability 8.9
GPU memory48 GB GDDR6 with ECC
Memory bandwidth864 GB/s
CUDA cores18,176
Tensor Cores568, fourth generation
RT Cores142, third generation
FP3291.6 TFLOPS
TF32 Tensor Core, dense / sparse183 / 366 TFLOPS
BF16 and FP16 Tensor Core, dense / sparse362 / 733 TFLOPS
FP8 Tensor Core, dense / sparse733 / 1,466 TFLOPS
INT8 Tensor Core, dense / sparse733 / 1,466 TOPS
FP4No
Host linkPCIe Gen4 x16, 64 GB/s bidirectional
NVLinkNo
Multi-Instance GPUNo
Video engines3 NVENC and 3 NVDEC, with AV1 encode and decode
Max power350 W
Form factorDual-slot, 4.4 x 10.5 in, passive cooling

Two rows decide most workloads. FP8 at 733 dense TFLOPS is the reason to pick the L40S over Ampere cards like the A100 and RTX A6000, which have no FP8 at all. In vLLM that speed applies to per-channel FP8 checkpoints, the format llm-compressor produces. Block-scaled ones, such as Qwen's own Qwen3-32B-FP8, run weight-only below Hopper. NVLink is absent, so multi-GPU jobs talk over PCIe Gen4. The video engines are a quieter advantage: NVIDIA's own H100 whitepaper notes that the H100 and A100 have no NVENC encoder, so a video pipeline on an L40S can generate and encode on the same GPU.

What fits in 48 GB#

The working budget for an LLM on one L40S is 41.4 GiB, because vLLM claims 92 percent of GPU memory by default and the card reports 46,068 MiB with ECC on, its default (our calculation: 46,068 / 1,024 x 0.92). ComfyUI does not reserve memory that way, and it can offload weights to system RAM when a model does not fit, at a cost in speed.

WorkloadMemory it needsOn one L40SBasis
SDXL 1.0Stability AI: "consumer GPUs with 8GB VRAM"YesPublished
FLUX.1 [dev] or [schnell]12.33 GB for BFL's FP8 file, about 23 GB for the full-precision file, plus text encodersYes, at either precisionPublished
Qwen-Image at fp820.4 GB model file plus a 9.38 GB fp8 text encoderYesPublished
Wan 2.2 TI2V-5B, 720p22.9 GB peak in Wan's own testYesPublished
Wan 2.2 A14B41.3 GB peak at 480p and 59.8 GB at 720p in Wan's single-GPU test480p yes. 720p only with ComfyUI's offloading, which is slowerPublished
Qwen3-8B at BF16, 32k context16.4 GB (15.3 GiB) of weights, plus 4.5 GiB of KV cache per 32k sequenceYes, nearly six full 32k sequences: (41.4 - 15.3) / 4.5 = 5.8Our calculation
Qwen3-32B at FP8, per-channel34.3 GB checkpointYes, with about 38,000 tokens of BF16 KV cache, or 76,000 with FP8 KV: (41.4 - 32.0) GiB / 256 KiB. That is before activations, so one full 32k request, which needs 8 GiB of BF16 KV, is borderlinePublished size, our calculation
Llama-3.3-70B at 4-bit (AWQ)39.8 GB checkpointTight: about 14,000 tokens of BF16 KV cache, (41.4 - 37.0) GiB / 320 KiBPublished size, our calculation
gpt-oss-20b"within 16GB of memory"YesPublished
gpt-oss-120b65.2 GB of weightsNot on 1x. Yes on 2x with pipeline parallelismOur calculation
QLoRA fine-tune, 32B model26 GB (Unsloth), 24 to 32 GB (Axolotl)YesPublished
QLoRA fine-tune, 70B model41 GB (Unsloth), 40 to 48 GB (Axolotl)Tight: Unsloth reached 12,106 tokens of context on 48 GBPublished

Check the licence before you build a product on an image model. FLUX.1 [dev] lets you use its outputs commercially, but using the model itself for revenue-generating work needs a licence from Black Forest Labs, while FLUX.1 [schnell] and Qwen-Image are Apache 2.0. The FLUX in ComfyUI guide covers the files and settings, ComfyUI GPU requirements lists VRAM by image and video model, and how much VRAM you need covers other models.

L40S vs L40, RTX 6000 Ada and A100#

The L40S sits between the other 48 GB cards and the 80 to 96 GB ones, and each neighbour wins on one line of the spec sheet.

GPUMemoryBandwidthFP8 TensorPick it over the L40S when
L4048 GB GDDR6864 GB/s362 dense, 724 sparse TFLOPSThe job is bound by bandwidth, like single-user LLM decoding, and the L40 costs less
RTX 6000 Ada48 GB GDDR6960 GB/s1,457 sparse TOPSIt costs less when you check: same generation, same memory, 11 percent more bandwidth
RTX A600048 GB GDDR6768 GB/sNoneFP8 does not matter to the job and the hourly price does
A100 80GB80 GB HBM2e1,935 GB/s (PCIe) or 2,039 GB/s (SXM4)NoneThe model needs more than 48 GB, or decode speed matters more than FP8
RTX PRO 6000 Blackwell96 GB GDDR71,597 or 1,792 GB/s, by editionYes, plus FP4You need up to 96 GB on one GPU, or FP4
GPUMemoryFromAvailable now
L40S48 GB$1.09/GPU-hrYes
L4048 GB$0.94/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes
RTX A600048 GB$0.48/GPU-hrYes
A100 SXM4 80GB80 GB$1.49/GPU-hrYes
RTX PRO 6000 Blackwell96 GB$2.39/GPU-hrYes

The honest comparison is with the RTX 6000 Ada. On paper it matches the L40S on FP8 within 1 percent (1,457 sparse TOPS against 1,466) and has 11 percent more memory bandwidth (our calculation: 960 / 864 = 1.11). If it is cheaper when you check, I would take it for a single GPU. The L40 has the same memory and bandwidth as the L40S with half the FP8 throughput (our calculation: 733 / 362 = 2.0), so my rule is simple: bandwidth-bound work goes to whichever of the two is cheaper, and compute-heavy work, meaning diffusion, prefill and large batches, goes to the L40S. The L40 vs L40S and L40S vs RTX 6000 Ada comparisons cover the rest of the differences.

Against the A100 the question is memory. The L40S has FP8 and the A100 does not, but the A100 has 80 GB and more than twice the bandwidth (our calculation: 1,935 / 864 = 2.2). The L40S vs A100 comparison puts both side by side, and L40S vs H100 does the same for the H100.

Two to eight L40S in one VM#

More L40S cards add memory, not a faster link. There is no NVLink, so the GPUs exchange data over PCIe Gen4, and vLLM's documentation names the L40S when it says GPUs without NVLink should use pipeline parallelism instead of tensor parallelism. In practice that means --pipeline-parallel-size 2 to fit gpt-oss-120b across two cards (our calculation: 2 x 41.4 = 82.8 GiB of budget against 60.8 GiB of weights), and one model copy per GPU whenever the model fits on one card. The 8x VM gives you 384 GB in total. It also takes longer to come up: across QuantaCloud, the median 8-GPU VM is running in about 10 minutes, against about 3 for a single GPU.

Launch it with a template#

The template decides what is running when the VM comes up, and it does not change the price. The deploy docs walk through the console steps.

TemplateWhat I would run on an L40SLaunch
ComfyUIFLUX.1, Qwen-Image at fp8 or Wan 2.2 5B video, following the ComfyUI guideLaunch ComfyUI
Open WebUI + OllamaA private chat on Qwen3-32B or gpt-oss-20bLaunch Open WebUI
PyTorch + JupyterQLoRA fine-tuning up to 32B, as in the Unsloth guideLaunch PyTorch + Jupyter
Bare Metal: Ubuntu 22.04, NVIDIA driver, DockervLLM in Docker serving per-channel FP8 checkpoints, which use Ada's FP8 Tensor CoresLaunch Bare Metal

SSH in as ubuntu on any template (connect over SSH). The app templates open at a private URL behind your QuantaCloud login. Download model weights and checkpoints after boot, and copy your outputs off before you stop. The templates docs list what each one contains, and the ComfyUI page has a GPU picker.

FAQ#

Is the L40S memory GDDR6 or GDDR6X?

GDDR6 with ECC: 48 GB at 864 GB/s, per NVIDIA's spec sheet. GDDR6X is the memory on GeForce cards such as the RTX 4090.

Should I rent the L40S or the L40?

Rent the L40S when the job leans on tensor math: diffusion models, prefill-heavy serving, big batches. It has twice the L40's FP8 throughput. For single-user LLM decoding, which is limited by the same 864 GB/s on both cards, I would take whichever is cheaper.

Should I rent the L40S or the A100?

The L40S when the model fits in 48 GB and you want FP8. The A100 when you need 80 GB on one GPU, or more than twice the memory bandwidth for decoding.

Neither. It connects over PCIe Gen4 x16, and NVIDIA lists MIG as unsupported. For multi-GPU inference, use pipeline parallelism or one model per GPU.

Can I get eight L40S in one VM?

Yes. The 1x, 2x, 4x and 8x configurations were all listed on 2026-09-27, and the 8x VM has 384 GB of GPU memory.

Is it a VM or a bare-metal server?

A VM. Every QuantaCloud instance runs Ubuntu 22.04 with the NVIDIA driver and Docker, and you connect over SSH as ubuntu. Bare Metal is the name of the plain Ubuntu template, not the hardware. If you need physical L40S servers to yourself, dedicated GPU servers are built to order: send a capacity brief.

Is it available right now?

The table at the top is live. On-demand capacity is not held for you, and there is no SLA. QuantaCloud runs in US regions only, and on 2026-09-27 every L40S configuration was in us-midwest-1.

What happens when I stop the instance?

Stopping terminates the VM and deletes its disk. There are no volumes or snapshots, so download generated images and trained adapters first. The unused seconds of the current hour are refunded to your balance. Pricing and the billing docs have the full rules.


The rule I follow for the L40S: rent it when the job fits in 48 GB and uses FP8 or heavy tensor math, like FLUX, Wan 2.2 5B or a 32B model at FP8. For plain decoding, check the L40 and RTX 6000 Ada prices in the GPU catalog first. When the job needs more than 48 GB on one GPU, go to the A100 or the RTX PRO 6000. Start with the ComfyUI template on a 1x VM, time one image at your own settings, and scale out from there. Launch ComfyUI on L40S

Keep building

Choose your next step.