GPU comparison

L40S vs RTX 6000 Ada

L40S vs RTX 6000 Ada: the same core counts and 48 GB, FP8 within 1%, but 11% more bandwidth on the RTX 6000 Ada. Specs, live prices and which to rent.

Faiz Ahmed9 min read
NVIDIA · Ada Lovelace

L40S

Cloud GPU with full SSH access

48 GBGDDR6 memory
Bandwidth
864 GB/s
FP8 compute
Native FP8
From / GPU-hr$1.09
Launch L40S
NVIDIA · Ada Lovelace

RTX 6000 Ada

Cloud GPU with full SSH access

48 GBGDDR6 memory
Bandwidth
960 GB/s
FP8 compute
Native FP8
From / GPU-hr$0.79
Launch RTX 6000 Ada

Prices checked

The RTX 6000 Ada is the one I would rent when both are listed: on paper it matches the L40S, and on 2026-09-27 it cost 28 percent less per GPU-hour. NVIDIA builds both on the Ada Lovelace architecture with the same core counts, 18,176 CUDA cores, 568 Tensor Cores and 142 RT Cores, and the same 48 GB of GDDR6 with ECC. Their rated FP8 throughput differs by 0.6 percent, 733 dense TFLOPS on the L40S against 728.5, and the RTX 6000 Ada reads memory 11 percent faster, 960 GB/s against 864. The prices behind that gap were $1.09 per GPU-hour for the L40S and $0.79 for the RTX 6000 Ada (our calculation: 1 - 0.79 / 1.09 = 0.28). What separates them is the board around the GPU: the L40S is a passively cooled data-center card allowed 350 W, and the RTX 6000 Ada a workstation card with a fan and a 300 W limit.

L40S and RTX 6000 Ada specs side by side#

The spec sheets match row for row on compute and split on memory speed, power and cooling.

SpecL40SRTX 6000 Ada Generation
Architecture (compute capability)Ada Lovelace (8.9)Ada Lovelace (8.9)
CUDA, Tensor and RT Cores18,176, 568 and 14218,176, 568 and 142
Boost clock2,520 MHz2,505 MHz
GPU memory48 GB GDDR6 with ECC48 GB GDDR6 with ECC
Memory bandwidth864 GB/s960 GB/s
FP3291.6 TFLOPS91.1 TFLOPS
BF16 and FP16 Tensor Core, dense362 TFLOPS364.2 TFLOPS
FP8 Tensor Core, dense733 TFLOPS728.5 TFLOPS
Max power350 W300 W
CoolingPassive, relies on server airflowFan
Form factorDual slot, 4.4 x 10.5 inDual slot, 4.4 x 10.5 in
Display outputs4x DisplayPort 1.4a, off by default4x DisplayPort 1.4, on by default
Video engines3 encode, 3 decode, with AV13 encode, 3 decode, with AV1
NVLink and MIGNeitherNeither
GPUs per QuantaCloud VM on 2026-09-271, 2, 4 or 81, 2, 4 or 8

The L40S figures come from NVIDIA's L40S page and its product brief, and the RTX 6000 Ada figures from NVIDIA's Ada professional architecture whitepaper and product page. Both headline their sparse FP8 figures, 1,466 TFLOPS for the L40S and 1,457 "AI TOPS" for the RTX 6000 Ada, and the dense FP8 figures in the table are half of those. The math rows are within 0.6 percent of each other, in both directions: the L40S leads in FP8 and FP32 by the same margin that separates the boost clocks, and the RTX 6000 Ada leads in NVIDIA's BF16 figure (our calculation: 733 / 728.5, 91.6 / 91.1, 2,520 / 2,505 and 364.2 / 362 all come to about 1.006). The rows that do differ are bandwidth, where the RTX 6000 Ada has 11 percent more (960 / 864), and power, where the L40S is allowed 50 W more.

The L40S also carries the features NVIDIA lists for data-center racks, secure boot with a hardware root of trust and NEBS Level 3 readiness. Its four DisplayPort outputs ship switched off, the mode NVIDIA's virtual GPU software needs, while the RTX 6000 Ada's are on by default. On QuantaCloud both run the same way: a VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu.

Price per job, and the VM around each card#

At the 2026-09-27 prices the L40S needed a 38 percent speed advantage to cost less per job, and the datasheets do not give it one.

GPUMemoryFromAvailable now
L40S48 GB$1.09/GPU-hrYes
RTX 6000 Ada48 GB$0.78/GPU-hrYes

Prices checked 5 Oct 2026, 21:30 UTC

On 2026-09-27 the L40S listed at $1.09 per GPU-hour in every configuration from 1x to 8x, and the RTX 6000 Ada at $0.79 for a 1x (today $1.09/GPU-hr and $0.78/GPU-hr). At those prices the L40S costs less per job only if it finishes in under 72.5 percent of the RTX 6000 Ada's time, which means running 1.38 times as fast (our calculation: 0.79 / 1.09 = 0.725, and 1.09 / 0.79 = 1.38). The math ratings are within 0.6 percent of each other, and the bandwidth favors the RTX 6000 Ada. For eight GPUs the gap was $2.51 an hour, $8.72 for the 8x L40S against $6.21 for the 8x RTX 6000 Ada (8.72 - 6.21).

The VMs around the two cards also differed on 2026-09-27. Every L40S configuration came with more local disk, while the RTX 6000 Ada's 2x, 4x and 8x came with more vCPUs, and its 8x with more RAM:

Configuration on 2026-09-27L40S VMRTX 6000 Ada VM
1x12 vCPUs, 72 GB RAM, 625 GB disk12 vCPUs, 72 GB RAM, 350 GB disk
2x24 vCPUs, 144 GB RAM, 1,250 GB disk26 vCPUs, 144 GB RAM, 700 GB disk
4x46 vCPUs, 288 GB RAM, 2,500 GB disk52 vCPUs, 288 GB RAM, 1,400 GB disk
8x94 vCPUs, 576 GB RAM, 5,000 GB disk104 vCPUs, 640 GB RAM, 2,800 GB disk
Regionsus-midwest-1us-midwest-1 and us-midwest-2, the 8x in us-midwest-1

The L40S VMs had 1.79 times the disk at every size (our calculation: 625 / 350 and 5,000 / 2,800). That matters for jobs that pull a large dataset or write many checkpoints to the VM. The disk is working space, not storage: stopping an instance terminates it and deletes the disk, and there are no volumes or snapshots, so copy results off first. Configurations change, so check the live tables on the L40S and RTX 6000 Ada pages before you launch. How billing works is on the pricing page.

What the L40S's extra 50 W buys#

On paper, the extra 50 W buys almost nothing. NVIDIA's figures for both cards follow their boost clocks, 2,520 MHz on the L40S and 2,505 on the RTX 6000 Ada, 15 MHz apart: the L40S's 91.6 TFLOPS of FP32 is its 18,176 cores at 2,520 MHz (our calculation: 18,176 x 2 x 2,520 MHz). What 50 W is worth in a long, heavy run depends on how close each card stays to its boost clock under load, and only a sustained test shows that. The L40S's product brief also lets a server maker program a lower power cap, so check the limit your VM reports with nvidia-smi -q -d POWER.

Which card for which job#

Most jobs point to the RTX 6000 Ada on paper, and a few point to the L40S.

JobWhat limits itAhead on paper
One user chatting with a model that fitsMemory bandwidthRTX 6000 Ada, 11 percent more
Batched FP8 serving, image generation, fine-tuningTensor Core mathA tie within 0.6 percent, with 50 W more power budget on the L40S
Datasets or checkpoints of several hundred GB on the VMLocal diskL40S: 625 GB per GPU against 350 GB on 2026-09-27
A model split across two to eight GPUsGPU count and PCIeA tie: both came in 1, 2, 4 and 8 on 2026-09-27, and neither has NVLink
Video encode and decodeNVENC and NVDECA tie: 3 and 3 on both, with AV1

Both cards run FP8 the same way, because both are compute capability 8.9. vLLM 0.30.0 runs per-channel FP8 checkpoints, the kind llm-compressor makes, with FP8 math on either card, and runs block-scaled ones, such as Qwen's own FP8 releases, weight-only on both. The per-channel build of Qwen3-32B, 34.3 GB, fits either card with about 38,000 tokens of BF16 KV cache at vLLM's default budget of 44.2 GB (our calculation: (44.2 - 34.3) GB / 256 KiB per token). The L40 vs L40S comparison lists what else fits in 48 GB, and the VRAM guide shows the arithmetic for other models.

FAQ#

Are the L40S and the RTX 6000 Ada the same GPU?

They share the architecture, the core counts and the 48 GB, but they are different cards. The L40S is a passively cooled data-center board allowed 350 W, with secure boot and NEBS Level 3 readiness on NVIDIA's spec list. The RTX 6000 Ada is a workstation board with a fan, a 300 W limit and faster memory: 960 GB/s against 864.

Which is faster for LLM inference?

On paper, the RTX 6000 Ada for one user, because token generation is memory-bound and it has 11 percent more bandwidth. For many concurrent users, where Tensor Core math decides, the datasheets call it a tie.

Can I move a job from one card to the other?

Yes, without changing it. Both are compute capability 8.9 with 48 GB, so the same container, checkpoint and vLLM or ComfyUI settings run on either, including FP8. Neither card has NVLink or MIG, so a multi-GPU job talks over PCIe on both, and vLLM's docs name the L40S when they recommend pipeline parallelism over tensor parallelism for such GPUs. What changes is the VM: the L40S came with more disk and ran only in us-midwest-1 on 2026-09-27, and a new VM starts without your files, so copy them across.


My rule for this pair: check both live prices and take the RTX 6000 Ada, unless the L40S is priced lower, the RTX 6000 Ada is not listed, or the job needs the L40S VM's bigger disk. RTX 6000 Ada vs RTX A6000 covers the older 48 GB card, L40S vs A100 and L40S vs H100 the step up to 80 GB, and how to rent a GPU the first launch.

Launch an RTX 6000 Ada Launch an L40S

Keep building

Choose your next step.