The RTX 6000 Ada is the one I would rent when both are listed: on paper it matches the L40S, and on 2026-09-27 it cost 28 percent less per GPU-hour. NVIDIA builds both on the Ada Lovelace architecture with the same core counts, 18,176 CUDA cores, 568 Tensor Cores and 142 RT Cores, and the same 48 GB of GDDR6 with ECC. Their rated FP8 throughput differs by 0.6 percent, 733 dense TFLOPS on the L40S against 728.5, and the RTX 6000 Ada reads memory 11 percent faster, 960 GB/s against 864. The prices behind that gap were $1.09 per GPU-hour for the L40S and $0.79 for the RTX 6000 Ada (our calculation: 1 - 0.79 / 1.09 = 0.28). What separates them is the board around the GPU: the L40S is a passively cooled data-center card allowed 350 W, and the RTX 6000 Ada a workstation card with a fan and a 300 W limit.
L40S and RTX 6000 Ada specs side by side#
The spec sheets match row for row on compute and split on memory speed, power and cooling.
| Spec | L40S | RTX 6000 Ada Generation |
|---|---|---|
| Architecture (compute capability) | Ada Lovelace (8.9) | Ada Lovelace (8.9) |
| CUDA, Tensor and RT Cores | 18,176, 568 and 142 | 18,176, 568 and 142 |
| Boost clock | 2,520 MHz | 2,505 MHz |
| GPU memory | 48 GB GDDR6 with ECC | 48 GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s | 960 GB/s |
| FP32 | 91.6 TFLOPS | 91.1 TFLOPS |
| BF16 and FP16 Tensor Core, dense | 362 TFLOPS | 364.2 TFLOPS |
| FP8 Tensor Core, dense | 733 TFLOPS | 728.5 TFLOPS |
| Max power | 350 W | 300 W |
| Cooling | Passive, relies on server airflow | Fan |
| Form factor | Dual slot, 4.4 x 10.5 in | Dual slot, 4.4 x 10.5 in |
| Display outputs | 4x DisplayPort 1.4a, off by default | 4x DisplayPort 1.4, on by default |
| Video engines | 3 encode, 3 decode, with AV1 | 3 encode, 3 decode, with AV1 |
| NVLink and MIG | Neither | Neither |
| GPUs per QuantaCloud VM on 2026-09-27 | 1, 2, 4 or 8 | 1, 2, 4 or 8 |
The L40S figures come from NVIDIA's L40S page and its product brief, and the RTX 6000 Ada figures from NVIDIA's Ada professional architecture whitepaper and product page. Both headline their sparse FP8 figures, 1,466 TFLOPS for the L40S and 1,457 "AI TOPS" for the RTX 6000 Ada, and the dense FP8 figures in the table are half of those. The math rows are within 0.6 percent of each other, in both directions: the L40S leads in FP8 and FP32 by the same margin that separates the boost clocks, and the RTX 6000 Ada leads in NVIDIA's BF16 figure (our calculation: 733 / 728.5, 91.6 / 91.1, 2,520 / 2,505 and 364.2 / 362 all come to about 1.006). The rows that do differ are bandwidth, where the RTX 6000 Ada has 11 percent more (960 / 864), and power, where the L40S is allowed 50 W more.
The L40S also carries the features NVIDIA lists for data-center racks, secure boot with a hardware root of trust and NEBS Level 3 readiness. Its four DisplayPort outputs ship switched off, the mode NVIDIA's virtual GPU software needs, while the RTX 6000 Ada's are on by default. On QuantaCloud both run the same way: a VM with Ubuntu 22.04, the NVIDIA driver and Docker, reached over SSH as ubuntu.
Price per job, and the VM around each card#
At the 2026-09-27 prices the L40S needed a 38 percent speed advantage to cost less per job, and the datasheets do not give it one.
| GPU | Memory | From | Available now |
|---|---|---|---|
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
Prices checked 5 Oct 2026, 21:30 UTC
On 2026-09-27 the L40S listed at $1.09 per GPU-hour in every configuration from 1x to 8x, and the RTX 6000 Ada at $0.79 for a 1x (today $1.09/GPU-hr and $0.78/GPU-hr). At those prices the L40S costs less per job only if it finishes in under 72.5 percent of the RTX 6000 Ada's time, which means running 1.38 times as fast (our calculation: 0.79 / 1.09 = 0.725, and 1.09 / 0.79 = 1.38). The math ratings are within 0.6 percent of each other, and the bandwidth favors the RTX 6000 Ada. For eight GPUs the gap was $2.51 an hour, $8.72 for the 8x L40S against $6.21 for the 8x RTX 6000 Ada (8.72 - 6.21).
The VMs around the two cards also differed on 2026-09-27. Every L40S configuration came with more local disk, while the RTX 6000 Ada's 2x, 4x and 8x came with more vCPUs, and its 8x with more RAM:
| Configuration on 2026-09-27 | L40S VM | RTX 6000 Ada VM |
|---|---|---|
| 1x | 12 vCPUs, 72 GB RAM, 625 GB disk | 12 vCPUs, 72 GB RAM, 350 GB disk |
| 2x | 24 vCPUs, 144 GB RAM, 1,250 GB disk | 26 vCPUs, 144 GB RAM, 700 GB disk |
| 4x | 46 vCPUs, 288 GB RAM, 2,500 GB disk | 52 vCPUs, 288 GB RAM, 1,400 GB disk |
| 8x | 94 vCPUs, 576 GB RAM, 5,000 GB disk | 104 vCPUs, 640 GB RAM, 2,800 GB disk |
| Regions | us-midwest-1 | us-midwest-1 and us-midwest-2, the 8x in us-midwest-1 |
The L40S VMs had 1.79 times the disk at every size (our calculation: 625 / 350 and 5,000 / 2,800). That matters for jobs that pull a large dataset or write many checkpoints to the VM. The disk is working space, not storage: stopping an instance terminates it and deletes the disk, and there are no volumes or snapshots, so copy results off first. Configurations change, so check the live tables on the L40S and RTX 6000 Ada pages before you launch. How billing works is on the pricing page.
What the L40S's extra 50 W buys#
On paper, the extra 50 W buys almost nothing. NVIDIA's figures for both cards follow their boost clocks, 2,520 MHz on the L40S and 2,505 on the RTX 6000 Ada, 15 MHz apart: the L40S's 91.6 TFLOPS of FP32 is its 18,176 cores at 2,520 MHz (our calculation: 18,176 x 2 x 2,520 MHz). What 50 W is worth in a long, heavy run depends on how close each card stays to its boost clock under load, and only a sustained test shows that. The L40S's product brief also lets a server maker program a lower power cap, so check the limit your VM reports with nvidia-smi -q -d POWER.
Which card for which job#
Most jobs point to the RTX 6000 Ada on paper, and a few point to the L40S.
| Job | What limits it | Ahead on paper |
|---|---|---|
| One user chatting with a model that fits | Memory bandwidth | RTX 6000 Ada, 11 percent more |
| Batched FP8 serving, image generation, fine-tuning | Tensor Core math | A tie within 0.6 percent, with 50 W more power budget on the L40S |
| Datasets or checkpoints of several hundred GB on the VM | Local disk | L40S: 625 GB per GPU against 350 GB on 2026-09-27 |
| A model split across two to eight GPUs | GPU count and PCIe | A tie: both came in 1, 2, 4 and 8 on 2026-09-27, and neither has NVLink |
| Video encode and decode | NVENC and NVDEC | A tie: 3 and 3 on both, with AV1 |
Both cards run FP8 the same way, because both are compute capability 8.9. vLLM 0.30.0 runs per-channel FP8 checkpoints, the kind llm-compressor makes, with FP8 math on either card, and runs block-scaled ones, such as Qwen's own FP8 releases, weight-only on both. The per-channel build of Qwen3-32B, 34.3 GB, fits either card with about 38,000 tokens of BF16 KV cache at vLLM's default budget of 44.2 GB (our calculation: (44.2 - 34.3) GB / 256 KiB per token). The L40 vs L40S comparison lists what else fits in 48 GB, and the VRAM guide shows the arithmetic for other models.
FAQ#
Are the L40S and the RTX 6000 Ada the same GPU?
They share the architecture, the core counts and the 48 GB, but they are different cards. The L40S is a passively cooled data-center board allowed 350 W, with secure boot and NEBS Level 3 readiness on NVIDIA's spec list. The RTX 6000 Ada is a workstation board with a fan, a 300 W limit and faster memory: 960 GB/s against 864.
Which is faster for LLM inference?
On paper, the RTX 6000 Ada for one user, because token generation is memory-bound and it has 11 percent more bandwidth. For many concurrent users, where Tensor Core math decides, the datasheets call it a tie.
Can I move a job from one card to the other?
Yes, without changing it. Both are compute capability 8.9 with 48 GB, so the same container, checkpoint and vLLM or ComfyUI settings run on either, including FP8. Neither card has NVLink or MIG, so a multi-GPU job talks over PCIe on both, and vLLM's docs name the L40S when they recommend pipeline parallelism over tensor parallelism for such GPUs. What changes is the VM: the L40S came with more disk and ran only in us-midwest-1 on 2026-09-27, and a new VM starts without your files, so copy them across.
My rule for this pair: check both live prices and take the RTX 6000 Ada, unless the L40S is priced lower, the RTX 6000 Ada is not listed, or the job needs the L40S VM's bigger disk. RTX 6000 Ada vs RTX A6000 covers the older 48 GB card, L40S vs A100 and L40S vs H100 the step up to 80 GB, and how to rent a GPU the first launch.
Launch an RTX 6000 Ada Launch an L40S