The short answer is: the L40S for work that keeps the GPU busy, the L40 for work that keeps it waiting. Both cards have 48 GB of GDDR6, 864 GB/s of memory bandwidth and the same core counts. The extra S buys you Tensor Core throughput: NVIDIA rates the L40S at twice the L40 in TF32, BF16, FP16 and FP8, and lets it draw 350 W instead of 300 W. On 2026-09-27 that cost 15 cents more per GPU-hour. I would pay it for fine-tuning, batch image generation and a busy inference endpoint. I would not pay it for one person chatting with a model, or for a notebook that sits idle while you think.
L40 and L40S specs side by side#
The spec sheets differ in two rows that matter: Tensor Core throughput and power. Memory size and bandwidth, the two numbers that decide what fits and how fast tokens stream out, are identical.
| Spec | NVIDIA L40 | NVIDIA L40S |
|---|---|---|
| Architecture | Ada Lovelace | Ada Lovelace |
| Compute capability | 8.9 | 8.9 |
| GPU memory | 48 GB GDDR6 with ECC | 48 GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s | 864 GB/s |
| CUDA, Tensor and RT cores | 18,176, 568 and 142 | 18,176, 568 and 142 |
| FP32 | 90.5 TFLOPS | 91.6 TFLOPS |
| TF32 Tensor Core (dense) | 90.5 TFLOPS | 183 TFLOPS |
| BF16 and FP16 Tensor Core (dense) | 181 TFLOPS | 362 TFLOPS |
| FP8 Tensor Core (dense) | 362 TFLOPS | 733 TFLOPS |
| Max power | 300 W | 350 W |
| NVLink | Not supported | Not supported |
| MIG | Not supported | Not supported |
| Cooling | Passive, dual slot | Passive, dual slot |
The figures come from NVIDIA's L40 datasheet and L40S product page, and compute capability from NVIDIA's CUDA GPU list. NVIDIA also quotes Tensor Core throughput with sparsity, at twice these values. I use the dense figures for both cards so the columns compare like with like. One caveat: NVIDIA still labels its L40 datasheet preliminary.
That difference pays off only when Tensor Core math is the bottleneck. NVIDIA's own guide to LLM inference draws the line: reading the prompt (prefill) "effectively saturates GPU utilization", while generating tokens one at a time is "a memory-bound operation". Training, fine-tuning and diffusion image models live mostly on the first side of that line. Single-user token generation lives on the second, where the two cards are equal on paper.
Prices per GPU-hour#
The L40S costs more per hour, so the question is whether it finishes soon enough to make up for it.
Prices checked 6 Oct 2026, 00:35 UTC
On 2026-09-27 the L40 listed at $0.94 per GPU-hour and the L40S at $1.09 (today $0.94/GPU-hr and $1.09/GPU-hr). At those prices the L40S is cheaper per job whenever it finishes in less than 86% of the L40's time, which means at least 16% faster (our calculation: 0.94 / 1.09 = 0.86). If a job runs twice as fast, as the datasheet ratio allows for math-bound work, the L40S does it for 58% of the L40's cost (1.09 x 0.5 / 0.94 = 0.58). If a job runs at the same speed on both, the L40 is 14% cheaper. On that date the catalog offered the L40 in 1, 2 and 4-GPU VMs and the L40S in 1, 2, 4 and 8-GPU VMs. The full configurations are on the L40 and L40S pages, and how billing works is on the pricing page.
Choose the L40S for fine-tuning, images and busy endpoints#
The L40S earns its premium when the Tensor Cores are the bottleneck, and three workloads fit that description.
Fine-tuning is the clearest case. A LoRA or QLoRA run is hours of BF16 matrix math, and the L40S is rated at 362 dense BF16 TFLOPS against the L40's 181. Memory does not separate them: Axolotl puts QLoRA on a 30 to 34B model at 24 to 32 GB, which fits either card. Over a run that lasts hours, a 16% speed edge is all the L40S needs to cost less at the 2026-09-27 prices.
Batch image generation is the second. FLUX.1 [dev] ships as a 12.3 GB FP8 file and SDXL runs on 8 GB cards, so on 48 GB the limit is how fast the GPU gets through each sampling step, not how much fits. If you render hundreds of images from a ComfyUI workflow, pay for the faster card.
A busy LLM endpoint is the third. Many concurrent users and long prompts push serving toward prefill and batched math. vLLM runs per-channel FP8 checkpoints, the format llm-compressor produces, as W8A8 on both cards, and there the L40S has 733 dense FP8 TFLOPS to the L40's 362. The L40S was also the only one of the two in an 8-GPU VM on 2026-09-27, which puts 384 GB of GPU memory in one machine. Neither card has NVLink, so vLLM's advice for splitting a model across them is pipeline parallelism rather than tensor parallelism.
Choose the L40 for one user at a time#
The L40 wins when the GPU spends its time reading memory or waiting on you.
Serving one user is the main case. Every generated token reads the model's weights from memory, and both cards read at 864 GB/s. A private chat assistant, a coding model you query yourself, or Open WebUI with Ollama for a couple of colleagues gets little from the L40S's extra Tensor Core rating during generation, and on 2026-09-27 the L40 cost 14% less per hour.
Interactive work is the other. A Jupyter notebook you are exploring, a ComfyUI graph you tweak one image at a time, or a training script you are still debugging leaves the GPU idle between runs. You pay for the hour whether the Tensor Cores are working or not, so the cheaper hour wins. Once the script works and the job becomes a long run, move it to the L40S.
What fits in 48 GB#
The same models fit on both cards, because both have 48 GB. By default vLLM claims 92% of GPU memory, a budget of about 44 GB for weights and KV cache together (our calculation: 48 x 0.92 = 44.2 GB). Both cards also have FP8 Tensor Cores, so a per-channel FP8 checkpoint is the easy way to fit a 32B model on either one and keep the FP8 speed.
| Workload | Memory it needs | Source | On one 48 GB card |
|---|---|---|---|
| SDXL 1.0 images | Runs on 8 GB cards | Stability AI | Fits with room |
| FLUX.1 [dev] images | About 23 GB file at full precision, 12.3 GB FP8 file | ComfyUI docs, Black Forest Labs | Fits at either precision |
| Wan2.2 A14B video | 41.3 GB peak at 480P, 59.8 GB at 720P (official script with offloading) | Wan-AI | 480P fits, 720P does not |
| Qwen3-8B, BF16, 32k context | 16.4 GB of weights plus 4.8 GB of KV cache | Our calculation from the model config | Fits with room |
| Qwen3-32B, per-channel FP8 | 34.3 GB checkpoint | Hugging Face (RedHatAI) | Fits, about 10 GB left for KV cache, with FP8 math |
| Qwen3-32B, BF16 | 65.5 GB of weights | Our calculation | Does not fit |
| Llama-3.3-70B, 4-bit AWQ | 39.8 GB checkpoint | Hugging Face | Loads, about 4.4 GiB left for KV cache, roughly 14,000 tokens |
| gpt-oss-120b | About 65 GB on disk | OpenAI | Does not fit on one card |
| QLoRA fine-tune, 70B | 40 to 48 GB | Axolotl | Tight |
ComfyUI's FP8 files and offloading run the video models in less memory than the official scripts, at the cost of speed. For a model that is not in the table, the VRAM guide walks through the arithmetic, and the fine-tuning guide covers LoRA and QLoRA sizing.
FAQ#
Do the L40 and L40S have NVLink?
No. Both NVIDIA spec sheets list NVLink as not supported, so the GPUs in a 2, 4 or 8-GPU VM talk over PCIe. vLLM's parallelism guide recommends pipeline parallelism over tensor parallelism on GPUs without NVLink, and it names the L40S as its example.
Can both cards run FP8 models?
Yes. FP8 Tensor Core math needs compute capability 8.9 or newer, and both cards are 8.9. In vLLM 0.30.0 that holds for per-channel FP8 checkpoints. Block-scaled ones, such as Qwen's own FP8 releases, get FP8-math kernels only on Hopper and Blackwell, so on these cards they run weight-only. Compute capability is the real dividing line against the older 48 GB and 80 GB cards: the RTX A6000 and the A100 are Ampere, with no FP8 Tensor Cores, which is most of the story in L40S vs A100.
Is there a cheaper 48 GB Ada card?
On 2026-09-27 there was: the RTX 6000 Ada, at $0.79 an hour for a 1x (today $0.78/GPU-hr). It has the same 48 GB, more bandwidth at 960 GB/s, and an FP8 rating of 1,457 AI TOPS with sparsity against the L40S's 1,466, also with sparsity. Check it before you settle on either card here. Its configurations are on the RTX 6000 Ada page. L40S vs RTX 6000 Ada compares it with the L40S.
What happens to my files when I stop the VM?
They are deleted. Stopping an instance terminates it and deletes its disk, and there are no volumes or snapshots, so copy checkpoints and outputs off the VM before you stop. Unused seconds of the current hour are refunded when you do, as the pricing page explains.
The rule I follow is to rent the L40S when the GPU will spend most of the hour doing math, and the L40 when it will spend most of the hour waiting on memory or on you. If you cannot tell which kind of job you have, run it on both and stop each one 20 minutes after launch. At the 2026-09-27 prices that comes to about 31 cents on the L40 and 36 cents on the L40S (our calculation: a third of an hour on each), because billing starts at launch and the unused seconds of the hour are refunded when you stop. The deploy guide covers launching and connecting.
Launch an L40S Launch an L40