The short answer is: the L40S for serving many users on a model that fits in 48 GB, the A100 for anything bigger and for one fast stream. The two cards pull in opposite directions. The L40S is an Ada card with FP8 Tensor Cores, rated at 733 dense FP8 TFLOPS, while the A100 is Ampere and has no FP8 at all. The A100 answers with 80 GB of HBM2e at up to 2,039 GB/s, against 48 GB of GDDR6 at 864 GB/s on the L40S. On 2026-09-27 a 1x L40S cost $1.09 an hour and a 1x A100 SXM4 $1.50. FP8, capacity and bandwidth each decide a different workload.
L40S and A100 specs side by side#
The spec sheets favor the L40S on math and the A100 on memory. I compare the L40S with both 80 GB A100s that QuantaCloud sells, the PCIe card and the SXM4 module.
| Spec | L40S | A100 80GB PCIe | A100 80GB SXM4 |
|---|---|---|---|
| Architecture (compute capability) | Ada Lovelace (8.9) | Ampere (8.0) | Ampere (8.0) |
| GPU memory | 48 GB GDDR6 with ECC | 80 GB HBM2e | 80 GB HBM2e |
| Memory bandwidth | 864 GB/s | 1,935 GB/s | 2,039 GB/s |
| FP32 | 91.6 TFLOPS | 19.5 TFLOPS | 19.5 TFLOPS |
| TF32 Tensor Core (dense) | 183 TFLOPS | 156 TFLOPS | 156 TFLOPS |
| BF16 and FP16 Tensor Core (dense) | 362 TFLOPS | 312 TFLOPS | 312 TFLOPS |
| FP8 Tensor Core (dense) | 733 TFLOPS | Not supported | Not supported |
| FP64 | Not listed by NVIDIA | 9.7 TFLOPS, 19.5 on Tensor Cores | 9.7 TFLOPS, 19.5 on Tensor Cores |
| NVLink on the card | Not supported | 2-GPU bridge, 600 GB/s | 600 GB/s |
| Max power | 350 W | 300 W | 400 W |
| Form factor and cooling | Dual-slot PCIe card, passive | Dual-slot PCIe card, passive | SXM4 module |
The L40S figures come from NVIDIA's L40S product page, the A100 figures from NVIDIA's A100 datasheet, and compute capability from NVIDIA's CUDA GPU list. NVIDIA also quotes Tensor Core throughput with sparsity, at twice these values. I use the dense figures for all three so the columns compare like with like.
The two A100s share every compute rating. The SXM4 module has 5% more bandwidth and a 400 W power limit instead of 300 W. On 2026-09-27 they cost about the same per GPU, $1.475 to $1.50, and the SXM4 came in 1 and 8-GPU VMs while the PCIe card came in 2 and 4-GPU VMs. The configurations are on the L40S page and the A100 page.
Prices per GPU-hour#
The L40S cost 27% less per GPU-hour on 2026-09-27, and the live table shows today's gap.
| GPU | Memory | From | Available now |
|---|---|---|---|
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.50/GPU-hr | Yes |
| A100 PCIe 80GB | 80 GB | $1.48/GPU-hr | Yes |
Prices checked 5 Oct 2026, 18:03 UTC
On 2026-09-27 the L40S listed at $1.09 per GPU-hour, a 1x A100 SXM4 at $1.50 and the A100 PCIe at $1.475 per GPU-hour (today $1.09/GPU-hr, $1.50/GPU-hr and $1.48/GPU-hr). At those prices the A100 SXM4 is cheaper per job whenever it runs at least 1.38 times as fast as the L40S (our calculation: 1.50 / 1.09 = 1.38). On paper it clears that bar for single-stream generation, where its bandwidth is 2.36 times the L40S's (2,039 / 864). It falls well short for FP8 math: the L40S is rated at 733 dense FP8 TFLOPS, while the A100 runs an FP8 checkpoint at its 312 dense BF16 TFLOPS, so the L40S has 2.35 times the rate (733 / 312). How billing works is on the pricing page.
Choose the L40S for FP8 serving and models up to 32B#
The L40S wins whenever the model fits in 48 GB and the GPU is kept busy with math.
Serving an FP8 model to many users is the clearest case. Take Qwen3-32B in the per-channel FP8 format that llm-compressor produces, a 34.3 GB checkpoint (RedHatAI/Qwen3-32B-FP8-dynamic). On the L40S, vLLM runs it as W8A8 on the FP8 Tensor Cores. On the A100, vLLM loads the same file weight-only (W8A16), so the weights still take 34.3 GB but the multiplications run at BF16 rates. Under a batch of concurrent requests that is 733 dense TFLOPS against 312, at 27% less per hour. The format matters: Qwen's own FP8 release uses 128x128 block scaling, and vLLM 0.30.0 has FP8-math kernels for that format only on Hopper and Blackwell GPUs, so it runs weight-only on both of these cards. The limit is memory. At vLLM's default of reserving 92% of GPU memory, the L40S has about 9.4 GiB left for KV cache, about 38,000 tokens across all requests at BF16, or twice that with --kv-cache-dtype fp8 (our calculation at 256 KiB per token from the 46,068 MiB the L40S reports). If your traffic needs more context in flight than that, the A100's 41.6 GiB of headroom wins instead.
Fine-tuning and image generation on models that fit follow the same logic with a smaller margin. The L40S is rated at 362 dense BF16 TFLOPS against the A100's 312 and cost 27% less per hour, so a QLoRA run on a model up to about 32B (24 to 32 GB, per Axolotl) or a batch of FLUX.1 images (a 12.3 GB FP8 file) should cost less on the L40S. It also has 4.7 times the A100's plain FP32 rate, 91.6 TFLOPS against 19.5, for code that does not use Tensor Cores.
Choose the A100 for 80 GB models and one fast stream#
The A100 wins on capacity first. The widely used AWQ build of Llama-3.3-70B is a 39.8 GB download. On the L40S, vLLM's default leaves about 4.4 GiB for KV cache, roughly 14,000 tokens shared by every request. On the A100 it leaves about 36.6 GiB, roughly 120,000 tokens (our calculation at 320 KiB of BF16 KV cache per token). Unsloth's benchmark shows the same gap for fine-tuning: QLoRA on Llama 3.3 70B reaches 12,106 tokens of context on a 48 GB GPU and 89,389 on 80 GB. gpt-oss-120b, about 65 GB, and Qwen3-32B in BF16, 65.5 GB, fit on one A100 and not on one L40S. The KV cache guide explains why context eats memory this fast.
One fast stream is the other A100 case. NVIDIA describes token-by-token generation as memory-bound, so for a single stream bandwidth is the number to compare, and the A100 SXM4 has 2.36 times the L40S's. If that stream runs continuously, as in an agent loop or a queue of sequential jobs, the A100 is faster and, on paper, cheaper per token: 1.38 times the price for up to 2.36 times the speed. If the stream is one person typing, the GPU idles between messages and the cheaper hour matters more than the faster token.
The A100 is also the card for double precision. NVIDIA rates it at 9.7 TFLOPS of FP64, and 19.5 on its Tensor Cores, and publishes no FP64 figure for the L40S.
What fits in 48 GB and in 80 GB#
Capacity decides more of these choices than speed does. vLLM reserves 92% of GPU memory by default, about 41.4 GiB on the L40S and 73.6 GiB on the A100 (our calculation from the 46,068 and 81,920 MiB that nvidia-smi reports).
| Workload | Memory it needs | Source | L40S, 48 GB | A100, 80 GB |
|---|---|---|---|---|
| FLUX.1 [dev] images | 12.3 GB FP8 file | Black Forest Labs | Fits | Fits |
| Qwen3-8B, BF16, 32k context | 16.4 GB of weights plus 4.8 GB of KV cache | Our calculation from the model config | Fits | Fits |
| Qwen3-32B, per-channel FP8 | 34.3 GB checkpoint | Hugging Face | Fits, about 9.4 GiB left, FP8 math | Fits, about 41.6 GiB left, 16-bit math |
| Qwen3-32B, BF16 | 65.5 GB of weights | Our calculation | Does not fit | Fits, about 12.6 GiB left |
| Llama-3.3-70B, 4-bit AWQ | 39.8 GB checkpoint | Hugging Face | Loads, about 4.4 GiB left | Fits, about 36.6 GiB left |
| gpt-oss-120b | About 65 GB on disk | OpenAI | Does not fit | Fits on one GPU |
| QLoRA, 30 to 34B model | 24 to 32 GB | Axolotl | Fits | Fits |
| QLoRA, 70B model | 40 to 48 GB | Axolotl | Tight | Fits |
The "left" figures are what remains for KV cache at vLLM's default, before activations. For any model not in the table, the VRAM guide walks through the arithmetic.
FAQ#
Does the A100 support FP8 at all?
Not in hardware. FP8 Tensor Core math needs compute capability 8.9, and the A100 is 8.0. vLLM still loads FP8 checkpoints on it as weight-only W8A16, which halves the memory of the weights but keeps the math in 16-bit. If you need FP8 math and 80 GB on one card, the H100 PCIe has both. It listed at $2.59 per GPU-hour on 2026-09-27, today $2.59/GPU-hr. L40S vs H100 compares it with the L40S.
Are two L40S better than one A100?
Only for FP8 models over 48 GB. The per-channel FP8 build of Llama-3.3-70B is a 72.7 GB checkpoint: on one A100 it leaves about 5.9 GiB of the 73.6 GiB budget for KV cache, about 19,000 tokens, and runs without FP8 math. Two L40S give you an 82.8 GiB budget, about 15.1 GiB of it left for KV cache, and FP8 Tensor Cores for that checkpoint (our calculation: 2 x 41.39 - 67.68). For a 70B model at 4-bit, one A100 is the better deal: $1.50 per hour against $2.18 for a 2-GPU L40S VM on 2026-09-27. The L40S has no NVLink, so a pair splits the model over PCIe, and vLLM recommends pipeline parallelism there rather than tensor parallelism.
What happens to my data when I stop the VM?
It is deleted. Stopping an instance terminates it and deletes its disk, and there are no volumes or snapshots, so copy checkpoints and outputs off first. Unused seconds of the current hour are refunded when you stop, as the pricing page explains.
The rule I follow is: if the model and its batch fit in 48 GB, rent the L40S and run a per-channel FP8 checkpoint, and if they do not, or if the speed of one stream is what you are paying for, rent the A100. Before you launch, add up the weights, the KV cache for your context and the number of requests you expect in flight. If the total sits under about 44 GB, the L40S is the better buy. The L40 vs L40S and RTX A6000 vs A100 comparisons cover the cheaper cards on each side, and the deploy guide covers launching and connecting.
Launch an L40S Launch an A100 80GB