GPU comparison

L40S vs H100 PCIe

NVIDIA L40S vs H100 PCIe: 48 against 80 GB, 864 against 2,000 GB/s, FP8 on both. Where 48 GB is enough, where it is not, and live hourly prices.

Faiz Ahmed12 min read
NVIDIA · Ada Lovelace

L40S

Cloud GPU with full SSH access

48 GBGDDR6 memory
Bandwidth
864 GB/s
FP8 compute
Native FP8
From / GPU-hr$1.09
Launch L40S
NVIDIA · Hopper

H100 PCIe

Cloud GPU with full SSH access

80 GBHBM2e memory
Bandwidth
2.0 TB/s
FP8 compute
Native FP8
From / GPU-hrSee console
Explore H100 PCIe

Prices checked

The L40S is the better buy for any job that fits in 48 GB, and the H100 PCIe for any job that does not or that someone waits on token by token. Both are PCIe cards with FP8 Tensor Cores and a 350 W limit. The H100 PCIe has 80 GB of HBM2e at 2,000 GB/s against 48 GB of GDDR6 at 864 GB/s, and about twice the L40S's dense Tensor Core rating. On 2026-09-27 it also cost 2.38 times as much per hour: $2.59 for a 1x H100 PCIe against $1.09 for a 1x L40S. No ratio on the spec sheet reaches 2.38, so the H100 has to earn its hour another way: with capacity, with the speed of a single stream, or with kernels that need Hopper.

L40S and H100 PCIe specs side by side#

The spec sheets favor the H100 on memory, bandwidth and Tensor Core math, and the L40S on plain FP32 and video encoding. QuantaCloud rents the H100 by the hour as the PCIe card, not the SXM module, so that is the version in the table.

SpecL40SH100 PCIe
Architecture (compute capability)Ada Lovelace (8.9)Hopper (9.0)
GPU memory48 GB GDDR6 with ECC80 GB HBM2e
Memory bandwidth864 GB/s2,000 GB/s
FP3291.6 TFLOPS51.2 TFLOPS
TF32 Tensor Core (dense)183 TFLOPS378 TFLOPS
BF16 and FP16 Tensor Core (dense)362 TFLOPS756 TFLOPS
FP8 Tensor Core (dense)733 TFLOPS1,513 TFLOPS
FP64Not listed by NVIDIA25.6 TFLOPS, 51.2 on Tensor Cores
NVLinkNot supportedBridge to one adjacent card, 600 GB/s
Multi-Instance GPUNot supportedUp to 7 x 10 GB
Video encoders (NVENC)3, with AV1None
Host linkPCIe Gen4 x16, 64 GB/sPCIe Gen5 x16, 128 GB/s
Max power350 W350 W
Form factor and coolingDual-slot PCIe card, passiveDual-slot PCIe card, passive

The L40S figures come from NVIDIA's L40S product page, the H100 PCIe figures from NVIDIA's H100 whitepaper and H100 PCIe product brief, and compute capability from NVIDIA's CUDA GPU list. Both spec sheets also print sparse Tensor Core figures, 1,466 FP8 TFLOPS for the L40S and 3,026 for the H100 PCIe, and the table keeps to the dense half on both sides.

Three rows decide most choices between these two. Memory decides what fits. Bandwidth decides how fast a single stream of tokens comes out, because NVIDIA describes token-by-token generation as memory-bound, and the H100 reads memory 2.31 times as fast (our calculation: 2,000 / 864). FP8 decides how much a busy server gets through, and the H100's dense FP8 rating is 2.06 times the L40S's (1,513 / 733). Both cards run per-channel FP8 checkpoints with FP8 math in vLLM 0.30.0. Block-scaled FP8 checkpoints, such as Qwen's own FP8 releases, get FP8-math kernels only on Hopper and Blackwell, so on the L40S they run weight-only.

Prices per GPU-hour#

The price gap is wider than any spec gap: on 2026-09-27 the H100 PCIe cost 2.38 times as much per GPU-hour as the L40S.

GPUMemoryFromAvailable now
L40S48 GB$1.09/GPU-hrYes
H100 PCIe-Not listedNo

Prices checked 5 Oct 2026, 20:19 UTC

That day a 1x L40S listed at $1.09 an hour and a 1x H100 PCIe at $2.59 (today $1.09/GPU-hr and the console price). For the H100 to cost less per job, it has to finish in 42% of the L40S's time, which means running 2.38 times as fast (our calculation: 1.09 / 2.59 = 0.42 and 2.59 / 1.09 = 2.38). Its bandwidth comes close, at 2.31 times. Its dense FP8 and BF16 rates fall further short, at 2.06 and 2.09 times. So on paper a job that fits in 48 GB and keeps the GPU busy costs a little less on the L40S and finishes sooner on the H100. The L40S came in 1, 2, 4 and 8-GPU VMs in us-midwest-1 and the H100 PCIe in 1 and 2-GPU VMs in us-midwest-2, and the L40S page and the H100 page list each configuration with its vCPUs, RAM and disk. The pricing page explains how a partly used hour is refunded.

Choose the L40S when the job fits in 48 GB#

The L40S wins whenever the model, its context and the batch fit in about 44 GB, because the H100's speed advantage is smaller than its price premium.

Serving a 32B model at FP8 is the clearest case. The per-channel FP8 build of Qwen3-32B (RedHatAI/Qwen3-32B-FP8-dynamic) is a 34.3 GB checkpoint, and vLLM runs it with FP8 math on both cards. At vLLM's default of reserving 92% of GPU memory, the L40S has about 9.4 GiB left for KV cache, roughly 38,000 tokens across all requests at BF16, or twice that with --kv-cache-dtype fp8 (our calculation at 256 KiB per token). That holds 16 requests of 2,000 tokens in flight at once (16 x 2,000 = 32,000 tokens). Pick the per-channel build: Qwen's own Qwen3-32B-FP8 is block-scaled, and on the L40S it loses the FP8 math.

Image and video generation is the second case. FLUX.1 [dev] ships as a 12.3 GB FP8 file and Wan 2.2's 5B model peaked at 22.9 GB at 720p in Wan's own test, so on 48 GB the limit is sampling speed, not memory. ComfyUI computes in FP8 on compute capability 8.9 and newer, which covers both cards. The L40S also has three NVENC encoders with AV1, and NVIDIA's H100 whitepaper notes that the H100 has no NVENC, so only the L40S can encode the finished video in hardware on the same GPU. Run ComfyUI on a cloud GPU covers the setup.

QLoRA fine-tuning up to 32B is the third. Axolotl puts QLoRA on a 30 to 34B model at 24 to 32 GB and Unsloth puts a 32B model at 26 GB, so the job fits either card, and over a run of many hours the cheaper hour decides the bill. The fine-tuning VRAM guide has the sizes for other models.

The L40S also scales wider in one VM. On 2026-09-27 it came in 8-GPU VMs with 384 GB in total, while the H100 PCIe stopped at two GPUs. For batch jobs that run one model copy per GPU, that is eight workers on one machine.

Choose the H100 PCIe for 80 GB, one fast stream or Hopper kernels#

The H100 wins on capacity first, and two model makers name it when they size their releases. Google says Gemma 4 31B in BF16, 62.5 GB of weights, will "fit efficiently on a single 80GB NVIDIA H100 GPU", and OpenAI sized gpt-oss-120b, about 65 GB, for a single 80 GB GPU such as the H100. Neither loads on one L40S, and neither does Qwen3-32B in BF16 at 65.5 GB. A 4-bit 70B model loads on both, but only the H100 has room to serve it: the AWQ build of Llama-3.3-70B (casperhansen/llama-3.3-70b-instruct-awq, 39.8 GB) leaves about 4.4 GiB for KV cache on the L40S, roughly 14,000 tokens shared by every request, and about 36.2 GiB on the H100, roughly 119,000 (our calculation at 320 KiB of BF16 KV cache per token). Training shows the same gap: Unsloth measured a maximum QLoRA context for Llama 3.3 70B of 12,106 tokens on a 48 GB GPU and 89,389 on 80 GB.

One fast stream is the second case. For a single request the H100's 2.31 times the bandwidth sets the ceiling, and on paper each token costs about the same on both cards: 2.38 times the price for 2.31 times the speed. So when a person or an agent loop waits on every response, the H100 buys back time at little extra cost per token.

Hopper kernels are the third. vLLM 0.30.0 runs block-scaled FP8 checkpoints with FP8 math on the H100, and FlashAttention-3 lists the H100 as a requirement. vLLM also picks bigger batch defaults on GPUs with at least 70 GiB of memory: 1,024 concurrent sequences and 8,192 batched tokens for its API server, against 256 and 2,048 on the L40S. The H100 is also the only one of the two for double precision, at 25.6 TFLOPS of FP64 and 51.2 on Tensor Cores, where NVIDIA lists no FP64 figure for the L40S.

What fits in 48 GB and in 80 GB#

The two working budgets are about 41.4 GiB and 73.3 GiB, since vLLM reserves 92% of GPU memory by default (our calculation from the 46,068 and 81,559 MiB that nvidia-smi reports, x 0.92).

WorkloadMemory it needsSourceL40S, 48 GBH100 PCIe, 80 GB
FLUX.1 [dev] images12.3 GB FP8 fileBlack Forest LabsFits, FP8 computeFits, FP8 compute
Qwen3-8B, BF16, 32k context16.4 GB of weights plus 4.8 GB of KV cacheOur calculation from the model configFitsFits
Qwen3-32B, per-channel FP834.3 GB checkpointHugging Face (RedHatAI)Fits, about 9.4 GiB left, FP8 mathFits, about 41.3 GiB left, FP8 math
Qwen3-32B, Qwen's block-scaled FP834.3 GB checkpointHugging Face (Qwen)Fits, weight-onlyFits, FP8 math
Gemma 4 31B, BF1662.5 GB of weightsGoogle, our calculationDoes not fitFits, per Google
Qwen3-32B, BF1665.5 GB of weightsOur calculationDoes not fitFits, about 12.3 GiB left
gpt-oss-120bAbout 65 GB on diskOpenAIDoes not fitFits, 2.8 full 131k-token sequences
Llama-3.3-70B, 4-bit AWQ39.8 GB checkpointHugging FaceLoads, about 4.4 GiB leftFits, about 36.2 GiB left
Llama-3.3-70B, per-channel FP872.7 GB checkpointHugging Face (RedHatAI)Does not fitLoads, about 5.6 GiB left
QLoRA, 30 to 34B model24 to 32 GBAxolotlFitsFits
QLoRA, Llama 3.3 70B40 to 48 GBAxolotl, UnslothTight, 12,106 tokens of contextFits, 89,389 tokens of context
Wan 2.2 A14B video, 720p59.8 GB peak (official script with offloading)Wan-AIOnly with ComfyUI offloading, slowerFits

The "left" figures are what each card has for KV cache after loading, before activations. The VRAM guide runs the same arithmetic for other models, and the KV cache guide shows why a long context fills 48 GB so quickly.

FAQ#

Are two L40S better than one H100 PCIe?

For a model too big for one L40S, sometimes, and for less money. A 2-GPU L40S VM listed at $2.18 an hour on 2026-09-27, 16% less than one H100 PCIe, with 96 GB against 80. That pair holds the per-channel FP8 build of Llama-3.3-70B, a 72.7 GB checkpoint that leaves about 5.6 GiB for KV cache on one H100 and about 15.1 GiB on two L40S (our calculation: 2 x 41.39 - 67.68). The catch is the link. The L40S has no NVLink, so the pair splits the model over PCIe, and vLLM recommends pipeline parallelism there, which adds capacity rather than speed for a single request. For a model that fits on one H100, the single card is the simpler setup and the faster stream.

How do I tell a per-channel FP8 checkpoint from a block-scaled one?

Open its config.json, because the file size will not tell you: both Qwen3-32B FP8 builds are 34.3 GB. Qwen's own Qwen3-32B-FP8 sets quant_method to fp8 with a weight_block_size of 128 x 128, which is block-scaled, so vLLM 0.30.0 runs it with FP8 math on the H100 and weight-only on the L40S. RedHatAI/Qwen3-32B-FP8-dynamic uses compressed-tensors with its weights quantized per channel, which gets FP8 math on both cards.

Only nvidia-smi topo -m on the VM can say. The H100 PCIe supports a bridge to one neighboring card, but QuantaCloud's catalog lists these offers as PCIe. If the GPU0 to GPU1 entry does not start with NV, the two H100s talk over PCIe, as a pair of L40S cards always does.

What happens to my files when I stop either VM?

They are deleted with the VM's local disk, which on 2026-09-27 was 625 GB on a 1x L40S and 1,250 GB on a 1x H100 PCIe. Stopping terminates the instance, and there are no volumes or snapshots, so copy QLoRA adapters, rendered video and benchmark logs off first. You get back the unused seconds of the current hour, as the pricing page explains.


My rule for this pair starts from the price: at 2.38 times the hourly rate, the H100 PCIe has to do something the L40S cannot. Add up the weights, the KV cache for your context and the requests you expect in flight. Under about 44 GB, rent the L40S, with a per-channel FP8 checkpoint if the model has one. Over it, or when a person waits on every token, rent the H100 PCIe. If the total passes about 74 GB, price the H200 NVL before you split the model across two H100s. L40S vs A100 covers the 80 GB card without FP8, RTX PRO 6000 vs H100 covers the 96 GB alternative, and the deploy guide covers launching and connecting.

Launch an L40S Launch an H100 PCIe

Keep building

Choose your next step.