The honest answer is about $30,000 to $34,000 for one H200 NVL card at US distributors in September 2026. Reuters reported a lower figure for China, around $27,000 per chip, and OEM-branded versions of the same card are advertised above $90,000, so check the part number before you compare prices. Renting an H200 NVL on QuantaCloud costs the console price. At the $3.43 per GPU-hour it cost on 2026-09-27, a card at $29,999.99 to $33,731.99 equals 8,746 to 9,834 hours of rental, or 12 to 13.5 months of round-the-clock use (our calculation, below). The SXM H200 ships on 4- and 8-GPU HGX boards inside servers, and on QuantaCloud it is a reserved build.
How much one H200 costs to buy#
The number I plan with is the stand-alone card price. Every figure below is dated and says what it covers.
| Date | Source | What the price covers | Price |
|---|---|---|---|
| December 2025 | Reuters, citing unnamed sources | H200 chips for Chinese customers | Around $27,000 per chip |
| December 2025 | Reuters, same report | Eight-chip H200 module for Chinese customers | Around 1.5 million yuan |
| 2026-09-28 | CDW listing, PNY kit NVH200NVLTCGPU-KIT | New H200 NVL card | $33,731.99 |
| 2026-09-28 | Central Computer listing, same kit | New H200 NVL card, out of stock | $29,999.99 |
| 2026-09-28 | CDW listing, Cisco part UCSC-GPU-H200-NVL= | H200 NVL sold as a Cisco server part | $91,644.99 |
| 2026-09-28 | CDW listing, HPE part S3U30C | H200 NVL sold as an HPE server part | $94,698.99 |
The same GPU lists at very different prices, and the reason is the part number. The PNY kit is the stand-alone card. The Cisco and HPE parts are the same H200 NVL sold as options for those vendors' servers, and on the same day at the same distributor they listed at 2.7 to 2.8 times the PNY card (our calculation: $91,644.99 / $33,731.99 = 2.72 and $94,698.99 / $33,731.99 = 2.81). The Reuters figure covers H200 variants priced for Chinese customers and comes from unnamed sources, so I do not use it for US planning. An OEM part price is a number to negotiate inside a server quote, not a card price.
H200 NVL vs H200 SXM#
The NVL and the SXM carry the same memory, so for one GPU the NVL gives up little.
| H200 NVL | H200 SXM | |
|---|---|---|
| Memory | 141 GB HBM3e | 141 GB HBM3e |
| Memory bandwidth | 4.8 TB/s | 4.8 TB/s |
| FP8 Tensor, with sparsity | 3,341 TFLOPS | 3,958 TFLOPS |
| Maximum power | Up to 600 W | Up to 700 W |
| GPU-to-GPU link | 2- or 4-way NVLink bridge, 900 GB/s per GPU | NVLink, 900 GB/s, NVSwitch on 8-GPU boards |
| Sold as | Dual-slot air-cooled PCIe card | Module on 4- or 8-GPU HGX H200 boards |
NVIDIA aimed the NVL at air-cooled enterprise racks when it launched in November 2024, citing a survey that roughly 70% of enterprise racks are 20 kW and below and use air cooling. The SXM version buys more compute per GPU and the NVSwitch fabric that ties eight GPUs into one node with 1.1 TB of HBM3e. If the model fits in one, two or four NVL cards, I would not pay for SXM: the NVL bridges up to four cards at the same 900 GB/s per GPU, and the smallest SXM purchase is a 4-GPU board. The KV cache guide shows how much of 141 GB a long-context model really uses, and the H200 page covers what fits on one card.
H100 vs H200 price#
The newer card was not the more expensive one on list price. On 2026-09-28 one distributor listed the PNY H200 NVL at $33,731.99 and another listed the PNY H100 PCIe at $35,315.93, out of stock, although the H200 NVL carries 141 GB against 80 GB, 76% more (our calculation: 141 / 80 = 1.76). By the hour, these are the live QuantaCloud prices for the two:
The H100 price guide has the full set of H100 purchase prices and sources.
HGX H200 and DGX H200 prices#
The HGX H200 has no US price I could verify. CDW's listing for HPE's HGX H200 8-GPU board shows it discontinued on June 9, 2026, with no price, and the only dated system figure I found is the Reuters report that an eight-chip module was expected to cost around 1.5 million yuan in China. A server quote adds CPUs, memory, storage, network cards and support to the board, so an HGX or DGX H200 has no single price.
If you are pricing an HGX H200 server, get a written number for a reserved build too. QuantaCloud builds HGX H200 servers to order, from a single server to InfiniBand clusters: we order and build the hardware to your spec, and the configuration, lead time and terms come back in writing. Send a capacity brief with the GPU count, network and term.
What renting an H200 costs on QuantaCloud#
An H200 NVL rents for the console price on QuantaCloud, billed by the hour. These are the live configurations:
No configuration is listed right now. See every GPU in the console.
Prices checked 6 Oct 2026, 01:25 UTC
On 2026-09-27 the single-GPU H200 NVL VM came with 16 vCPU, 180 GB of RAM and 750 GB of disk in Virginia (us-east-1), and a 2-GPU VM was listed as well. Each instance is an Ubuntu 22.04 VM with the NVIDIA driver and Docker. Billing works from a prepaid balance with a $5 minimum deposit: the first hour is charged at launch, later hours as each previous hour is used up, and the unused seconds of the current hour are refunded when you stop. There are no egress or ingress charges. Stopping terminates the VM and deletes its disk, so copy results off before you stop. The pricing page and the billing docs have the details. SXM H200 servers are not on demand: they are the reserved builds above.
Launch an H200 NVLBreak-even: how many rented hours one H200 buys#
The rule I follow: divide the card price by the hourly price, then by the hours a year the GPU will really be busy. At $3.43 per GPU-hour, the H200 NVL price on 2026-09-27 (live now: the console price), our calculation gives:
| Card price used | Rental hours it equals | Years at 100% use | Years at 50% use | Years at 25% use |
|---|---|---|---|---|
| $29,999.99 (PNY kit, Central Computer, out of stock) | 8,746 | 1.0 | 2.0 | 4.0 |
| $33,731.99 (PNY kit, CDW) | 9,834 | 1.1 | 2.2 | 4.5 |
| $91,644.99 (Cisco part, CDW) | 26,719 | 3.1 | 6.1 | 12.2 |
The arithmetic for the CDW row: $33,731.99 / $3.43 = 9,834 hours. A year has 8,760 hours, so that is 1.1 years at 100% use, 2.2 years at 50% and 4.5 years at 25%. If the live price differs from $3.43, divide by the live number instead.
The OEM row shows why the part number matters: bought at that advertised price, the card needs 3.1 years of round-the-clock use to match the rent. Every row is the price of the card alone, and the costs in the next section push the real break-even later.
What the card price leaves out#
The card price covers 141 GB of memory on a PCIe board and nothing else.
| Cost | Buying an H200 NVL card | Renting on QuantaCloud |
|---|---|---|
| Host server with PCIe Gen5 slots, CPUs, RAM, disk | A separate purchase | In the hourly price |
| Power | Up to 600 W for the card alone, 5,256 kWh a year at full load | In the hourly price |
| Cooling and rack space | Yours: the card is passive and needs server airflow | In the hourly price |
| Network, setup and staff time | Yours | In the hourly price, no egress charges |
| Depreciation and resale | Yours | None |
The 5,256 kWh is our calculation: 600 W x 8,760 hours, for the card at its maximum configurable power. Multiply by your electricity rate, then add the host server and the cooling. I could not find a reputable, dated source for used H200 prices, so resale value stays out of the sums.
Buy, rent or reserve#
My rule is to rent until the busy hours are measured, then choose between buying and reserving. Renting fits a model that needs 141 GB on one GPU for weeks rather than years, and it fits uncertain demand, because unused seconds are refunded when you stop.
Buying fits when you already run air-cooled racks with the power, space and people for 600 W cards, and your measured use keeps each card busy past the break-even above.
Reserving fits sustained work on dedicated hardware you do not want to own, and it is the only way to get SXM H200 servers or several nodes on InfiniBand from QuantaCloud. The reserved capacity page explains the process, and GPU clusters covers multi-node builds. If you are planning a multi-year commitment, price the Blackwell generation as well: the B200 and B300 pages cover them as reserved builds. H200 vs B200 compares the two for training and inference.
The H200 decision comes down to the same measured number as every other GPU: busy hours a year. Rent an H200 NVL for a week of real work, count them, and put them into the break-even table. If the numbers point to owning the hardware, send a capacity brief and compare the written quote with the server quote before you sign either.