H200 NVL vs RTX PRO 6000 vs L40S: where each wins
Three PCIe cards, three price bands. A side-by-side of memory, bandwidth, power and cooling, and the workloads where each is the right answer rather than the expensive one.

If you are building or expanding a PCIe-based GPU server in 2026, the shortlist almost always comes down to the same three cards: NVIDIA H200 NVL, RTX PRO 6000 Blackwell Server Edition and L40S. They span roughly a 4× price range, all fit standard dual-slot FHFL bays and all are passively cooled. The differences that matter are memory type, interconnect and what the software stack expects from them.
The numbers side by side
| H200 NVL | RTX PRO 6000 Blackwell SE | L40S | |
|---|---|---|---|
| Architecture | Hopper | Blackwell | Ada Lovelace |
| Memory | 141 GB HBM3e | 96 GB GDDR7 (ECC) | 48 GB GDDR6 (ECC) |
| Memory bandwidth | 4.8 TB/s | ~1.6 TB/s | 864 GB/s |
| Host interface | PCIe 5.0 x16 | PCIe 5.0 x16 | PCIe 4.0 x16 |
| GPU-to-GPU | NVLink bridge, 2- or 4-way, 900 GB/s | PCIe only | PCIe only |
| Lowest precision | FP8 | FP4 (NVFP4) | FP8 |
| TDP | Up to 600 W (configurable) | 400-600 W (configurable) | 350 W |
| Form factor | Dual-slot FHFL | Dual-slot FHFL | Dual-slot FHFL |
| Cooling | Passive | Passive | Passive |
| MIG | Up to 7 instances | Up to 4 instances | No |
| Display / video engines | None | NVENC/NVDEC, no display | NVENC/NVDEC/AV1, no display |
| Typical use | Large-model inference, fine-tuning, HPC | LLM/VLM inference, rendering, mixed AI | Mid-size inference, media, VDI, graphics |
Confirm exact figures against the current datasheet for the SKU you are quoted: vendors ship several TDP and firmware variants of both the H200 NVL and the RTX PRO 6000, and server OEMs may lock a lower power cap.
H200 NVL: capacity and bandwidth without the SXM node
H200 NVL exists for one reason: to bring Hopper HBM3e into ordinary PCIe servers. 141 GB on one card means a 70B model at FP8 fits with a large KV-cache, and a 4-way NVLink bridge turns four cards into a 564 GB pool with 900 GB/s between them, which is what makes it viable for full fine-tuning of mid-size models outside an HGX chassis.
It wins when latency at low batch matters (bandwidth per token), when the model or context does not fit in 96 GB, and when you need MIG to carve one card into isolated slices for several tenants. It loses on cost per token for small models and on FP4 support, which it does not have.
RTX PRO 6000 Blackwell Server Edition: the volume inference card
The Server Edition of RTX PRO 6000 is the card most hosting providers are standardizing on for open-model inference in 2026. 96 GB of GDDR7 with ECC is enough for 70B-class models at FP8 or 120B-class at FP4, Blackwell tensor cores provide native FP4, and the price sits far below HBM parts. Eight cards in a 4U server give a dense, air-cooled 768 GB inference node.
Its limits are the ones you would expect from GDDR: ~1.6 TB/s is a third of the H200 NVL, so single-stream decode is slower, and there is no NVLink, so tensor-parallel setups across cards lean on PCIe 5.0. It also carries media engines and full graphics support, which makes it the better choice for mixed rendering, simulation and AI fleets.
L40S: still the pragmatic middle
L40S launched in 2023 and remains in volume supply. 48 GB and 864 GB/s are enough for 7B-30B models at FP8, embeddings, rerankers, diffusion models, transcoding and VDI. At 350 W it runs in servers that cannot feed 600 W cards, and it is the least demanding on airflow of the three.
It does not do FP4, MIG or NVLink. For a new build the RTX PRO 6000 usually offers a better cost per token, but L40S wins where the model is small, the chassis is fixed or the budget per slot is.
Scenario verdicts
- 70B model, FP8, 20-50 concurrent users
- RTX PRO 6000 (1 card). H200 NVL if p50 latency at batch 1 is the KPI.
- 70B model, BF16 or 32K+ context
- H200 NVL, 2 cards with bridge.
- Multi-tenant GPU slices (MIG)
- H200 NVL (7 slices) or RTX PRO 6000 (4 slices).
- Fine-tuning up to 70B with LoRA
- RTX PRO 6000 or H200 NVL; both work, H200 NVL is faster per step.
- Full fine-tuning 13B-30B
- 4× H200 NVL with NVLink bridge.
- Embeddings, rerankers, sub-30B chat
- L40S, or RTX PRO 6000 if FP4 headroom is wanted.
- Rendering + AI on one fleet
- RTX PRO 6000 Server Edition.
- Media pipelines, VDI, 350 W servers
- L40S.
Server compatibility notes
- All three are passive: the chassis must provide front-to-back airflow rated for the card TDP. Workstation-style cases will not work.
- PCIe 5.0 hosts (EPYC 9004/9005, Xeon 5th/6th gen) are needed to get full bandwidth from H200 NVL and RTX PRO 6000; L40S is fine on PCIe 4.0.
- NVLink bridges for H200 NVL are a separate part number and require adjacent slots with the correct spacing; check the server qualified-configuration list.
- Plan power at the configured TDP plus 10-15%: eight 600 W cards need a 4U server with at least 2× 3000 W PSUs and 30+ kW rack capacity.
Nodeforge keeps all three cards in Dubai and Hong Kong stock in most weeks, alongside qualified Supermicro, Dell and HPE servers to host them. Send us the model list and the number of users, and we will return a BOM within one business day.
Next step
Need help choosing the hardware?
Send the model and the workload. An engineer replies with a BOM, availability in Dubai and Hong Kong, and lead times.
Related


