Location: USA Shipping: Worldwide Shipping Call +1 872 293 4196 WhatsApp supply@allchipsupply.com

Data Center GPU Buying Guide: H100 vs H200 vs B200 vs L40S

By All Chip Supply · Published

Choosing a data center GPU is mostly a question of matching the card to the workload and to the servers you already run. The NVIDIA H100, H200, B200 and L40S are often compared side by side, but they are built for different jobs and different systems. This guide compares them using the figures NVIDIA publishes on its product pages and datasheets, and lists the questions worth answering before you request a quote.

All Tensor Core figures below are NVIDIA's numbers with sparsity unless marked otherwise. Dense throughput is half the sparse figure. Figures differ between SXM modules, PCIe cards and system configurations.

The short version

Side-by-side specifications

NVIDIA L40S (Ada Lovelace, PCIe Gen4 x16, dual slot, passive cooling)

NVIDIA H100 SXM (Hopper, SXM module)

NVIDIA H100 NVL (Hopper, PCIe dual-slot, air-cooled)

NVIDIA H200 SXM (Hopper, SXM module)

NVIDIA H200 NVL (Hopper, PCIe dual-slot, air-cooled)

NVIDIA B200 (Blackwell, SXM module in HGX B200) NVIDIA publishes B200 figures for the 8-GPU HGX B200 board. Divided by eight, that is roughly:

What the numbers mean in practice

Memory capacity decides what fits. A model's weights, the KV cache for inference and the activations for training all need GPU memory. The jump from 80GB (H100 SXM) to 141GB (H200) to around 180GB (B200) can decide whether a model runs on one GPU or must be split across several. The L40S, with 48GB, suits smaller models, fine-tuning and inference at modest batch sizes.

Bandwidth often matters as much as FLOPS. Large language model inference is frequently limited by how fast weights can be read from memory. The H200 has the same compute figures as the H100 SXM but 4.8TB/s of bandwidth against 3.35TB/s, which is why it is often chosen as an inference upgrade within Hopper systems.

Precision support. All four support FP8 Tensor Core math. B200 adds FP4 and FP6. Using these formats depends on your software stack, so check framework and library support before planning around them.

Double precision for HPC. H100 and H200 SXM list 67 TFLOPS of FP64 Tensor Core throughput. The L40S is aimed at AI and graphics and NVIDIA does not list FP64 Tensor Core figures for it, so it is not the right choice for FP64-heavy simulation.

Form factor, power and the server you have

This is often the deciding factor. SXM modules (H100 SXM, H200 SXM, B200) are not standalone cards: they go into HGX baseboards in servers designed for them, with matching power delivery and cooling. They are usually bought as part of complete systems.

PCIe cards (L40S, H100 NVL, H200 NVL, H100 PCIe) fit qualified PCIe servers, but you still need to check:

Questions to answer before you request a quote

  1. Which workloads will run: training, inference, rendering, HPC, or a mix?
  2. What is the largest model or dataset that must fit in memory on one GPU?
  3. Are you buying modules for an HGX system, cards for existing PCIe servers, or complete servers?
  4. How many GPUs, and do they need NVLink between them?
  5. What power and cooling are available per rack?
  6. What delivery timeline and documentation do you need?

Next steps

You can compare listings for the NVIDIA L40S 48GB, H100 NVL 94GB, H100 SXM5 80GB, H100 80GB PCIe, H200 NVL 141GB, H200 SXM 141GB and B200 180GB SXM in our AI and data center category. Availability and lead times change quickly for these parts, so send the details above through a chip request and we will reply with current pricing and timing.

Specifications are taken from NVIDIA's published product pages and datasheets and may change. Always confirm against the latest NVIDIA documentation for your exact part number.

Chat with Us on WhatsApp