H100 vs H200 vs B200: NVIDIA GPU Buyer’s Guide
The short answer: choose the H100 when 80GB of memory per GPU is enough and you want the best price on proven Hopper silicon; step up to the H200 when memory capacity and bandwidth are the bottleneck for inference or long-context models; and move to the B200 when you need the highest per-GPU throughput and are building new, liquid-cooled capacity where the added power and cost are justified. All three are current, buyable parts as of 2026-08, and all three sell through OEM and integrator channels rather than at a fixed price.
This guide compares NVIDIA’s H100, H200 and B200 on the specifications that actually drive procurement decisions: HBM capacity and bandwidth, thermal design power, form factor and interconnect, and indicative market pricing. The goal is to match each accelerator to a workload profile so your AI GPU and accelerator budget lands where it earns the most. Figures below are indicative ranges, not quotes, and should be confirmed by RFQ because HBM supply constraints keep contract prices moving.
Memory and bandwidth: the real dividing line
Memory is the single clearest difference across this lineup. The H100 SXM5 carries 80GB of HBM3 at 3.35 TB/s. The H200 keeps the same Hopper compute but upgrades to 141GB of HBM3e at 4.8 TB/s, a large jump that directly benefits memory-bound inference, larger batch sizes and longer context windows. The B200, built on the newer Blackwell architecture, pushes to roughly 180–192GB of HBM3e at about 8 TB/s, close to double the H200’s bandwidth.
For training and inference of large language models, capacity determines how much of a model fits per GPU and how few GPUs you need for a given context length, while bandwidth governs how fast tokens move. If your models spill out of 80GB or you are paying for extra GPUs purely to hold weights and KV cache, the H200 and B200 change the math.
Power, form factor and interconnect
Higher performance comes with higher power. The H100 and H200 SXM parts both sit at 700W, so an existing HGX H100 platform can often host H200 modules within a similar thermal envelope. The B200 raises TDP to 1000W per GPU, which is why dense 8-GPU B200 systems are generally direct-liquid-cooled and demand facility planning before purchase.
Interconnect scales in step. H100 and H200 use NVLink at 900 GB/s per GPU, while the B200 doubles NVLink to 1.8 TB/s, improving all-reduce performance across large training jobs. An H100 also ships in a lower-power 350W PCIe variant with reduced bandwidth (roughly 2.0 TB/s on HBM2e), useful for mixed or air-cooled servers where SXM density is not required.
Specification comparison
Figures are representative SXM parts as of 2026-08. Prices are indicative per-GPU market ranges, confirmed by RFQ.
| Spec | H100 SXM5 | H200 SXM | B200 SXM |
|---|---|---|---|
| Architecture | Hopper | Hopper | Blackwell |
| HBM capacity | 80GB HBM3 | 141GB HBM3e | 180–192GB HBM3e |
| Memory bandwidth | 3.35 TB/s | 4.8 TB/s | ~8 TB/s |
| TDP | 700W | 700W | 1000W |
| Form factor | SXM5 | SXM | SXM |
| NVLink | 900 GB/s | 900 GB/s | 1.8 TB/s |
| Cooling | Air / liquid | Air / liquid | Typically liquid |
| Indicative price | ~$35,000–40,000 | ~$32,000–50,000 | ~$40,000–80,000 |
Matching GPU to workload
For inference and fine-tuning on models that fit in 80GB, the H100 remains the value pick; used and refurbished units widen the discount further. For memory-heavy inference, retrieval-augmented pipelines and long-context serving, the H200’s 141GB and higher bandwidth reduce GPU count and often lower total cost per served token. For frontier-scale training and the densest new build-outs, the B200 delivers the highest throughput per GPU and the fastest interconnect, provided you can supply the power and liquid cooling.
At the node level, an 8-GPU HGX H100 system runs roughly $310,000–370,000, an HGX H200 node roughly $370,000–430,000, and an HGX B200 system roughly $430,000–520,000 as of 2026-08. Complete nodes are typically quoted as bundles including CPU, memory, storage, networking and cooling, so per-GPU math only goes so far. Our accelerator desk can model both per-GPU and full-node options against live availability.
FAQ
Is the H200 just an H100 with more memory? Largely, yes. The H200 keeps the Hopper compute engine of the H100 but swaps in 141GB of HBM3e at 4.8 TB/s, so it mainly helps memory-bound inference and large-context work rather than raw compute.
Do I need liquid cooling for the B200? At 1000W per SXM GPU, dense 8-GPU B200 nodes are typically direct-liquid-cooled. Air-cooled configurations exist but are rarer and lower density. Plan facility cooling before ordering.
Should I still buy H100 in 2026? For many inference and fine-tuning jobs, yes. H100 supply is loosening as Blackwell ships, and used units trade well below new, making it cost-effective where 80GB per GPU is enough.
Why is there no list price for these GPUs? NVIDIA sets no MSRP for datacenter GPUs. They sell through OEMs, ODMs and integrators under negotiated, volume-dependent terms, so prices are ranges confirmed by RFQ.
Can I mix H100 and H200 in one cluster? You can run them in the same fabric, but a single training job generally wants matched GPUs. Mixed fleets are common when H200 nodes are added to expand an existing H100 estate.
Bottom line
The H100, H200 and B200 form a clear ladder: proven value, more memory, or maximum throughput. Pick the rung that clears your workload’s memory and power constraints without overspending on capability you will not use. Browse the full AI GPUs and accelerators category for current options, and request a quote with your model sizes and node count for an indicative price against live stock.