NVIDIA HGX B200
Rent NVIDIA Blackwell GPUs on demand or reserve a cluster, with 180GB HBM3e and native FP4 for frontier-scale training and inference.
NVIDIA HGX B200 Pricing & Specs
Every cluster is dedicated bare metal, orchestrated with Slurm or Kubernetes, and tuned by the research team that built FlashAttention. Your data and model weights stay yours, backed by SOC 2 Type II controls.
3x
vs H100
15x
vs H100
12x
vs H100
Pricing
Rent HGX B200 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.
Plan | Price |
|---|---|
On-demand (pay as you go) | $8.19 |
Reserved 7-30 days | $7.99 |
Reserved 31-90 days | $7.79 |
Reserved 91-180 days | $6.79 |
Reserved 181+ days | Contact us |
Prices as of July 2026.
See pricing for every GPU on the GPU cluster pricing page.
Technical specification
- Model data NVIDIA B200 (HGX)
- Architecture NVIDIA Blackwell (dual-die, 208B transistors)
- GPU memory (VRAM) 180GB HBM3e
- Memory bandwidth 7.7TB/s
- FP4 Tensor Core (per GPU) 18 PFLOPS
- FP8 / FP6 Tensor Core (per GPU) 9 PFLOPS
- INT8 Tensor Core 9 POPS
- FP16 / BF16 Tensor Core 4.5 PFLOPS
- TF32 Tensor Core 2.2 PFLOPS
- FP32 75 TFLOPS
- FP64 / FP64 Tensor Core 37 TFLOPS
- Total fast memory (node) Up to 1.4TB
- Total memory bandwidth (node) Up to 62TB/s
- Total NVLink bandwidth (node) 14.4TB/s
- NVLink bandwidth (per GPU) 1.8TB/s (fifth-generation)
- Multi-Instance GPU (MIG) Up to 7
- Decoders 7 NVDEC, 7 nvJPEG
- Max thermal design power (TDP) Up to 1,000W
- Interconnect 5th Gen NVLink 1.8TB/s, PCIe Gen5 128GB/s
- Form factor HGX B200, 8 GPUs
Why Rent NVIDIA HGX B200 on Together
The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.
Train faster on 8-GPU HGX nodes
Each system includes eight Blackwell GPUs with second-gen Transformer Engine and FP8 precision, optimized for maximum throughput.
Optimized NVLink topologies
Configured 1.8TB/s NVSwitch fabrics per node, extending to spine-leaf or Clos topologies for dense LLMs or sparsely activated MoE workloads.
Delivery in 4–6 weeks, no NVIDIA lottery required
Full-rack clusters ship with thousands of GPUs available; no backorder delays.
Run by researchers who train models
Team actively tunes training workloads on NVIDIA GB200 systems.
What our customers are saying
Rent NVIDIA B200 GPUs on Together GPU Clusters
Spin up a dedicated B200 cluster with real-time availability, from a single 8-GPU node to multi-node deployments over non-blocking InfiniBand and fifth-generation NVLink.
Every cluster is dedicated bare metal with liquid cooling, orchestrated with Slurm or Kubernetes, and tuned by the research team that built FlashAttention. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Ready to put HGX B200 to work?
See GPU cluster pricing for full details.
Infrastructure you can trust at scale.
Production-grade security.
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022
Regions and availability zones
Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.
- USA2GW+ in the portfolio with 600MW of near-term capacity in US.
- Europe150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
- Asia & Middle EastOptions available based on the scale of the projects in Asia and the Middle East.
FAQ
What is the NVIDIA B200?
The NVIDIA B200 is the flagship GPU of the Blackwell architecture, with 180GB of HBM3e memory, 7.7TB/s of bandwidth, and fifth-generation Tensor Cores with native FP4, built for frontier-scale AI training and inference.
How much does it cost to rent an NVIDIA B200?
On Together GPU Clusters, B200 GPUs are $8.19 per GPU per hour on demand, with reserved rates from $6.79 per GPU per hour for longer commitments. See the pricing section above or contact sales for volume pricing.
Can I rent B200 GPUs by the hour?
Yes. Launch an on-demand B200 cluster and pay per GPU, or reserve dedicated capacity for a defined training window.
How much memory does the B200 have?
Each B200 has 180GB of HBM3e memory.
What is the B200 memory bandwidth?
The B200 delivers 7.7TB/s of memory bandwidth.
What is the power consumption of the B200?
The B200 has a thermal design power of up to 1,000W and is liquid-cooled in most deployments.
Is the B200 available on demand or only reserved?
Both. Together offers on-demand B200 capacity with real-time availability and reserved clusters for longer terms.
Which regions are B200 clusters available in?
Across the US, Europe (UK, Spain, France, Portugal, Iceland), and select Asia and Middle East regions, supporting data-residency and compliance needs.
What is the NVIDIA B200 used for?
The B200 is used for frontier-scale model training, high-throughput and FP4 inference for 70B to 100B+ parameter models, and memory-bound workloads that exceed a single Hopper GPU.
Does the B200 support FP4?
Yes. The B200's fifth-generation Tensor Cores and second-generation Transformer Engine add native FP4 precision, roughly doubling inference throughput over FP8.
What is the B200 SXM?
B200 SXM is the socketed module form factor used in HGX B200 and DGX B200 systems, connected over fifth-generation NVLink for multi-GPU scaling.
Browse all NVIDIA GPUs
Self-serve GPUs with transparent billing.





