NVIDIA GB200 NVL72
Rent 72 Blackwell GPUs and 36 Grace CPUs in one NVLink domain, built for trillion-parameter training and inference.
NVIDIA GB200 NVL72 Pricing & Specs
Every cluster is dedicated bare metal with liquid cooling, orchestrated with Slurm or Kubernetes, and tuned by the research team that built FlashAttention. Your data and model weights stay yours, backed by SOC 2 Type II controls.
4x
vs H100
30x
vs H100
25x
vs H100
Pricing
Rent GB200 NVL72 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.
Plan | Price |
|---|---|
On-demand (pay as you go) | — |
Reserved 7-30 days | Contact us |
Reserved 31-90 days | |
Reserved 91-180 days | |
Reserved 181+ days |
Prices as of July 2026.
See pricing for every GPU on the GPU cluster pricing page.
Technical specification
- Model data NVIDIA GB200 NVL72
- Architecture NVIDIA Grace Blackwell
- Configuration 36 Grace CPUs : 72 Blackwell GPUs
- GPU memory (VRAM) 186GB HBM3e per GPU
- Total GPU memory Up to 13.4TB HBM3e
- Total fast memory Up to 30TB (with Grace LPDDR5X)
- Total memory bandwidth Up to 576TB/s
- FP4 Tensor Core (total) 1,440 PFLOPS (1.4 exaFLOPS)
- FP8 / FP6 Tensor Core (total) 720 PFLOPS
- INT8 Tensor Core (total) 720 POPS
- FP16 / BF16 Tensor Core (total) 360 PFLOPS
- TF32 Tensor Core (total) 180 PFLOPS
- FP32 (total) 5,760 TFLOPS
- FP64 (total) 2,880 TFLOPS
- NVLink bandwidth 130TB/s total, 1.8TB/s per GPU (fifth-generation)
- CPU core count 2,592 Arm Neoverse V2 cores
- CPU memory Up to 17TB LPDDR5X, up to 18.4TB/s
- Cooling Liquid-cooled
- Form factor NVL72 rack (72 GPUs, 36 Grace CPUs)
Why Rent NVIDIA GB200 NVL72 on Together
The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.
Train trillion-parameter models
Combine 72 Blackwell GPUs and 36 Grace CPUs into one liquid-cooled, memory-coherent rack — enabling tightly synchronized, low-latency training.
Custom networking over 130TB/s
Close topologies for dense LLMs & oversubscription for MoEs, built on InfiniBand or high-speed Ethernet.
AI-native shared storage
VAST and Weka for high-throughput, parallel access to massive datasets and model state.
Expert support
Engineers co-develop Blackwell optimizations; continually tune workloads and publish breakthroughs.
What our customers are saying
Rent NVIDIA GB200 GPUs on Together GPU Clusters
Rent GB200 sized to your run, from a single node to a full 72-GPU rack or multiple racks, all in one NVLink domain, so the largest models train and infer without leaving the rack.
Together provisions, tunes, and operates the liquid-cooled system with high-performance storage, monitoring, and direct support from the research team. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Ready to put GB200 NVL72 to work?
See GPU cluster pricing for full details.
Infrastructure you can trust at scale.
Production-grade security.
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022
Regions and availability zones
Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.
- USA2GW+ in the portfolio with 600MW of near-term capacity in US.
- Europe150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
- Asia & Middle EastOptions available based on the scale of the projects in Asia and the Middle East.
FAQ
What is the NVIDIA GB200 NVL72?
The GB200 NVL72 is a liquid-cooled system that connects 72 NVIDIA Blackwell GPUs and 36 Grace CPUs into a single NVLink domain, delivering up to 1.4 exaFLOPS of FP4 compute for trillion-parameter AI.
How much does it cost to rent a GB200 NVL72?
GB200 NVL72 is offered to rent as reserved capacity, priced per GPU or per rack by quantity, term, and region. Contact sales for a current quote.
Can I rent GB200 GPUs by the hour?
GB200 NVL72 is available to rent as reserved capacity rather than by the hour. You can rent any count, from a single node up to a full rack or multiple racks. Contact sales to get started.
How much memory does the GB200 have?
Each Blackwell GPU in the GB200 has 186GB of HBM3e. A full NVL72 rack holds 13.4TB of HBM3e, and up to 30TB of fast memory including the Grace CPUs.
What is the GB200 memory bandwidth?
A full GB200 NVL72 rack provides up to 576TB/s of aggregate memory bandwidth.
What is the power consumption of the GB200 NVL72?
The GB200 NVL72 is a liquid-cooled system; full-rack power is on the order of tens of kilowatts and is handled by Together's data center infrastructure.
Is the GB200 NVL72 available on demand or only reserved?
It is offered to rent as reserved capacity. Contact sales to check availability.
Which regions are GB200 NVL72 clusters available in?
Across the US, Europe, and select Asia and Middle East regions, supporting data-residency and compliance needs.
How many GPUs are in a GB200 NVL72?
A full NVL72 rack contains 72 Blackwell GPUs and 36 Grace CPUs.
What CPU does the GB200 use?
The GB200 pairs Blackwell GPUs with the NVIDIA Grace CPU, an Arm Neoverse V2 design, over a coherent NVLink-C2C link.
Browse all NVIDIA GPUs
Self-serve GPUs with transparent billing.





