GPU / NVIDIA GB200 NVL72

NVIDIA GB200 NVL72

Rent 72 Blackwell GPUs and 36 Grace CPUs in one NVLink domain, built for trillion-parameter training and inference.

NVIDIA GB200 NVL72 Pricing & Specs

Every cluster is dedicated bare metal with liquid cooling, orchestrated with Slurm or Kubernetes, and tuned by the research team that built FlashAttention. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Performance
FASTER TRAINING

4x

vs H100

FASTER INFERENCE

30x

vs H100

ENERGY EFFICIENCY

25x

vs H100

Pricing

Rent GB200 NVL72 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.

Plan

Price

On-demand (pay as you go)

Reserved 7-30 days

Contact us

Reserved 31-90 days

Reserved 91-180 days

Reserved 181+ days

Prices as of July 2026.

See pricing for every GPU on the GPU cluster pricing page.

Technical specification

  • Model data NVIDIA GB200 NVL72
  • Architecture NVIDIA Grace Blackwell
  • Configuration 36 Grace CPUs : 72 Blackwell GPUs
  • GPU memory (VRAM) 186GB HBM3e per GPU
  • Total GPU memory Up to 13.4TB HBM3e
  • Total fast memory Up to 30TB (with Grace LPDDR5X)
  • Total memory bandwidth Up to 576TB/s
  • FP4 Tensor Core (total) 1,440 PFLOPS (1.4 exaFLOPS)
  • FP8 / FP6 Tensor Core (total) 720 PFLOPS
  • INT8 Tensor Core (total) 720 POPS
  • FP16 / BF16 Tensor Core (total) 360 PFLOPS
  • TF32 Tensor Core (total) 180 PFLOPS
  • FP32 (total) 5,760 TFLOPS
  • FP64 (total) 2,880 TFLOPS
  • NVLink bandwidth 130TB/s total, 1.8TB/s per GPU (fifth-generation)
  • CPU core count 2,592 Arm Neoverse V2 cores
  • CPU memory Up to 17TB LPDDR5X, up to 18.4TB/s
  • Cooling Liquid-cooled
  • Form factor NVL72 rack (72 GPUs, 36 Grace CPUs)

Why Rent NVIDIA GB200 NVL72 on Together

The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.

Train trillion-parameter models

Combine 72 Blackwell GPUs and 36 Grace CPUs into one liquid-cooled, memory-coherent rack — enabling tightly synchronized, low-latency training.

Custom networking over 130TB/s

Close topologies for dense LLMs & oversubscription for MoEs, built on InfiniBand or high-speed Ethernet.

AI-native shared storage

VAST and Weka for high-throughput, parallel access to massive datasets and model state.

Expert support

Engineers co-develop Blackwell optimizations; continually tune workloads and publish breakthroughs.

What our customers are saying

Man with glasses and beard smiling against a background with glowing circuit-like lines.

"Delivering competitive pricing, strong reliability and a properly set up cluster is the bulk of the value differentiation for most AI clouds. The only differentiated value we have seen outside this set is from a Neocloud called Together AI, where the inventor of FlashAttention, Tri Dao, works. We don't believe the value created by Together can be replicated elsewhere."

Dylan Patel

Founder, SemiAnalysis

Smiling man with short dark hair wearing a blue blazer and white shirt in an outdoor corridor.

"Training our omnimodal Character-3 model required infrastructure designed for large-scale AI. The Together Frontier AI Factory delivered the performance we needed to push the boundaries of multimodal video generation. Together AI understands what builders need — and that made all the difference."

Michael Lingelbach

CEO, Hedra

"Together GPU Clusters provided a combination of amazing training performance, expert support, and the ability to scale to meet our rapid growth to help us serve our growing community of AI creators."

Young woman with long dark hair smiling outdoors wearing a white turtleneck and statement earrings.

Demi Guo

CEO, Pika

“Together AI provides the performance and reliability we need for real-time, high-quality image and video generation at scale. We value that Together AI is much more than an infrastructure provider — they're a true innovation partner, enabling us to push creative boundaries without compromise.”

Young man wearing a cap sprays graffiti on a wall with a spray paint can in black and white.

Victor Perez

Co-Founder, Krea

Rent NVIDIA GB200 GPUs on Together GPU Clusters

  • Scale from one node to a full rack

    Rent GB200 sized to your run, from a single node to a full 72-GPU rack or multiple racks, all in one NVLink domain, so the largest models train and infer without leaving the rack.

  • Fully managed, your data stays yours

    Together provisions, tunes, and operates the liquid-cooled system with high-performance storage, monitoring, and direct support from the research team. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Ready to put GB200 NVL72 to work?

See GPU cluster pricing for full details.

Infrastructure you can trust at scale.
Production-grade security.

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Learn More

As an NVIDIA Cloud Partner, Together builds and operates clusters on NVIDIA NCP reference architectures for predictable performance and faster time to production. Your data and models remain under your control with strict privacy safeguards and SOC 2–compliant security practices.

  • NVIDIA logo with text Preferred Partner on a black background.
    preferred partner
  • SOC 2 Type II
  • ISO 27001:2022

Regions and availability zones

Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.

  • USA
    2GW+ in the portfolio with 600MW of near-term capacity in US.
  • Europe
    150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
  • Asia & Middle East
    Options available based on the scale of the projects in Asia and the Middle East.

FAQ

What is the NVIDIA GB200 NVL72?

The GB200 NVL72 is a liquid-cooled system that connects 72 NVIDIA Blackwell GPUs and 36 Grace CPUs into a single NVLink domain, delivering up to 1.4 exaFLOPS of FP4 compute for trillion-parameter AI.

How much does it cost to rent a GB200 NVL72?

GB200 NVL72 is offered to rent as reserved capacity, priced per GPU or per rack by quantity, term, and region. Contact sales for a current quote.

Can I rent GB200 GPUs by the hour?

GB200 NVL72 is available to rent as reserved capacity rather than by the hour. You can rent any count, from a single node up to a full rack or multiple racks. Contact sales to get started.

How much memory does the GB200 have?

Each Blackwell GPU in the GB200 has 186GB of HBM3e. A full NVL72 rack holds 13.4TB of HBM3e, and up to 30TB of fast memory including the Grace CPUs.

What is the GB200 memory bandwidth?

A full GB200 NVL72 rack provides up to 576TB/s of aggregate memory bandwidth.

What is the power consumption of the GB200 NVL72?

The GB200 NVL72 is a liquid-cooled system; full-rack power is on the order of tens of kilowatts and is handled by Together's data center infrastructure.

Is the GB200 NVL72 available on demand or only reserved?

It is offered to rent as reserved capacity. Contact sales to check availability.

Which regions are GB200 NVL72 clusters available in?

Across the US, Europe, and select Asia and Middle East regions, supporting data-residency and compliance needs.

How many GPUs are in a GB200 NVL72?

A full NVL72 rack contains 72 Blackwell GPUs and 36 Grace CPUs.

What CPU does the GB200 use?

The GB200 pairs Blackwell GPUs with the NVIDIA Grace CPU, an Arm Neoverse V2 design, over a coherent NVLink-C2C link.

Browse all NVIDIA GPUs

Self-serve GPUs with transparent billing.

H100
Hardware

NVIDIA H100 (80 GB)

On-demand

$3.99/hr per GPU

Reserved

Starting at $3.19/hr per GPU

Scale

8 to 256 GPUs

Create cluster
H200
Hardware

NVIDIA H200 (140 GB)

On-demand

$5.99/hr per GPU

Reserved

Starting at $3.99/hr per GPU

Scale

256 to 1,000 GPUs

Create cluster
HGX B200
Hardware

NVIDIA HGX B200 (180 GB)

On-demand

$8.19/hr per GPU

Reserved

Starting at $6.79/hr per GPU

Scale

256 to 1,000+ GPUs

Create cluster
HGX B300
Hardware

NVIDIA HGX B300 (270 GB)

Reserved

Contact us for pricing

Contact sales
GB200 NVL72
Hardware

NVIDIA GB200 NVL72 (186 GB)

Reserved

Contact us for pricing

Scale

512 to 1,000+ GPUs

Contact sales
GB300 NVL72
Hardware

NVIDIA GB300 NVL72 (288 GB)

Reserved

Contact us for pricing

Contact sales