GPU / NVIDIA GB300 NVL72

NVIDIA GB300 NVL72

Rent 72 Blackwell Ultra GPUs and 36 Grace CPUs in one NVLink domain, built for AI reasoning and trillion-parameter inference.

NVIDIA GB300 NVL72 Pricing & Specs

Together provisions, tunes, and operates the liquid-cooled system with high-performance storage, monitoring, and direct support from the research team. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Performance
Faster inference

10x

vs Hopper GPUs

Faster inference

5x

vs Hopper GPUs

Better efficiency

1.5x

vs GB200 NVL72

Pricing

Rent GB300 NVL72 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.

Plan

Price

On-demand (pay as you go)

Reserved 7-30 days

Contact us

Reserved 31-90 days

Reserved 91-180 days

Reserved 181+ days

Prices as of July 2026.

See pricing for every GPU on the GPU cluster pricing page.

Technical specification

  • Configuration 36 Grace CPU: 72 Blackwell Ultra GPUs
  • FP4 Tensor Core 1,400 PFLOPS (with sparsity) | 1,100 PFLOPS (dense)
  • FP8/FP6 Tensor Core 720 PFLOPS
  • INT8 Tensor Core 23 PFLOPS
  • FP16/BF16 Tensor Core 360 PFLOPS
  • TF32 Tensor Core 180 PFLOPS
  • FP32 6 PFLOPS
  • FP64 / FP64 Tensor Core 100 TFLOPS
  • GPU Memory | Bandwidth Up to 21 TB | Up to 576 TB/s
  • NVLink Bandwidth 130 TB/s
  • CPU Core Count 2,592 Arm® Neoverse V2 cores
  • CPU Memory | Bandwidth Up to 18 TB SOCAMM with LPDDR5X | Up to 14.3 TB/s

Why Rent NVIDIA GB300 NVL72 on Together

The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.

A single 72-GPU NVLink domain for reasoning at scale

GB300 unifies 72 Blackwell Ultra GPUs and 36 Grace CPUs in one platform, optimized for test-time scaling inference.

Purpose-built for AI reasoning throughput

Compared to Hopper, GB300 delivers 10x higher TPS per user and 5x higher TPS per megawatt, combining to 50x higher AI-factory output.

Massive on-rack memory for frontier contexts

Run long-context LLMs and agentic workloads with up to 21 TB of aggregate GPU HBM (up to 576 TB/s bandwidth) and up to 40 TB fast memory.

Run by the same people pushing the Blackwell stack forward

Work with engineers who co-develop Blackwell optimizations; our team continually tunes workloads and publishes cutting-edge training breakthroughs.

What our customers are saying

Man with glasses and beard smiling against a background with glowing circuit-like lines.

"Delivering competitive pricing, strong reliability and a properly set up cluster is the bulk of the value differentiation for most AI clouds. The only differentiated value we have seen outside this set is from a Neocloud called Together AI, where the inventor of FlashAttention, Tri Dao, works. We don't believe the value created by Together can be replicated elsewhere."

Dylan Patel

Founder, SemiAnalysis

Smiling man with short dark hair wearing a blue blazer and white shirt in an outdoor corridor.

"Training our omnimodal Character-3 model required infrastructure designed for large-scale AI. The Together Frontier AI Factory delivered the performance we needed to push the boundaries of multimodal video generation. Together AI understands what builders need — and that made all the difference."

Michael Lingelbach

CEO, Hedra

"Together GPU Clusters provided a combination of amazing training performance, expert support, and the ability to scale to meet our rapid growth to help us serve our growing community of AI creators."

Young woman with long dark hair smiling outdoors wearing a white turtleneck and statement earrings.

Demi Guo

CEO, Pika

“Together AI provides the performance and reliability we need for real-time, high-quality image and video generation at scale. We value that Together AI is much more than an infrastructure provider — they're a true innovation partner, enabling us to push creative boundaries without compromise.”

Young man wearing a cap sprays graffiti on a wall with a spray paint can in black and white.

Victor Perez

Co-Founder, Krea

Rent NVIDIA GB300 NVL72 GPUs on Together GPU Clusters

  • Scale from one node to a full rack

    Rent GB300 NVL72 sized to your run, from a single node to a full 72-GPU rack or multiple racks, all in one NVLink domain.

  • Fully managed, your data stays yours

    Together provisions, tunes, and operates the liquid-cooled system with high-performance storage, monitoring, and direct support from the research team. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Ready to put GB300 NVL72 to work?

See GPU cluster pricing for full details.

Infrastructure you can trust at scale.
Production-grade security.

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Learn More

As an NVIDIA Cloud Partner, Together builds and operates clusters on NVIDIA NCP reference architectures for predictable performance and faster time to production. Your data and models remain under your control with strict privacy safeguards and SOC 2–compliant security practices.

  • NVIDIA logo with text Preferred Partner on a black background.
    preferred partner
  • SOC 2 Type II
  • ISO 27001:2022

Regions and availability zones

Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.

  • USA
    2GW+ in the portfolio with 600MW of near-term capacity in US.
  • Europe
    150 MW+ available in Europe: UK, Spain, France, Portugal, and Iceland also.
  • Asia & Middle East
    Options available based on the scale of the projects in Asia and the Middle East.

FAQ

What is the NVIDIA GB300 NVL72?

The GB300 NVL72 is the Blackwell Ultra system, connecting 72 Blackwell Ultra GPUs and 36 Grace CPUs into one liquid-cooled NVLink domain, with 288GB of HBM3e per GPU and around 1.1 exaFLOPS of FP4 compute, built for AI reasoning and trillion-parameter inference.

How much does it cost to rent a GB300 NVL72?

GB300 NVL72 is offered to rent as reserved capacity, priced per GPU or per rack by quantity, term, and region. Contact sales for a current quote.

Can I rent GB300 GPUs by the hour?

GB300 NVL72 is available to rent as reserved capacity rather than by the hour. You can rent any count, from a single node up to a full rack or multiple racks. Contact sales to get started.

How much memory does the GB300 have?

Each Blackwell Ultra GPU has 288GB of HBM3e. A full NVL72 rack holds up to 21TB of HBM3e, and up to 40TB of fast memory including the Grace CPUs.

What is the GB300 memory bandwidth?

Each Blackwell Ultra GPU provides 8TB/s of memory bandwidth, with 130TB/s of NVLink bandwidth across the rack.

What is the power consumption of the GB300 NVL72?

The GB300 NVL72 is a 100% liquid-cooled system drawing roughly 120kW per rack, handled by Together's data center infrastructure.

Is the GB300 NVL72 available on demand or only reserved?

It is offered to rent as reserved capacity. Contact sales to check availability.

Which regions are GB300 NVL72 clusters available in?

Across the US, Europe, and select Asia and Middle East regions, supporting data-residency and compliance needs.

How is the GB300 different from the GB200?

The GB300 uses Blackwell Ultra GPUs with 288GB of HBM3e (versus 186GB on the GB200) and adds roughly 1.5x the FP4 compute, with a focus on AI reasoning and test-time scaling inference.

How many GPUs are in a GB300 NVL72?

A full NVL72 rack contains 72 Blackwell Ultra GPUs and 36 Grace CPUs.

Browse other NVIDIA GPUs

Self-serve GPUs with transparent billing.

HGX H100
Hardware

NVIDIA H100 (80 GB)

On-demand

$3.99/hr per GPU

Reserved

Starting at $3.09/hr per GPU

I am interested
HGX H200
Hardware

NVIDIA H200 (140GB)

On-demand

$5.99/hr per GPU

Reserved

Starting at $3.99/hr per GPU

I am interested
HGX B200
Hardware

NVIDIA HGX B200 (180GB)

On-demand

$8.19/hr per GPU

Reserved

Starting at $6.79/hr per GPU

I am interested
GB200 NVL72
Hardware

NVIDIA GB200 NVL72 ()

Reserved

Contact us for pricing

I am interested