GPU / NVIDIA H200

NVIDIA H200

Rent NVIDIA H200 GPUs on demand or reserve a cluster, with 141GB HBM3e and 4.8TB/s for memory-bound LLM inference and training.

NVIDIA H200 Pricing & Specs

The NVIDIA H200 is the memory-upgraded Hopper GPU, the first with HBM3e: 141GB at 4.8TB/s on the same compute as the H100. That capacity suits memory-bound work like long-context inference and large-batch serving. Rent it on Together on demand or as a dedicated cluster.

Performance
Faster training

2x

vs. H100

Faster inference

110x

higher

Better efficiency

1.4x

vs. H100

Pricing

Rent HGX H200 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.

Plan

Price

On-demand (pay as you go)

$5.99

Reserved 7-30 days

$4.99

Reserved 31-90 days

$4.15

Reserved 91-180 days

$3.99

Reserved 181+ days

Contact us

Prices as of July 2026.

See pricing for every GPU on the GPU cluster pricing page.

Technical specification

  • Model data NVIDIA H200 (SXM)
  • Architecture NVIDIA Hopper
  • GPU memory (VRAM) 141GB HBM3e
  • Memory bandwidth 4.8TB/s
  • FP64 34 TFLOPS
  • FP64 Tensor Core 67 TFLOPS
  • FP32 67 TFLOPS
  • TF32 Tensor Core 989 TFLOPS
  • FP16 / BF16 Tensor Core 1,979 TFLOPS
  • FP8 Tensor Core 3,958 TFLOPS
  • INT8 Tensor Core 3,958 TOPS
  • NVLink bandwidth 900GB/s
  • Interconnect NVLink 900GB/s, PCIe Gen5 128GB/s
  • Multi-Instance GPU (MIG) Up to 7 instances
  • Max thermal design power (TDP) Up to 700W (configurable)
  • Form factor HGX H200, 4 or 8 GPUs

Why Rent NVIDIA H200 on Together

The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.

2x performance over H100

Each H200 GPU cluster offers double the inference throughput compared to H100, ideal for deploying LLMs at unprecedented scale.

Enhanced memory bandwidth

With 141GB HBM3e GPU memory and 4.8TB/s bandwidth, H200 significantly accelerates memory-intensive generative AI workloads and HPC applications.

Maximum efficiency and TCO savings

Achieve higher performance within the same power profile as previous-gen GPUs, drastically reducing energy consumption and total cost of ownership.

Run by researchers who train models

Our research team actively runs and tunes training workloads on NVIDIA H200 systems for edge-of-possibility expertise.

What our customers are saying

Man with glasses and beard smiling against a background with glowing circuit-like lines.

"Delivering competitive pricing, strong reliability and a properly set up cluster is the bulk of the value differentiation for most AI clouds. The only differentiated value we have seen outside this set is from a Neocloud called Together AI, where the inventor of FlashAttention, Tri Dao, works. We don't believe the value created by Together can be replicated elsewhere."

Dylan Patel

Founder, SemiAnalysis

Smiling man with short dark hair wearing a blue blazer and white shirt in an outdoor corridor.

"Training our omnimodal Character-3 model required infrastructure designed for large-scale AI. The Together Frontier AI Factory delivered the performance we needed to push the boundaries of multimodal video generation. Together AI understands what builders need — and that made all the difference."

Michael Lingelbach

CEO, Hedra

"Together GPU Clusters provided a combination of amazing training performance, expert support, and the ability to scale to meet our rapid growth to help us serve our growing community of AI creators."

Young woman with long dark hair smiling outdoors wearing a white turtleneck and statement earrings.

Demi Guo

CEO, Pika

“Together AI provides the performance and reliability we need for real-time, high-quality image and video generation at scale. We value that Together AI is much more than an infrastructure provider — they're a true innovation partner, enabling us to push creative boundaries without compromise.”

Young man wearing a cap sprays graffiti on a wall with a spray paint can in black and white.

Victor Perez

Co-Founder, Krea

Rent NVIDIA H200 GPUs on Together GPU Clusters

  • Scale from one node to thousands

    Spin up a dedicated H200 cluster in minutes with real-time availability, from a single 8-GPU node to thousands of GPUs over non-blocking InfiniBand and NVLink.

  • Fully managed, your data stays yours

    Every cluster is dedicated bare metal, orchestrated with Slurm or Kubernetes, and tuned by the research team that built FlashAttention. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Ready to put HGX H200 to work?

See GPU cluster pricing for full details.

Infrastructure you can trust at scale.
Production-grade security.

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Learn More

As an NVIDIA Cloud Partner, Together builds and operates clusters on NVIDIA NCP reference architectures for predictable performance and faster time to production. Your data and models remain under your control with strict privacy safeguards and SOC 2–compliant security practices.

  • NVIDIA logo with text Preferred Partner on a black background.
    preferred partner
  • SOC 2 Type II
  • ISO 27001:2022

Regions and availability zones

Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.

  • USA
    2GW+ in the portfolio with 600MW of near-term capacity in US.
  • Europe
    150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
  • Asia & Middle East
    Options available based on the scale of the projects in Asia and the Middle East.

FAQ

What is the NVIDIA H200?

The NVIDIA H200 is a Hopper-architecture data center GPU with 141GB of HBM3e memory and 4.8TB/s of bandwidth, the first GPU with HBM3e, built for memory-bound LLM inference and large-scale training.

How much does it cost to rent an NVIDIA H200?

On Together GPU Clusters, H200 GPUs are $5.99 per GPU per hour on demand, with reserved rates from $3.99 per GPU per hour for longer commitments. See the pricing section above or contact sales for volume pricing.

Can I rent H200 GPUs by the hour?

Yes. Launch an on-demand H200 cluster in minutes and pay per GPU, or reserve dedicated capacity for a defined training window.

How much memory does the H200 have?

Each H200 has 141GB of HBM3e memory, up from 80GB on the H100.

What is the H200 memory bandwidth?

The H200 delivers 4.8TB/s of memory bandwidth, roughly 1.4x the H100.

What is the power consumption of the H200?

The H200 SXM has a configurable thermal design power of up to 700W.

Is the H200 available on demand or only reserved?

Both. Together offers on-demand H200 capacity with real-time availability and reserved clusters for longer terms.

Which regions are H200 clusters available in?

Across the US, Europe (UK, Spain, France, Portugal, Iceland), and select Asia and Middle East regions, supporting data-residency and compliance needs.

What is the NVIDIA H200 used for?

The H200 is used for long-context and large-batch LLM inference, serving models that exceed 80GB at full precision, and large-scale training where memory capacity and bandwidth are the binding constraint.

Is the H200 a Blackwell GPU?

No. The H200 is a Hopper-architecture GPU. Blackwell is the next generation, which includes the B200 and the GB200 and GB300 NVL72 systems.

Browse other NVIDIA GPUs

Self-serve GPUs with transparent billing.

HGX H100
Hardware

NVIDIA H100 (80 GB)

On-demand

$3.99/hr per GPU

Reserved

Starting at $3.19/hr per GPU

Scale

8 to 256 GPUs

I am interested
HGX B200
Hardware

NVIDIA HGX B200 (180GB)

On-demand

$8.19/hr per GPU

Reserved

Starting at $6.79/hr per GPU

Scale

256 to 1,000+ GPUs

I am interested
GB300 NVL72
Hardware

NVIDIA GB300 NVL72 ()

Reserved

Contact us for pricing

I am interested
GB200 NVL72
Hardware

NVIDIA GB200 NVL72 ()

Reserved

Contact us for pricing

Scale

512 to 1,000+ GPUs

I am interested