NVIDIA H100
Rent NVIDIA H100 GPUs on demand or reserve a cluster, with 80GB HBM3 and fourth-generation Tensor Cores for large-scale training and inference.
NVIDIA H100 Pricing & Specs
The NVIDIA H100 is NVIDIA's Hopper-generation GPU for large-scale AI, with 80GB of HBM3 and an FP8 Transformer Engine. It's the workhorse for training and serving large language models. Rent it on Together on demand or as a dedicated cluster.
4x
vs A100
30x
on Megatron 530B
7x
higher
Pricing
Rent HGX H100 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.
Plan | Price |
|---|---|
On-demand (pay as you go) | $3.99 |
Reserved 7-30 days | $3.69 |
Reserved 31-90 days | $3.45 |
Reserved 91-180 days | $3.19 |
Reserved 181+ days | Contact us |
Prices as of July 2026.
See pricing for every GPU on the GPU cluster pricing page.
Technical specification
- Model data NVIDIA H100 (SXM)
- Architecture NVIDIA Hopper
- GPU memory (VRAM) 80GB HBM3
- Memory bandwidth 3.35TB/s
- FP64 34 TFLOPS
- FP64 Tensor Core 67 TFLOPS
- FP32 67 TFLOPS
- TF32 Tensor Core 989 TFLOPS
- FP16 / BF16 Tensor Core 1,979 TFLOPS
- FP8 Tensor Core 3,958 TFLOPS
- INT8 Tensor Core 3,958 TOPS
- NVLink bandwidth 900GB/s
- Interconnect NVLink 900GB/s, PCIe Gen5 128GB/s
- Multi-Instance GPU (MIG) Up to 7 instances @ 10GB each
- Decoders 7 NVDEC, 7 JPEG
- Max thermal design power (TDP) Up to 700W (configurable)
- Form factor HGX H100, 4 or 8 GPUs
Why Rent NVIDIA H100 on Together
The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.
Efficient training with NVIDIA Hopper architecture
Each H100 cluster leverages fourth-generation Tensor Cores and the Transformer Engine with FP8 precision, enabling fast training for GPT-scale models.
Advanced multi-GPU connectivity
We deploy NVIDIA's NVLink Switch System for 900GB/s bidirectional bandwidth per GPU, providing unparalleled scalability for multi-node AI and HPC workloads.
Secure, multi-instance GPU configurations
Second-generation MIG technology securely partitions GPUs into isolated instances, maximizing resource utilization and quality of service across diverse teams.
Run by researchers who train models
Our research team actively runs and tunes training workloads on NVIDIA H100 systems. You're not just getting hardware — you're working with experts at the edge of what's possible.
What our customers are saying
Rent NVIDIA H100 GPUs on Together GPU Clusters
Spin up a dedicated H100 cluster in minutes with real-time availability, from a single 8-GPU node to thousands of GPUs over non-blocking InfiniBand and NVLink.
Every cluster is dedicated bare metal, orchestrated with Slurm or Kubernetes, and tuned by the research team that built FlashAttention. Your data and model weights stay yours, backed by SOC 2 Type II controls.

Ready to put HGX H100 to work?
See GPU cluster pricing for full details.
Infrastructure you can trust at scale.
Production-grade security.
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022
Regions and availability zones
Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.
- USA2GW+ in the portfolio with 600MW of near-term capacity in US.
- Europe150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
- Asia & Middle EastOptions available based on the scale of the projects in Asia and the Middle East.
FAQ
What is the NVIDIA H100?
The NVIDIA H100 is a Hopper-architecture data center GPU with 80GB of HBM3 memory, 3.35TB/s of bandwidth, and fourth-generation Tensor Cores with FP8 precision, built for large-scale AI training, LLM inference, and HPC.
How much does it cost to rent an NVIDIA H100?
On Together GPU Clusters, H100 GPUs are $3.99 per GPU per hour on demand, with reserved rates from $3.09 per GPU per hour for longer commitments. See the pricing section above or contact sales for volume pricing.
Can I rent H100 GPUs by the hour?
Yes. Launch an on-demand H100 cluster in minutes and pay per GPU, or reserve dedicated capacity for a defined training window.
How much memory does the H100 have?
Each H100 SXM has 80GB of HBM3 memory.
What is the H100 memory bandwidth?
The H100 SXM delivers 3.35TB/s of memory bandwidth.
What is the power consumption of the H100?
The H100 SXM has a configurable thermal design power of up to 700W.
Is the H100 available on demand or only reserved?
Both. Together offers on-demand H100 capacity with real-time availability and reserved clusters for longer terms.
Which regions are H100 clusters available in?
Across the US, Europe (UK, Spain, France, Portugal, Iceland), and select Asia and Middle East regions, supporting data-residency and compliance needs.
What is the NVIDIA H100 used for?
The H100 is used for training and fine-tuning large language models, high-throughput inference, and high-performance computing such as simulation and scientific research.
When was the NVIDIA H100 released?
The H100 launched in 2022 on the NVIDIA Hopper architecture and remains one of the most widely deployed AI data center GPUs.
Why is the NVIDIA H100 expensive?
Demand for H100 capacity has consistently outstripped supply, and each GPU carries costly HBM3 memory and advanced packaging. Renting on Together avoids the upfront hardware cost and long procurement lead times.
How many CUDA cores does the H100 have?
The H100 SXM has 16,896 CUDA cores and 528 fourth-generation Tensor Cores.
Browse other NVIDIA GPUs
Self-serve GPUs with transparent billing.





