NVIDIA HGX B300
Rent NVIDIA B300 GPUs on demand or reserve a cluster, with 288GB HBM3e and enhanced FP4 for reasoning-scale training and inference.
NVIDIA HGX B300 Pricing & Specs
The NVIDIA B300 is NVIDIA's Blackwell Ultra GPU, with 288GB of HBM3e, 8TB/s of bandwidth, and enhanced FP4 built for the age of reasoning. It's a large step up over the B200 in memory and inference throughput. Rent it on Together on demand or as a dedicated cluster.
10x
vs HGX B200
2x
vs HGX B200
30x
vs Hopper
Pricing
Rent HGX B300 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.
Plan | Price |
|---|---|
On-demand (pay as you go) | — |
Reserved 7-30 days | Contact us |
Reserved 31-90 days | |
Reserved 91-180 days | |
Reserved 181+ days |
Prices as of July 2026.
See pricing for every GPU on the GPU cluster pricing page.
Technical specification
- Architecture NVIDIA Blackwell Ultra (dual-die)
- GPU Memory (VRAM) 288GB HBM3e
- Memory Bandwidth 8TB/s
- Blackwell Ultra GPUs 8 GPUs
- Total FP4 Tensor Core 144 PFLOPS
- Total FP8/FP6 Tensor Core 72 PFLOPS
- Total Fast Memory 2.1 TB
- Total Memory Bandwidth 64 TB/s
- Total NVLink Bandwidth 14.4 TB/s
- FP4 Tensor Core (per GPU) 14 PFLOPS
- FP8/FP6 Tensor Core (per GPU) 9 PFLOPS
- INT8 Tensor Core 307 TOPS
- FP16/BF16 Tensor Core 4.5 PFLOPS
- TF32 Tensor Core 2.2 PFLOPS
- FP32 75 TFLOPS
- FP64 / FP64 Tensor Core 1.2 TFLOPS
- Multi-Instance GPU (MIG) 7
- Decompression Engine Yes
- Decoders 7 NVDEC, 7 nvJPEG
- Max Thermal Design Power (TDP) Configurable up to 1,400 W
- Interconnect 5th Gen NVLink: 1.8 TB/s, PCIe Gen6: 256 GB/s
- Networking ConnectX-8 SuperNIC, 800Gb/s per GPU
- Form factor HGX B300, 8 GPUs
Why Rent NVIDIA HGX B300 on Together
The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.
Native FP4 for reasoning
Blackwell Ultra's enhanced Transformer Engine and native FP4 are tuned for test-time scaling and long reasoning, with roughly 1.5x the dense FP4 compute of the B200.
288GB HBM3e at 8TB/s
50% more memory than the B200, so trillion-parameter and long-context models fit with fewer nodes and less offloading.
Scale from one node up with fifth-gen NVLink
Spin up with real-time availability, from a single 8-GPU node to multi-node clusters over non-blocking InfiniBand and 1.8TB/s fifth-generation NVLink.
Run by researchers who train models
Team actively tunes training workloads on NVIDIA Blackwell systems.
What our customers are saying
Infrastructure you can trust at scale.
Production-grade security.
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022
Regions and availability zones
Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.
- USA2GW+ in the portfolio with 600MW of near-term capacity in US.
- Europe150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
- Asia & Middle EastOptions available based on the scale of the projects in Asia and the Middle East.
FAQ
What is the NVIDIA B300?
The NVIDIA B300 is NVIDIA's Blackwell Ultra GPU, with 288GB of HBM3e, 8TB/s of bandwidth, and enhanced FP4, built for AI reasoning and large-scale inference and training.
How much does it cost to rent an NVIDIA B300?
Contact sales for current B300 pricing. See the pricing section above or the GPU cluster pricing page.
Can I rent B300 GPUs by the hour?
Contact sales to check current availability and terms for on-demand and reserved B300 capacity.
How much memory does the B300 have?
Each B300 has 288GB of HBM3e, 50% more than the B200.
What is the B300 memory bandwidth?
8TB/s of HBM3e bandwidth per GPU. That feeds the FP4 compute during large-batch inference and long-context reasoning so throughput isn't held back by memory.
What is the power consumption of the B300?
The B300 draws up to 1,400W per GPU at NVIDIA's maximum spec.
How is the B300 different from the B200?
The B300 (Blackwell Ultra) carries 288GB of HBM3e versus 192GB on the B200, delivers roughly 1.5x the dense FP4 compute, and doubles attention throughput. It's tuned specifically for reasoning workloads.
Does the B300 support FP4?
Yes. It uses a second-generation Transformer Engine with native FP4, which shrinks the memory footprint of large models while holding accuracy for inference.
What is the difference between the B300 and the GB300 NVL72?
The B300 is a single GPU, deployed eight to a node in an HGX system. The GB300 NVL72 is a 72-GPU Grace Blackwell Ultra system that links all 72 GPUs into one NVLink domain for the largest training and inference jobs.
Is the B300 available on demand or reserved?
Both, subject to current availability. Contact sales to check what's open in your target region.
Which regions are B300 clusters available in?
Currently across the US, Europe (UK, Spain, France, Portugal, Iceland), and select Asia and Middle East regions.
Browse all NVIDIA GPUs
Self-serve GPUs with transparent billing.





