GPU / NVIDIA HGX B300

NVIDIA HGX B300

Rent NVIDIA B300 GPUs on demand or reserve a cluster, with 288GB HBM3e and enhanced FP4 for reasoning-scale training and inference.

NVIDIA HGX B300 Pricing & Specs

The NVIDIA B300 is NVIDIA's Blackwell Ultra GPU, with 288GB of HBM3e, 8TB/s of bandwidth, and enhanced FP4 built for the age of reasoning. It's a large step up over the B200 in memory and inference throughput. Rent it on Together on demand or as a dedicated cluster.

Performance
More dense FP4

10x

vs HGX B200

Attention performance

2x

vs HGX B200

AI factory output

30x

vs Hopper

Pricing

Rent HGX B300 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.

Plan

Price

On-demand (pay as you go)

Reserved 7-30 days

Contact us

Reserved 31-90 days

Reserved 91-180 days

Reserved 181+ days

Prices as of July 2026.

See pricing for every GPU on the GPU cluster pricing page.

Technical specification

  • Architecture NVIDIA Blackwell Ultra (dual-die)
  • GPU Memory (VRAM) 288GB HBM3e
  • Memory Bandwidth 8TB/s
  • Blackwell Ultra GPUs 8 GPUs
  • Total FP4 Tensor Core 144 PFLOPS
  • Total FP8/FP6 Tensor Core 72 PFLOPS
  • Total Fast Memory 2.1 TB
  • Total Memory Bandwidth 64 TB/s
  • Total NVLink Bandwidth 14.4 TB/s
  • FP4 Tensor Core (per GPU) 14 PFLOPS
  • FP8/FP6 Tensor Core (per GPU) 9 PFLOPS
  • INT8 Tensor Core 307 TOPS
  • FP16/BF16 Tensor Core 4.5 PFLOPS
  • TF32 Tensor Core 2.2 PFLOPS
  • FP32 75 TFLOPS
  • FP64 / FP64 Tensor Core 1.2 TFLOPS
  • Multi-Instance GPU (MIG) 7
  • Decompression Engine Yes
  • Decoders 7 NVDEC, 7 nvJPEG
  • Max Thermal Design Power (TDP) Configurable up to 1,400 W
  • Interconnect 5th Gen NVLink: 1.8 TB/s, PCIe Gen6: 256 GB/s
  • Networking ConnectX-8 SuperNIC, 800Gb/s per GPU
  • Form factor HGX B300, 8 GPUs

Why Rent NVIDIA HGX B300 on Together

The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.

Native FP4 for reasoning

Blackwell Ultra's enhanced Transformer Engine and native FP4 are tuned for test-time scaling and long reasoning, with roughly 1.5x the dense FP4 compute of the B200.

288GB HBM3e at 8TB/s

50% more memory than the B200, so trillion-parameter and long-context models fit with fewer nodes and less offloading.

Scale from one node up with fifth-gen NVLink

Spin up with real-time availability, from a single 8-GPU node to multi-node clusters over non-blocking InfiniBand and 1.8TB/s fifth-generation NVLink.

Run by researchers who train models

Team actively tunes training workloads on NVIDIA Blackwell systems.

What our customers are saying

Man with glasses and beard smiling against a background with glowing circuit-like lines.

"Delivering competitive pricing, strong reliability and a properly set up cluster is the bulk of the value differentiation for most AI clouds. The only differentiated value we have seen outside this set is from a Neocloud called Together AI, where the inventor of FlashAttention, Tri Dao, works. We don't believe the value created by Together can be replicated elsewhere."

Dylan Patel

Founder, SemiAnalysis

Smiling man with short dark hair wearing a blue blazer and white shirt in an outdoor corridor.

"Training our omnimodal Character-3 model required infrastructure designed for large-scale AI. The Together Frontier AI Factory delivered the performance we needed to push the boundaries of multimodal video generation. Together AI understands what builders need — and that made all the difference."

Michael Lingelbach

CEO, Hedra

"Together GPU Clusters provided a combination of amazing training performance, expert support, and the ability to scale to meet our rapid growth to help us serve our growing community of AI creators."

Young woman with long dark hair smiling outdoors wearing a white turtleneck and statement earrings.

Demi Guo

CEO, Pika

“Together AI provides the performance and reliability we need for real-time, high-quality image and video generation at scale. We value that Together AI is much more than an infrastructure provider — they're a true innovation partner, enabling us to push creative boundaries without compromise.”

Young man wearing a cap sprays graffiti on a wall with a spray paint can in black and white.

Victor Perez

Co-Founder, Krea

Infrastructure you can trust at scale.
Production-grade security.

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Learn More

As an NVIDIA Cloud Partner, Together builds and operates clusters on NVIDIA NCP reference architectures for predictable performance and faster time to production. Your data and models remain under your control with strict privacy safeguards and SOC 2–compliant security practices.

  • NVIDIA logo with text Preferred Partner on a black background.
    preferred partner
  • SOC 2 Type II
  • ISO 27001:2022

Regions and availability zones

Choose from global regions to meet data residency and compliance requirements—HIPAA for healthcare, GDPR for Europe, or banking regulations.

  • USA
    2GW+ in the portfolio with 600MW of near-term capacity in US.
  • Europe
    150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
  • Asia & Middle East
    Options available based on the scale of the projects in Asia and the Middle East.

FAQ

What is the NVIDIA B300?

The NVIDIA B300 is NVIDIA's Blackwell Ultra GPU, with 288GB of HBM3e, 8TB/s of bandwidth, and enhanced FP4, built for AI reasoning and large-scale inference and training.

How much does it cost to rent an NVIDIA B300?

Contact sales for current B300 pricing. See the pricing section above or the GPU cluster pricing page.

Can I rent B300 GPUs by the hour?

Contact sales to check current availability and terms for on-demand and reserved B300 capacity.

How much memory does the B300 have?

Each B300 has 288GB of HBM3e, 50% more than the B200.

What is the B300 memory bandwidth?

8TB/s of HBM3e bandwidth per GPU. That feeds the FP4 compute during large-batch inference and long-context reasoning so throughput isn't held back by memory.

What is the power consumption of the B300?

The B300 draws up to 1,400W per GPU at NVIDIA's maximum spec.

How is the B300 different from the B200?

The B300 (Blackwell Ultra) carries 288GB of HBM3e versus 192GB on the B200, delivers roughly 1.5x the dense FP4 compute, and doubles attention throughput. It's tuned specifically for reasoning workloads.

Does the B300 support FP4?

Yes. It uses a second-generation Transformer Engine with native FP4, which shrinks the memory footprint of large models while holding accuracy for inference.

What is the difference between the B300 and the GB300 NVL72?

The B300 is a single GPU, deployed eight to a node in an HGX system. The GB300 NVL72 is a 72-GPU Grace Blackwell Ultra system that links all 72 GPUs into one NVLink domain for the largest training and inference jobs.

Is the B300 available on demand or reserved?

Both, subject to current availability. Contact sales to check what's open in your target region.

Which regions are B300 clusters available in?

Currently across the US, Europe (UK, Spain, France, Portugal, Iceland), and select Asia and Middle East regions.

Browse all NVIDIA GPUs

Self-serve GPUs with transparent billing.

H100
Hardware

NVIDIA H100 (80 GB)

On-demand

$3.99/hr per GPU

Reserved

Starting at $3.19/hr per GPU

Scale

8 to 256 GPUs

Create cluster
H200
Hardware

NVIDIA H200 (140 GB)

On-demand

$5.99/hr per GPU

Reserved

Starting at $3.99/hr per GPU

Scale

256 to 1,000 GPUs

Create cluster
HGX B200
Hardware

NVIDIA HGX B200 (180 GB)

On-demand

$8.19/hr per GPU

Reserved

Starting at $6.79/hr per GPU

Scale

256 to 1,000+ GPUs

Create cluster
HGX B300
Hardware

NVIDIA HGX B300 (270 GB)

Reserved

Contact us for pricing

Contact sales
GB200 NVL72
Hardware

NVIDIA GB200 NVL72 (186 GB)

Reserved

Contact us for pricing

Scale

512 to 1,000+ GPUs

Contact sales
GB300 NVL72
Hardware

NVIDIA GB300 NVL72 (288 GB)

Reserved

Contact us for pricing

Contact sales