Two leaders. One platform.

Together AI and NVIDIA combine Together AI’s fastest, research-optimized AI platform with the NVIDIA full-stack AI Factory platform to help teams move from experimentation to production at scale.

The AI Native Cloud, built on the NVIDIA accelerated computing platform

The most advanced AI workloads demand more than bare-metal compute alone. Together AI brings inference optimization, developer experience, and enterprise trust directly onto NVIDIA accelerated computing to help your team ship faster — without managing the underlying infrastructure.

Together AI

Purpose-built for AI-native companies that need reliability, speed, and control at every stage of the model lifecycle.

Optimized inference with FlashAttention, speculative decoding & continuous batching

Full model lifecycle: pre-training → fine-tuning → production inference

SOC II Type 2 certified, Zero Data Retention policy

Dedicated clusters and private deployments for regulated industries

OpenAI-compatible APIs to migrate in minutes

NVIDIA

NVIDIA provides the best full-stack AI factory platform spanning accelerated computing, open models, networking, storage, reference architectures, and software — designed to deliver high performance and efficient AI training and inference at scale.

NVIDIA Blackwell platform

NVIDIA NVLink™ & NVIDIA Quantum-X800 InfiniBand for scale-up and scale-out connectivity

NVIDIA TensorRT for optimized generative AI inference

Optimizations on Dynamo and CUDA ecosystem

NVIDIA Nemotron open models

NVIDIA AI Enterprise software and support

Everything your AI team needs to move faster

Through extreme co-design, Together AI and NVIDIA provide the lowest cost-per-token on leading open models, unlocking higher performance on the NVIDIA accelerated computing platform.

H100
Hardware

NVIDIA H100 (80 GB)

On-demand

$3.99/hr per GPU

Reserved

Starting at $3.19/hr per GPU

Scale

8 to 256 GPUs

Create cluster
H200
Hardware

NVIDIA H200 (140 GB)

On-demand

$5.99/hr per GPU

Reserved

Starting at $3.99/hr per GPU

Scale

256 to 1,000 GPUs

Create cluster
HGX B200
Hardware

NVIDIA HGX B200 (180 GB)

On-demand

$8.19/hr per GPU

Reserved

Starting at $6.79/hr per GPU

Scale

256 to 1,000+ GPUs

Create cluster
HGX B300
Hardware

NVIDIA HGX B300 (270 GB)

Reserved

Contact us for pricing

Contact sales
GB200 NVL72
Hardware

NVIDIA GB200 NVL72 (186 GB)

Reserved

Contact us for pricing

Scale

512 to 1,000+ GPUs

Contact sales
GB300 NVL72
Hardware

NVIDIA GB300 NVL72 (288 GB)

Reserved

Contact us for pricing

Contact sales

High-performance inference

Together Inference runs on latest-gen NVIDIA accelerated compute with custom kernels, adaptive speculative decoding, and intelligent quantization to deliver industry-leading tokens/sec at scale.

Private & dedicated deployments

Run models on dedicated NVIDIA AI infrastructure in Together's global data centers. Zero Data Retention policy and SOC II compliance for privacy-sensitive production workloads.

200+ models, one API

Access the broadest catalog of open models — including Nemotron, DeepSeek, Kimi, MiniMax, Qwen, GLM, Gemma, GPT OSS, and more — all running on NVIDIA AI infrastructure via a single OpenAI-compatible endpoint.

Built for every AI workload.

    • Production-grade inference at any scale

      Inference

      Serve 10T+ tokens per day with sub-100ms time to first token (TTFT). Using our cutting-edge research paired with NVIDIA inference software, NVIDIA TensorRT, and NVIDIA Dynamo to maximize throughput and achieve lower cost per token.

    • Build with NVIDIA Nemotron models

      NVIDIA Nemotron

      Access NVIDIA Nemotron models for reasoning, agentic AI applications, and speech recognition — optimized and served via Together AI’s inference API.

    • Built for what you are building

      Model Shaping

      Fine-tune your models to your domain. Domain-specific fine-tuning, reinforcement learning (RLHF), and direct preference optimization (DPO) on the NVIDIA accelerated computing platform. From LoRA adapters to full-model updates, you retain control of your weights and deployment.

Production ready from day one

Privacy-sensitive U.S. organizations and global enterprises choose Together AI because security and compliance aren’t afterthoughts, they’re built into every layer of the stack.

  • SOC II Type 2 compliant

    Independently audited security controls covering availability, confidentiality, and processing integrity.

  • Zero Data Retention

    Prompts and completions are never stored, logged, or used for training. Your data stays yours.

  • Dedicated infrastructure

    Isolated NVIDIA AI infrastructure for customers who need multi-tenant environments, network isolation, and custom SLAs.

  • Global data centers

    Multi-region deployments across North America, Europe, and Asia with data residency and zero-trust controls.

Customer stories

Young man with black hair wearing a dark jacket and sunglasses standing near a waterfall.
  • cost reduction

  • <400ms

    p95 model latency

  • Weekly

    model deployments

"Low latency is especially important for voice because there’s a much higher UX bar. Together helped us push latency down by optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner — proactive about risks and fast when issues come up."

Max Lu

Head of Research, Decagon