Dedicated Model Inference

Dedicated inference, tuned for production

Deploy any open model in minutes. Roll out safely, scale to meet demand, and keep full control.
No platform team required.

Why Dedicated Inference
with Together AI?

Designed for production workloads that need 
consistent performance and operational control.

Production control plane

Safe rollouts, autoscaling, fast cold starts, auto-rollback, and multi-region failover.

Best-in-market economics

Faster inference and more tokens per GPU deliver closed-model quality at a lower cost.

Research in production

Frontier research ships into the product continuously, so you're always on the leading edge.

Build with leading models

Explore top-performing models across text, image, video, code, and voice.

new

Chat

DeepSeek V4 Pro

new

Chat

MiniMax M3

new

Chat

GLM-5.2

new

Chat

Kimi K3

new

Chat

Inkling Small

new

Chat

DeepSeek V4 Flash 0731

new

Chat

Qwen3.8 Max

new

Chat

NVIDIA Nemotron 3.5 Lightning

new

Chat

Muse Glimmer

new

Video

ByteDance Seedance 2.5

new

Chat

Gemma 4 31B

new

Chat

NVIDIA Nemotron 3 Ultra

new

Chat

Kimi K2.7 Code

new

Chat

Qwen3.7-Plus

new

Image

GPT Image 2

Video

ByteDance Seedance 2.0

free

Code

PrismMLTernary Bonsai 27B

new

Code

Inkling

new

LFM2.5-8B-A1B

new

Video

FLUX 3

Have your own model?

Deploy custom containers on Together’s managed GPU infrastructure with automatic scaling, job queues, and built-in observability.

Key capabilities, purpose built for AI natives

Bring any open-weight model, deploy it to a dedicated endpoint in minutes, and run it in production with full control and optimized performance.

    • Deploy in minutes

      NO DEVOPS REQUIRED
      LIVE IN MINUTES
      SIMPLE CONFIGURATION

      Select a target model and hardware config and be live in minutes. Production-ready endpoints, no deep infra expertise

    • Adaptive speculative decoding

      Faster Outputs
      Learns in production
      Lossless quality

      Cut latency on dedicated infrastructure with ATLAS — Together's AdapTive-LeArning Speculator System. Predict and validate multiple tokens per step to accelerate workloads continuously. No decoding bottlenecks.

    • Bring your own language model

      BRING ANY MODEL
      DEPLOY IN MINUTES
      UI OR CLI

      Deploy custom models directly from Hugging Face or S3 onto dedicated endpoints via the UI or CLI. Maintain complete ownership while offloading infrastructure management.

Production-grade deployment, built in

The deployment safety a platform team would build — already built and managed.

  • Zero-downtime rollouts

    Canary, rolling, blue/green & automatic rollback on your thresholds.

  • Shadow & A/B

    Mirror live traffic to a candidate model; zero user impact.

  • Multi-region failover

    Declare preferred regions; traffic shifts automatically.

  • SLO-driven autoscaling

    Scale on TTFT, latency, and throughput, not demo loads.

  • Advanced routing

    Least-loaded, session affinity, prefix-cache-aware.

  • Coming soon
    Multi-LoRA serving

    Serve multiple LoRA adapters from a single deployment.

Research that ships

Our research team doesn't just publish. They build the optimizations that power every inference request.

  • Atlas
  • CPD
  • Megakernel
  • ThunderKittens
  • Performance on DeepSeek V3.1 (Arena Hard)

    • Atlas
    • Static Speculator
    • No Speculator

    ATLAS performance

    3.18x faster

    ATLAS, our AdapTive-LeArning Speculator System, continuously learns from live traffic — outperforming static speculators and specialized hardware.

    learn more
  • CPD improves sustainable QPS by 35-40%

    • CPD
    • Baseline

    Together AI CPD vs 2P1D

    +40% throughput

    Long-context inference without the latency penalty. CPD (cache-aware prefill-decode disaggregation) separates warm and cold requests, cutting time-to-first-token and boosting throughput by up to 40%.

    learn more
  • Time to first 64 tokens

    • Megakernel (H100)
    • Baseline (B200)

    Megakernel vs baseline

    Up to 3.6x faster

    Megakernel fuses an entire model's forward pass into a single GPU kernel. Made using the ThunderKittens framework, Megakernel eliminates the idle gaps between operations that rob GPUs of their full potential.

    learn more
  • BF16 all-reduce sum performance (on 8x NVIDIA B200s)

    • PK
    • NCCL

    ParallelKittens vs NCCL

    Up to 1.79x faster

    ParallelKittens—an extension to ThunderKittens for multi-GPU workloads developed in collaboration with Stanford's Hazy Lab—cuts the synchronization overhead that large multi-GPU models pay on every single forward pass.

    learn more

Deployment options

Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.

  • Serverless Inference

  • Provisioned 
Throughput

  • Dedicated Model 
Inference

  • Dedicated Container 
Inference

Serverless Inference

A fully managed real-time or batch inference API with access to dozens of the most popular AI models.

Best for

Variable or unpredictable traffic

Rapid prototyping and iteration

Cost-sensitive or early-stage production workloads

Provisioned 
Throughput

Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit. 

Best for

Production workloads

Reliability guarantees

Predictable pricing

Dedicated Model 
Inference

An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.

Best for

Predictable or steady traffic

Latency-sensitive applications

High-throughput production workloads

Dedicated Container 
Inference

Run inference with your own engine and model on fully-managed, scalable infrastructure.

Best for

Generative media models

Non-standard runtimes

Custom inference pipelines

Single-tenant
security and data privacy

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Learn More

Every dedicated deployment runs on single-tenant, isolated GPUs, so your traffic and data are never shared. Your prompts and model weights stay under your control, and Together never trains on your data. Choose your deployment region for data residency, backed by SOC 2 Type II and ISO 27001.

  • NVIDIA logo with text Preferred Partner on a black background.
    preferred partner
  • SOC 2 Type II
  • ISO 27001:2022

Customers running inference in production

Young man with black hair wearing a dark jacket and sunglasses standing near a waterfall.
  • cost reduction

  • <400ms

    p95 model latency

  • Weekly

    model deployments

"Low latency is especially important for voice because there’s a much higher UX bar. Together helped us push latency down by optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner — proactive about risks and fast when issues come up."

Max Lu

Head of Research, Decagon

Smiling young man with light brown hair wearing a blue patterned shirt in a softly blurred indoor setting.
  • ~30%

    Cost savings

"Together has helped us deploy VyUI, our state-of-the-art computer AI model. We had multiple in-depth meetings where we brainstormed how we could satisfy our model's custom technical requirements while still leveraging Together's infrastructure for efficient, load-balanced inference."

Luca Weihs

Co-founder, Vercept

Smiling man with short dark hair wearing a black shirt and dark gray blazer against a light background.

"Together AI offers optimized performance at scale, and at a lower cost than closed-source providers – all while maintaining strict privacy standards."

Vineet Khosla

CTO, The Washington Post

Inference FAQ

If you can't find the answer you were looking for, feel free to contact our team — we’re here to help.

What is dedicated inference?

Dedicated inference means running a model on GPUs reserved exclusively for your workload, instead of sharing capacity through a pay-per-token API. You get consistent latency, predictable cost, and full control over the model and hardware.

How is dedicated inference different from serverless on Together AI?

Serverless bills per token on shared, best-effort capacity and suits spiky or early-stage traffic. Dedicated inference reserves GPUs at a fixed GPU-hour rate for predictable performance and cost on steady production workloads.

Can I move from a serverless prototype to a dedicated endpoint without re-architecting?

Yes. Serverless and dedicated share the same OpenAI-compatible API, so you can prototype on serverless and move the same code to a dedicated endpoint by pointing at the new endpoint.

How much does dedicated inference cost?

Dedicated inference is billed per GPU-hour based on the GPU type and number of replicas you provision. See our GPU cluster pricing for current rates.

Can I deploy my own model?

Yes. Deploy custom or fine-tuned models from Hugging Face or your own S3 storage via the UI or CLI, while Together AI manages the infrastructure.

Can I update a deployed model without downtime?

Yes. You can create a new deployment with an updated configuration (e.g. a different model version, fine-tuned weights, GPU type, or speculative decoding) within the same endpoint, then run shadow traffic or A/B testing against it to confirm performance holds up before fully switching over. Together AI manages this rollout without taking the serving endpoint down, so production traffic continues uninterrupted during the update.

Is dedicated inference single-tenant and private?

Yes. Each deployment runs on single-tenant, isolated GPUs reserved for you. Your data and weights stay under your control, and Together does not train on your data. SOC 2 Type II and ISO 27001 certified.