Provisioned Throughput

Committed inference capacity, 
guaranteed to perform.

Provisioned Throughput gives your workloads token-based pricing, reserved capacity, 
and a production-grade SLA.

Together AI interface showing MiniMax M3 model details, code snippet, pricing, and endpoints status offline.

Why Provisioned Throughput 
with Together AI?

Production inference with the reliability your team can build on.

Production performance, guaranteed

Reserved capacity means your traffic doesn't compete. 99% uptime SLA. Guaranteed tokens per minute.

Frontier models at a fraction of the cost

Benchmark-topping quality with dramatically better tokenomics than proprietary APIs. Priced in tokens, just like the APIs you're moving from.

Seamless migration from proprietary APIs

Drop-in API compatibility. No infrastructure to manage. Send your traffic and let the SLA do its job.

Key capabilities, built for inference at scale

Committed throughput, hard SLAs, and defined overage behavior. Everything production teams need without any infrastructure overhead.

    • PTU-based reserved capacity

      GUARANTEED CAPACITY
      Lower latency
      PREDICTABLE COST

      Buy Provisioned Throughput Units for the model you're running. Each PTU represents a fixed amount of normalized token throughput per minute (TPM). Your reserved capacity is held exclusively for you.

    • Hard performance guarantees

      RELIABILITY SLA
      GUARANTEED THROUGHPUT
      NO BEST-EFFORT

      Every PTU purchase comes with a capacity SLA, not best-effort targets. Provision the throughput you need, and we guarantee it's there when your traffic demands it.

    • Predictable overage behavior

      AUTOMATIC OVERFLOW
      LIST-PRICE BILLING
      SERVERLESS FALLBACK

      When traffic exceeds your purchased capacity, we'll do our best to route overflow to our serverless fleet at list price.

Research that ships

Our research team doesn't just publish. They build the optimizations that power every inference request.

  • Atlas
  • CPD
  • Megakernel
  • ThunderKittens
  • Performance on DeepSeek V3.1 (Arena Hard)

    • Atlas
    • Static Speculator
    • No Speculator

    ATLAS performance

    3.18x faster

    ATLAS, our AdapTive-LeArning Speculator System, continuously learns from live traffic — outperforming static speculators and specialized hardware.

    learn more
  • CPD improves sustainable QPS by 35-40%

    • CPD
    • Baseline

    Together AI CPD vs 2P1D

    +40% throughput

    Long-context inference without the latency penalty. CPD (cache-aware prefill-decode disaggregation) separates warm and cold requests, cutting time-to-first-token and boosting throughput by up to 40%.

    learn more
  • Time to first 64 tokens

    • Megakernel (H100)
    • Baseline (B200)

    Megakernel vs baseline

    Up to 3.6x faster

    Megakernel fuses an entire model's forward pass into a single GPU kernel. Made using the ThunderKittens framework, Megakernel eliminates the idle gaps between operations that rob GPUs of their full potential.

    learn more
  • BF16 all-reduce sum performance (on 8x NVIDIA B200s)

    • PK
    • NCCL

    ParallelKittens vs NCCL

    Up to 1.79x faster

    ParallelKittens—an extension to ThunderKittens for multi-GPU workloads developed in collaboration with Stanford's Hazy Lab—cuts the synchronization overhead that large multi-GPU models pay on every single forward pass.

    learn more

Provisioned Throughput

Reserve dedicated capacity in throughput units (PTUs). Each PTU represents fixed capacity. The tokens-per-minute it delivers depends on the model and the token type.

Estimate your PTUs & cost
%
-$7 236 83% lower

Compute costs

Input
TPM/PTU

Cached
TPM/PTU

Output
TPM/PTU

Price
PTU/MIN

MiniMax M3

138,840

694,200

23,140

$0.05

GLM-5.2

35,731

192,400

9,620

$0.05

Savings compare Together PTU cost against the selected commercial model's published list price ($/1M tokens) on the same traffic profile. Estimates assume continuous 24/7 provisioning (~43,800 min/mo).

Deployment options

Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.

  • Serverless Inference

  • Provisioned 
Throughput

  • Dedicated Model 
Inference

  • Dedicated Container 
Inference

Serverless Inference

A fully managed real-time or batch inference API with access to dozens of the most popular AI models.

Best for

Variable or unpredictable traffic

Rapid prototyping and iteration

Cost-sensitive or early-stage production workloads

Provisioned 
Throughput

Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit. 

Best for

Production workloads

Reliability guarantees

Predictable pricing

Dedicated Model 
Inference

An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.

Best for

Predictable or steady traffic

Latency-sensitive applications

High-throughput production workloads

Dedicated Container 
Inference

Run inference with your own engine and model on fully-managed, scalable infrastructure.

Best for

Generative media models

Non-standard runtimes

Custom inference pipelines

Production-grade
security and data privacy

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Learn More

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

  • NVIDIA logo with text Preferred Partner on a black background.
    preferred partner
  • SOC 2 Type II
  • ISO 27001:2022

Customers running inference in production

Young man with black hair wearing a dark jacket and sunglasses standing near a waterfall.
  • cost reduction

  • <400ms

    p95 model latency

  • Weekly

    model deployments

"Low latency is especially important for voice because there’s a much higher UX bar. Together helped us push latency down by optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner — proactive about risks and fast when issues come up."

Max Lu

Head of Research, Decagon

Smiling young man with light brown hair wearing a blue patterned shirt in a softly blurred indoor setting.
  • ~30%

    Cost savings

"Together has helped us deploy VyUI, our state-of-the-art computer AI model. We had multiple in-depth meetings where we brainstormed how we could satisfy our model's custom technical requirements while still leveraging Together's infrastructure for efficient, load-balanced inference."

Luca Weihs

Co-founder, Vercept

Smiling man with short dark hair wearing a black shirt and dark gray blazer against a light background.

"Together AI offers optimized performance at scale, and at a lower cost than closed-source providers – all while maintaining strict privacy standards."

Vineet Khosla

CTO, The Washington Post

Inference FAQ

If you can't find the answer you were looking for, feel free to contact our team — we’re here to help.

What is a PTU and how is it different from a token?

A PTU (Provisioned Throughput Unit) is a unit of reserved model capacity. When you buy PTUs, you are buying throughput: a guaranteed rate of inference your application can rely on.

Tokens are still how usage is measured and billed underneath. But instead of paying per token as you go, PTUs commit you to a certain level of throughput . Your actual traffic (input tokens, output tokens, cache reads) is converted into normalized PTU consumption using model-specific ratios of input, output, and cached tokens, so a single PTU commitment covers whatever mix of request types you send.

What happens if my traffic exceeds my purchased PTU capacity?

Traffic that exceeds your committed PTU capacity is handled on a best-effort basis, routed to serverless capacity where available. Requests handled this way are not covered by your PTU SLA, and serverless pricing applies.

If you consistently hit your capacity ceiling, the right move is to purchase additional PTUs. Your account team can help you size based on your actual traffic patterns.

What SLAs does Provisioned Throughput include?

Provisioned Throughput includes two customer-facing commitments for eligible traffic within your purchased PTU capacity:

Max normalized TPM. You can send up to your contracted normalized TPM rate for the covered model. Your actual traffic (input tokens, output tokens, cache reads) is converted into normalized TPM using model-specific ratios, so a single PTU commitment covers whatever mix of request types you send. Traffic shape does not change the SLA. It changes how quickly you burn through your PTU capacity.

Availability. Together targets a 99% monthly success rate for eligible requests. Eligible requests are those sent to the contracted model through the PTU endpoint, within your purchased capacity, and within standard product limits. Failures caused by Together systems count against the availability target. Client errors, invalid requests, cancelled requests, and traffic above contracted capacity do not.

SLAs do not apply to overage traffic. If your traffic exceeds your purchased PTU capacity, excess requests may fall back to serverless on a best-effort basis, with no availability guarantee.

Which models are available on Provisioned Throughput?

Provisioned Throughput is currently available for MiniMax M3 and GLM 5.2. PTUs are scoped to a specific model, so capacity purchased for one model does not carry over to another.

We are expanding model availability. Contact your account team for the current roadmap and capacity availability.

Are there any minimum commitments for Provisioned Throughput?

We currently offer a one-month minimum term for Provisioned Throughput. Discounts are available at higher levels of commitment.

Where is Provisioned Throughput available?

We currently have Provisioned Throughput capacity for workloads served in the US and Canada. Please contact us to discuss inference outside of the US.

How is Provisioned Throughput different from Dedicated Inference?

The key difference is what you are buying and how much control you get over the serving environment.

Provisioned Throughput gives you reserved capacity for base models, expressed as PTUs, with a standard SLA. You get stronger guarantees than serverless and token-based pricing you can plan around, without managing GPU infrastructure. It is designed for production workloads where reliability matters but the model itself does not need to be customized.

Dedicated Inference is the right fit when you need a post-trained or custom model, or when your workload requires tighter control over the serving environment. If you are running a standard open-weight model in production, Provisioned Throughput is the faster, simpler path.