Provisioned Throughput
Committed inference capacity, guaranteed to perform.
Provisioned Throughput gives your workloads token-based pricing, reserved capacity, and a production-grade SLA.

Why Provisioned Throughput with Together AI?
Production inference with the reliability your team can build on.
Production performance, guaranteed
Reserved capacity means your traffic doesn't compete. 99% uptime SLA. Guaranteed tokens per minute.
Frontier models at a fraction of the cost
Benchmark-topping quality with dramatically better tokenomics than proprietary APIs. Priced in tokens, just like the APIs you're moving from.
Seamless migration from proprietary APIs
Drop-in API compatibility. No infrastructure to manage. Send your traffic and let the SLA do its job.
Key capabilities, built for inference at scale
Committed throughput, hard SLAs, and defined overage behavior. Everything production teams need without any infrastructure overhead.
Buy Provisioned Throughput Units for the model you're running. Each PTU represents a fixed amount of normalized token throughput per minute (TPM). Your reserved capacity is held exclusively for you.
Every PTU purchase comes with a capacity SLA, not best-effort targets. Provision the throughput you need, and we guarantee it's there when your traffic demands it.
When traffic exceeds your purchased capacity, we'll do our best to route overflow to our serverless fleet at list price.



Research that ships
Our research team doesn't just publish. They build the optimizations that power every inference request.
Performance on DeepSeek V3.1 (Arena Hard)
- Atlas
- Static Speculator
- No Speculator
ATLAS performance
3.18x faster
ATLAS, our AdapTive-LeArning Speculator System, continuously learns from live traffic — outperforming static speculators and specialized hardware.
learn moreCPD improves sustainable QPS by 35-40%
- CPD
- Baseline
Together AI CPD vs 2P1D
+40% throughput
Long-context inference without the latency penalty. CPD (cache-aware prefill-decode disaggregation) separates warm and cold requests, cutting time-to-first-token and boosting throughput by up to 40%.
learn moreTime to first 64 tokens
- Megakernel (H100)
- Baseline (B200)
Megakernel vs baseline
Up to 3.6x faster
Megakernel fuses an entire model's forward pass into a single GPU kernel. Made using the ThunderKittens framework, Megakernel eliminates the idle gaps between operations that rob GPUs of their full potential.
learn moreBF16 all-reduce sum performance (on 8x NVIDIA B200s)
- PK
- NCCL
ParallelKittens vs NCCL
Up to 1.79x faster
ParallelKittens—an extension to ThunderKittens for multi-GPU workloads developed in collaboration with Stanford's Hazy Lab—cuts the synchronization overhead that large multi-GPU models pay on every single forward pass.
learn more
Provisioned Throughput
Reserve dedicated capacity in throughput units (PTUs). Each PTU represents fixed capacity. The tokens-per-minute it delivers depends on the model and the token type.
Compute costs | Input | Cached | Output | Price |
|---|---|---|---|---|
MiniMax M3 | 138,840 | 694,200 | 23,140 | $0.05 |
GLM-5.2 | 35,731 | 192,400 | 9,620 | $0.05 |
Savings compare Together PTU cost against the selected commercial model's published list price ($/1M tokens) on the same traffic profile. Estimates assume continuous 24/7 provisioning (~43,800 min/mo).
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
A fully managed real-time or batch inference API with access to dozens of the most popular AI models.
Best for
Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit.
Best for
An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.
Best for
Run inference with your own engine and model on fully-managed, scalable infrastructure.
Best for
Production-grade
security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022
Customers running inference in production
Inference FAQ
If you can't find the answer you were looking for, feel free to contact our team — we’re here to help.
What is a PTU and how is it different from a token?
A PTU (Provisioned Throughput Unit) is a unit of reserved model capacity. When you buy PTUs, you are buying throughput: a guaranteed rate of inference your application can rely on.
Tokens are still how usage is measured and billed underneath. But instead of paying per token as you go, PTUs commit you to a certain level of throughput . Your actual traffic (input tokens, output tokens, cache reads) is converted into normalized PTU consumption using model-specific ratios of input, output, and cached tokens, so a single PTU commitment covers whatever mix of request types you send.
What happens if my traffic exceeds my purchased PTU capacity?
Traffic that exceeds your committed PTU capacity is handled on a best-effort basis, routed to serverless capacity where available. Requests handled this way are not covered by your PTU SLA, and serverless pricing applies.
If you consistently hit your capacity ceiling, the right move is to purchase additional PTUs. Your account team can help you size based on your actual traffic patterns.
What SLAs does Provisioned Throughput include?
Provisioned Throughput includes two customer-facing commitments for eligible traffic within your purchased PTU capacity:
Max normalized TPM. You can send up to your contracted normalized TPM rate for the covered model. Your actual traffic (input tokens, output tokens, cache reads) is converted into normalized TPM using model-specific ratios, so a single PTU commitment covers whatever mix of request types you send. Traffic shape does not change the SLA. It changes how quickly you burn through your PTU capacity.
Availability. Together targets a 99% monthly success rate for eligible requests. Eligible requests are those sent to the contracted model through the PTU endpoint, within your purchased capacity, and within standard product limits. Failures caused by Together systems count against the availability target. Client errors, invalid requests, cancelled requests, and traffic above contracted capacity do not.
SLAs do not apply to overage traffic. If your traffic exceeds your purchased PTU capacity, excess requests may fall back to serverless on a best-effort basis, with no availability guarantee.
Which models are available on Provisioned Throughput?
Provisioned Throughput is currently available for MiniMax M3 and GLM 5.2. PTUs are scoped to a specific model, so capacity purchased for one model does not carry over to another.
We are expanding model availability. Contact your account team for the current roadmap and capacity availability.
Are there any minimum commitments for Provisioned Throughput?
We currently offer a one-month minimum term for Provisioned Throughput. Discounts are available at higher levels of commitment.
Where is Provisioned Throughput available?
We currently have Provisioned Throughput capacity for workloads served in the US and Canada. Please contact us to discuss inference outside of the US.
How is Provisioned Throughput different from Dedicated Inference?
The key difference is what you are buying and how much control you get over the serving environment.
Provisioned Throughput gives you reserved capacity for base models, expressed as PTUs, with a standard SLA. You get stronger guarantees than serverless and token-based pricing you can plan around, without managing GPU infrastructure. It is designed for production workloads where reliability matters but the model itself does not need to be customized.
Dedicated Inference is the right fit when you need a post-trained or custom model, or when your workload requires tighter control over the serving environment. If you are running a standard open-weight model in production, Provisioned Throughput is the faster, simpler path.


