Reinforcement Learning - now in beta
Post-training, all the way to production
RL and SFT, full-weight and LoRA, on frontier open models. The highest-quality custom models through fast, large-scale experimentation.

Why Together Custom Training
One platform that takes a model from first experiment to production, without leaving the stack.
State-of-the-art quality
Full-weight training at frontier scale, native expert training for MoE architectures, and reinforcement learning validated at the token level. Custom models engineered the way frontier labs build them.
Cost-efficient at scale
Guaranteed throughput on dedicated infrastructure, with on-demand compute when you need it. Concurrent LoRA training and purpose-built stacks keep every GPU working, for industry-leading price-performance.
End-to-end platform
Run concurrent reinforcement learning experiments, validate with high-scale multi-LoRA inference, and merge and quantize for serving at scale. The fastest path from experiment to production, on one platform.
Choose your training method
Both run through a granular Python SDK, with high configurability over every training job.
- Reinforcement learning
Optimize your model against a reward.
GRPO and custom lossesReward and KL tracking per runFrontier open models, full-weight or LoRA - Supervised fine-tuning
Teach the model from your own demonstrations.
Your data, your formatsInstruction and conversation tuningSame SDK, same deployment path
Everything you need to reach production
Your code defines the run. Full configurability, dedicated capacity, and one path from training to serving.
Run many LoRA adapter experiments in parallel inside one training deployment. Converge on your best model in one fast test loop.
We match the computations between training and inference and minimize any discrepancies, keeping training stable even for the largest runs. Reach the strongest version of your model.
Your experiments run on capacity that is completely yours. Predictable performance and predictable cost, with your code defining the run.
Iterate on experiments, deploy multiple versions, and promote the best to production, with a dashboard tracking every run. One stack carries the model straight to serving.




Backed by frontier research
Every run sits on Together's own training and inference research.
Accurary (%) on DeepMath reasoning benchmark
RARO vs VERIFIER-FREE
+10% accuracy
RARO, our Relativistic Adversarial Reasoning Optimization method, learns strong reasoning from expert demonstrations, without verifiers. It outperforms the best verifier-free methods in domains where ground truth doesn't exist, and scales like verifier-based RL.
Learn moreContext parallelism approaches on long-context training
- Together AI (DCT)
- Baseline (LD)
UPipe vs other SOTA Approaches
82.5% less memory
Long-context training hits a memory wall at the attention layer. UPipe processes attention heads in smaller chunks, cutting peak activation memory by up to 82.5% — enabling 5M token context lengths on a single 8×H100 node.
Learn moreContext parallelism approaches on long-context training
- Together AI (DCT)
- Baseline (LD)
FFT Optimizer results
25% less memory
Fine-tuning large models is memory-hungry. Our FFT-based optimizer replaces expensive SVD projections with fast Fourier transforms, reducing optimizer memory by up to 25% with no loss in training quality.
learn more
Production-grade
security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022