Running open models in production: a live walkthrough of our new inference platform
Open-weight models give you real control over quality, performance, and cost. Getting that control into production usually means building a platform team's worth of infrastructure first.
- Deploying a model or adapter from Hugging Face, S3, or local, with ~4x faster warm starts
- Safe rollouts with canary, blue-green, and automatic rollback
- Testing new versions on live traffic with shadow requests and A/B tests
- SLO-driven autoscaling and multi-region deployment behind one stable endpoint
- Live Q&A with the team building it

Nikitha Suryadevara
Staff Product Manager
Together AI

Zain Hasan
Staff AI/ML Engineer
Together AI
