ZAI
Deploy the latest GLM models on Together AI. Open-weight frontier coding, configurable thinking effort, and OpenAI-compatible APIs.
Why ZAI on Together AI?
Designed for production workloads that need consistent performance and operational control.
Frontier agentic coding, open weights
Z.ai's GLM models rank among the top open-source systems for agentic coding and multi-step reasoning. They ship under the MIT license with open weights, so you get frontier performance with full commercial freedom.
Tune compute to the task
Configurable thinking-effort levels let you trade reasoning depth against throughput depending on task complexity. Dial up for complex refactors or down for fast, cost-sensitive calls, all through one OpenAI-compatible API.
Enterprise-ready from day one
SOC 2 Type II certified, HIPAA compliant, and deployed on US-based infrastructure. Full model ownership with no data retention by default.
Meet the ZAI family
Explore top-performing models across text, image, video, code, and voice.
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
A fully managed real-time or batch inference API with access to dozens of the most popular AI models.
Best for
Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit.
Best for
An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.
Best for
Run inference with your own engine and model on fully-managed, scalable infrastructure.
Best for