Models / Moonshot AI

Moonshot AI

Deploy the latest Kimi models on Together AI. A million-token context, native vision, and frontier agentic coding in open models built for production.

Why Moonshot AI on Together AI?

Designed for production workloads that need 
consistent performance and operational control.

Built for long-horizon agentic work

Kimi sustains multi-step tasks with minimal supervision and navigates large codebases across a million-token context window. Its reasoning modes support reliable multi-step execution.

Native multimodal, drop-in API

Kimi models support vision inputs and are compatible with the OpenAI SDK, so most teams switch with a one-line model change. Move from prototype to production without re-plumbing your stack.

Enterprise-ready from day one

SOC 2 Type II certified, HIPAA compliant, and deployed on US-based infrastructure. Full model ownership with no data retention by default.

Meet the Moonshot AI family

Explore top-performing models across text, image, video, code, and voice.

Chat

Kimi K3

Chat

Kimi K2.7 Code

Chat

Kimi K2.6

Chat

Kimi K2 Thinking

Chat

Kimi K2 Instruct

Chat

Kimi K2 Instruct-0905

Chat

Kimi K2.5

Deployment options

Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.

  • Serverless Inference

  • Provisioned 
Throughput

  • Dedicated Model 
Inference

  • Dedicated Container 
Inference

Serverless Inference

A fully managed real-time or batch inference API with access to dozens of the most popular AI models.

Best for

Variable or unpredictable traffic

Rapid prototyping and iteration

Cost-sensitive or early-stage production workloads

Provisioned 
Throughput

Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit. 

Best for

Production workloads

Reliability guarantees

Predictable pricing

Dedicated Model 
Inference

An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.

Best for

Predictable or steady traffic

Latency-sensitive applications

High-throughput production workloads

Dedicated Container 
Inference

Run inference with your own engine and model on fully-managed, scalable infrastructure.

Best for

Generative media models

Non-standard runtimes

Custom inference pipelines