Kimi K3
Open 3T-class model for long-horizon coding and knowledge work
Pay-as-you-go — no credit card to start
Drop-in OpenAI-compatible API
Match closed model quality at 80% lower cost
Drop-in OpenAI-compatible API
Match closed model quality at 80% lower cost
Served in the EU: your data stays in-region
Endpoint:








Why Kimi K3 on Together AI
We build for production workloads — not just demos
Leading price performance
Full-stack optimizations deliver higher throughput and lower latency than the leading open source implementations on identical hardware.
Run it your way
Serverless scales to your traffic with per token pricing. Dedicated gives you isolated GPUs tuned to your workload.
Backed by research that ships
Inference performance is driven by continuous optimization across kernels, scheduling, and runtime systems.
Kimi K3
Endpoint:
About
Kimi K3 is Moonshot AI's most capable model and the first open model at the 3-trillion-parameter class, with 2.8T total parameters activating 16 of 896 experts per token. It is built on Kimi Delta Attention and Attention Residuals, two architectural changes to how information flows across sequence length and model depth, with native vision and a 1-million-token context window. The model is designed for long-horizon work: sustaining extended engineering sessions across large repositories, carrying multi-step research and document tasks end to end, and reading screenshots, charts, and documents inside the same model. It runs at maximum thinking effort at launch, and weights are released under an open license. Available on Together AI.
Pricing
- Input$3.00/ 1M tokens
- output$15.00/ 1M tokens
- cache$0.30/ 1M cached
Performance benchmark
- FeaturesJSON ModeFunction Calling
- Parameters2.8T
- Context length1.05M
- Input modalitiesTextImage
- Output modalitiesText
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
A fully managed real-time or batch inference API with access to dozens of the most popular AI models.
Best for
Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit.
Best for
An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.
Best for
Run inference with your own engine and model on fully-managed, scalable infrastructure.
Best for
Customers running inference in production
Production-grade
security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022


