DeepSeek-V3-0324
Mixture-of-Experts model challenging top AI models at much lower cost. Updated on March 24th, 2025.
Pay-as-you-go — no credit card to start
Drop-in OpenAI-compatible API
Match closed model quality at 80% lower cost
Drop-in OpenAI-compatible API
Match closed model quality at 80% lower cost
Served in the EU: your data stays in-region








Why DeepSeek-V3-0324 on Together AI
We build for production workloads — not just demos
Leading price performance
Full-stack optimizations deliver higher throughput and lower latency than the leading open source implementations on identical hardware.
Run it your way
Serverless scales to your traffic with per token pricing. Dedicated gives you isolated GPUs tuned to your workload.
Backed by research that ships
Inference performance is driven by continuous optimization across kernels, scheduling, and runtime systems.
DeepSeek-V3-0324
Endpoint:
About
DeepSeek-V3-0324 is a strong Mixture-of-Experts language model with 671B parameters, 37B activated per token, designed for efficient inference and cost-effective training. It excels in performance, outpacing other open-source models and rivaling leading closed-source models. Suitable for applications requiring high-quality language understanding and generation.
This endpoint was updated on March 24th, 2025 to use the weights of the improved DeepSeek-V3-0324 model.
Pricing
- Input$1.25/ 1M tokens
- output$1.25/ 1M tokens
Performance benchmark
- Main use casesChatFunction Calling
- FeaturesFunction CallingJSON Mode
- Parameters671B
- Context length131K
- Input modalitiesText
- Output modalitiesText
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
A fully managed real-time or batch inference API with access to dozens of the most popular AI models.
Best for
Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit.
Best for
An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.
Best for
Run inference with your own engine and model on fully-managed, scalable infrastructure.
Best for
Customers running inference in production
Production-grade
security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022


