DeepSeek V4 Flash 0731
Efficient 1M-context model for agentic coding with adjustable reasoning effort
Pay-as-you-go — no credit card to start
Drop-in OpenAI-compatible API
Match closed model quality at 80% lower cost
Drop-in OpenAI-compatible API
Match closed model quality at 80% lower cost
Served in the EU: your data stays in-region
Endpoint:








Why DeepSeek V4 Flash 0731 on Together AI
We build for production workloads — not just demos
Leading price performance
Full-stack optimizations deliver higher throughput and lower latency than the leading open source implementations on identical hardware.
Run it your way
Serverless scales to your traffic with per token pricing. Dedicated gives you isolated GPUs tuned to your workload.
Backed by research that ships
Inference performance is driven by continuous optimization across kernels, scheduling, and runtime systems.
DeepSeek V4 Flash 0731
Endpoint:
About
DeepSeek V4 Flash 0731 is the official release of DeepSeek's efficiency-focused V4 Flash model, superseding the April preview with substantially stronger agentic capability. It is a Mixture-of-Experts model with 284B total parameters and 13B active per token, supporting a 1M-token context window on a hybrid attention design built to keep ultra-long-context inference cheap, and it ships with a speculative decoding module attached for faster generation. Reasoning effort is adjustable per request across low, high, and max levels, letting the same deployment serve quick responses and deep deliberation. DeepSeek reports the release outperforms the much larger V4 Pro preview on its published agentic benchmark set despite far fewer activated parameters. Released under the MIT license. Available on Together AI.
Pricing
- Input$0.14/ 1M tokens
- output$0.28/ 1M tokens
- cache$0.03/ 1M cached
Performance benchmark
- Parameters284B
- Context length1M
- Input modalitiesText
- Output modalitiesText
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
A fully managed real-time or batch inference API with access to dozens of the most popular AI models.
Best for
Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit.
Best for
An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.
Best for
Run inference with your own engine and model on fully-managed, scalable infrastructure.
Best for
Customers running inference in production
Production-grade
security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
SOC 2 Type II
ISO 27001:2022


