Qwen3 235B A22B Instruct 2507 FP8 Throughput
235B MoE model with 22B activation featuring enhanced instruction following, reasoning, and 262K context for cost-efficient high-throughput inference.
About model
Enhanced Qwen3 model optimized for serverless inference with superior price-performance.
Performance benchmarks
Model | AIME 2025 | GPQA Diamond | HLE | LiveCodeBench | MATH500 | SWE-bench verified |
|---|---|---|---|---|---|---|
Qwen3 235B A22B Instruct 2507 FP8 Throughput | 65.9% | Related open-source models | Competitor closed-source models | |||
92.6% | 53% | |||||
93.2% | 53% | |||||
94.1% | 47% | |||||
93.1% | 40% |
API usage
Endpoint:
Related models
- TypeChatReasoning
- Main use casesChatSmall & FastMedium General PurposeFunction Calling
- FeaturesFunction CallingJSON Mode
- DeploymentDedicatedServerless
- Parameters235B
- Context length262K
- Input price
$0.20 / 1M tokens
- Output price
$0.60 / 1M tokens
- Input modalitiesText
- Output modalitiesText
- ReleasedJuly 22, 2025
- Quantization levelFP8
- External link
- CategoryChat