Qwen3.8 27B
Compact vision-language model for agentic work with flexible thinking control

About model
Qwen3.8 27B is the compact dense model of Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, it is a native vision-language model with substantial gains across coding, professional work, research, and long-horizon agentic tasks, designed to carry complex multi-step work through to completion with greater reliability. Thinking is on by default and can be disabled per request, reasoning depth is tunable across three effort levels, and reasoning context from prior messages is retained by default. The 262K-token native context extends up to 1M tokens. Released under Apache 2.0, with fine-tuning supported on Together AI. Available on Together AI.
27B
Compact vision-language model with full capability on every token
262K
Extensible up to 1M tokens for long-horizon tasks
3
xhigh, medium, and low, with thinking on by default and per-request disable
- Agent Execution: Stronger autonomous planning and better handling of environment feedback for reliable end-to-end task completion
- Vision-Language Understanding: Native image and document understanding, from STEM diagrams and charts to full document layouts
- Flexible Thinking Control: Thinking on by default with per-request disable, three reasoning effort levels, and preserved reasoning context across turns
- Production-Ready Infrastructure: 99.9% SLA, available on Together AI dedicated infrastructure with fine-tuning support
Model | FrontierMath Tier 4 | GPQA Diamond | HLE | SciCode | GDPval-AA | Terminal-Bench 2.1 | Agent Arena | FrontierCode | DeepSWE |
|---|---|---|---|---|---|---|---|---|---|
Qwen3.8 27B | 89.2 | 30.8 | Related open-source models | Competitor closed-source models | |||||
87.8% | 92.6% | 53% | 60% | 62% | 85% | 53.5% | 70% | ||
73.2% | 93.2% | 53% | 56% | 68% | 89% | +16.4pp | 53.4% | 74% | |
82.9% | 94.1% | 47% | 56% | 61% | 88% | 47.5% | 73% | ||
24.4% | 93.1% | 40% | 54% | 51% | 82% | +4.2pp | 42.4% | 54% | |
61.0% | 91.1% | 37% | 53% | 54% | 81% | 39.8% | 67% |
Model card
Architecture Overview:
• 27B parameter dense causal language model with a vision encoder
• 64 layers in a hybrid layout: 16 blocks of three Gated DeltaNet layers followed by one gated attention layer, each paired with feed-forward networks
• Gated DeltaNet with 48 linear attention heads for V and 16 for QK; gated attention with 24 query heads and 4 KV heads
• Multi-Token Prediction trained with multiple steps; 248K vocabulary
• 262,144-token native context, extensible up to 1,000,000 tokens
Training Methodology:
• Pre-training and post-training built on the Qwen3.5 architectural foundation
• Broader compatibility with popular agent harnesses and development tools than prior generations
• Preserved thinking retains reasoning blocks across historical messages, improving decision consistency and cache utilization in agent scenarios
Performance Characteristics:
• Terminal coding (Terminal Bench 2.1, completing real tasks in a command-line environment): 73.0, up from 63.4 for Qwen3.6-27B
• Agentic coding (SWE-bench Pro: 61.7; DeepSWE 1.1: 42.2, up from 13.3; NL2Repo repository generation: 42.3)
• Long-horizon professional work (CoWorkBench, Qwen in-house, spanning finance, law, medical, and other domains: 70.7; JobBench professional tasks: 33.4)
• Reasoning and knowledge (GPQA Diamond: 89.2; Humanity's Last Exam: 30.8; LiveCodeBench v6: 90.3; IFBench instruction following: 79.5)
• Computer, browser, and mobile use (OSWorld-Verified: 84.3, up from 63.9; WebArena-Verified: 64.8; AndroidWorld: 81.9)
• Multimodal engineering and analysis (SWE-MM: 38.6; Vision2Web: 62.9; MathVision: 90.0, or 94.6 with code interpreter; CharXiv RQ chart reasoning: 83.7; OmniDocBench document parsing: 91.1; RealWorldQA: 85.9)
Prompting
Together AI API Access:
• Access Qwen3.8 27B via Together AI APIs using the endpoint Qwen/Qwen3.8-27B
• Authenticate using your Together AI API key in request headers
• Thinking is on by default and can be disabled per request for direct instruct-style responses
• Set reasoning_effort to xhigh (default), medium, or low; preserved thinking retains reasoning context across turns by default
• Recommended sampling: temperature 1.0 with top-p 0.95 in thinking mode; temperature 0.7 with top-p 0.8 in non-thinking mode
• For agentic tasks, allow generous budgets: up to 262K reasoning tokens and 131K final-response tokens
• Available on Together AI dedicated infrastructure with fine-tuning support
Applications & use cases
Agentic Software Engineering:
• Run terminal and repository agents that carry multi-step engineering tasks to completion
• Tune reasoning effort per step: low for routine calls, xhigh for hard planning
• Fine-tune on Together AI to specialize the model for your codebase and workflows
Computer & Device Automation:
• Automate desktop, browser, and mobile workflows grounded in screenshots of real interfaces
• Build UI agents that read application state visually and act on it
• Recreate and test application flows across platforms
Document & Visual Analysis:
• Parse documents, charts, and STEM diagrams alongside text in one request
• Run visual math and chart reasoning within analysis pipelines
• Combine document intelligence with long context for large mixed working sets
- TypeChatReasoningVision
- Fine tuningSupported
- DeploymentDedicated
- Parameters27B
- Context length262.1K
- Input modalitiesTextImage
- Output modalitiesText
- External link
- CategoryChat
