Models / Qwen
Chat
Reasoning
Vision

Qwen3.8 27B

Compact vision-language model for agentic work with flexible thinking control

About model

Qwen3.8 27B is the compact dense model of Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, it is a native vision-language model with substantial gains across coding, professional work, research, and long-horizon agentic tasks, designed to carry complex multi-step work through to completion with greater reliability. Thinking is on by default and can be disabled per request, reasoning depth is tunable across three effort levels, and reasoning context from prior messages is retained by default. The 262K-token native context extends up to 1M tokens. Released under Apache 2.0, with fine-tuning supported on Together AI. Available on Together AI.

Dense Parameters

27B

Compact vision-language model with full capability on every token

Native Context

262K

Extensible up to 1M tokens for long-horizon tasks

Reasoning Effort Levels

3

xhigh, medium, and low, with thinking on by default and per-request disable

Model key capabilities
  • Agent Execution: Stronger autonomous planning and better handling of environment feedback for reliable end-to-end task completion
  • Vision-Language Understanding: Native image and document understanding, from STEM diagrams and charts to full document layouts
  • Flexible Thinking Control: Thinking on by default with per-request disable, three reasoning effort levels, and preserved reasoning context across turns
  • Production-Ready Infrastructure: 99.9% SLA, available on Together AI dedicated infrastructure with fine-tuning support
Performance benchmarks

Model

FrontierMath Tier 4

GPQA Diamond

HLE

SciCode

GDPval-AA

Terminal-Bench 2.1

Agent Arena

FrontierCode

DeepSWE

89.2

30.8

Related open-source models

Competitor closed-source models

Claude Fable 5

87.8%

92.6%

53%

60%

62%

85%

53.5%

70%

Claude Opus 5

73.2%

93.2%

53%

56%

68%

89%

+16.4pp

53.4%

74%

GPT-5.6 Sol

82.9%

94.1%

47%

56%

61%

88%

47.5%

73%

Grok 4.5

24.4%

93.1%

40%

54%

51%

82%

+4.2pp

42.4%

54%

GPT-5.6 Luna

61.0%

91.1%

37%

53%

54%

81%

39.8%

67%

  • Model card

    Architecture Overview:
    • 27B parameter dense causal language model with a vision encoder
    • 64 layers in a hybrid layout: 16 blocks of three Gated DeltaNet layers followed by one gated attention layer, each paired with feed-forward networks
    • Gated DeltaNet with 48 linear attention heads for V and 16 for QK; gated attention with 24 query heads and 4 KV heads
    • Multi-Token Prediction trained with multiple steps; 248K vocabulary
    • 262,144-token native context, extensible up to 1,000,000 tokens

    Training Methodology:
    • Pre-training and post-training built on the Qwen3.5 architectural foundation
    • Broader compatibility with popular agent harnesses and development tools than prior generations
    • Preserved thinking retains reasoning blocks across historical messages, improving decision consistency and cache utilization in agent scenarios

    Performance Characteristics:
    • Terminal coding (Terminal Bench 2.1, completing real tasks in a command-line environment): 73.0, up from 63.4 for Qwen3.6-27B
    • Agentic coding (SWE-bench Pro: 61.7; DeepSWE 1.1: 42.2, up from 13.3; NL2Repo repository generation: 42.3)
    • Long-horizon professional work (CoWorkBench, Qwen in-house, spanning finance, law, medical, and other domains: 70.7; JobBench professional tasks: 33.4)
    • Reasoning and knowledge (GPQA Diamond: 89.2; Humanity's Last Exam: 30.8; LiveCodeBench v6: 90.3; IFBench instruction following: 79.5)
    • Computer, browser, and mobile use (OSWorld-Verified: 84.3, up from 63.9; WebArena-Verified: 64.8; AndroidWorld: 81.9)
    • Multimodal engineering and analysis (SWE-MM: 38.6; Vision2Web: 62.9; MathVision: 90.0, or 94.6 with code interpreter; CharXiv RQ chart reasoning: 83.7; OmniDocBench document parsing: 91.1; RealWorldQA: 85.9)

  • Prompting

    Together AI API Access:
    • Access Qwen3.8 27B via Together AI APIs using the endpoint Qwen/Qwen3.8-27B
    • Authenticate using your Together AI API key in request headers
    • Thinking is on by default and can be disabled per request for direct instruct-style responses
    • Set reasoning_effort to xhigh (default), medium, or low; preserved thinking retains reasoning context across turns by default
    • Recommended sampling: temperature 1.0 with top-p 0.95 in thinking mode; temperature 0.7 with top-p 0.8 in non-thinking mode
    • For agentic tasks, allow generous budgets: up to 262K reasoning tokens and 131K final-response tokens
    • Available on Together AI dedicated infrastructure with fine-tuning support

  • Applications & use cases

    Agentic Software Engineering:
    • Run terminal and repository agents that carry multi-step engineering tasks to completion
    • Tune reasoning effort per step: low for routine calls, xhigh for hard planning
    • Fine-tune on Together AI to specialize the model for your codebase and workflows

    Computer & Device Automation:
    • Automate desktop, browser, and mobile workflows grounded in screenshots of real interfaces
    • Build UI agents that read application state visually and act on it
    • Recreate and test application flows across platforms

    Document & Visual Analysis:
    • Parse documents, charts, and STEM diagrams alongside text in one request
    • Run visual math and chart reasoning within analysis pipelines
    • Combine document intelligence with long context for large mixed working sets

Related models
  • Model provider
    Qwen
  • Type
    Chat
    Reasoning
    Vision
  • Fine tuning
    Supported
  • Deployment
    Dedicated
  • Parameters
    27B
  • Context length
    262.1K
  • Input modalities
    Text
    Image
  • Output modalities
    Text