Models / MiniMax AI
LLM
Reasoning
Chat

MiniMax M1 40K

456B-parameter hybrid MoE reasoning model with 40K thinking budget, lightning attention, and 1M token context for efficient reasoning and problem-solving tasks.

About model

MiniMax M1 40K is a large-scale hybrid-attention reasoning model suitable for complex tasks requiring long input processing and extensive thinking. It supports a context length of 1 million tokens and enables efficient scaling of test-time compute. Ideal for users needing advanced language modeling capabilities for tasks like software engineering and mathematical reasoning.

To run this model you first need to deploy it on a Dedicated Endpoint.

Performance benchmarks

Model

FrontierMath Tier 4

GPQA Diamond

HLE

SciCode

GDPval-AA

Terminal-Bench 2.1

Agent Arena

FrontierCode

DeepSWE

70.0%

Related open-source models

Competitor closed-source models

Claude Fable 5

87.8%

92.6%

53%

60%

62%

85%

53.5%

70%

Claude Opus 5

73.2%

93.2%

53%

56%

68%

89%

+16.4pp

53.4%

74%

GPT-5.6 Sol

82.9%

94.1%

47%

56%

61%

88%

47.5%

73%

Grok 4.5

24.4%

93.1%

40%

54%

51%

82%

+4.2pp

42.4%

54%

GPT-5.6 Luna

61.0%

91.1%

37%

53%

54%

81%

39.8%

67%

  • Model card

    Architecture Overview:
    • Hybrid Mixture-of-Experts with 456 billion total parameters and 45.9 billion activated per token
    • Lightning attention mechanism for efficient test-time compute scaling
    • 1 million token context window for extensive document processing and analysis
    • Optimized hybrid attention design balancing performance with computational efficiency

    Training Methodology:
    • Large-scale reinforcement learning on diverse reasoning and engineering problems
    • CISPO algorithm optimization for efficient importance sampling weight management
    • 40K thinking budget providing balanced reasoning capabilities with computational efficiency
    • Trained on diverse problems from mathematical reasoning to real-world software engineering

    Performance Characteristics:
    • Efficient test-time compute scaling with lightning attention mechanism
    • Strong performance on AIME 2024 (83.3), SWE-bench Verified (55.6), and coding benchmarks
    • Superior efficiency compared to larger reasoning models while maintaining quality
    • Optimized for tasks requiring substantial reasoning with moderate computational budgets

  • Prompting

    Reasoning Capabilities:
    • Advanced reasoning model with 40K thinking budget for efficient problem-solving
    • System/user/assistant format optimized for reasoning chains and complex tasks
    • Lightning attention enables efficient scaling while maintaining reasoning quality
    • Balanced approach to extensive reasoning with computational efficiency considerations

    Optimization Settings:
    • Temperature 1.0 and top_p 0.95 for creativity and logical coherence balance
    • General scenarios: "You are a helpful assistant" for broad applications
    • Mathematical reasoning: Step-by-step reasoning with structured output formatting
    • Code generation: Comprehensive web development and engineering assistance

    Efficiency Features:
    • Function calling capabilities for structured external integrations
    • Efficient reasoning budget allocation for cost-effective complex problem-solving
    • Strong performance across diverse domains with moderate computational requirements
    • Optimal balance between reasoning capability and resource utilization

  • Applications & use cases

    Efficient Reasoning Applications:
    • Mathematical problem-solving and competition-level tasks with budget efficiency
    • Software engineering and coding assistance requiring moderate reasoning depth
    • Long-context document analysis with 1M token processing capability
    • Multi-step reasoning tasks with computational efficiency requirements

    Business & Development:
    • Cost-effective reasoning applications for business problem-solving
    • Development environments requiring advanced AI assistance with budget considerations
    • Educational applications requiring step-by-step reasoning and explanation
    • Research and analysis tasks with moderate complexity and reasoning requirements

    Balanced Applications:
    • Applications requiring advanced reasoning capabilities without premium computational costs
    • Complex problem-solving scenarios with efficiency and performance balance
    • Next-generation AI agents for real-world challenges with resource optimization
    • Advanced decision-making systems requiring substantial reasoning with cost-effectiveness

Related models
  • Model provider
    MiniMax AI
  • Type
    LLM
    Reasoning
    Chat
  • Main use cases
    Chat
    Reasoning
  • Deployment
    Dedicated
  • Parameters
    456B
  • Activated parameters
    45.9B
  • Context length
    1M
  • Input modalities
    Text
  • Output modalities
    Text