🚀 DeepSeek V4 Pro 0813 vs. GPT-5.6 Sol on DeepSWE →

🤝 Together AI & Y Combinator announce partnership to deliver the first dedicated YC GPU cluster →

⚡ On-demand B200s now available on Together GPU Clusters →

🚀 Now serving MiniMax-M3 for efficient inference →

  • Inference

    • Serverless Inference

      High-performance inference as APIs

    • Batch Inference

      Inference for batch workloads

    • Provisioned Throughput

      Token-based capacity with SLAs

    • Dedicated Model Inference

      Inference on custom hardware

    • Dedicated Container Inference

      Inference for custom models

    MiniMax M3
    Gemma 4 31B
    DeepSeek V4 Pro
    GLM-5.2
    White circular shape with uneven edges and three extended finger-like projections on a black background.
    kimi K2.7 Code
    OpenAI logo with a symmetrical abstract geometric knot design.
    gpt-oss-120B

    Model library

    Explore the top open-source models

  • Compute

    Accelerated Compute

    • GPU Clusters

      Reliable GPU clusters at scale

    • AI Factory

      Custom infrastructure at frontier scale

    Developer Environments

    • Sandbox

      Build development environments for AI

    Storage

    • Managed Storage

      Store model weights & data securely

    • GB300

    • GB200

    • B200

    • H200

    • H100

  • Model Shaping

    • Custom Training

      From reinforcement learning to full control

    • Fine-Tuning

      Shape models with your data

    • Evaluations

      Measure model quality

    kimi K2.7 Code
    Gemma 4 31B-it FP8
    GLM 5.1 FP4
    OpenAI logo with a symmetrical abstract geometric knot design.
    gpt-oss-120b
    Qwen3.5 397B A17b
    Llama 4 Maverick

    Model library

    Fine-tune top open-source models

  • Research

    • Research

      Systems research for production AI

    • Research blog

      All our research publications

    Featured publications

    • FlashAttention

    • ATLAS

    • Kernel Collection

    • ThunderKittens

    • DSGym

    Show all
  • Developers

    • Documentation

      Technical docs for Together AI

    • Demos

      Our open-source demo apps

    • Cookbooks

      Practical implementation guides

    • Voice Agents

      Build voice agents for production

    • Open-source AI

      Build better with open models

    • Model Library

    • Playground

    • Together Chat

    • Which LLM to use

    • Open-source ROI calculator

  • Company

    Resources

    • Customer stories

      Testimonials from AI Natives

    • Startup accelerator

      Build and scale your startup

    • Customer support

      Find answers to your questions

    • Blog

      Our latest news & blog posts

    • Events

      Explore our events calendar

    Company

    • About

      Get to know us

    • Careers

      Join our mission

    • Press

      Together in the news

  • Pricing

    • Serverless Inference

      High-performance inference as APIs

    • Batch Inference

      Inference for batch workloads

    • Provisioned Throughput

      Token-based capacity with SLAs

    • Dedicated Model Inference

      Inference on custom hardware

    • Dedicated Container Inference

      Inference for custom models

    MiniMax M3
    Gemma 4 31B
    DeepSeek V4 Pro
    GLM-5.2
    White circular shape with uneven edges and three extended finger-like projections on a black background.
    kimi K2.7 Code
    OpenAI logo with a symmetrical abstract geometric knot design.
    gpt-oss-120B

    Model library

    Explore the top open-source models

  • Accelerated Compute

    • GPU Clusters

      Reliable GPU clusters at scale

    • AI Factory

      Custom infrastructure at frontier scale

    Developer Environments

    • Sandbox

      Build development environments for AI

    Storage

    • Managed Storage

      Store model weights & data securely

    • GB300

    • GB200

    • B200

    • H200

    • H100

    • Custom Training

      From reinforcement learning to full control

    • Fine-Tuning

      Shape models with your data

    • Evaluations

      Measure model quality

    kimi K2.7 Code
    Gemma 4 31B-it FP8
    GLM 5.1 FP4
    OpenAI logo with a symmetrical abstract geometric knot design.
    gpt-oss-120b
    Qwen3.5 397B A17b
    Llama 4 Maverick

    Model library

    Fine-tune top open-source models

    • Research

      Systems research for production AI

    • Research blog

      All our research publications

    Featured publications

    • FlashAttention

    • ATLAS

    • Kernel Collection

    • ThunderKittens

    • DSGym

    Show all
    • Documentation

      Technical docs for Together AI

    • Demos

      Our open-source demo apps

    • Cookbooks

      Practical implementation guides

    • Voice Agents

      Build voice agents for production

    • Open-source AI

      Build better with open models

    • Model Library

    • Playground

    • Together Chat

    • Which LLM to use

    • Open-source ROI calculator

  • Resources

    • Customer stories

      Testimonials from AI Natives

    • Startup accelerator

      Build and scale your startup

    • Customer support

      Find answers to your questions

    • Blog

      Our latest news & blog posts

    • Events

      Explore our events calendar

    Company

    • About

      Get to know us

    • Careers

      Join our mission

    • Press

      Together in the news

Contact sales
Contact sales
Sign in
Explore Research

Research blog

All
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Inference
DeepCoder: A Fully Open-Source 14B Coder at O3-mini Level

Michael Luo*, Sijun Tan*, Roy Huang*, Ameen Patel*, Alpay Ariyak*, Qingyang Wu*, Xiaoxiang Shi, Rachel Xin, Colin Cai, Maurice Weber, Ce Zhang, Li Erran Li, Raluca Ada Popa, Ion Stoica

Text on dark background: together.ai DeepCoder: A Fully Open-Source 14B Coder at O3-mini Level.
Kernels
ThunderKittens Now Optimized for NVIDIA Blackwell GPUs

Benjamin Spector, Aaryan Singhal, Dan Fu, Chris Ré

Inference
Minions: embracing small LMs, shifting compute on-device, and cutting cloud costs in the process

Avanika Narayan*, Dan Biderman*, Sabri Eyuboglu*, Avner May, Scott Linderman, James Zou, Christopher Ré

Graph showing accuracy vs remote cost per problem for local Llama and remote GPT-4o models with cost and accuracy labels.
Model Shaping
Long Context Fine-Tuning: A Technical Deep Dive

George Grigorev, Zain Hasan, Max Ryabinin

together.ai Long Context Fine-Tuning: A Technical Deep Dive on dark background with blue dots.
Previous
Load more
10 / 20

No search result

Try expanding your search or changing the filters.

Be at the forefront of AI innovation

From optimized training and model shaping to large-scale production inference

See open roles
  • Products

    • Accelerated Compute

    • Serverless Inference

    • Provisioned Throughput

    • Dedicated Inference

    • Fine-Tuning

    • Sandbox

    • Evaluations

  • Models

    See all models

    DeepSeek

    Meta

    Qwen

    Google

    OpenAI

    Mistral AI

    Custom models

  • Developers

    • Research

    • Docs

    • Open-source AI

    • OSS ROI calculator

    Pricing

    • Pricing overview

    • Inference

    • Fine-Tuning

    • GPU Clusters

  • Resources

    • Blog

    • About us

    • Careers

    • Customer Stories

    • Support

  • Privacy Policy

  • Terms of service

  • Cookie Policy

  • Consent Preferences

© 2026 Together AI. All Rights Reserved.