Llama 4 Scout
SOTA 109B model with 17B active params & large context, excelling at multi-document analysis, codebase reasoning, and personalized tasks.
About model
Together AI offers day 1 support for the new Llama 4 multilingual vision models that can analyze multiple images and respond to queries about them.
Register for a Together AI account to get an API key. New accounts come with free credits to start. Install the Together AI library for your preferred language.
Model | FrontierMath Tier 4 | GPQA Diamond | HLE | SciCode | GDPval-AA | Terminal-Bench 2.1 | Agent Arena | FrontierCode | DeepSWE |
|---|---|---|---|---|---|---|---|---|---|
Llama 4 Scout | 51.8% | Related open-source models | Competitor closed-source models | ||||||
87.8% | 92.6% | 53% | 60% | 62% | 85% | 53.5% | 70% | ||
73.2% | 93.2% | 53% | 56% | 68% | 89% | +16.4pp | 53.4% | 74% | |
82.9% | 94.1% | 47% | 56% | 61% | 88% | 47.5% | 73% | ||
24.4% | 93.1% | 40% | 54% | 51% | 82% | +4.2pp | 42.4% | 54% | |
61.0% | 91.1% | 37% | 53% | 54% | 81% | 39.8% | 67% |
How to use model
Input
Output
Function Calling
Input
Output
Query models with multiple images
Currently this model supports 5 images as input.
Input
Output
Model card
- Model String: meta-llama/Llama-4-Scout-17B-16E-Instruct
- Specs:
- 17B active parameters (109B total)
- 16-expert MoE architecture
- 327,680 context length (will be increased to 10M)
- Support for 12 languages: Arabic, English, French, German, Hindi, Indonesian, Italian, Portuguese, Spanish, Tagalog, Thai, and Vietnamese
- Multimodal capabilities (text + images)
- Support Function Calling
- Best for: Multi-document analysis, codebase reasoning, and personalized tasks
- Knowledge Cutoff: August 2024
Applications & use cases
- Multi-document summarization for legal/financial analysis: Analyze multiple legal contracts or financial statements simultaneously, identifying key terms, inconsistencies, and patterns across documents to generate comprehensive summaries and risk assessments.
- Personalized task automation using years of user data: Create tailored automation workflows by analyzing an individual's historical data patterns, communication style, and preferences, enabling highly personalized digital assistants that adapt to specific user needs.
- Efficient image parsing for multimodal applications: Process and understand image content in conjunction with text to power applications like visual search, content moderation, and accessibility features that require understanding the relationship between visual and textual elements.
- TypeChatVision
- Main use casesChatFunction CallingVision
- FeaturesFunction Calling
- Fine tuningSupported
- DeploymentDedicated
- Parameters109B
- Context length1M
- Input price
$0.18 / 1M tokens
- Output price
$0.59 / 1M tokens
- Input modalitiesTextImage
- Output modalitiesText
- ReleasedApril 2, 2025
- Last updatedFebruary 5, 2026
- Quantization levelFP16
- External link
- CategoryChat
