Llama 4 Scout

4.55
Meta open-source LLM allowing commercial use, local execution, and fine-tuning.
Advertisement 728 × 90
CompanyMeta
Context192K
Released2025-04
Updated2026-08-06

Llama 4 Scout Overview

Llama 4 Scout is a natively multimodal large language model released by Meta on April 5, 2025. It employs a Mixture-of-Experts (MoE) architecture with 109 billion total parameters, activating 17 billion parameters per inference across 16 experts. The model features a 10 million token context window and can process up to 5 images simultaneously. Scout supports 12 languages: Arabic, English, French, German, Hindi, Indonesian, Italian, Portuguese, Spanish, Tagalog, Thai, and Vietnamese. With Int4 quantization, it runs efficiently on a single NVIDIA H100 GPU, making it ideal for multi-document summarization, codebase reasoning, and visual understanding tasks.

Llama 4 Scout Pricing

PlanPriceDescription
API via DeepInfra Input: $0.08 / Output: $0.30 per 1M tokens Hosted through DeepInfra with support for a 10M-token context window. Offers highly competitive pricing, making it ideal for large-scale AI applications and cost-efficient deployments.
API via OpenRouter Input: $0.10 / Output: $0.30 per 1M tokens Access Llama 4 Scout through the OpenRouter platform with standard token-based pricing. Provides a convenient way to integrate the model into various AI applications.
API via Groq Input: $0.11 / Output: $0.34 per 1M tokens Powered by Groq’s high-performance inference platform, delivering speeds of up to approximately 446 tokens/s. Ideal for real-time applications requiring fast responses and low latency.
Self-Hosted Deployment (Open Weights) Free (hardware required) Available under the Meta Llama 4 Community License. Uses a 109B-parameter MoE architecture with 17B active parameters, enabling private deployments and customized enterprise solutions.

Llama 4 Scout Key Features

1. Native Multimodality: Natively supports integrated text and image understanding and reasoning
2. Ultra-Long Context: 10 million token context window for processing massive documents or multi-turn conversations
3. MoE Architecture: 16 experts, 109B total parameters, only 17B activated per inference for optimal performance and efficiency
4. Multilingual Support: Covers 12 languages with consistent cross-lingual performance
5. Efficient Deployment: Runs on a single H100 GPU with Int4 quantization
6. Instruction-Tuned: Optimized for multilingual chat, captioning, and image understanding tasks

Summary

1.Developers & Researchers: Individuals or teams seeking high-performance open-source models with limited GPU resources
2. Enterprise Users: Organizations needing multi-document summarization, codebase analysis, and multilingual customer support
3.RAG System Developers: Developers building RAG pipelines that require large-scale knowledge base retrieval and processing
4.AI Product Teams: Teams aiming for rapid prototyping and production deployment

Comments (0)

Leave a comment

Advertisement 728 × 90

Meta Model Comparison

Model Context Pricing API Released Global Heat
Llama 4 Scout
192K Freemium YES 2025-04
91/100
1M YES 2026-08
89/100
1M YES 2026-08
1M YES 2026-07
88/100

Similar Models

Muse Spark 1.2
89
Muse Spark 1.2 landed just a month after version 1.1. This release is more focused. Meta is putting more weight behind code generation, software engineering, and its new terminal-based agent, Muse Code.
Meta
Muse Spark 1.1
88
Meta’s multimodal reasoning model optimized for AI Agents, coding, tool use, and complex task automation.
Meta
Muse Spark 1.2 Contributor
The first thing that caught my attention about Muse Spark 1.2 Contributor wasn’t the low price. It was the question behind it: why is Meta willing to make it this cheap? The answer is in the terms. If you let Meta use your interactions to improve its models, you get much cheaper API access.
Meta
Claude Opus 5
100
Anthropic’s flagship Claude model built for complex reasoning, AI Agents, software development, and enterprise knowledge workflows.
Anthropic
GPT-5.5 Pro
99
I’ve been using GPT-5.5 Pro with a few colleagues for the past several weeks, mostly on the kinds of jobs where regular chatbots tend to fall apart: long documents, messy research questions, code debugging, and tasks that need more than one or two steps of reasoning.
OpenAI
GPT-5.5
99
OpenAI flagship model, a multimodal AI natively supporting text, images, and audio.
OpenAI
GPT-5.6 Luna
99
An efficient GPT-5.6 series model optimized for speed, cost efficiency, and intelligent performance across large-scale AI applications.
OpenAI
Claude Opus 4.8
98
Claude Opus 4.8 is Anthropic's most powerful Opus-series model, featuring multimodal input, reasoning, and a 1M-token context window, excelling in complex reasoning and coding.
Anthropic

Related News