GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท llama-4-scout

Llama 4 Scout

v2025-04

Meta

Modelopen-sourceefficientedge-deploymentlow-latency
85
Strong
About This Model

Meta's efficient Llama 4 model (released April 5, 2025): a natively multimodal mixture-of-experts with 109B total / 17B active parameters (16 experts) and an industry-leading 10M-token context window, deployable on a single H100-class GPU. Optimized for speed and cost-sensitive applications requiring open-weight flexibility. Now a legacy line: Meta has shipped no new open weights since Scout/Maverick and pivoted to closed models with Muse Spark (April 2026).

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

Efficient performance optimized for speed and resource usage. Good balance for edge deployment and cost-sensitive applications.

task accuracy code

Industry-standard coding benchmarks

Evidence
HumanEval Benchmark โ€” 42% pass rate (estimated)
mediumVerified: 2026-07-09
task accuracy reasoning

Mathematical reasoning benchmarks

Evidence
MATH Benchmark โ€” 52% on mathematical reasoning tasks
mediumVerified: 2026-07-09
task accuracy general

Knowledge testing benchmarks

Evidence
MMLU Benchmark โ€” 57.2% on multitask language understanding
highVerified: 2026-07-09
output consistency

Internal testing with repeated prompts

Evidence
Meta Internal Testing โ€” Good consistency for typical tasks
mediumVerified: 2026-07-09
latency p50

Median latency on recommended hardware

Evidence
Community benchmarking โ€” ~0.6s on standard hardware
highVerified: 2026-07-09
latency p95

95th percentile response time

Evidence
Community benchmarking โ€” p95 latency ~1.2s
highVerified: 2026-07-09
context window

Official specification

Evidence
Meta AI Blog - The Llama 4 herd โ€” Industry-leading 10M-token context window
Hugging Face - meta-llama/Llama-4-Scout-17B-16E โ€” 10M context confirmed on model card; hosted API providers typically expose far less (128K-1M)
highVerified: 2026-07-09
uptime

User-controlled deployment

Evidence
Self-hosted model โ€” Uptime depends on hosting infrastructure
mediumVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Good baseline security with self-hosted deployment providing full control. Smaller model may have slightly lower resistance than Behemoth.

prompt injection resistance

Testing against prompt injection attacks

Evidence
Meta Safety Testing โ€” Good baseline resistance, additional safeguards recommended
mediumVerified: 2026-07-09
jailbreak resistance

Testing against adversarial prompts

Evidence
Meta Safety Evaluations โ€” Built-in safety mechanisms
mediumVerified: 2026-07-09
data leakage prevention

Analysis of deployment model

Evidence
Self-hosted deployment โ€” Full control over data in self-hosted deployments
highVerified: 2026-07-09
output safety

Safety testing

Evidence
Meta Safety Benchmarks โ€” Safety training applied
mediumVerified: 2026-07-09
api security

Review of deployment practices

Evidence
Deployment documentation โ€” Security depends on deployment
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Exceptional privacy with self-hosted deployment. Full control over all data aspects.

data residency

Analysis of deployment model

Evidence
Open-source model โ€” Full control over data location
highVerified: 2026-07-09
training data optout

Analysis of data flow

Evidence
Self-hosted model โ€” No data sent to Meta
highVerified: 2026-07-09
data retention

Analysis of deployment model

Evidence
Self-hosted deployment โ€” Full control over retention
highVerified: 2026-07-09
pii handling

Review of deployment architecture

Evidence
Self-hosted deployment โ€” Full PII control
highVerified: 2026-07-09
compliance certifications

Review of deployment options

Evidence
Self-hosted model โ€” Compliance through deployment infrastructure
highVerified: 2026-07-09
zero data retention

Analysis of deployment model

Evidence
Self-hosted deployment โ€” Complete control over data
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Strong transparency as open-source model. Good documentation and customizable guardrails.

explainability

Evaluation of reasoning transparency

Evidence
Model Behavior โ€” Good explanations for typical tasks
mediumVerified: 2026-07-09
hallucination rate

Community evaluation

Evidence
Community Testing โ€” Moderate hallucination rate
mediumVerified: 2026-07-09
bias fairness

Evaluation on bias benchmarks

Evidence
Meta Responsible AI Report โ€” Bias testing applied
mediumVerified: 2026-07-09
uncertainty quantification

Qualitative assessment

Evidence
Model Behavior โ€” Reasonable uncertainty expression
mediumVerified: 2026-07-09
model card quality

Review of documentation

Evidence
Meta Model Card โ€” Comprehensive model card
highVerified: 2026-07-09
training data transparency

Review of technical documentation

Evidence
Meta Technical Report โ€” Good transparency on training
highVerified: 2026-07-09
guardrails

Review of safety systems

Evidence
Open-source implementation โ€” Transparent, customizable safety
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Good operational maturity with strong ecosystem. Easier to deploy than Behemoth due to smaller size.

api design quality

Review of API design

Evidence
Meta Documentation โ€” Standard inference API
highVerified: 2026-07-09
sdk quality

Review of SDKs

Evidence
Meta GitHub โ€” Official libraries and community tools
highVerified: 2026-07-09
versioning policy

Review of versioning

Evidence
Meta Release Policy โ€” Clear versioning
highVerified: 2026-07-09
monitoring observability

Review of monitoring tools

Evidence
Community tools โ€” Depends on deployment stack
mediumVerified: 2026-07-09
support quality

Assessment of support

Evidence
Community Support โ€” Active community support
mediumVerified: 2026-07-09
ecosystem maturity

Analysis of ecosystem

Evidence
Open-source ecosystem โ€” Mature ecosystem
highVerified: 2026-07-09
license terms

Review of license

Evidence
Meta Llama License โ€” Permissive commercial license
highVerified: 2026-07-09
Strengths
  • +Fast inference (~0.6s p50) suitable for real-time applications
  • +Fits on a single H100-class GPU (17B active parameters)
  • +Industry-leading 10M-token native context window
  • +Natively multimodal (text + image input) via early fusion
  • +Complete data sovereignty with self-hosted deployment โ€” no data retention or sharing concerns
  • +Open weights with full transparency
  • +Cost-effective for high-volume workloads
Limitations
  • !Moderate accuracy compared to larger models (Maverick, frontier proprietary)
  • !Limited coding capabilities relative to coding-specialized models
  • !Native 10M context rarely exposed by hosted providers (typically capped at 128K-1M)
  • !Requires infrastructure for deployment
  • !Less capable for complex reasoning tasks
  • !No managed API service from Meta
  • !Legacy status: Meta has shipped no new open weights since Llama 4 Scout/Maverick (April 2025) and pivoted to closed models (Muse Spark, April 2026)
Metadata
pricing
input: Self-hosted (infrastructure costs)
output: Self-hosted (infrastructure costs)
notes: Open-weight model. Typically $0.10-0.50 per 1M tokens via optimized deployment or third-party hosts as of July 2026.
last verified: 2026-07-09
context window: 10000000
languages
0: English
1: Spanish
2: French
3: German
4: Italian
5: Portuguese
6: Japanese
7: Korean
8: Chinese
9: 100+ languages
modalities
0: text
1: image (input)
api endpoint: Self-hosted or third-party hosts (Together, Groq, etc.)
open source: true
architecture: Mixture-of-experts (16 experts), natively multimodal via early fusion
parameters: 109B total / 17B active

Use Case Ratings

code generation

Adequate for basic coding tasks. Fast inference makes it suitable for development tools.

customer support

Well-suited for customer support with fast response times and privacy benefits.

content creation

Good for content creation with balanced quality and speed.

data analysis

Adequate for basic data analysis. Not suitable for complex mathematical tasks.

research assistant

Good for basic research tasks. 57.2% MMLU shows solid general knowledge.

legal compliance

Good for basic legal tasks with data sovereignty benefits.

healthcare

Good for healthcare with self-hosted HIPAA compliance. Basic clinical tasks.

financial analysis

Adequate for basic financial tasks. Not suitable for complex modeling.

education

Good for educational content. Fast inference suitable for interactive learning.

creative writing

Adequate creative writing for typical use cases.