Model ComparisonMeta

Llama 4 Maverick

Llama 4 Maverick — Meta's most capable open-weight model with 1M token context.

VisionTool usePrompt cachingReasoningFast
Pricing (verified 2026-08-19)
TypePrice / 1M tokensNotes
Input$0.27Per 1M input tokens
Output$0.85Per 1M output tokens
Source: Meta official pricing page · Last verified 2026-08-19
Context Window
1M tokens
750K words · 3 pages of text

Llama 4 Maverick competes with GPT-5.4 on benchmarks at a fraction of the cost. At $0.27/M input via Together AI, it's the strongest open-weight model for high-quality generation.

Best for

  • High-volume production inference needing GPT-4+ quality
  • Teams with GPU infrastructure for self-hosting
  • Long-context (up to 1M token) research tasks
  • Fine-tuning on proprietary data

Not ideal for

  • Real-time latency-critical apps (MoE routing adds overhead)
  • Teams without the capacity to manage open-weight deployment

Real cost scenarios

ScenarioEst. monthly costBreakdown
Content generation — 100k docs/month at 3k tokens each$81300M input at $0.27/M = $81
Self-hosted on 8× A100 80GB cluster (RunPod)$14,600$2.49/hr × 730hr × 8 GPUs = unlimited tokens

Estimates with typical token distributions. Use the Prompt Cost Calculator for your exact workload.

How to cut costs on Llama 4 Maverick

  1. At $0.27/M, Llama 4 Maverick is 9× cheaper than GPT-5.4 on input. For GPT-4-class tasks at scale, this is the first model to evaluate.

Calculate your Llama 4 Maverick costs

Prompt Cost Calculator
Exact cost for your prompt
Agent Cost Calculator
Multi-step agent loops
Prompt Caching Calculator
Caching savings
Conversation Cost
Chat history strategies

Other Meta models

Llama 4 Scout$0.18/MLlama 3.3 70B$0.23/MLlama 3.1 405B$1.00/M

Compare Llama 4 Maverick

deepseek v3 vs llama 4 scoutllama 4 scout vs gpt 5 mini

Frequently Asked Questions

What is Llama 4 Maverick vs Llama 4 Scout?

Maverick (400B total, 17B active) is larger and more capable than Scout (109B total, 17B active). Both use MoE architecture so they run at similar speeds. Maverick benchmarks closer to GPT-5.4; Scout is closer to GPT-5 mini — but both are dramatically cheaper.

Pricing verified 2026-08-19 from Meta official pricing. All prices in USD per 1M tokens. See methodology and price changelog.