Llama 4 Maverick (17B-128E)
FP8 repo meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8. Bedrock max output 8K. No first-party pay-as-you-go pricing verified. Superseded at Meta by closed Muse Spark and open Muse Glimmer.
- Context window
- 1,000,000 tokens
- Knowledge cutoff
- 2024-08
- Input
- text, image
- Output
- text
- License
- llama4-community
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | — | huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct | — |
| AWS Bedrock | meta.llama4-maverick-17b-instruct-v1:0 | — | docs |
| OpenRouter | meta-llama/llama-4-maverick | openrouter.ai/meta-llama/llama-4-maverick | — |
| Web app | — | meta.ai | — |
Notable capabilities (3)
- FIRST First natively multimodal Llama (early fusion): Llama 4 were the first Llama models with native multimodality via early fusion of text and vision tokens. source
- 400B-total MoE on one H100 host: 17B active / 128 experts / ~400B total; runs on a single H100 host. source
- LMArena experimental-variant controversy (found after launch): Launch LMArena Elo 1417 came from an experimental chat-tuned variant, not the released weights, drawing criticism. source
Open-weight MoE multimodal model; still widely hosted and cheap on third-party providers.
curl https://openrouter.ai/api/v1/chat/completions -H "Authorization: Bearer $OPENROUTER_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"meta-llama/llama-4-maverick","messages":[{"role":"user","content":"Hello"}]}'
Sources: https://ai.meta.com/blog/llama-4-multimodal-intelligence/ , https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-maverick-17b-instruct.html
Other Meta models
Muse Voice Transcribe 1.0 · Muse Spark 1.3 · Muse Glimmer 30B · Omnilingual ASR · Llama 4 Scout (17B-16E)