Llama 4 Scout (17B-16E)
Context: 10M per Meta; provider limits vary (not listed as context_window). Knowledge cutoff Aug 2024 per Meta model card (not re-verified today). No first-party pricing verified.
- Knowledge cutoff
- 2024-08
- Input
- text, image
- Output
- text
- License
- llama4-community
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | — | huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct | — |
| AWS Bedrock | meta.llama4-scout-17b-instruct-v1:0 | — | docs |
| OpenRouter | meta-llama/llama-4-scout | openrouter.ai/meta-llama/llama-4-scout | — |
| Web app | — | meta.ai | — |
Notable capabilities (2)
- FIRST 10M-token context (claimed): Meta advertised an 'industry-leading' 10M-token context via the iRoPE architecture; hosted providers typically serve far less (e.g. ~1.3M on OpenRouter). source
- Single-H100 multimodal MoE: 17B active / 16 experts / 109B total; fits one H100 with Int4 quantization. source
Small open-weight multimodal MoE for long-context and on-prem use.
curl https://openrouter.ai/api/v1/chat/completions -H "Authorization: Bearer $OPENROUTER_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"meta-llama/llama-4-scout","messages":[{"role":"user","content":"Hello"}]}'
Sources: https://ai.meta.com/blog/llama-4-multimodal-intelligence/ , https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct
Other Meta models
Muse Voice Transcribe 1.0 · Muse Spark 1.3 · Muse Glimmer 30B · Omnilingual ASR · Llama 4 Maverick (17B-128E)