vLLM
vLLM v0.8.3 supports the Llama 4 herd.
This reference covers the gated Maverick instruction checkpoint and separates sibling claims from this artifact. It treats provider output limits as third-party listings.
Llama 4 Maverick is the vendor reference for meta-llama/Llama-4-Maverick-17B-128E-Instruct, as of Aug 12, 2026. Model scale: 17B active, 128 experts, 400B total; Context length: 1M, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | meta-llama/Llama-4-Maverick-17B-128E-Instruct | huggingface.coRetrieved |
| Identity | Herd release date | April 5, 2025 | ai.meta.comRetrieved |
| Identity | Maverick identity | Llama 4 Maverick, a 17 billion active parameter model with 128 experts | ai.meta.comRetrieved |
| Identity | Scout sibling | Llama 4 Scout, a 17 billion active parameter model with 16 experts and 10M context | ai.meta.comRetrieved |
| Identity | Behemoth announcement status | The release announcement said: While we're not yet releasing Llama 4 Behemoth as it is still training. | ai.meta.comRetrieved |
| Architecture | Model scale | 17B active, 128 experts, 400B total | huggingface.coRetrieved |
| Architecture | Multimodal fusion | natively multimodal models with early fusion to seamlessly integrate text and vision tokens | ai.meta.comRetrieved |
| Architecture | Training scale | more than 30 trillion tokens | ai.meta.comRetrieved |
| Context | Context length | 1M | huggingface.coRetrieved |
| Context | OpenRouter provider output limits | Provider limits range from 8K to 32K. | openrouter.aiRetrieved |
| Modalities | Inputs and output | text and image input; text output | huggingface.coRetrieved |
| Modalities | Languages | 12 supported languages | huggingface.coRetrieved |
| License | Weights license | Llama 4 Community License Agreement | developer.meta.comRetrieved |
| License | Monthly-active-user clause | If, on the Llama 4 version release date, the monthly active users of the products or services made available by or for Licensee...is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta | developer.meta.comRetrieved |
| License | Branding clause | prominently display 'Built with Llama'... you shall also include 'Llama' at the beginning of any such AI model name. | developer.meta.comRetrieved |
| Pricing | Meta API price: OpenRouter listing | OpenRouter lists $0.20 input and $0.696 output per 1M tokens. | openrouter.aiRetrieved |
| Availability | Repository access | The Hugging Face repository is gated; access requires Meta's form, and the raw configuration is not publicly fetchable without access. | huggingface.coRetrieved |
| Availability | Vendor deployment guidance | The FP8 quantized weights fit on a single H100 DGX host while still maintaining quality. | huggingface.coRetrieved |
| Availability | Gated instruction weights | meta-llama/Llama-4-Maverick-17B-128E-Instruct | huggingface.coRetrieved |
"the first open-weight natively multimodal models with unprecedented context length support and our first built using a mixture-of-experts (MoE) architecture."
ai.meta.comRetrieved
"Llama 4 Scout dramatically increases the supported context length from 128K in Llama 3 to an industry leading 10 million tokens."
ai.meta.comRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| MMLU Pro | 80.5 | huggingface.coRetrieved |
| GPQA Diamond | 69.8 | huggingface.coRetrieved |
| LiveCodeBench | 43.4 | huggingface.coRetrieved |
| LMArena experimental chat version | an experimental chat version scoring ELO of 1417 on LMArena | ai.meta.comRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
vLLM v0.8.3 supports the Llama 4 herd.
SGLang v0.4.5 includes Llama 4 support from merged pull request 5092.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: meta-llama/Llama-4-Maverick-17B-128E-Instruct.
As of Aug 12, 2026, Repository access: The Hugging Face repository is gated; access requires Meta's form, and the raw configuration is not publicly fetchable without access.
As of Aug 12, 2026, Vendor deployment guidance: The FP8 quantized weights fit on a single H100 DGX host while still maintaining quality.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs