vLLM
A merged pull request implements Gemma 4 architecture support for MoE, multimodal input, reasoning, and tool use.
This reference covers the dense flagship checkpoint and keeps family-only capabilities separate. It distinguishes model-vendor facts from third-party endpoint pricing.
Gemma 4 31B is the vendor reference for google/gemma-4-31B-it, as of Aug 12, 2026. Attention architecture: hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global; Context length: 256K tokens for the 12B, 26B, and 31B models, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | google/gemma-4-31B-it | huggingface.coRetrieved |
| Identity | Family release date | April 2, 2026 | blog.googleRetrieved |
| Identity | Flagship variant | 31B Dense; 30.7B parameters | ai.google.devRetrieved |
| Identity | Family variants | E2B, E4B, 12B, 26B MoE, and 31B Dense; the 26B MoE has 25.2B total, 3.8B active, 8 active experts, 128 total experts, and 1 shared expert. | ai.google.devRetrieved |
| Architecture | Attention architecture | hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global | ai.google.devRetrieved |
| Architecture | Sliding windows | 512-1024 tokens | ai.google.devRetrieved |
| Context | Context length | 256K tokens for the 12B, 26B, and 31B models | ai.google.devRetrieved |
| Context | Smaller-family context | 128K tokens for the E2B and E4B models | ai.google.devRetrieved |
| Context | Maximum output | Not stated by Google for this artifact. | ai.google.devRetrieved |
| Modalities | Inputs | All models natively process video and images, supporting variable resolutions. | blog.googleRetrieved |
| Modalities | Output | text output only | ai.google.devRetrieved |
| Modalities | Agentic workflows | Native support for function-calling, structured JSON output | blog.googleRetrieved |
| Modalities | Languages | 140+ languages | ai.google.devRetrieved |
| Modalities | Audio scope | Audio input is not listed for the 31B model. | ai.google.devRetrieved |
| License | Weights license | released under a commercially permissive Apache 2.0 license | blog.googleRetrieved |
| License | Repository access | Apache 2.0 metadata; ungated repository | huggingface.coRetrieved |
| License | Use-policy pointer | The vendor model card links a Prohibited use policy page. | ai.google.devRetrieved |
| Pricing | Google API price: OpenRouter listing | $0.08 input and $0.35 output per 1M tokens | openrouter.aiRetrieved |
| Availability | Repository family | Verified repositories include the 31B instruction and base checkpoints, the 26B MoE instruction checkpoint, and the 12B, E4B, and E2B instruction checkpoints, with additional base and QAT variants. | huggingface.coRetrieved |
| Availability | Instruction weights | google/gemma-4-31B-it | huggingface.coRetrieved |
| Availability | Base weights | google/gemma-4-31B | huggingface.coRetrieved |
| Availability | Mixture instruction weights | google/gemma-4-26B-A4B-it | huggingface.coRetrieved |
| Availability | Instruction weights | google/gemma-4-12B-it | huggingface.coRetrieved |
| Availability | Efficient instruction weights | google/gemma-4-E4B-it | huggingface.coRetrieved |
| Availability | Efficient instruction weights | google/gemma-4-E2B-it | huggingface.coRetrieved |
"our most intelligent open models to date. Purpose-built for advanced reasoning"
blog.googleRetrieved
"breakthrough capabilities made widely accessible under an Apache 2.0 license"
blog.googleRetrieved
"intelligence-per-parameter means achieving frontier-level capabilities with less hardware"
blog.googleRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| Arena AI text leaderboard | 31B: #3 open model in the world on the industry-standard Arena AI text leaderboard | blog.googleRetrieved |
| MMLU Pro | 85.2% | ai.google.devRetrieved |
| AIME 2026, no tools | 89.2% | ai.google.devRetrieved |
| Codeforces ELO | 2150 | ai.google.devRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
A merged pull request implements Gemma 4 architecture support for MoE, multimodal input, reasoning, and tool use.
The Gemma 4 cookbook states that all Gemma 4 models require the Triton attention backend for bidirectional image-token attention.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: google/gemma-4-31B-it.
As of Aug 12, 2026, Repository family: Verified repositories include the 31B instruction and base checkpoints, the 26B MoE instruction checkpoint, and the 12B, E4B, and E2B instruction checkpoints, with additional base and QAT variants.
As of Aug 12, 2026, Repository family: Verified repositories include the 31B instruction and base checkpoints, the 26B MoE instruction checkpoint, and the 12B, E4B, and E2B instruction checkpoints, with additional base and QAT variants.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs