vLLM
vLLM now supports gpt-oss on NVIDIA Blackwell and Hopper GPUs, as well as AMD MI300x and MI355x GPUs.
This reference covers the larger open reasoning checkpoint and its required response format. It keeps self-hosting guidance separate from measured RunInfra evidence.
gpt-oss-120b is the vendor reference for openai/gpt-oss-120b, as of Aug 12, 2026. Model scale: 116.8B total parameters and 5.1B 'active' parameters per token per forward pass; Context length: 131,072, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | openai/gpt-oss-120b | huggingface.coRetrieved |
| Identity | Paper submission date | August 8, 2025 | arxiv.orgRetrieved |
| Identity | Sibling checkpoint | openai/gpt-oss-20b | huggingface.coRetrieved |
| Architecture | Model scale | 116.8B total parameters and 5.1B 'active' parameters per token per forward pass | arxiv.orgRetrieved |
| Architecture | Mixture configuration | 36 layers; 128 experts; top-4 routing | huggingface.coRetrieved |
| Architecture | Attention pattern | banded window and fully dense patterns alternate, with 128-token windows and grouped-query attention using 64 attention heads and 8 key-value heads | arxiv.orgRetrieved |
| Architecture | Weights format | MoE weights quantized to MXFP4 format (4.25 bits per parameter) | arxiv.orgRetrieved |
| Architecture | Position scaling | YaRN | huggingface.coRetrieved |
| Context | Context length | 131,072 | developers.openai.comRetrieved |
| Context | Maximum output | 131,072 | developers.openai.comRetrieved |
| Modalities | Modality | Tagged text-generation on the model hub listing. | huggingface.coRetrieved |
| Modalities | Reasoning effort | Configurable reasoning effort: Easily adjust the reasoning effort (low, medium, high) | huggingface.coRetrieved |
| Modalities | Agentic capabilities | Use the models' native capabilities for function calling, web browsing, Python code execution, and Structured Outputs | huggingface.coRetrieved |
| License | Weights license | Apache 2.0 | huggingface.coRetrieved |
| License | License description | Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk | huggingface.coRetrieved |
| Pricing | OpenAI API price: OpenRouter listing | OpenRouter lists $0.03 per million input tokens and $0.17 per million output tokens. | openrouter.aiRetrieved |
| Availability | Repository access | The openai/gpt-oss-120b repository is ungated. | huggingface.coRetrieved |
| Availability | Required response format | Both models were trained using our harmony response format and should only be used with this format; otherwise, they will not work correctly. | huggingface.coRetrieved |
| Availability | Vendor deployment guidance | for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) | huggingface.coRetrieved |
| Availability | Larger open reasoning weights | openai/gpt-oss-120b | huggingface.coRetrieved |
| Availability | Smaller open reasoning weights | openai/gpt-oss-20b | huggingface.coRetrieved |
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| AIME 2025, with tools | 97.9% | arxiv.orgRetrieved |
| GPQA Diamond, with tools | 80.9% | arxiv.orgRetrieved |
| MMLU | 90.0% | arxiv.orgRetrieved |
| SWE-bench Verified | 62.4% | arxiv.orgRetrieved |
| Codeforces, with tools | 2622 Elo | arxiv.orgRetrieved |
| Paper comparison | gpt-oss-120b surpasses OpenAI o3-mini and approaches OpenAI o4-mini accuracy | arxiv.orgRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
vLLM now supports gpt-oss on NVIDIA Blackwell and Hopper GPUs, as well as AMD MI300x and MI355x GPUs.
The project tracked day-zero gpt-oss support.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: openai/gpt-oss-120b.
As of Aug 12, 2026, Repository access: The openai/gpt-oss-120b repository is ungated.
As of Aug 12, 2026, Vendor deployment guidance: for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X).
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs