vLLM
A merged pull request enables Qwen3.8 on additional hardware, implying in-tree model support; no release note names it yet.
This reference covers the open text artifact, not its hosted multimodal sibling. It separates artifact facts from hosted service claims and leaves absent pricing unstated.
Qwen3.8-2.4T-A95B is the vendor reference for Qwen/Qwen3.8-2.4T-A95B, as of Aug 12, 2026. Scale: 2.4T in total and 95B activated; Context: 262,144 natively and extensible up to 1,010,000 tokens, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | Qwen/Qwen3.8-2.4T-A95B | huggingface.coRetrieved |
| Identity | FP8 variant | Qwen/Qwen3.8-2.4T-A95B-FP8 | huggingface.coRetrieved |
| Identity | Artifact scope | Qwen3.8 Max is the vendor's hosted multimodal API sibling; this page covers only the open-weights text model, and no Max API fact about vision, video, or API pricing applies to this artifact. | qwen.aiRetrieved |
| Identity | Repository timing | in the week before 2026-08-12 | huggingface.coRetrieved |
| Architecture | Scale | 2.4T in total and 95B activated | huggingface.coRetrieved |
| Architecture | Experts | 512 experts; 10 Routed + 1 Shared activated | huggingface.coRetrieved |
| Architecture | Layer pattern | 92 layers in the pattern 23 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE)) | huggingface.coRetrieved |
| Architecture | Hybrid linear attention | Gated DeltaNet: 128 linear attention V heads, 16 QK; Gated Attention: 64 Q heads, 4 KV heads, head dim 256; hidden dim 8192 | huggingface.coRetrieved |
| Context | Context | 262,144 natively and extensible up to 1,010,000 tokens | huggingface.coRetrieved |
| Context | Reasoning content budget | Reasoning Content: 262,144 tokens | huggingface.coRetrieved |
| Context | Final response budget | Final Response: 131,072 tokens | huggingface.coRetrieved |
| Modalities | Modality | text only for THIS artifact | huggingface.coRetrieved |
| Modalities | Thinking mode | requires thinking mode for all interactions | huggingface.coRetrieved |
| Modalities | Reasoning format | reasoning in <think> tags | huggingface.coRetrieved |
| Modalities | Thinking control | Flexible Thinking Control via reasoning_effort levels xhigh (default), medium, low | huggingface.coRetrieved |
| Modalities | Agentic support | tool calling and agentic use supported | huggingface.coRetrieved |
| Modalities | Recommended sampling | temperature 1.0, top_p 0.95, top_k 20 | huggingface.coRetrieved |
| License | License | Qwen3.8-Max License | huggingface.coRetrieved |
| License | License distinction | Not Apache 2.0. | huggingface.coRetrieved |
| License | Attribution clause | above >100,000,000 monthly active users or US$ 20,000,000 ... monthly revenue, respective model name must be prominently displayed on the user interface | huggingface.coRetrieved |
| License | Model service clause | If the licensee or any of its affiliates conducts a Model as a Service or AI Work Assistant business, and the aggregate revenue ... exceeds US$50,000,000 ... during any consecutive twelve (12) months, the licensee shall obtain a separate license from Qwen before Using the Software or its derivative works for any commercial purpose. | huggingface.coRetrieved |
| Pricing | Qwen API price: Open artifact pricing | Not published for the open artifact. | huggingface.coRetrieved |
| Availability | Reference weights | Qwen/Qwen3.8-2.4T-A95B | huggingface.coRetrieved |
| Availability | FP8 weights | Qwen/Qwen3.8-2.4T-A95B-FP8 | huggingface.coRetrieved |
| Availability | Vendor-recommended engines | SGLang, vLLM | huggingface.coRetrieved |
| Availability | Reference weights | Qwen/Qwen3.8-2.4T-A95B | huggingface.coRetrieved |
| Availability | Quantized weights | Qwen/Qwen3.8-2.4T-A95B-FP8 | huggingface.coRetrieved |
"the most capable generation in the Qwen open-model family to date"
huggingface.coRetrieved
"the first time we will open-source the weights of a Qwen-Max-class model"
qwen.aiRetrieved
"Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains"
huggingface.coRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| PaperBench | 93.0 | huggingface.coRetrieved |
| Terminal Bench 2.1 | 86.6 | huggingface.coRetrieved |
| GPQA Diamond | 92.6 | huggingface.coRetrieved |
| SWE-bench Pro | 67.7 | huggingface.coRetrieved |
| HLE | 43.6 | huggingface.coRetrieved |
| FrontierSWE | 73.5 | huggingface.coRetrieved |
| WideSearch | 81.9 | huggingface.coRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
A merged pull request enables Qwen3.8 on additional hardware, implying in-tree model support; no release note names it yet.
Merged documentation adds a Qwen3.8 cookbook and serving configs while the core support pull request was still open.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: Qwen/Qwen3.8-2.4T-A95B.
As of Aug 12, 2026, Reference weights: Qwen/Qwen3.8-2.4T-A95B.
As of Aug 12, 2026, Vendor-recommended engines: SGLang, vLLM.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs