vLLM
The model card states a vLLM v0.23.0+ serving floor for GLM-5.2.
This reference covers the open flagship checkpoint and preserves conflicting source claims side by side. It separates vendor deployment floors from independent serving support.
GLM-5.2 is the vendor reference for zai-org/GLM-5.2, as of Aug 12, 2026. Vendor-stated scale: 744B parameters (40B active); Context length: Solid 1M-token context that stably sustains long-horizon work, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | zai-org/GLM-5.2 | huggingface.coRetrieved |
| Identity | API model id | glm-5.2 | docs.z.aiRetrieved |
| Identity | Family position | Newest GLM flagship listed by the organization at retrieval. | huggingface.coRetrieved |
| Identity | Hugging Face card date | June 17, 2026 | huggingface.coRetrieved |
| Architecture | Vendor-stated scale | 744B parameters (40B active) | github.comRetrieved |
| Architecture | Repository counter | The Hugging Face automatic parameter counter shows about 753B, while the vendor states 744B. | huggingface.coRetrieved |
| Architecture | Sparse attention | DSA sparse attention with IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length | github.comRetrieved |
| Architecture | Speculative decoding | MTP layer for speculative decoding, increasing the acceptance length by up to 20% | github.comRetrieved |
| Architecture | Configuration | 78 layers, 256 routed experts, 1 shared expert, 8 experts per token, max_position_embeddings 1048576, model_type glm_moe_dsa | huggingface.coRetrieved |
| Context | Context length | Solid 1M-token context that stably sustains long-horizon work | huggingface.coRetrieved |
| Context | Maximum output | 128K | docs.z.aiRetrieved |
| Context | Maximum output | maximum generation length of 163,840 tokens | huggingface.coRetrieved |
| Modalities | Modality | text | docs.z.aiRetrieved |
| Modalities | Thinking control | Thinking toggle with reasoning_effort including max | docs.z.aiRetrieved |
| Modalities | Agentic support | Tool calling and MCP integration | docs.z.aiRetrieved |
| License | Weights license | MIT; Copyright (c) 2026 Zhipu AI | huggingface.coRetrieved |
| Pricing | Zhipu AI API price: Official input | $1.40 per 1M tokens | docs.z.aiRetrieved |
| Pricing | Zhipu AI API price: Official cached input | $0.26 per 1M tokens | docs.z.aiRetrieved |
| Pricing | Zhipu AI API price: Official output | $4.40 per 1M tokens | docs.z.aiRetrieved |
| Availability | Repositories | zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings. | huggingface.coRetrieved |
| Availability | Quantized repository downloads | zai-org/GLM-5.2-FP8 had 2.03M downloads at retrieval. | huggingface.coRetrieved |
| Availability | Vendor serving floors | vLLM v0.23.0+ and SGLang v0.5.13.post1+ | huggingface.coRetrieved |
| Availability | Reference weights | zai-org/GLM-5.2 | huggingface.coRetrieved |
| Availability | Quantized weights | zai-org/GLM-5.2-FP8 | huggingface.coRetrieved |
| Availability | Earlier sibling weights | zai-org/GLM-5.1 | huggingface.coRetrieved |
| Availability | Earlier sibling weights | zai-org/GLM-5 | huggingface.coRetrieved |
"IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length"
github.comRetrieved
"for speculative decoding, increasing the acceptance length by up to 20%"
github.comRetrieved
"Stronger coding capabilities with multiple thinking effort levels"
huggingface.coRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| Terminal-Bench 2.1 | 81.0 | docs.z.aiRetrieved |
| SWE-bench Pro | 62.1 | docs.z.aiRetrieved |
| GPQA-Diamond | 91.2 | huggingface.coRetrieved |
| AIME 2026 | 99.2 | huggingface.coRetrieved |
| HLE | 40.5 | huggingface.coRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
The model card states a vLLM v0.23.0+ serving floor for GLM-5.2.
The SGLang project publishes a GLM-5.2 serving cookbook.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: zai-org/GLM-5.2.
As of Aug 12, 2026, Repositories: zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings.
As of Aug 12, 2026, Repositories: zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs