SGLang
The [Feature] Xiaomi MiMo-V2.5 day0 support pull request was merged on April 30, 2026.
This reference covers the open omnimodal mixture checkpoint and keeps sibling scale claims separate. It records the vendor's permissive license statement without implying that a repository license file exists.
MiMo-V2.5 is the vendor reference for XiaomiMiMo/MiMo-V2.5, as of Aug 12, 2026. Model scale: Sparse MoE (Mixture of Experts), 310B total / 15B activated parameters; Context length: Up to 1M tokens, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | XiaomiMiMo/MiMo-V2.5 | huggingface.coRetrieved |
| Identity | Vendor description | MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. | huggingface.coRetrieved |
| Identity | OpenRouter release date | Apr 22, 2026 | openrouter.aiRetrieved |
| Identity | Open-source announcement | Today, we officially open source the Xiaomi MiMo-V2.5 series | mimo.mi.comRetrieved |
| Identity | Public testing date | Public testing began April 23, 2026. | mimo.mi.comRetrieved |
| Architecture | Model scale | Sparse MoE (Mixture of Experts), 310B total / 15B activated parameters | huggingface.coRetrieved |
| Architecture | Expert routing | 256 routed experts; 8 per token | huggingface.coRetrieved |
| Architecture | Layer configuration | 48 layers (1 dense + 47 MoE) | huggingface.coRetrieved |
| Architecture | Attention layers | 9 full-attention + 39 SWA | huggingface.coRetrieved |
| Architecture | Attention interleaving | interleaving Sliding Window Attention (SWA) and Global Attention (GA) with a 5:1 ratio and 128 sliding window. This reduces KV-cache storage by nearly 6x while maintaining long-context performance via learnable attention sink bias. | huggingface.coRetrieved |
| Architecture | Vision encoder | 729M-param ViT (28 layers: 24 SWA + 4 Full) | huggingface.coRetrieved |
| Architecture | Audio encoder | 261M-param Audio Transformer | huggingface.coRetrieved |
| Architecture | Multi-token prediction | 329M parameters, 3 layers, for speculative decoding | huggingface.coRetrieved |
| Architecture | Training | Trained on a total of ~48T tokens using FP8 mixed precision. | huggingface.coRetrieved |
| Architecture | Backbone lineage | inherits from the MiMo-V2-Flash architecture | huggingface.coRetrieved |
| Context | Context length | Up to 1M tokens | huggingface.coRetrieved |
| Context | Context extension schedule | 32K -> 256K -> 1M | huggingface.coRetrieved |
| Modalities | Input understanding | text, image, video, and audio | huggingface.coRetrieved |
| Modalities | Output | text | openrouter.aiRetrieved |
| License | Vendor license statement | uses the MIT license, supports commercial inference deployment and secondary training, and requires no additional authorization. | mimo.mi.comRetrieved |
| License | Repository metadata | license: mit | huggingface.coRetrieved |
| Pricing | Xiaomi API price: Xiaomi cache-hit input | $0.0028 / MTok | mimo.mi.comRetrieved |
| Pricing | Xiaomi API price: Xiaomi cache-miss input | $0.14 / MTok | mimo.mi.comRetrieved |
| Pricing | Xiaomi API price: Xiaomi output | $0.28 / MTok | mimo.mi.comRetrieved |
| Pricing | Xiaomi API price: OpenRouter listing | $0.14 input and $0.28 output per 1M tokens | openrouter.aiRetrieved |
| Availability | Repository access | Ungated repository. | huggingface.coRetrieved |
| Availability | Tensor types | F32, BF16, F8_E4M3 | huggingface.coRetrieved |
| Availability | Configuration refresh | The config.json and tokenizer_config.json files in this repository have been updated since the initial release... Using the outdated config may lead to degraded model performance. | huggingface.coRetrieved |
| Availability | Recommended sampling | temperature=1.0, top_p=0.95 | huggingface.coRetrieved |
| Availability | Base sibling | MiMo-V2.5-Base; 256K context | huggingface.coRetrieved |
| Availability | Pro sibling | MiMo-V2.5-Pro; 1.02T total / 42B activated | huggingface.coRetrieved |
| Availability | vLLM recipe floor | vLLM 0.21.0+ | recipes.vllm.aiRetrieved |
| Availability | Main weights | XiaomiMiMo/MiMo-V2.5 | huggingface.coRetrieved |
"Today, we are releasing MiMo-V2.5, a major step forward in agentic capability and multimodal understanding. With native visual and audio understanding, MiMo-V2.5 reasons seamlessly across modalities, surpasses MiMo-V2-Pro in agentic performance, and supports up to 1 million tokens of context."
mimo.xiaomi.comRetrieved
"Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD), achieving strong performance on agentic tasks and multimodal understanding benchmarks."
huggingface.coRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| Claw-Eval, general subset | 62.3; placing it at the Pareto frontier of performance and efficiency | mimo.xiaomi.comRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
The [Feature] Xiaomi MiMo-V2.5 day0 support pull request was merged on April 30, 2026.
The [Model] Add MiMo-V2.5 support pull request was merged on April 27, 2026.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: XiaomiMiMo/MiMo-V2.5.
As of Aug 12, 2026, Repository access: Ungated repository.
As of Aug 12, 2026, Repository access: Ungated repository.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs