vLLM
vLLM's own recipe page serves Mistral Medium 3.5 on nightly builds; architecture support had not yet shipped in a stable release.
This reference covers the open merged flagship checkpoint and its restrictive license terms. It preserves each vendor date and currency without reconciling them.
Mistral Medium 3.5 is the vendor reference for mistralai/Mistral-Medium-3.5-128B, as of Aug 12, 2026. Model architecture: Dense 128B model with a 256k context window, handling instruction-following, reasoning, and coding; Context length: 256k, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | mistralai/Mistral-Medium-3.5-128B | huggingface.coRetrieved |
| Identity | Vendor description | our first flagship merged model | mistral.aiRetrieved |
| Identity | Announcement date | May 22, 2026 | mistral.aiRetrieved |
| Identity | Documentation card date | April 28, 2026 | docs.mistral.aiRetrieved |
| Identity | Repository timing | The repository was created March 31, 2026. | huggingface.coRetrieved |
| Architecture | Model architecture | Dense 128B model with a 256k context window, handling instruction-following, reasoning, and coding | huggingface.coRetrieved |
| Architecture | Vision encoder | vision encoder trained from scratch | mistral.aiRetrieved |
| Context | Context length | 256k | huggingface.coRetrieved |
| Modalities | Inputs and output | Accepts both text and image input, with text output | huggingface.coRetrieved |
| Modalities | Tool calling | Agentic: Best-in-class agentic capabilities with native function calling and JSON output | huggingface.coRetrieved |
| Modalities | Reasoning control | Opt-in reasoning_effort with [THINK] blocks | recipes.vllm.aiRetrieved |
| License | Weights license | Modified MIT License | huggingface.coRetrieved |
| License | Revenue clause | You are not authorized to exercise any rights under this license if the global consolidated monthly revenue of your company (or that of your employer) exceeds $20 million (or its equivalent in another currency) for the preceding month. | huggingface.coRetrieved |
| License | Derivative restriction | The revenue restriction extends to derivatives; a commercial license is available through Mistral. | huggingface.coRetrieved |
| License | Vendor license framing | Released as open weights, under a modified MIT license | mistral.aiRetrieved |
| Pricing | Mistral AI API price: Documentation input price | EUR 1.25 per 1M tokens | docs.mistral.aiRetrieved |
| Pricing | Mistral AI API price: Documentation output price | EUR 6.40 per 1M tokens | docs.mistral.aiRetrieved |
| Pricing | Mistral AI API price: Announcement input price | $1.50 per 1M tokens | mistral.aiRetrieved |
| Pricing | Mistral AI API price: Announcement output price | $7.50 per 1M tokens | mistral.aiRetrieved |
| Availability | Vendor hardware guidance | self-hosting possible on as few as four GPUs | mistral.aiRetrieved |
| Availability | Merged flagship weights | mistralai/Mistral-Medium-3.5-128B | huggingface.coRetrieved |
"our first flagship merged model"
mistral.aiRetrieved
"replaces Devstral 2, Magistral, and Medium 3.1"
huggingface.coRetrieved
"Built for long-horizon tasks, calling multiple tools reliably"
mistral.aiRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
No vendor-claimed benchmark results for this exact build were present in the cited material as of Aug 12, 2026.
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
vLLM's own recipe page serves Mistral Medium 3.5 on nightly builds; architecture support had not yet shipped in a stable release.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: mistralai/Mistral-Medium-3.5-128B.
As of Aug 12, 2026, Vendor hardware guidance: self-hosting possible on as few as four GPUs.
As of Aug 12, 2026, Vendor hardware guidance: self-hosting possible on as few as four GPUs.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs