vLLM
The DeepSeek V4 Rebased pull request adding initial DeepSeek V4 support was merged on April 27, 2026.
This reference restores an archived open checkpoint without implying a current measured package. A future published package with the same identity will replace this view automatically.
DeepSeek-V4-Flash-0731 is the vendor reference for deepseek-ai/DeepSeek-V4-Flash-0731, as of Aug 12, 2026. Release architecture: DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained.; Context length: 1M, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. | huggingface.coRetrieved |
| Identity |
"DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained."
api-docs.deepseek.comRetrieved
"The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex."
api-docs.deepseek.comRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| Terminal Bench 2.1 | 82.7 | api-docs.deepseek.comRetrieved |
| NL2Repo |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
The DeepSeek V4 Rebased pull request adding initial DeepSeek V4 support was merged on April 27, 2026.
Day-zero support covers the open DeepSeek V4 family.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
As of Aug 12, 2026, Open repository: deepseek-ai/DeepSeek-V4-Flash-0731 exists with about 1.05M downloads at retrieval.
As of Aug 12, 2026, Vendor deployment guidance: The model card provides vLLM expert-parallel serving and SGLang DSPARK speculative-serving commands.
| API model id |
| deepseek-v4-flash |
api-docs.deepseek.comRetrieved |
| Identity | Release date | 2026-07-31 | api-docs.deepseek.comRetrieved |
|---|
| Architecture | Release architecture | DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained. | api-docs.deepseek.comRetrieved |
|---|
| Architecture | Vendor-stated scale | Total 284B, Activated 13B | huggingface.coRetrieved |
|---|
| Architecture | Repository counter | The Hugging Face repository size badge reads 304B from the safetensors count, while the vendor-stated model total is 284B. | huggingface.coRetrieved |
|---|
| Architecture | Configuration | 43 layers, 64 attention heads, 256 routed experts, 1 shared expert, 6 experts per token, max_position_embeddings 1048576, bfloat16 with fp8 e4m3 | huggingface.coRetrieved |
|---|
| Context | Context length | 1M | api-docs.deepseek.comRetrieved |
|---|
| Context | Maximum output | MAX OUTPUT: MAXIMUM: 384K | api-docs.deepseek.comRetrieved |
|---|
| Modalities | Modality | text | huggingface.coRetrieved |
|---|
| Modalities | Reasoning effort | The reasoning_effort parameter now supports three levels - low, high, and max - which control how much deliberation the model spends before answering. | huggingface.coRetrieved |
|---|
| Modalities | Tool calling | Tool calling is supported. | huggingface.coRetrieved |
|---|
| License | Weights license | MIT; Copyright (c) 2023 DeepSeek | huggingface.coRetrieved |
|---|
| Pricing | DeepSeek API price: DeepSeek API input cache hit | $0.0028 per 1M tokens | api-docs.deepseek.comRetrieved |
|---|
| Pricing | DeepSeek API price: DeepSeek API input cache miss | $0.14 per 1M tokens | api-docs.deepseek.comRetrieved |
|---|
| Pricing | DeepSeek API price: DeepSeek API output | $0.28 per 1M tokens | api-docs.deepseek.comRetrieved |
|---|
| Pricing | DeepSeek API price: DeepSeek pricing warning | We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. | api-docs.deepseek.comRetrieved |
|---|
| Pricing | DeepSeek API price: OpenRouter listing | $0.08 input and $0.25 output per 1M tokens | openrouter.aiRetrieved |
|---|
| Availability | Open repository | deepseek-ai/DeepSeek-V4-Flash-0731 exists with about 1.05M downloads at retrieval. | huggingface.coRetrieved |
|---|
| Availability | Open lineage repositories | DeepSeek-V4-Flash preview, DeepSeek-V4-Flash-Base, and DeepSeek-V4-Flash-DSpark repositories exist. | huggingface.coRetrieved |
|---|
| Availability | Dated-build variants | No -0731-Base or -0731-DSpark repository is listed for the dated build. | huggingface.coRetrieved |
|---|
| Availability | Vendor deployment guidance | The model card provides vLLM expert-parallel serving and SGLang DSPARK speculative-serving commands. | huggingface.coRetrieved |
|---|
| Availability | Dated instruction weights | deepseek-ai/DeepSeek-V4-Flash-0731 | huggingface.coRetrieved |
|---|
| Availability | Preview instruction weights | deepseek-ai/DeepSeek-V4-Flash | huggingface.coRetrieved |
|---|
| Availability | Base weights | deepseek-ai/DeepSeek-V4-Flash-Base | huggingface.coRetrieved |
|---|
| Availability | Speculative decoding draft variant | deepseek-ai/DeepSeek-V4-Flash-DSpark | huggingface.coRetrieved |
|---|
| 54.2 |
api-docs.deepseek.comRetrieved |
| Cybergym | 76.7 | api-docs.deepseek.comRetrieved |
|---|
| DeepSWE | 54.4 | api-docs.deepseek.comRetrieved |
|---|
| Toolathlon verified | 70.3 | api-docs.deepseek.comRetrieved |
|---|
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs