vLLM
Support DeepseekV4 was merged; following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass in v0.23.0. This applies to the open V4 lineage because the 0813 weights are not public.
This reference tracks an API build documented separately from an earlier open lineage. It keeps current API facts distinct from lineage details and treats unpublished weights as unavailable.
DeepSeek-V4-Pro-0813, vendor specifications as of Aug 12, 2026. Open lineage scale: 1.6T parameters (49B activated); Context length: 1M, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | DeepSeek-V4-Pro-0813 | api-docs.deepseek.comRetrieved |
| Identity | API model id | deepseek-v4-pro | api-docs.deepseek.comRetrieved |
| Identity | Availability | GA of DeepSeek V4 Pro; appeared 2026-08-12 | openrouter.aiRetrieved |
| Architecture | Open lineage scale | 1.6T parameters (49B activated) | huggingface.coRetrieved |
| Architecture | Open lineage configuration | 384 routed experts, 6 experts per token, 1 shared expert, 61 layers, 128 attention heads, hidden 7168, vocab 129280, DeepseekV4ForCausalLM | huggingface.coRetrieved |
| Architecture | Attention | Novel Attention: Token-wise compression + DSA (DeepSeek Sparse Attention) | api-docs.deepseek.comRetrieved |
| Architecture | Current-build architecture disclosure | The vendor has not stated whether the 0813 build changes the April open-weights architecture. | api-docs.deepseek.comRetrieved |
| Context | Context length | 1M | api-docs.deepseek.comRetrieved |
| Context | Maximum output | 384K | api-docs.deepseek.comRetrieved |
| Context | Open lineage position limit | max_position_embeddings 1048576 | huggingface.coRetrieved |
| Modalities | Modality | text input and output | openrouter.aiRetrieved |
| Modalities | Thinking support | thinking and non-thinking modes supported | api-docs.deepseek.comRetrieved |
| Modalities | Open lineage mode names | Non-think, Think High, Think Max | huggingface.coRetrieved |
| Modalities | Vision | Not stated by the vendor. | api-docs.deepseek.comRetrieved |
| License | April open lineage | MIT | huggingface.coRetrieved |
| Pricing | DeepSeek API price: Input, cache hit | $0.003625 per 1M tokens | api-docs.deepseek.comRetrieved |
| Pricing | DeepSeek API price: Input, cache miss | $0.435 per 1M tokens | api-docs.deepseek.comRetrieved |
| Pricing | DeepSeek API price: Output | $0.87 per 1M tokens | api-docs.deepseek.comRetrieved |
| Availability | Current-build public weights | No public weights for DeepSeek-V4-Pro-0813. | huggingface.coRetrieved |
| Availability | Open lineage | DeepSeek-V4-Pro open weights are available. | huggingface.coRetrieved |
| Availability | Change log | No 0813 entry was present in the DeepSeek change log. | api-docs.deepseek.comRetrieved |
| Availability | Open lineage weights | deepseek-ai/DeepSeek-V4-Pro | huggingface.coRetrieved |
| Availability | Base weights | deepseek-ai/DeepSeek-V4-Pro-Base | huggingface.coRetrieved |
| Availability | Speculative decoding draft variant | deepseek-ai/DeepSeek-V4-Pro-DSpark | huggingface.coRetrieved |
"The `deepseek-v4-flash` model has been updated to DeepSeek-V4-Flash-0731, and the `deepseek-v4-pro` model has been updated to DeepSeek-V4-Pro-0813."
api-docs.deepseek.comRetrieved
"Novel Attention: Token-wise compression + DSA (DeepSeek Sparse Attention)"
api-docs.deepseek.comRetrieved
"In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2."
huggingface.coRetrieved
"We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected."
api-docs.deepseek.comRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
No vendor-claimed benchmark results for this exact build were present in the cited material as of Aug 12, 2026.
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
Support DeepseekV4 was merged; following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass in v0.23.0. This applies to the open V4 lineage because the 0813 weights are not public.
Day-0 support is documented for the open DeepSeek-V4 lineage. This applies to the open lineage because the 0813 weights are not public.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: DeepSeek-V4-Pro-0813.
DeepSeek API price: As of Aug 12, 2026, Input, cache hit: $0.003625 per 1M tokens. Input, cache miss: $0.435 per 1M tokens. Output: $0.87 per 1M tokens.
No, API only. As of Aug 12, 2026, Current-build public weights: No public weights for DeepSeek-V4-Pro-0813.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs