On this page
Key takeaways
- In the August 1, 2026 test, v4turbo measured 362.86 output tokens per second single stream for DeepSeek V4 Flash 0731; a stock deployment on the same four B200 GPUs in the same window measured 113.37.
- The median 1024-token request drops from 9248.33 ms to 3095.38 ms, a platform-computed 2.98x, for about 13 ms of added time to first token.
- The August 2, 2026 provider comparison is not like for like: our number is single stream on an idle node, theirs are production endpoints, and both halves are printed together. It does not establish a current provider ranking.
- The announcement described both configurations with a measured operating rule: accelerated at tested concurrency 1 and 8, standard at concurrency 32, where it held 397.74 output tokens per second per GPU. Results apply to those tested operating points.
- The article preserves the measured operating rule and reproduction context for the tested four-GPU configuration.
Status note, September 21, 2026: this article preserves the August 2, 2026 report for the DeepSeek V4 Flash 0731 checkpoint. No current measured optimization package is published for this checkpoint. See the model reference and the Model APIs library for current availability. The measurements below describe the historical four-GPU configuration, not a current product offer.
The rest of this post preserves the measured pair, the serving curve including the point where the package loses, and the market context and package scope reported on August 2, 2026.
The claim, with its conditions
As of August 2, 2026, Artificial Analysis listed no provider speed benchmarks for the 0731 release; its provider page said benchmarks were not available and named DeepSeek's own API as the sole provider. On the preview-era board for the non-reasoning variant, the fastest listed provider was Makora at 241.4 output tokens per second. Our measured number was 362.86, a single-stream measurement on an idle dedicated node with the exact request profile below, while leaderboard providers were measured under production load. These different workloads do not establish a like-for-like provider ranking.
The protocol behind our number: 1024 input tokens, 1024 output tokens, temperature 0, streaming, 8 timed requests per arm, both arms back to back in one window in one container on the same four B200 cards, request profiles verified identical before any number rendered. One arm is a stock deployment of the same model on the same node. The other is the v4turbo package. The only changed variable is the serving configuration. The August 2, 2026 announcement described the raw artifacts as committed.
The headline pair
362.86 tok/s
p50 with the v4turbo package. The stock deployment measured 113.37 tok/s in the same window.
3.1 s
Down from 9.2 s. Exact medians 3095.38 ms and 9248.33 ms for 1024 tokens in, 1024 out.
2.98x
Computed by the catalog platform from the median pair, never authored by hand.
232.93 ms
Against 220.21 ms stock. The package pays about 13 ms at the first token.
| run | stock deployment | v4turbo package |
|---|---|---|
| single stream, p50 | 113.37 | 362.86 |
The serving curve, including where we lose
Single-stream speed is one operating point. The serving curve is the rest. We measured 32768-token prompts with 2048-token outputs at concurrency 1, 8, and 32, per-GPU output tokens per second, both arms.
| concurrency | v4turbo package | stock deployment |
|---|---|---|
| 1 | 74.31 | 27.73 |
| 8 | 235.91 | 176.58 |
| 32 | 370.43 | 397.74 |
The measured recipe used the accelerated configuration at concurrency 1 and 8. At concurrency 32 the standard configuration held the higher aggregate, 397.74 per GPU against 370.43, or 1590.94 output tokens per second for the node. The published rule follows those measured points; it does not establish which configuration wins on every workload.
Market context on August 2, 2026
The claim section above carries the conditions; this chart is the picture. The three provider bars are what Artificial Analysis lists for the preview-era variant of this model. The highlighted bar is our measured single-stream number.
| who | our measurement, idle node, single stream | AA listing, production endpoints under load |
|---|---|---|
| v4turbo, 4x B200 | 362.86 | |
| Makora | 241.4 | |
| DeepSeek API | 107.2 | |
| CoreWeave | 40.8 |
Where the speed matters
- Coding agents and tool-use loops. Agent steps are serial, each one waits for the last, so per-user tokens per second sets wall-clock time. A 1024-token step streams in about 3.1 seconds instead of 9.2.
- Interactive assistants. A person watching tokens stream reads the difference between 113 and 363 tokens per second directly, and it costs about 13 ms at the first token.
- High-volume batch. In the measured run, the standard configuration held 397.74 output tokens per second per GPU at concurrency 32 (the accelerated configuration measured 370.43 at that point). The August 2, 2026 announcement described both configurations and the measured curve. Measure your own operating point before you commit.
Own your AI
Every closed API in the table below is a rented dependency: rented latency, rented rate limits, rented deprecation schedules. The table is what list prices look like beside a model you can own.
| API | Input $/1M | Output $/1M | Source note |
|---|---|---|---|
| Claude Fable 5 | $10.00 | $50.00 | Anthropic pricing page |
| Claude Opus 5 | $5.00 | $25.00 | Anthropic pricing page |
| Claude Sonnet 5 | $2.00 | $10.00 | Introductory through 2026-08-31, then $3.00 / $15.00 |
| gpt-5.6-sol | $5.00 | $30.00 | OpenAI Standard tier |
| gpt-5.6-terra | $2.00 | $12.00 | OpenAI Standard tier |
| gpt-5.6-luna | $0.20 | $1.20 | OpenAI Standard tier |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | Prompts up to 200k tokens; $4.00 / $18.00 above |
| DeepSeek V4 Flash, first party | $0.14 | $0.28 | Cache miss; cache-hit input $0.0028. 2x peak-hour pricing announced, effective date pending |
Serving it yourself means renting the node instead. The measured target for this package is four B200 GPUs, and three providers publish on-demand rates for exactly that.
| Provider | $ per GPU-hour | $ per 4-GPU node-hour |
|---|---|---|
| RunPod | $5.89 | $23.56 |
| Lambda | $6.79 | $27.16 |
| Nebius | $7.15 | $28.60 |
One piece of arithmetic, with its basis stated. At the measured concurrency 32 operating point one four-GPU node aggregates 1590.94 output tokens per second in the stock configuration and 1481.71 with the package. Batch is where stock wins, so the batch arithmetic uses the stock number: about 5.7 million output tokens per hour. Divide the node rent by that and a fully busy node lands between $4.11 and $4.99 per 1M output tokens, and each of those requests also carried a 32768-token prompt the same dollars paid to process. The assumptions are visible: the node stays busy, and your traffic resembles that workload. Redo this arithmetic against your own measured curve before you commit.
As of August 2, 2026, DeepSeek's own API listed $0.28 per 1M output tokens, far below that node arithmetic, so the case for self-hosting this model was not raw price per token. It is the list that follows.
- Your prompts and outputs stay on hardware you control.
- There are no rate limits except the ones you set.
- There are no model deprecations; the weights you validated today are the weights you serve next year.
- Latency is yours to control, down to where the node physically sits.
- The model is MIT licensed, so it is yours to keep. We read the license from the upstream repository on 2026-08-01.
What was announced
The August 2, 2026 announcement described v4turbo as a finished serving configuration for DeepSeek V4 Flash 0731 on a four B200 node. No current measured optimization package is published for it.
The measured serving results and reproduction context are documented in this article.
Sources
- Measured performance numbers: RunInfra measurement session of 2026-08-01, one 4x B200 node, with the measurement conditions documented in this article.
- Artificial Analysis, deepseek-v4-flash provider board (no provider speed benchmarks; sole listed provider DeepSeek at $0.14 in / $0.28 out): https://artificialanalysis.ai/models/deepseek-v4-flash/providers, accessed 2026-08-02.
- Artificial Analysis, deepseek-v4-flash-non-reasoning provider board (Makora 241.4 tok/s, DeepSeek 107.2 tok/s, CoreWeave 40.8 tok/s): https://artificialanalysis.ai/models/deepseek-v4-flash-non-reasoning/providers, accessed 2026-08-02.
- Anthropic pricing (Claude Fable 5 $10.00 in / $50.00 out, Claude Opus 5 $5.00 / $25.00, Claude Sonnet 5 $2.00 / $10.00 introductory through 2026-08-31, then $3.00 / $15.00): https://platform.claude.com/docs/en/about-claude/pricing, accessed 2026-08-02.
- OpenAI pricing, Standard tier (gpt-5.6-sol $5.00 in / $30.00 out, gpt-5.6-terra $2.00 / $12.00, gpt-5.6-luna $0.20 / $1.20): https://developers.openai.com/api/docs/pricing, accessed 2026-08-02.
- Google pricing (Gemini 3.1 Pro Preview $2.00 in / $12.00 out for prompts up to 200k tokens, $4.00 / $18.00 above): https://ai.google.dev/gemini-api/docs/pricing, accessed 2026-08-02.
- DeepSeek API pricing (deepseek-v4-flash $0.14 in on cache miss, $0.0028 on cache hit, $0.28 out; 2x peak-hour pricing announced, effective date pending): https://api-docs.deepseek.com/quick_start/pricing, accessed 2026-08-02.
- 4x B200 on-demand rental: RunPod $5.89 per GPU-hour (https://www.runpod.io/pricing), Lambda $6.79 per GPU-hour (https://lambda.ai/service/gpu-cloud/pricing), Nebius $7.15 per GPU-hour (https://nebius.com/prices), all accessed 2026-08-02.
- DeepSeek-V4-Flash-0731 weights and MIT license: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731, license read 2026-08-01.
References
- 01 Artificial Analysis, deepseek-v4-flash provider board
- 02 Artificial Analysis, deepseek-v4-flash-non-reasoning provider board
- 03 Anthropic API pricing
- 04 OpenAI API pricing
- 05 Google Gemini API pricing
- 06 DeepSeek API pricing
- 07 Lambda GPU cloud pricing
- 08 RunPod pricing
- 09 Nebius pricing
- 10 DeepSeek-V4-Flash-0731 weights and MIT license
Written by

