On this page
Key takeaways
- On 2026-07-29 no released vLLM build served Kimi K3. Reaching a first token required resolving several independent failures on 8x B300.
- A one-variable, concurrency-1 comparison measured 55.28 and 118.22 output tokens per second p50 per request across the package's batch and interactive profiles, a displayed 2.1x result. These are not per-GPU values or node aggregates.
- The interactive profile's advantage measured +83.07 percent per GPU at concurrency 1 and +20.25 percent at concurrency 8, then minus 9.46 percent at concurrency 32. The earlier batch-profile switching rule is retired in version 4.
- There is no stock-vLLM speed multiplier because no released build served the model. Accuracy recovery was NOT measured, and output parity is stated by construction because no weight byte changes.
- The July 2026 kit was offered at $2,300, was 2.5 MB, contained zero weight bytes, and pinned about 1.56 TB across 96 shards. Its earlier deploy files still require the measured serve command to be transcribed before this version is complete.
Status note, September 26, 2026: package sales are closed, and kimi-k3-b300x8-k3turbo is no longer sold. For hosted inference, see Model APIs. The July 2026 measurements and package details below are historical. Version 4 retired the earlier batch-profile switching rule.
On 2026-07-29, no released vLLM build could serve Kimi K3. We checked the installable options, then worked through several distinct failures before the model produced a first token on an 8x B300 node. The result was hard won, and the proof is publishable even though the implementation recipe is not.
That work was a catalog package. This post explains the result, the measurement protocol, the operating limits, what the buyer received, and what we still refuse to claim.
The package, as offered in July
Every card below reads its GPU, engine, and measurements from the source definition at build time. If a measurement is retracted or re-run, the card changes with it.
Kimi K3
2.12x faster
Verified Jul 31, 2026.
Baseline 19,090 ms; optimized 8,978 ms.
Baseline and run conditions
Single-stream decode with one variable between the two columns of this comparison. Both columns were measured on 2026-07-31 in one window, one container, one engine install, on the same 8x B300 cards, run back to back with an idle drain between them, and the comparison verified the request profiles identical across the arms, including per-point prompt fingerprints, before rendering any number. The baseline column is the batch profile and the optimized column is the interactive profile. The throughput figures are per-request decode rates at concurrency 1, the speed one user experiences, not per-GPU figures and not node aggregates: 55.15 output tokens per second p50 for the batch profile and 119.84 for the interactive profile, which is 2.17x. Dividing either figure by eight describes nothing real because all eight cards cooperate on one stream, and multiplying by eight would invent a many-stream aggregate this profile did not measure. The profile, stated exactly: single-stream sequential, concurrency 1, 8 timed requests per arm, 1024 prompt tokens target, 1024 output tokens requested per request, temperature 0.0, ignore_eos on, streaming, a fixed seed, and 8 of 8 requests ok on each arm. The latency stats beside each column are end-to-end request times from the same runs: the interactive profile median request finishes in 9.0 seconds instead of 19.1, its time to first token is better in this window (429.30 ms against 535.06), and its time per output token is higher (23.10 ms against 18.13). Know the boundary of this result before you deploy on it: the 2.17x is a concurrency-1 fact, and the same window measured how the advantage narrows with load on 32768-token-prompt workloads, +89.59 percent per GPU at concurrency 1, +17.83 percent at concurrency 8, and flat at concurrency 32 (minus 0.83 percent, inside the measurement noise floor) where time to first token still improves. Version 3 published a different tuning of the interactive profile whose advantage INVERTED at concurrency 32, and told you to switch to the batch profile above roughly concurrency 8; a one-variable window then measured the two tunings against each other and the version 4 tuning won at every point, so that rule is retired. This is a 1024-token profile, not the 32K repo-scale profile version 2 published, and version 2 per-GPU throughput figure must not be compared with these per-request rates.
These numbers were measured on an unreleased vLLM 0.23.1 build. The kit does not ship the engine binary. It ships deploy files and a build definition, and that build definition now describes the exact build these numbers were measured on, which earlier versions named as owed work. It has not been executed on a compatible build machine, so treat your first build from it as a validation run. The deploy files still carry an earlier validated serve configuration rather than the measured one, and transcribing the measured serve command into them remains a named prerequisite. No released or nightly vLLM build served this model when these measurements were recorded.
AREX-Turbo
1.2x faster
Verified Jul 27, 2026.
Baseline 666 ms; optimized 551 ms.
Baseline and run conditions
chat-128
Baseline runtime disclosure not published.
Qwen3.6 27B
1.28x faster
Verified Jul 25, 2026.
Baseline 2,857 ms; optimized 2,215 ms.
Baseline and run conditions
chat-128
Baseline runtime disclosure not published.
Qwythos-9B-Claude-Mythos-5-1M
1.29x faster
Verified Jul 25, 2026.
Baseline 1,058 ms; optimized 815 ms.
Baseline and run conditions
chat-128
Baseline runtime disclosure not published.
Why this was hard
The failures were independent and surfaced at different stages of bring-up. Fixing one exposed another, and no public installable build provided a working starting point. That is the useful public fact: this was not a missing toggle or a one-line launch command.
We are not publishing the failure sequence, runtime internals, or launch settings. Those were the paid work product. We are publishing the measurements and limitations because buyers could judge the result before they bought it.
What we measured
Initial validation used repo-scale coding prompts of 33 to 36K tokens as reported by the server, 2048 output tokens per request, and 16 measured requests after 1 warmup. Every published result is scoped to text generation on a single 8x B300 node.
The headline comparison was a one-variable test on 2026-07-30. Both arms ran back to back in one measurement window on the same eight cards and the same vLLM installation. Each arm used 8 timed requests at concurrency 1, a 1024-token prompt target, 1024 requested output tokens, temperature 0.0, streaming, a fixed seed, and generation through the requested output target. All 8 requests completed successfully in each arm. The only changed variable was the selected operating profile.
| Metric | Batch profile | Interactive profile |
|---|---|---|
| Output tokens per second p50 | 55.28 | 118.22 |
| End-to-end latency p50 | 18,729.50 ms | 8,907.00 ms |
| End-to-end latency p95 | 18,755.05 ms | 10,168.25 ms |
| End-to-end latency p99 | 18,757.54 ms | 10,332.88 ms |
| Time to first token p50 | 226.22 ms | 244.87 ms |
| Time per output token p50 | 18.09 ms | 24.94 ms |
The serving curve
The repo-scale serving curve used 32768 prompt tokens and 2048 requested output tokens per request, temperature 0.0, streaming, a fixed seed, and one warmup request per point. It measured concurrency 1 with 4 requests, concurrency 8 with 16 requests, and concurrency 32 with 64 requests. Every request completed successfully. Concurrency 128 is not characterized on this version.
- Concurrency 1: 11.13 output tokens per second per GPU, 89.04 aggregate, time to first token p50 5707.65 ms, time per output token p50 26.46 ms, end-to-end latency p50 22629.34 ms, 4 of 4 requests completed.
- Concurrency 8: 28.74 output tokens per second per GPU, 229.93 aggregate, time to first token p50 6900.96 ms, time per output token p50 99.59 ms, end-to-end latency p50 69479.02 ms, 16 of 16 requests completed.
- Concurrency 32: 34.10 output tokens per second per GPU, 272.77 aggregate, time to first token p50 114549.13 ms, time per output token p50 175.66 ms, end-to-end latency p50 212831.87 ms, 64 of 64 requests completed. Request rate was 0.133 requests per second.
The profile advantage changes with load. The interactive profile measured +83.07 percent per GPU at concurrency 1 and +20.25 percent at concurrency 8, then minus 9.46 percent at concurrency 32. At concurrency 32, the batch profile measured 37.66 output tokens per second per GPU and 301.28 aggregate, with time to first token p50 of 7.4 seconds, versus 34.10 per GPU, 272.77 aggregate, and 114.5 seconds for the interactive profile.
What we did not measure
Accuracy recovery was NOT measured. Output parity is by construction because the package changes no weight byte. We do not claim token-level identity across operating profiles, and we publish no cross-server accuracy score. A control between two identical server instances measured 78.12 percent top-1 agreement, so that method could not distinguish a server from itself and certifies nothing.
- We did not measure image input. Every number here is text only.
- We did not evaluate refusal, toxicity, or jailbreak behaviour. The weights are unmodified, so we do not expect a shift, but we did not check.
- We did not measure anywhere near the 1 million token context the model supports.
- Peak VRAM is not published. The measured hardware was 8x B300 with 288 GB per GPU and 2304 GB across the node.
- The 8-request single-stream sample supports a median, but it is not a strong tail sample. Its p95 and p99 sit between the two slowest requests.
Why it was offered as a package
With a catalog package, the work already ran. What buyers received was the finished artifact, the measured proof, and the verifier. Buyers deployed it on their own hardware, subject to the named prerequisite below, and the receipt states exactly what was measured.
The catalog answered one question: someone already solved this model on this hardware, could a buyer purchase the measured work product? K3 is a strong example because the difficulty was the independent failure chain and the validation burden, not one isolated setting.
What the package contained
- Deploy files for the measured 8x B300 target and the interactive and batch operating profiles, with the prerequisite below stated plainly.
- The exact runtime build record and measured serve command in the signed buyer receipt delivered after purchase.
- A benchmark receipt signed with Ed25519, so the numbers you were sold are the numbers in the artifact.
- A verifier that checks the bytes you downloaded against what we measured.
- Zero weight bytes. The kit pins an immutable upstream revision and your node pulls the model from its source. We do not store, mirror, or redistribute the weights.
The K3 serving recipe was 2.5 MB. The model behind it was about 1.56 TB across 96 shards. That asymmetry mattered: the artifact captured serving engineering rather than duplicating the weights.
We did not quantize this model. The upstream weights arrived untouched, byte for byte. Claiming otherwise would be false, so the package stated that no weight byte changed and tied the download to measured checksums.
The original enterprise tier
The July 2026 package was offered at $2,300, plus $200 per month for maintenance and version updates, and required a signed enterprise licence before purchase. The buyer named the licensed entity, the signer, and their authority to bind it, then signed four separate affirmations covering redistribution, resale, and traceability. The signature was recorded append-only with the exact terms text and its hash, and it was bound to the purchase inside the same database transaction that moved the money. An entitlement under those terms could not exist without a signature.
The upstream weights remain governed by the Kimi K3 License. The package licence governs the paid serving work product. There is no technical measure that stops a determined buyer from copying configuration and documentation, and we will not pretend otherwise. The original gate provided a specific, evidenced agreement rather than an implied one.
Where this goes
Frontier open models are getting larger faster than the tooling around them. K3 is 2.8 trillion parameters total with about 104 billion active per token, across 92 mixture-of-experts layers. Day-zero support in a released inference engine is now the exception, not the rule. The window between a model release and a measured production path is where the work lives.
We saw that window as a product: a verified artifact buyers owned and ran on their own hardware, rather than a consulting engagement or an ongoing hosted rental. Enterprises that could not send weights to a third party needed someone to have already done the difficult bring-up and measurement work.
Buyers receive subsequent package versions through the maintenance subscription. Each version remains tied to its own signed receipt, measurements, checksums, known limits, and deployment prerequisites rather than silently inheriting claims from an earlier run.
Written by

