NVIDIA GPUs for open model inference
Under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026, Kimi K3: 2.12x faster; 19,090 ms at baseline and 8,978 ms optimized. This page covers 1 published package on this GPU and keeps every claim tied to its recorded engine, concurrency, and date. Other GPU, engine, and unlisted-concurrency measurements are not published.
We measured 1 published catalog package on this GPU. The rows below keep every claim tied to its recorded engine, concurrency, and verified date.
View Model APIsThe one B300 package is a serving recipe on unchanged weights, and the specification rows below carry the vendor's published maximums, never our measurements.
GPU time is not currently sold. Historical rates are reference values, not an offer. Active-tier rates refer to reserved deployments. No historical rate is published for this GPU.
Memory and bandwidth are the vendor's published figures in the vendor's own wording, retrieved Sep 21, 2026, source. They are not our measurements.
Packages are no longer sold. For hosted inference, see Model APIs.
| Model | Technique | P50 latency | Throughput | Speed verdict | Accuracy |
|---|---|---|---|---|---|
| Kimi K3 moonshotai/Kimi-K3 Engine: vLLM 0.23.1 Concurrency 1 License: Kimi K3 License Verified: Jul 31, 2026 | TechniqueNone (weights unchanged) | P50 latency 19,090 ms baseline 8,978 ms optimized P99 latency 19,377 ms baseline 9,595 ms optimized | Throughput 55.2 tokens/s baseline 120 tokens/s optimized | Speed verdict2.12x faster | Accuracy Same weights, accuracy unchanged |
Published open-model inference measurements on NVIDIA B300, with baseline and optimized values only where the catalog carries both.
Measured by RunInfra on rented B300 hardware with vLLM 0.23.1; each package kit includes a signed benchmark receipt.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Published package measurements and signed kit records. https://runinfra.ai/gpu.
Measurement policy: Methodology.
Other GPU measurements for these listed packages not published.
Other engine measurements for these listed packages not published.
Measurements for concurrency not listed in the table for these packages not published.
Kimi K3 p50 latency measured 19,090 ms at baseline and 8,978 ms optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.
Kimi K3 p99 latency measured 19,377 ms at baseline and 9,595 ms optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.
Kimi K3 throughput measured 55.2 tokens/s at baseline and 120 tokens/s optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs