vLLM 0.23.0 on NVIDIA H100 80GB BF16: 5,333 output tokens/s and 2,221 ms TTFT p50 at concurrency 256 (Jun 20, 2026). Llama 3.1 8B Instruct, three timed repeats, confidence intervals unpublished. Throughput is the mean; TTFT is p50. Sweep, cost, prefix-cache results and published packages follow.
RunInfra measured vLLM 0.23.0 on Llama 3.1 8B Instruct with NVIDIA H100 80GB at BF16 across 6 published concurrency points, as of Jun 20, 2026. The page reports measured output throughput and TTFT, keeps derived cost on its separate source basis, and keeps prefix-cache results tied to their published prefix conditions. 4 published packages appear below, each with its own engine version, concurrency, and verified date.
vLLM changes position across load and cache shape, so each decision stays tied to its measured condition.
Source article and methodMeasured throughput and TTFT remain separate from derived cost. Every comparison ratio shows both source values and its full condition.
Scroll horizontally for all columns.
| Concurrency | Measured throughput | Measured TTFT p50 | Full condition |
|---|---|---|---|
| 1 | 158 output tokens/s | 35 ms | Concurrency 1, shared sweep condition above |
| 8 | 1,096 output tokens/s | 177 ms | Concurrency 8, shared sweep condition above |
| 32 | 2,921 output tokens/s | 514 ms | Concurrency 32, shared sweep condition above |
| 64 | 4,040 output tokens/s | 790 ms | Concurrency 64, shared sweep condition above |
| 128 | 4,943 output tokens/s | 1,658 ms | Concurrency 128, shared sweep condition above |
| 256 | 5,333 output tokens/s | 2,221 ms | Concurrency 256, shared sweep condition above |
output throughput: vLLM 5,333, SGLang 5,235 output tokens/s. 1.8% gap; significance unknown.
Recorded means for output throughput: vLLM recorded 5,333 output tokens/s versus SGLang at 5,235 output tokens/s, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 1.8% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off
TTFT p50: vLLM 2,221, SGLang 2,543 ms. 12.7% gap; significance unknown.
Recorded means for TTFT p50: vLLM recorded 2,221 ms versus SGLang at 2,543 ms, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 12.7% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.14x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column
output throughput: vLLM 5,333, TensorRT-LLM 4,813 output tokens/s. 9.8% gap; significance unknown.
Recorded means for output throughput: vLLM recorded 5,333 output tokens/s versus TensorRT-LLM at 4,813 output tokens/s, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 9.8% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.11x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off
TTFT p50: vLLM 2,221, TensorRT-LLM 2,848 ms. 22.0% gap; significance unknown.
Recorded means for TTFT p50: vLLM recorded 2,221 ms versus TensorRT-LLM at 2,848 ms, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 22.0% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.28x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column
Scroll horizontally for all columns.
| Configuration | Published value or absence | Full condition |
|---|---|---|
| NVIDIA H100 80GB BF16 | $0.206 USD per 1M output tokens, derived | NVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026 |
| NVIDIA H100 80GB FP8 | $0.158 USD per 1M output tokens, derived | NVIDIA H100 80GB, FP8, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026 |
| NVIDIA L40S 48GB BF16 | $0.425 USD per 1M output tokens, derived | NVIDIA L40S 48GB, BF16, source-selected cost concurrency 128, Llama 3.1 8B Instruct, as of Jun 20, 2026 |
A single shared prefix is the easy case that any block-level cache handles well. It is not the workload SGLang's RadixAttention is built for. At zero hits, cache-on throughput exceeds cache-off by 9.8% for vLLM (999 versus 910 output tokens per second) and 6.6% for SGLang (937 versus 879 output tokens per second). This gap is unexplained by the published data; it is not evidence of cache reuse.
Scroll horizontally for all columns.
| Cache state | Hit rate | Measured throughput | Full condition |
|---|---|---|---|
| Cache on | 0 percent | 999 output tokens/s | One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026. |
| Cache off | 0 percent | 910 output tokens/s | One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026. |
| Cache on | 50 percent | 1,559 output tokens/s | One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026. |
| Cache off | 50 percent | 912 output tokens/s | One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026. |
| Cache on | 90 percent | 2,855 output tokens/s | One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026. |
| Cache off | 90 percent | 910 output tokens/s | One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026. |
The highest published hit rate compares vLLM, cache on, 90 percent hit rate at 2,855 output tokens/s with vLLM, cache on, zero hit rate at 999 output tokens/s, under One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026.. The render-derived ratio is 2.86x.derived at render time for vLLM from its own cache-on throughput at a 90 percent hit rate against its cache-on throughput at a zero hit rate, output tokens per second, mean of three timed repeats after warmup, cache state per row
RadixAttention may pull ahead with longer prefixes, deeper trees, or heavier eviction pressure than we tested. On our test it did not, and we are not going to claim otherwise.
Scroll horizontally for all columns.
| Distinct prefixes | Measured throughput | Full condition |
|---|---|---|
| 8 | 2,677 output tokens/s | 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. NVIDIA H100 80GB. Precision was not published. As of Jun 20, 2026. |
| 32 | 2,277 output tokens/s | 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. NVIDIA H100 80GB. Precision was not published. As of Jun 20, 2026. |
| 128 | 1,467 output tokens/s | 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. NVIDIA H100 80GB. Precision was not published. As of Jun 20, 2026. |
One client drives all three engines with identical request streams, so the timing definitions are the same everywhere. Time to first token is the time to the first streamed token. Throughput is output tokens per second. Warm up, then three timed repeats per operating point, charted as the mean. Raw repeats and confidence intervals are not publicly available. Prefix caching is off for the unique-prompt sweeps. Concurrency swept from 1 to 256 at 1,024 input and 256 output tokens, unique prompts, prefix caching off. Same weights and same request stream across engines, single GPU, tensor-parallel size 1. 4 published package rows retain their recorded engine version, concurrency, and verified date.
Where package measurements are listed: We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Source article and method, Jun 20, 2026..
These rows come from the published catalog, not the comparison sweep. Each package retains its own model, engine version, concurrency, and verified date.
Packages are no longer sold. For hosted inference, see Model APIs.
Scroll horizontally for all columns.
| Model | Published facts | P50 latency | P99 latency | Throughput | Recorded conditions |
|---|---|---|---|---|---|
| AREX-TurboBAAI/AREX-Turbo |
| 666 ms baseline 551 ms optimized | 684 ms baseline 579 ms optimized | 1,534 tokens/s baseline 1,847 tokens/s optimized | vLLM 0.25.1 Concurrency 8 Jul 27, 2026 1.2x faster gsm8k no measurable accuracy change, passed Channelwise FP8 weights, dynamic per-token activations |
| Kimi K3moonshotai/Kimi-K3 |
| 19,090 ms baseline 8,978 ms optimized | 19,377 ms baseline 9,595 ms optimized | 55.2 tokens/s baseline 120 tokens/s optimized | vLLM 0.23.1 (unverified, no PyPI release found September 21, 2026). Source: PyPI release index. Concurrency 1 Jul 31, 2026 2.12x faster Parity by construction Not published |
| Qwen3.6 27BQwen/Qwen3.6-27B |
| 2,857 ms baseline 2,215 ms optimized | 2,881 ms baseline 2,230 ms optimized | 361 tokens/s baseline 465 tokens/s optimized | vLLM 0.25.1 Concurrency 8 Jul 25, 2026 1.28x faster gsm8k 99.87% recovery, passed Channelwise FP8, measured selective-layer recipe |
| Qwythos-9B-Claude-Mythos-5-1Mempero-ai/Qwythos-9B-Claude-Mythos-5-1M |
| 1,058 ms baseline 815 ms optimized | 1,083 ms baseline 861 ms optimized | 974 tokens/s baseline 1,255 tokens/s optimized | vLLM 0.25.1 Concurrency 8 Jul 25, 2026 1.29x faster gsm8k 99.35% recovery, passed Channelwise FP8 weights, dynamic per-token activations |
Recorded means for output throughput. vLLM recorded 5,333 output tokens/s; SGLang recorded 5,235 output tokens/s. The difference is 1.8% of the larger mean. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
Recorded means for TTFT p50. vLLM recorded 2,221 ms; TensorRT-LLM recorded 2,848 ms. The difference is 22.0% of the larger mean. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
We report these derived costs. vLLM was $0.206; SGLang was $0.21 USD per 1M output tokens. Condition: NVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. Both values are derived from the recorded GPU rate and measured saturation throughput. Rounded cost differences do not resolve a statistically significant ranking.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs