H100 BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026: vLLM 5,333 versus SGLang 5,235 output tokens/s, a 1.8% gap relative to the larger mean. Three timed repeats per concurrency without published confidence intervals do not establish a statistically resolved ordering or a universal ranking.
H100 BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026: vLLM 5,333 versus SGLang 5,235 output tokens/s, a 1.8% gap relative to the larger mean. Same condition: vLLM 2,221 versus SGLang 2,543 ms TTFT p50. Derived cost, same condition: vLLM $0.206 versus SGLang $0.21 per million output tokens. From three repeats, confidence intervals unpublished. No universal winner is claimed.
Throughput and TTFT were measured on NVIDIA H100 80GB at BF16. Derived cost compared on H100 BF16, H100 FP8, and L40S BF16 only. The comparison uses Llama 3.1 8B Instruct, vLLM 0.23.0 and SGLang 0.5.13, as of Jun 20, 2026. RunInfra's later pages record vLLM 0.25.1 in published model packages and SGLang 0.5.16 in B200 article (read Sep 21, 2026). Those newer versions were not compared in this June sweep. Sources: published model packages, B200 article.
The published vLLM and SGLang throughput means at the highest measured concurrency are close; three repeats without published confidence intervals do not establish a statistically resolved ordering.
We compare measured throughput and TTFT p50 at every published concurrency, then show derived cost cells at each published configuration. Every ratio sits beside both source values.
Llama 3.1 8B Instruct, NVIDIA H100 80GB, BF16, Jun 20, 2026. Three timed repeats; confidence intervals unpublished. Throughput: mean of repeats. TTFT: p50. Ratios are arithmetic comparisons, not statistical significance.
Concurrency (log scale). vLLM: solid; SGLang: dashed. Zero-based vertical scale.
Concurrency (log scale). vLLM: solid; SGLang: dashed. Zero-based vertical scale.
Measured throughput and TTFT use ratios from the two absolute values. Cost is derived from the recorded GPU rate; no cross-engine cost ratio is defined.
| Metric and condition | vLLM | SGLang | Conditional verdict |
|---|---|---|---|
| Highest measured throughput on NVIDIA H100 80GB at BF16Shared sweep condition belowoutput tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off | vLLM 5,333 output tokens/s | SGLang 5,235 output tokens/s | vLLM 1.8% higher; significance unknown Evidence and calculationvLLM recorded 5,333 output tokens/s versus SGLang at 5,235 output tokens/s, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 98 output tokens/s (1.8% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| TTFT p50 at concurrency 1Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | vLLM 35 ms | SGLang 41 ms | vLLM 14.6% lower; significance unknown Evidence and calculationvLLM recorded 35 ms versus SGLang at 41 ms, under NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 6 ms (14.6% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| TTFT p50 at concurrency 8Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | vLLM 177 ms | SGLang 221 ms | vLLM 19.9% lower; significance unknown Evidence and calculationvLLM recorded 177 ms versus SGLang at 221 ms, under NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 44 ms (19.9% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.25x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 8, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 32Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | vLLM 514 ms | SGLang 597 ms | vLLM 13.9% lower; significance unknown Evidence and calculationvLLM recorded 514 ms versus SGLang at 597 ms, under NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 83 ms (13.9% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.16x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 32, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 64Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | vLLM 790 ms | SGLang 997 ms | vLLM 20.8% lower; significance unknown Evidence and calculationvLLM recorded 790 ms versus SGLang at 997 ms, under NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 207 ms (20.8% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.26x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 64, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 128Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | vLLM 1,658 ms | SGLang 1,761 ms | vLLM 5.8% lower; significance unknown Evidence and calculationvLLM recorded 1,658 ms versus SGLang at 1,761 ms, under NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 103 ms (5.8% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.06x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 128, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 256Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | vLLM 2,221 ms | SGLang 2,543 ms | vLLM 12.7% lower; significance unknown Evidence and calculationvLLM recorded 2,221 ms versus SGLang at 2,543 ms, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 322 ms (12.7% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.14x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, BF16, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026Derived from the recorded GPU rate and measured throughput at the recorded cost concurrency, not measured directly. | vLLM $0.206 BF16, derived | SGLang $0.21 BF16, derived | 1.9% rounded cost gap Cost basisPublished derived costs: vLLM $0.206 versus SGLang $0.21, under NVIDIA H100 80GB, BF16, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The gap between these rounded derived costs is 1.9% of the larger cost. The published throughput inputs differ by 1.8% of the larger mean. The throughput method uses three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, FP8, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026Derived from the recorded GPU rate and measured throughput at the recorded cost concurrency, not measured directly. | vLLM $0.158 FP8, derived | SGLang $0.17 FP8, derived | 7.1% rounded cost gap Cost basisPublished derived costs: vLLM $0.158 versus SGLang $0.17, under NVIDIA H100 80GB, FP8, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The gap between these rounded derived costs is 7.1% of the larger cost. The underlying throughput rows for this configuration are not published. The throughput method uses three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA L40S 48GB, BF16, recorded cost concurrency 128, Llama 3.1 8B Instruct, as of Jun 20, 2026Derived from the recorded GPU rate and measured throughput at the recorded cost concurrency, not measured directly. | vLLM $0.425 BF16, derived | SGLang $0.438 BF16, derived | 3.0% rounded cost gap Cost basisPublished derived costs: vLLM $0.425 versus SGLang $0.438, under NVIDIA L40S 48GB, BF16, recorded cost concurrency 128, Llama 3.1 8B Instruct, as of Jun 20, 2026. The gap between these rounded derived costs is 3.0% of the larger cost. The underlying throughput rows for this configuration are not published. The throughput method uses three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
We keep measured throughput on its own scale. Derived cost never shares this chart.
We separate one shared prefix from many distinct prefixes because the source scopes those workloads differently.
A single shared prefix is the easy case that any block-level cache handles well. It is not the workload SGLang's RadixAttention is built for. At zero hits, cache-on throughput exceeds cache-off by 9.8% for vLLM (999 versus 910 output tokens per second) and 6.6% for SGLang (937 versus 879 output tokens per second). This gap is unexplained by the published data; it is not evidence of cache reuse.
We measured higher cache-on throughput for vLLM. vLLM, cache on, 90 percent hit rate recorded 2,855 output tokens/s versus vLLM, cache on, zero hit rate at 999 output tokens/s, under One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026.. The ratio is 2.86x.derived at render time for vLLM from its own cache-on throughput at a 90 percent hit rate against its cache-on throughput at a zero hit rate, output tokens per second, mean of three timed repeats after warmup, cache state per row
We measured higher cache-on throughput for SGLang. SGLang, cache on, 90 percent hit rate recorded 2,488 output tokens/s versus SGLang, cache on, zero hit rate at 937 output tokens/s, under One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of Jun 20, 2026.. The ratio is 2.66x.derived at render time for SGLang from its own cache-on throughput at a 90 percent hit rate against its cache-on throughput at a zero hit rate, output tokens per second, mean of three timed repeats after warmup, cache state per row
RadixAttention may pull ahead with longer prefixes, deeper trees, or heavier eviction pressure than we tested. On our test it did not, and we are not going to claim otherwise.
We measured higher throughput on the many-prefix workload. vLLM recorded 2,677 output tokens/s versus SGLang at 2,493 output tokens/s, under 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. Distinct prefixes: 8. NVIDIA H100 80GB. Precision was not published. As of Jun 20, 2026.. The ratio is 1.07x.derived at render time at 8 distinct prefixes, output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on
We measured higher throughput on the many-prefix workload. vLLM recorded 2,277 output tokens/s versus SGLang at 2,086 output tokens/s, under 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. Distinct prefixes: 32. NVIDIA H100 80GB. Precision was not published. As of Jun 20, 2026.. The ratio is 1.09x.derived at render time at 32 distinct prefixes, output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on
We measured higher throughput on the many-prefix workload. vLLM recorded 1,467 output tokens/s versus SGLang at 1,376 output tokens/s, under 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. Distinct prefixes: 128. NVIDIA H100 80GB. Precision was not published. As of Jun 20, 2026.. The ratio is 1.07x.derived at render time at 128 distinct prefixes, output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on
We measured throughput and TTFT only for Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16. Derived cost is compared on H100 BF16, H100 FP8, and L40S BF16 only. Single-engine cost rows do not establish a comparison. We did not test engine versions newer than vLLM 0.23.0 and SGLang 0.5.13 on the Jun 20, 2026 as-of date.
The higher recorded throughput mean at the highest measured point was for vLLM. vLLM recorded 5,333 output tokens/s; SGLang recorded 5,235 output tokens/s. The difference is 1.8% of the larger mean. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
The lower recorded TTFT p50 mean was for vLLM. vLLM recorded 2,221 ms; SGLang recorded 2,543 ms. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
Published derived costs: vLLM $0.206; SGLang $0.21 USD per 1M output tokens. Condition: NVIDIA H100 80GB, BF16, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. Both values are derived from the recorded GPU rate and measured throughput at the recorded cost concurrency. The gap between these rounded derived costs is 1.9% of the larger cost. The published throughput inputs differ by 1.8% of the larger mean. The throughput method uses three timed repeats; confidence intervals are unpublished, so statistical significance is unknown.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs