H100 BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026: SGLang 5,235 versus TensorRT-LLM 4,813 output tokens/s. Three timed repeats per concurrency without published confidence intervals do not establish a statistically resolved ordering or a universal ranking.
H100 BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026: SGLang 5,235 versus TensorRT-LLM 4,813 output tokens/s. TensorRT-LLM recorded lower TTFT at five of six concurrencies (1, 8, 32, 64, and 128), SGLang at 256. Same condition: SGLang 2,543 versus TensorRT-LLM 2,848 ms TTFT p50. Derived cost, same condition: SGLang $0.21 versus TensorRT-LLM $0.228 per million output tokens. From three repeats, confidence intervals unpublished. No universal winner is claimed.
Throughput and TTFT were measured on NVIDIA H100 80GB at BF16. Derived cost compared on H100 BF16 only. The comparison uses Llama 3.1 8B Instruct, SGLang 0.5.13 and TensorRT-LLM 1.2.1, as of Jun 20, 2026. RunInfra's later pages record SGLang 0.5.16 in B200 article (read Sep 21, 2026). Those newer versions were not compared in this June sweep. Sources: B200 article.
The recorded ordering changes by metric and load; three repeats without published confidence intervals do not establish a universal ranking.
We compare measured throughput and TTFT p50 at every published concurrency, then show derived cost cells at each published configuration. Every ratio sits beside both source values.
Llama 3.1 8B Instruct, NVIDIA H100 80GB, BF16, Jun 20, 2026. Three timed repeats; confidence intervals unpublished. Throughput: mean of repeats. TTFT: p50. Ratios are arithmetic comparisons, not statistical significance.
Concurrency (log scale). SGLang: solid; TensorRT-LLM: dashed. Zero-based vertical scale.
Concurrency (log scale). SGLang: solid; TensorRT-LLM: dashed. Zero-based vertical scale.
Measured throughput and TTFT use ratios from the two absolute values. Cost is derived from the recorded GPU rate; no cross-engine cost ratio is defined.
| Metric and condition | SGLang | TensorRT-LLM | Conditional verdict |
|---|---|---|---|
| Highest measured throughput on NVIDIA H100 80GB at BF16Shared sweep condition belowoutput tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off | SGLang 5,235 output tokens/s | TensorRT-LLM 4,813 output tokens/s | SGLang 8.1% higher; significance unknown Evidence and calculationSGLang recorded 5,235 output tokens/s versus TensorRT-LLM at 4,813 output tokens/s, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 422 output tokens/s (8.1% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.09x, an arithmetic comparison only.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off |
| TTFT p50 at concurrency 1Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | SGLang 41 ms | TensorRT-LLM 32 ms | TensorRT-LLM 22.0% lower; significance unknown Evidence and calculationSGLang recorded 41 ms versus TensorRT-LLM at 32 ms, under NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 9 ms (22.0% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.28x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 1, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 8Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | SGLang 221 ms | TensorRT-LLM 56 ms | TensorRT-LLM 74.7% lower; significance unknown Evidence and calculationSGLang recorded 221 ms versus TensorRT-LLM at 56 ms, under NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 165 ms (74.7% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 3.95x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 8, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 32Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | SGLang 597 ms | TensorRT-LLM 235 ms | TensorRT-LLM 60.6% lower; significance unknown Evidence and calculationSGLang recorded 597 ms versus TensorRT-LLM at 235 ms, under NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 362 ms (60.6% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 2.54x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 32, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 64Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | SGLang 997 ms | TensorRT-LLM 504 ms | TensorRT-LLM 49.4% lower; significance unknown Evidence and calculationSGLang recorded 997 ms versus TensorRT-LLM at 504 ms, under NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 493 ms (49.4% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.98x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 64, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 128Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | SGLang 1,761 ms | TensorRT-LLM 1,691 ms | TensorRT-LLM 4.0% lower; significance unknown Evidence and calculationSGLang recorded 1,761 ms versus TensorRT-LLM at 1,691 ms, under NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 70 ms (4.0% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.04x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 128, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 256Shared sweep condition belowp50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | SGLang 2,543 ms | TensorRT-LLM 2,848 ms | SGLang 10.7% lower; significance unknown Evidence and calculationSGLang recorded 2,543 ms versus TensorRT-LLM at 2,848 ms, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The recorded difference is 305 ms (10.7% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.12x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, BF16, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026Derived from the recorded GPU rate and measured throughput at the recorded cost concurrency, not measured directly. | SGLang $0.21 BF16, derived | TensorRT-LLM $0.228 BF16, derived | 7.9% rounded cost gap Cost basisPublished derived costs: SGLang $0.21 versus TensorRT-LLM $0.228, under NVIDIA H100 80GB, BF16, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The gap between these rounded derived costs is 7.9% of the larger cost. The published throughput inputs differ by 8.1% of the larger mean. The throughput method uses three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, FP8, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026Derived from the recorded GPU rate and measured throughput at the recorded cost concurrency, not measured directly. | SGLang $0.17 FP8, derived | TensorRT-LLM TensorRT-LLM at fp8 was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | Comparison unavailable Cost basisWe do not calculate a ratio because one published cost cell is absent. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA L40S 48GB, BF16, recorded cost concurrency 128, Llama 3.1 8B Instruct, as of Jun 20, 2026Derived from the recorded GPU rate and measured throughput at the recorded cost concurrency, not measured directly. | SGLang $0.438 BF16, derived | TensorRT-LLM TensorRT-LLM on the L40S was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | Comparison unavailable Cost basisWe do not calculate a ratio because one published cost cell is absent. |
We keep measured throughput on its own scale. Derived cost never shares this chart.
We separate one shared prefix from many distinct prefixes because the source scopes those workloads differently.
We did not measure TensorRT-LLM's prefix cache. Both prefix experiments ran vLLM and SGLang only.
TensorRT-LLM 1.2.1 ran the PyTorch backend, its only execution backend. NVIDIA's TensorRT-LLM 1.2 release notes document removal of the TensorRT backend and engine-build CLI.
We measured throughput and TTFT only for Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16. Derived cost is compared on H100 BF16 only. Single-engine cost rows do not establish a comparison. We did not test engine versions newer than SGLang 0.5.13 and TensorRT-LLM 1.2.1 on the Jun 20, 2026 as-of date.
The higher recorded throughput mean at the highest measured point was for SGLang. SGLang recorded 5,235 output tokens/s; TensorRT-LLM recorded 4,813 output tokens/s. The difference is 8.1% of the larger mean. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
The lower recorded TTFT p50 mean was for SGLang. SGLang recorded 2,543 ms; TensorRT-LLM recorded 2,848 ms. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
Published derived costs: SGLang $0.21; TensorRT-LLM $0.228 USD per 1M output tokens. Condition: NVIDIA H100 80GB, BF16, recorded cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. Both values are derived from the recorded GPU rate and measured throughput at the recorded cost concurrency. The gap between these rounded derived costs is 7.9% of the larger cost. The published throughput inputs differ by 8.1% of the larger mean. The throughput method uses three timed repeats; confidence intervals are unpublished, so statistical significance is unknown.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs