TensorRT-LLM 1.2.1 on NVIDIA H100 80GB BF16: 4,813 output tokens/s and 2,848 ms TTFT p50 at concurrency 256 (Jun 20, 2026). Llama 3.1 8B Instruct, three timed repeats, confidence intervals unpublished. Throughput is the mean; TTFT is p50. Sweep, cost, prefix-cache results and published packages follow.
RunInfra measured TensorRT-LLM 1.2.1 on Llama 3.1 8B Instruct with NVIDIA H100 80GB at BF16 across 6 published concurrency points, as of Jun 20, 2026. The page reports measured output throughput and TTFT, keeps derived cost on its separate source basis, and keeps published conditions attached to every row. No published catalog package currently serves on this engine, so comparison evidence is not presented as package availability.
TensorRT-LLM's PyTorch-backend run makes backend scope part of every throughput and latency reading.
Source article and methodMeasured throughput and TTFT remain separate from derived cost. Every comparison ratio shows both source values and its full condition.
Scroll horizontally for all columns.
| Concurrency | Measured throughput | Measured TTFT p50 | Full condition |
|---|---|---|---|
| 1 | 152 output tokens/s | 32 ms | Concurrency 1, shared sweep condition above |
| 8 | 1,027 output tokens/s | 56 ms | Concurrency 8, shared sweep condition above |
| 32 | 2,978 output tokens/s | 235 ms | Concurrency 32, shared sweep condition above |
| 64 | 4,029 output tokens/s | 504 ms | Concurrency 64, shared sweep condition above |
| 128 | 4,774 output tokens/s | 1,691 ms | Concurrency 128, shared sweep condition above |
| 256 | 4,813 output tokens/s | 2,848 ms | Concurrency 256, shared sweep condition above |
output throughput: TensorRT-LLM 4,813, vLLM 5,333 output tokens/s. 9.8% gap; significance unknown.
Recorded means for output throughput: vLLM recorded 5,333 output tokens/s versus TensorRT-LLM at 4,813 output tokens/s, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 9.8% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.11x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off
TTFT p50: TensorRT-LLM 2,848, vLLM 2,221 ms. 22.0% gap; significance unknown.
Recorded means for TTFT p50: vLLM recorded 2,221 ms versus TensorRT-LLM at 2,848 ms, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 22.0% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.28x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column
output throughput: TensorRT-LLM 4,813, SGLang 5,235 output tokens/s. 8.1% gap; significance unknown.
Recorded means for output throughput: SGLang recorded 5,235 output tokens/s versus TensorRT-LLM at 4,813 output tokens/s, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 8.1% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.09x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off
TTFT p50: TensorRT-LLM 2,848, SGLang 2,543 ms. 10.7% gap; significance unknown.
Recorded means for TTFT p50: SGLang recorded 2,543 ms versus TensorRT-LLM at 2,848 ms, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. The difference is 10.7% of the larger mean, from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The ratio is 1.12x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column
Scroll horizontally for all columns.
| Configuration | Published value or absence | Full condition |
|---|---|---|
| NVIDIA H100 80GB BF16 | $0.228 USD per 1M output tokens, derived | NVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026 |
| NVIDIA H100 80GB FP8 | TensorRT-LLM at fp8 was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | NVIDIA H100 80GB, FP8, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026 |
| NVIDIA L40S 48GB BF16 | TensorRT-LLM on the L40S was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | NVIDIA L40S 48GB, BF16, source-selected cost concurrency 128, Llama 3.1 8B Instruct, as of Jun 20, 2026 |
We did not measure TensorRT-LLM's prefix cache. Both prefix experiments ran vLLM and SGLang only.
One client drives all three engines with identical request streams, so the timing definitions are the same everywhere. Time to first token is the time to the first streamed token. Throughput is output tokens per second. Warm up, then three timed repeats per operating point, charted as the mean. Raw repeats and confidence intervals are not publicly available. Prefix caching is off for the unique-prompt sweeps. Concurrency swept from 1 to 256 at 1,024 input and 256 output tokens, unique prompts, prefix caching off. Same weights and same request stream across engines, single GPU, tensor-parallel size 1. 0 published package rows retain their recorded engine version, concurrency, and verified date.
Where package measurements are listed: We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Source article and method, Jun 20, 2026..
These rows come from the published catalog, not the comparison sweep. Each package retains its own model, engine version, concurrency, and verified date.
No published package currently serves on this engine.
Recorded means for output throughput. TensorRT-LLM recorded 4,813 output tokens/s; vLLM recorded 5,333 output tokens/s. The difference is 9.8% of the larger mean. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
Recorded means for TTFT p50. TensorRT-LLM recorded 2,848 ms; SGLang recorded 2,543 ms. The difference is 10.7% of the larger mean. Three timed repeats, confidence intervals unpublished; statistical significance is unknown. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026.
We report these derived costs. TensorRT-LLM was $0.228; vLLM was $0.206 USD per 1M output tokens. Condition: NVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of Jun 20, 2026. Both values are derived from the recorded GPU rate and measured saturation throughput. Rounded cost differences do not resolve a statistically significant ranking.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs