vLLM vs SGLang
vLLM
5,333 output tokens/s
2,221 ms TTFT p50
SGLang
5,235 output tokens/s
2,543 ms TTFT p50
Conditions and versions
Measured sweep Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16; derived cost compared on H100 BF16, H100 FP8, and L40S BF16 only; engine versions vLLM 0.23.0 and SGLang 0.5.13; as of Jun 20, 2026.