H100, Channelwise FP8 weights, dynamic per-token activationsarex-turbo-fp8cd-h100-vllm
Published Jul 27, 2026. Proof verified Jul 27, 2026.
How this was measured
chat-128
vLLM 0.25.1. Concurrency 8. Proof verified Jul 27, 2026.
1.2x faster
Median latency measured 666 ms at baseline and 551 ms optimized.
Quality evidence
gsm8k no measurable accuracy change, passed.