Measured speed and accuracy on open models
Under H100, vLLM 0.25.1, request profile chat-128, concurrency 8, measured basis, verified Jul 25, 2026, Qwythos-9B-Claude-Mythos-5-1M with Channelwise FP8 weights, dynamic per-token activations: 1.29x faster; 1,058 ms at baseline and 815 ms optimized. This page covers 2 published packages using this technique and keeps every claim tied to its recorded run. Other tasks, GPUs, and serving engines are not published.
Channelwise FP8 stores weights in eight-bit floating point and scales activations for each token at runtime.
We measured 2 published packages across 1 GPU target, using the recorded baseline and optimized conditions.
What made these builds shippable is not the FP8 itself but the measured accuracy verdict beside it, taken on the same protocol as the baseline.
We show absolute baseline and optimized values before the shared verdict. We keep each condition attached to its result.
Packages are no longer sold. For hosted inference, see Model APIs.
Scroll horizontally for all columns.
| Model | Target | Throughput | Median latency | Verdict and conditions | Accuracy | Verified |
|---|---|---|---|---|---|---|
| AREX-TurboBAAI/AREX-Turbo | H100vLLM 0.25.1 | 1,534 tokens/s baseline1,847 tokens/s optimized | 666 ms p50 baseline551 ms p50 optimized | 1.2x faster.Both runs: chat-128, concurrency 8, measured. | Baseline 0.3821, optimized 0.4094.gsm8k no measurable accuracy change, passedStandard errors: baseline +/- 0.0134; optimized +/- 0.0135. | Verified Jul 27, 2026 |
| Qwythos-9B-Claude-Mythos-5-1Mempero-ai/Qwythos-9B-Claude-Mythos-5-1M | H100vLLM 0.25.1 | 974 tokens/s baseline1,255 tokens/s optimized | 1,058 ms p50 baseline815 ms p50 optimized | 1.29x faster.Both runs: chat-128, concurrency 8, measured. | Baseline 0.8378, optimized 0.8324.gsm8k 99.35% recovery, passedStandard errors not published. | Verified Jul 25, 2026 |
We restate the published throughput pairs. The neutral fills do not declare a winner.
1.20x throughput.
1.28x throughput.
We do not call a score change an improvement when the gap is smaller than about twice the combined standard errors.
AREX-Turbo measured 0.3821 at baseline and 0.4094 optimized on gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2). No measurable change in accuracy. The difference between the baseline and optimized scores is smaller than the combined measurement error, so it is within measurement noise. Standard errors: baseline +/- 0.0134; optimized +/- 0.0135.
Qwythos-9B-Claude-Mythos-5-1M measured 0.8378 at baseline and 0.8324 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. The optimized score measured below the baseline, at 99.35% of it. Neither score published a standard error, so this difference cannot be separated from sampling noise. The difference is not stated as a regression. Standard errors not published.
AREX-Turbo measured 0.3821 at baseline and 0.4094 optimized on gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2). No measurable change in accuracy. Standard errors: baseline +/- 0.0134; optimized +/- 0.0135. Qwythos-9B-Claude-Mythos-5-1M measured 0.8378 at baseline and 0.8324 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published.
AREX-Turbo: Throughput measured 1,534 tokens/s at baseline and 1,847 tokens/s optimized. P50 latency measured 666 ms at baseline and 551 ms optimized. Both runs: chat-128, concurrency 8, measured. Qwythos-9B-Claude-Mythos-5-1M: Throughput measured 974 tokens/s at baseline and 1,255 tokens/s optimized. P50 latency measured 1,058 ms at baseline and 815 ms optimized. Both runs: chat-128, concurrency 8, measured.
AREX-Turbo and Qwythos-9B-Claude-Mythos-5-1M were measured on H100 using vLLM 0.25.1.
Tasks beyond gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2) and gsm8k exact_match strict (N=1319, completion protocol) not published. GPUs beyond H100 not published. Serving engines beyond vLLM 0.25.1 not published.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs