NVIDIA GPUs for open model inference
Under concurrency 8, vLLM 0.25.1, measurement verified Jul 25, 2026, Qwythos-9B-Claude-Mythos-5-1M: 1.29x faster; 1,058 ms at baseline and 815 ms optimized. This page covers 3 published packages on this GPU and keeps every claim tied to its recorded engine, concurrency, and date. Other GPU, engine, and unlisted-concurrency measurements are not published.
We measured 3 published catalog packages on this GPU. The rows below keep every claim tied to its recorded engine, concurrency, and verified date.
View Model APIsAll three H100 packages use channelwise FP8. Qwen3.6 27B and Qwythos stay within one percent of BF16 accuracy; AREX-Turbo recorded a higher gsm8k score. The paired scores below show the difference, without establishing statistical significance.
GPU time is not currently sold. Historical rates are reference values, not an offer. Active-tier rates refer to reserved deployments. Rate source: GPU rate catalog, captured Aug 9, 2026.
Memory and bandwidth are the vendor's published figures in the vendor's own wording, retrieved Aug 8, 2026, source. They are not our measurements.
Packages are no longer sold. For hosted inference, see Model APIs.
| Model | Technique | P50 latency | Throughput | Speed verdict | Accuracy |
|---|---|---|---|---|---|
| AREX-Turbo BAAI/AREX-Turbo Engine: vLLM 0.25.1 Concurrency 8 License: Apache-2.0 Verified: Jul 27, 2026 | TechniqueChannelwise FP8 weights, dynamic per-token activations | P50 latency 666 ms baseline 551 ms optimized P99 latency 684 ms baseline 579 ms optimized | Throughput 1,534 tokens/s baseline 1,847 tokens/s optimized | Speed verdict1.2x faster | Accuracy gsm8k no measurable accuracy change, passed gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2): 0.3821 baseline to 0.4094 optimized |
| Qwen3.6 27B Qwen/Qwen3.6-27B Engine: vLLM 0.25.1 Concurrency 8 License: Apache-2.0 Verified: Jul 25, 2026 | TechniqueChannelwise FP8, measured selective-layer recipe | P50 latency 2,857 ms baseline 2,215 ms optimized P99 latency 2,881 ms baseline 2,230 ms optimized | Throughput 361 tokens/s baseline 465 tokens/s optimized | Speed verdict1.28x faster | Accuracy gsm8k 99.87% recovery, passed gsm8k exact_match strict (N=1319, completion protocol): 0.5754 baseline to 0.5747 optimized |
| Qwythos-9B-Claude-Mythos-5-1M empero-ai/Qwythos-9B-Claude-Mythos-5-1M Engine: vLLM 0.25.1 Concurrency 8 License: Apache-2.0 Verified: Jul 25, 2026 | TechniqueChannelwise FP8 weights, dynamic per-token activations | P50 latency 1,058 ms baseline 815 ms optimized P99 latency 1,083 ms baseline 861 ms optimized | Throughput 974 tokens/s baseline 1,255 tokens/s optimized | Speed verdict1.29x faster | Accuracy gsm8k 99.35% recovery, passed gsm8k exact_match strict (N=1319, completion protocol): 0.8378 baseline to 0.8324 optimized |
Published open-model inference measurements on NVIDIA H100, with baseline and optimized values only where the catalog carries both.
Measured by RunInfra on rented H100 hardware with vLLM 0.25.1; each package kit includes a signed benchmark receipt.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Published package measurements and signed kit records. https://runinfra.ai/gpu.
Measurement policy: Methodology.
Other GPU measurements for these listed packages not published.
Other engine measurements for these listed packages not published.
Measurements for concurrency not listed in the table for these packages not published.
AREX-Turbo p50 latency measured 666 ms at baseline and 551 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 27, 2026. Package verdict: 1.2x faster.
Qwen3.6 27B p50 latency measured 2,857 ms at baseline and 2,215 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 25, 2026. Package verdict: 1.28x faster.
Qwythos-9B-Claude-Mythos-5-1M p50 latency measured 1,058 ms at baseline and 815 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 25, 2026. Package verdict: 1.29x faster.
AREX-Turbo p99 latency measured 684 ms at baseline and 579 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 27, 2026. Package verdict: 1.2x faster.
AREX-Turbo throughput measured 1,534 tokens/s at baseline and 1,847 tokens/s optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 27, 2026. Package verdict: 1.2x faster.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs