Channelwise FP8 weights, dynamic per-token activations
llm-compressor 0.12.0, data-free: no calibration corpus is read at any point, so the quantization cannot have seen evaluation data
Open models with serving tuned on real hardware. Buy once, run anywhere.
Quantized and serving-tuned for H100 on vLLM 0.25.1, with the benchmark receipt included.
You get a self-contained kit: the optimized weights, the exact serving configuration, and the measured proof.
Core plan required
We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.
Request profile
chat-128
Run the model on your infrastructure, without deprecations, silent swaps, or provider restrictions.
Buy once instead of paying per token; your deployment cost follows the GPUs you run.
Prompts and outputs stay on your network, with nothing retained elsewhere.
Delivery contract
Download
./models/kit.zipArchive root
arex-turbo-fp8cd-h100-vllm/Model directory
arex-turbo-fp8cd-h100-vllm/weights/Included
Kit contains
aca1170bd762...64f114a83895llm-compressor 0.12.0, data-free: no calibration corpus is read at any point, so the quantization cannot have seen evaluation data
the gated linear-attention (GDN) path and the vision tower are excluded from quantization and remain at BF16, which is the recipe measured lossless on this architecture family before it was reused here
the exact engine build and arguments every receipt number was produced on: vLLM 0.25.1, max_num_seqs 128, max_num_batched_tokens 16384, max-model-len 4096, chosen because it beat the wider 256/32768 anchor on both latency and throughput for this model rather than inherited from another package
Measured accuracy
(N=1319 full set, chat-template, deliberation-aware extraction v2)
Passed accuracy gate
No measurable change in accuracy
The difference between the baseline and optimized scores is smaller than the combined measurement error, so it is within measurement noise.
Methodology
Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.
RunInfra Package License v1.1 (2026-07-25)
Buy once. No meter, no expiry, nothing calling home. The full terms ship inside the kit.
This package is licensed to the purchasing workspace, one time, for the exact version purchased. Everyone in that workspace may run it in production without limits: unlimited inference, on any hardware or cloud the workspace controls, for any lawful commercial purpose. You may modify the configuration and tooling for your own use.
Everything the model produces for you belongs to you. You may use, sell, and build products on the model's outputs without restriction or royalty. Serving the model to your own customers as part of your product is use, not redistribution, and is fully allowed.
You may not redistribute, resell, sublicense, rent, publish, or otherwise make the package or its artifacts available to any third party. That covers the optimized weights, serving configuration, scripts, kit archive, and benchmark receipts, whole or in part, modified or not. The license belongs to the purchasing workspace and cannot be transferred separately from it.
The underlying model remains under its upstream open-source license, which is included in this kit with attribution and a statement of RunInfra's modifications. Nothing in this license restricts rights the upstream license grants you for the ORIGINAL model; the restrictions above apply to RunInfra's optimized package.
You purchased this exact version, proven against the exact engine version named in the receipt. It stays downloadable to your workspace and never expires. The benchmark receipt describes measurements taken at verification time on the named hardware; RunInfra does not promise future updates to this version, and later package versions are separate purchases.
If the workspace redistributes the package or its artifacts, this license terminates for that workspace. Sections about your outputs survive termination for outputs already produced.
Apache-2.0, as published by the model author on Hugging Face. The full license text ships inside the kit.
Company
© 2026 RunInfra. All rights reserved.