RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading the optimized model catalog

Optimized models, measured on real GPUs

Open models with serving tuned on real hardware. Buy once, run anywhere.

Optimized models, measured on real GPUs

Open models with serving tuned on real hardware. Buy once, run anywhere.

FiltersAll models
Type
Serving engine
GPU
DeepSeek V4 Flash
Technique not publishedvLLM 0.25.0B200
2.98×faster
9248 ms to 3095 ms
Concurrency 1
Parity by construction
Peak VRAM not published
$60
one-time
→
Kimi K3
Technique not publishedvLLM 0.23.1B300
2.12×faster
19090 ms to 8978 ms
Concurrency 1
Parity by construction
Peak VRAM not published
$100
one-time
→
AREX-Turbo
Channelwise FP8 weights, dynamic per-token activationsvLLM 0.25.1H100
1.2×faster
666 ms to 551 ms
Concurrency 8
gsm8k no measurable accuracy change, passed
Peak VRAM not published
$15
one-time
→
Qwythos-9B-Claude-Mythos-5-1M
Channelwise FP8 weights, dynamic per-token activationsvLLM 0.25.1H100
1.29×faster
1058 ms to 815 ms
Concurrency 8
gsm8k 99.35% recovery, passed
Peak VRAM not published
$20
one-time
→
Qwen3.6 27B
Channelwise FP8, measured selective-layer recipevLLM 0.25.1H100
1.28×faster
2857 ms to 2215 ms
Concurrency 8
gsm8k 99.87% recovery, passed
Peak VRAM not published
$40
one-time
→
Do not see your model?
Or describe your use case and the agent optimizes it for you→
Type
Serving engine
GPU

Common questions

Can't find what you're looking for? Get in touch

What exactly do I get when I buy a package?

There are two package kinds. An optimized-weights package is a self-contained kit with the optimized model weights, exact serving configuration, verifier, and measured benchmark receipt. It calls nothing home. A recipe package contains the serving configuration, verifier, and receipt, but no weights. When its source pin is published, you download the weights separately from Hugging Face at the pinned commit named on the model page. If the source pin is unavailable, the model page says so instead of showing a partial fetch instruction. Every model page states which kind you are buying. Both kits are yours to keep and do not expire.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy