RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading the optimized model catalog

Optimized models, measured on real GPUs

Open models with serving tuned on real hardware. Buy once, run anywhere.

Loading the model package
← All optimized models

Qwythos-9B-Claude-Mythos-5-1M

empero-ai/Qwythos-9B-Claude-Mythos-5-1M

Quantized and serving-tuned for H100 on vLLM 0.25.1, with the benchmark receipt included.

You get a self-contained kit: the optimized weights, the exact serving configuration, and the measured proof.

$20
one-time, yours forever
+ $20/mo continuous optimization
Get this model

Core plan required

$20/mo

We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.

On. Confirmed at checkout.

Speed result
1.29x faster
Median latency measured 1058 ms at baseline and 815 ms optimized. Median latency (p50), concurrency 8.The full benchmark methodology and receipt ship inside the kit.
Throughput
1255tokens/s, concurrency 8
Latency p50
815ms
Latency p99
861ms
Accuracy
99.35%
gsm8k exact_match strict (N=1319, completion protocol) recoveryPassed accuracy gate
GPU target
H100
Base model license
Apache-2.0
Proof verified
Jul 25, 2026
Published
Jul 25, 2026

Request profile

chat-128

Latency p50lower is better↓ 23%
Baseline1058 ms
Optimized815 ms
010002000
Latency p99lower is better↓ 21%
Baseline1083 ms
Optimized861 ms
010002000
Throughputhigher is better↑ 29%
Baseline974 tokens/s
Optimized1255 tokens/s
010002000

Runs on any GPU provider

Docker Compose
Kubernetes
RunPod
Modal
Your own hardware

This is what it means to own your AI

You control the intelligence

Run the model on your infrastructure, without deprecations, silent swaps, or provider restrictions.

You keep the economics

Buy once instead of paying per token; your deployment cost follows the GPUs you run.

You keep the data

Prompts and outputs stay on your network, with nothing retained elsewhere.

Technical implementation

In this panel

  1. 01What is in the kit
  2. 02Technical specifications
  3. 03Optimization techniques
  4. 04Measured accuracy

Delivery contract

What is in the kit

Download

./models/kit.zip

Archive root

qwythos-9b-fp8cd-gdnvis-h100-vllm/

Model directory

qwythos-9b-fp8cd-gdnvis-h100-vllm/weights/

Included

Kit contains

  • Serving configuration
  • Benchmark receipt
  • Verifier
  • Optimized weights
Kit size
13.5 GB
Kit SHA-25649fde441b83f...8d8ae3c61e54

Technical specifications

Will it run here?

GPU target
H100
Parameters
9B

What is it built from?

Engine
vLLM 0.25.1
Quantization
Channelwise FP8 weights, dynamic per-token activations
Modality
Text
Base model license
Apache-2.0
Package version
v1

How was it measured?

Measurement basis
Measured
Request profile
chat-128
Concurrency
8
Proof verified
Jul 25, 2026
Published
Jul 25, 2026

Optimization techniques

Channelwise FP8 weights, dynamic per-token activations

llm-compressor 0.12.0, data-free: no calibration corpus is read at any point

Gated-DeltaNet path and vision configuration held at BF16

the ignore list is re:.*visual.* and re:.*linear_attn.*, carried over from the measured per-layer sensitivity map for this qwen3_5 hybrid family; this is the third model built on that map, and it was validated here by the full-set gsm8k run rather than re-searched layer by layer

Pinned engine and measured serving configuration

the exact engine build and arguments the receipt numbers were produced on: vLLM 0.25.1, max-model-len 4096, max-num-seqs 64, gpu-memory-utilization 0.90, speculative decoding off

Measured accuracy

gsm8k exact_match strict

(N=1319, completion protocol)

Passed accuracy gate

Baseline
0.8378
Optimized
0.8324
Recovery
99.35%
Delta
-0.64%

Measured difference, significance not published

The optimized score measured below the baseline, at 99.35% of it. Neither score published a standard error, so this difference cannot be separated from sampling noise. The difference is not stated as a regression.

Methodology

STATISTICALLY INDISTINGUISHABLE FROM THE BF16 BASELINE, neither an improvement nor a demonstrated loss:
0.8324 +/- 0.0103 against a 0.8378 +/- 0.0102 BF16 baseline is a 0.0053 gap against a 0.0145 combined standard error on the difference, about 0.37 sigma, and the two confidence intervals overlap.
Both framings, because the numbers carry the claim and not the adjective: that is 99.4 percent recovery, equivalently a 0.6 percent relative difference in the baseline's favor, too small for this measurement to resolve. lm-eval 0.4.12, vLLM backend, gsm8k full set N=1319, exact_match strict-match, completion protocol with no chat template, seed 1234, identical settings, same H100 and same engine build on both sides.

License

Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.

Package license

RunInfra Package License v1.1 (2026-07-25)

Buy once. No meter, no expiry, nothing calling home. The full terms ship inside the kit.

What you are licensed to do

This package is licensed to the purchasing workspace, one time, for the exact version purchased. Everyone in that workspace may run it in production without limits: unlimited inference, on any hardware or cloud the workspace controls, for any lawful commercial purpose. You may modify the configuration and tooling for your own use.

Your outputs are yours

Everything the model produces for you belongs to you. You may use, sell, and build products on the model's outputs without restriction or royalty. Serving the model to your own customers as part of your product is use, not redistribution, and is fully allowed.

What you may not do

You may not redistribute, resell, sublicense, rent, publish, or otherwise make the package or its artifacts available to any third party. That covers the optimized weights, serving configuration, scripts, kit archive, and benchmark receipts, whole or in part, modified or not. The license belongs to the purchasing workspace and cannot be transferred separately from it.

The base model keeps its own license

The underlying model remains under its upstream open-source license, which is included in this kit with attribution and a statement of RunInfra's modifications. Nothing in this license restricts rights the upstream license grants you for the ORIGINAL model; the restrictions above apply to RunInfra's optimized package.

Version-pinned, as measured

You purchased this exact version, proven against the exact engine version named in the receipt. It stays downloadable to your workspace and never expires. The benchmark receipt describes measurements taken at verification time on the named hardware; RunInfra does not promise future updates to this version, and later package versions are separate purchases.

Breach ends the license

If the workspace redistributes the package or its artifacts, this license terminates for that workspace. Sections about your outputs survive termination for outputs already produced.

Base model license

Apache-2.0, as published by the model author on Hugging Face. The full license text ships inside the kit.

Company

Who you are buying from

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy