RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading the optimized model catalog

Optimized models, measured on real GPUs

Open models with serving tuned on real hardware. Buy once, run anywhere.

Loading the model package
← All optimized models

Kimi K3

moonshotai/Kimi-K3

Serving-tuned for B300 on vLLM 0.23.1.

$100
one-time, yours forever
+ $200/mo continuous optimization
Get this model

Core plan required

$200/mo

We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.

On. Confirmed at checkout.

Speed result
2.12x faster
Median latency measured 19090 ms at baseline and 8978 ms optimized. Median latency (p50), concurrency 1.Measured with unshipped runtime. The full benchmark methodology and receipt ship inside the kit.
Throughput
120tokens/s, concurrency 1
Latency p50
8978ms
Latency p99
9595ms
Accuracy
Served-output parity with the pinned upstream weights and engine version (by construction; no measured evaluation)Accuracy recovery not measuredOutput parity by construction
GPU target
B300
Base model license
Kimi K3 License
Proof verified
Jul 31, 2026
Published
Jul 31, 2026
Request profile, the exact conditions measured

Single-stream decode with one variable between the two columns of this comparison.

Both columns were measured on 2026-07-31 in one window, one container, one engine install, on the same 8x B300 cards, run back to back with an idle drain between them, and the comparison verified the request profiles identical across the arms, including per-point prompt fingerprints, before rendering any number.

The baseline column is the batch profile and the optimized column is the interactive profile.

The throughput figures are per-request decode rates at concurrency 1, the speed one user experiences, not per-GPU figures and not node aggregates: 55.15 output tokens per second p50 for the batch profile and 119.84 for the interactive profile, which is 2.17x.

Dividing either figure by eight describes nothing real because all eight cards cooperate on one stream, and multiplying by eight would invent a many-stream aggregate this profile did not measure.

The profile, stated exactly: single-stream sequential, concurrency 1, 8 timed requests per arm, 1024 prompt tokens target, 1024 output tokens requested per request, temperature 0.0, ignore_eos on, streaming, a fixed seed, and 8 of 8 requests ok on each arm.

The latency stats beside each column are end-to-end request times from the same runs: the interactive profile median request finishes in 9.0 seconds instead of 19.1, its time to first token is better in this window (429.30 ms against 535.06), and its time per output token is higher (23.10 ms against 18.13).

Know the boundary of this result before you deploy on it: the 2.17x is a concurrency-1 fact, and the same window measured how the advantage narrows with load on 32768-token-prompt workloads, +89.59 percent per GPU at concurrency 1, +17.83 percent at concurrency 8, and flat at concurrency 32 (minus 0.83 percent, inside the measurement noise floor) where time to first token still improves.

Version 3 published a different tuning of the interactive profile whose advantage INVERTED at concurrency 32, and told you to switch to the batch profile above roughly concurrency 8; a one-variable window then measured the two tunings against each other and the version 4 tuning won at every point, so that rule is retired.

This is a 1024-token profile, not the 32K repo-scale profile version 2 published, and version 2 per-GPU throughput figure must not be compared with these per-request rates.

Latency p50lower is better↓ 53%
Baseline19090 ms
Optimized8978 ms
01000020000
Latency p99lower is better↓ 50%
Baseline19377 ms
Optimized9595 ms
01000020000
Throughputhigher is better↑ 117%
Baseline55.2 tokens/s
Optimized120 tokens/s
0100200

Measured serving curve

Repo-scale coding profile: 32768 prompt tokens target and 2048 output tokens requested per request, temperature 0.0, ignore_eos on, streaming, a fixed seed, measured at concurrency 1 (4 requests), 8 (16 requests) and 32 (64 requests) with one warmup request per point, every request ok at every point. This curve is the interactive profile. Request rate at concurrency 32 measured 0.133 requests per second on this profile. Concurrency 128 is not characterized on this configuration.

Concurrency 111.5 tokens/s per GPU
Aggregate 92.4 tokens/sTTFT p50 4836 msTPOT p50 24 mse2e p50 21363 ms
Concurrency 830.3 tokens/s per GPU
Aggregate 242 tokens/sTTFT p50 7008 msTPOT p50 78 mse2e p50 64709 ms
Concurrency 3242.7 tokens/s per GPU
Aggregate 341 tokens/sTTFT p50 29435 msTPOT p50 206 mse2e p50 170409 ms
02550

Deployment guidance: request concurrency is a knob you set on your own deployment, and each point states the measured trade at that setting, not a package speedup.

Measured 2026-07-31 on the same 8x B300 cards and in the same one-container window as the headline comparison. Read the curve with its shape in view: the interactive profile advantage narrows as load rises, +89.59 percent per GPU at concurrency 1, +17.83 percent at concurrency 8, and flat at concurrency 32 (minus 0.83 percent, inside the measurement noise floor), while time to first token is better with the interactive profile at every measured point, 35.5 s against 29.4 s at concurrency 32. Version 3 published a different tuning whose advantage inverted at concurrency 32 with a 114.5 s time to first token, and told you to switch to the batch profile above roughly concurrency 8. Version 4 retires that rule: there is no measured operating point on this curve where the batch profile wins throughput outside noise, and latency prefers the interactive profile everywhere.

Runs on any GPU provider

Docker Compose
Kubernetes
RunPod
Your own hardware

This is what it means to own your AI

You control the intelligence

Run the model on your infrastructure, without deprecations, silent swaps, or provider restrictions.

You keep the economics

Buy once instead of paying per token; your deployment cost follows the GPUs you run.

You keep the data

Inference runs on your infrastructure, keeping prompts and outputs on your network.

Technical implementation

In this panel

  1. 01What is in the kit
  2. 02Technical specifications
  3. 03Optimization techniques
  4. 04Measured accuracy

Delivery contract

What is in the kit

Download

./models/kit.zip

Archive root

kimi-k3-b300x8-k3turbo/

Model directory

kimi-k3-b300x8-k3turbo/weights/

Filled by --weights

Kit contains

  • Serving configuration
  • Benchmark receipt
  • Verifier

Fetched separately

  • Pinned source weightsmoonshotai/Kimi-K3 @ 9f62e4e9fffb
Kit size
2.9 MB
Weights download
1.56 TB
Total download
1.56 TB
Kit SHA-256a7352bdfd181...084c738647fd

Technical specifications

Will it run here?

GPU target
B300
Compute capability
Sm 103
Parameters
2800B

What is it built from?

Engine
vLLM 0.23.1
Quantization
None
Modality
Text
Base model license
Kimi K3 License
Package version
v4

How was it measured?

Measurement basis
Measured with unshipped runtime
Request profile
Stated in full on the Overview tab
Concurrency
1
Proof verified
Jul 31, 2026
Published
Jul 31, 2026

Optimization techniques

Verified prefill execution path

A required prefill execution path was active and verified in both arms of the measuring window. This is upstream work, not ours. What this package contributes is carrying the validated path, verifying that it was active, and measuring the complete serving profiles. No speedup figure is attached to this entry, because a dedicated one-variable measurement with and without this path has not been run.

Interactive profile with measured low-concurrency gain

The interactive profile was measured against the batch profile with one variable changed, in one container on the same eight cards, with request profiles verified identical.

Validated expert execution path with a portability limit

The routed expert layers used the only execution path available in the dependencies used for validation. This was a compatibility constraint, not a tuning claim. A potentially faster path was unavailable, so this package does not claim the fastest possible implementation and attaches no speedup figure to this entry.

Request-reuse support in the measured configuration

A request-reuse optimization was active in the configuration these numbers were measured on. No contribution is attributed to it here: the measured profile issues 8 distinct single-stream requests, which is not a workload designed to exercise repeated input, and no one-variable pair has been measured. No speedup figure is attached to this entry.

Validated engine package and measured profiles

The part of this package that is genuinely ours is the validated vLLM 0.23.1 package, the delivered interactive and batch profiles, the verification record, and the measured proof. One honest gap remains: the deploy files in this kit carry an earlier validated configuration rather than the configuration used for the 2026-07-30 measurements, and transcribing the measured serve command into them is a named prerequisite of this version. No stock configuration exists to compare against, so this package is sold as the validated way to serve the model and is never counted into a speedup multiplier.

Measured accuracy

Served-output parity with the pinned upstream weights and engine version

(by construction; no measured evaluation)

Output parity by construction

Baseline
Not reported
Optimized
Not reported
Recovery
Not reported
Delta
Not reported

Accuracy recovery not measured

This package does not publish both a baseline and an optimized score on this metric, so there is no recovery figure to state.

Methodology

By construction, not a measured evaluation, and stated with its limits.
This package changes no weight byte: the tree you fetch is verified at deploy time against the measured digest ebf0b1a75a29ab5e01f319252302c418d0e8f9340e402950497094a6fc64aced.
It also changes no engine code.
The served model is therefore the pinned upstream MXFP4 weights served by vLLM 0.23.1, and its outputs are that engine version's outputs under the delivered profiles.
What is NOT claimed: that one serving profile produces the same tokens as another profile or a different engine build.
Token-level identity across configurations is unmeasured rather than asserted.
No cross-server token-agreement score is published either, because a control between two identical server instances measured 78.12 percent top-1 agreement, so cross-server comparison cannot certify accuracy and is informational only.
Accuracy recovery NOT measured.

License

Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.

Package license

RunInfra Package License v1.1 (2026-07-25)

Buy once. No meter and no expiry. No RunInfra service is required at runtime. The full terms ship inside the kit.

What you are licensed to do

This package is licensed to the purchasing workspace, one time, for the exact version purchased. Everyone in that workspace may run it in production without limits: unlimited inference, on any hardware or cloud the workspace controls, for any lawful commercial purpose. You may modify the configuration and tooling for your own use.

Your outputs are yours

Everything the model produces for you belongs to you. You may use, sell, and build products on the model's outputs without restriction or royalty. Serving the model to your own customers as part of your product is use, not redistribution, and is fully allowed.

What you may not do

You may not redistribute, resell, sublicense, rent, publish, or otherwise make the package or its artifacts available to any third party. That covers the serving configuration, scripts, kit archive, and benchmark receipts, whole or in part, modified or not. The upstream model weights are not included in this kit. You fetch them separately from Hugging Face at the pinned revision, and they remain governed by the model's own license, not this package license. The license belongs to the purchasing workspace and cannot be transferred separately from it.

The base model keeps its own license

The underlying model remains under its upstream open-source license, which is included in this kit with attribution and a statement of RunInfra's modifications. Nothing in this license restricts rights the upstream license grants you for the ORIGINAL model; the restrictions above apply to RunInfra's optimized package.

Version-pinned, as measured

You purchased this exact version, proven against the exact engine version named in the receipt. It stays downloadable to your workspace and never expires. The benchmark receipt describes measurements taken at verification time on the named hardware; RunInfra does not promise future updates to this version, and later package versions are separate purchases.

Breach ends the license

If the workspace redistributes the package or its artifacts, this license terminates for that workspace. Sections about your outputs survive termination for outputs already produced.

Base model license

Kimi K3 License, as published by the model author on Hugging Face. The full license text ships inside the kit.

Company

Who you are buying from

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy