RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading the optimized model catalog

Optimized models, measured on real GPUs

Open models with serving tuned on real hardware. Buy once, run anywhere.

Loading the model package
← All optimized models

DeepSeek V4 Flash

deepseek-ai/DeepSeek-V4-Flash-0731

Serving-tuned for B200 on vLLM 0.25.0.

$60
one-time, yours forever
+ $150/mo continuous optimization
Get this model

Core plan required

$150/mo

We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.

On. Confirmed at checkout.

Speed result

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsDocs
2.98x faster
Median latency measured 9248 ms at baseline and 3095 ms optimized. Median latency (p50), concurrency 1.The full benchmark methodology and receipt ship inside the kit.
Throughput
363tokens/s, concurrency 1
Latency p50
3095ms
Latency p99
4306ms
Accuracy
output parityAccuracy recovery not measuredOutput parity by construction
GPU target
B200
Base model license
MIT
Proof verified
Aug 1, 2026
Published
Aug 1, 2026
Request profile, the exact conditions measured

Single-stream decode with one variable between the two columns of this comparison.

Both columns were measured on 2026-08-01 in one window, one container, one engine install, on the same 4x B200 cards, run back to back with an idle VRAM drain between them, and the comparison tool verified the request profiles identical across the arms before rendering any number.

The baseline column is this same configuration with speculative decoding off, not a different engine and not a failed serve.

The optimized column enables the DSpark speculative decoding that ships fused inside the DeepSeek-V4-Flash-0731 checkpoint, which the released engine does not enable on its own.

The throughput figures are per-request decode rates at concurrency 1, the speed one user experiences, not per-GPU figures and not node aggregates: 113.37 output tokens per second p50 with speculation off and 362.86 with it on.

The latency stats beside each column are end-to-end request times from the same runs: the optimized median request finishes in 3.1 seconds instead of 9.2, and its time to first token is 232.93 ms against 220.21, a cost of about 13 ms.

The profile, stated exactly: single-stream sequential, concurrency 1, 8 timed requests per arm, 1024 prompt tokens target, 1024 output tokens requested per request, temperature 0.0, ignore_eos on, streaming, a fixed seed, and 8 of 8 requests ok on each arm.

Know the boundary of this result before you deploy on it: the advantage is a low-concurrency fact, and the same window measured how it narrows with load on 32768-token-prompt workloads, plus 168 percent per GPU at concurrency 1, plus 33.6 percent at concurrency 8, and minus 6.9 percent at concurrency 32 where the median end-to-end time stays flat.

The optimized configuration also ran in the measurement slot that this campaign's A/A control series showed to be systematically disadvantaged, so the comparison understates the effect rather than overstating it.

Latency p50lower is better↓ 67%
Baseline9248 ms
Optimized3095 ms
0500010000
Latency p99lower is better↓ 55%
Baseline9548 ms
Optimized4306 ms
0500010000
Throughputhigher is better↑ 220%
Baseline113 tokens/s
Optimized363 tokens/s
0250500

Measured serving curve

Workload profile 32768 prompt tokens and 2048 output tokens per request, repo-scale coding shape, closed loop, 8, 16 and 64 timed requests at concurrency 1, 8 and 32, temperature 0.0, fixed seed, measured on the optimized configuration in the same one-container window as the headline pair.

Concurrency 174.3 tokens/s per GPU
Aggregate 297 tokens/sTTFT p50 175 msTPOT p50 14 mse2e p50 5440 ms
Concurrency 8236 tokens/s per GPU
Aggregate 944 tokens/sTTFT p50 2388 msTPOT p50 30 mse2e p50 15008 ms
Concurrency 32370 tokens/s per GPU
Aggregate 1482 tokens/sTTFT p50 4647 msTPOT p50 84 mse2e p50 41958 ms
0250500

Deployment guidance: request concurrency is a knob you set on your own deployment, and each point states the measured trade at that setting, not a package speedup.

Deployment guidance, not a speedup claim: the concurrency knob belongs to the buyer. The same window measured the speculation advantage narrowing with load, and at concurrency 32 the optimized configuration's aggregate throughput sat 6.9 percent BELOW its own no-speculation baseline while median end-to-end time stayed flat, with the optimized arm in the measurement slot the A/A control series showed to be disadvantaged. Batch-first deployments above concurrency 8 should benchmark their own operating point; interactive and low-concurrency serving is where this recipe earns its multiplier.

Runs on any GPU provider

Docker Compose
Kubernetes
RunPod
Your own hardware

This is what it means to own your AI

You control the intelligence

Run the model on your infrastructure, without deprecations, silent swaps, or provider restrictions.

You keep the economics

Buy once instead of paying per token; your deployment cost follows the GPUs you run.

You keep the data

Inference runs on your infrastructure, keeping prompts and outputs on your network.

Technical implementation

In this panel

  1. 01What is in the kit
  2. 02Technical specifications
  3. 03Optimization techniques
  4. 04Measured accuracy

Delivery contract

What is in the kit

Download

./models/kit.zip

Archive root

deepseek-v4-flash-b200x4-v4turbo/

Model directory

deepseek-v4-flash-b200x4-v4turbo/weights/

Filled by --weights

Kit contains

  • Serving configuration
  • Benchmark receipt
  • Verifier

Fetched separately

  • Pinned source weightsdeepseek-ai/DeepSeek-V4-Flash-0731 @ 7872f01b1d1f
Kit size
2.8 MB
Weights download
166.9 GB
Total download
166.9 GB
Kit SHA-256f9304227d189...6d67a51f3bfc

Technical specifications

Will it run here?

GPU target
B200
Compute capability
Sm 100
Parameters
284B

What is it built from?

Engine
vLLM 0.25.0
Quantization
None
Modality
Text
Base model license
MIT
Package version
v1

How was it measured?

Measurement basis
Measured

Optimization techniques

Fused DSpark speculative decoding, enabled and tuned

The checkpoint ships its own trained draft module fused into the weights; the released engine leaves it off. The draft module is DeepSeek's work; the enablement, configuration and measured operating envelope are this package's.

Validated 4x B200 serving configuration

Data-parallel 4 with expert parallelism, FP8 KV cache, 256-token attention blocks, the model-specific tokenizer and reasoning parser flags the fused checkpoint requires, an engine readiness timeout sized to this model's real cold-start, and a 131072-token context sized to the measured workloads. Every flag carries its reason in the shipped configuration notes.

Measured accuracy

output parity

Output parity by construction

Baseline
Not reported
Optimized
Not reported
Recovery
Not reported
Delta
Not reported

Accuracy recovery not measured

This package does not publish both a baseline and an optimized score on this metric, so there is no recovery figure to state.

Methodology

Configuration-only recipe on the released engine's own kernels with weights served untouched at a pinned revision.

License

Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.

Package license

RunInfra Package License v1.1 (2026-07-25)

Buy once. No meter and no expiry. No RunInfra service is required at runtime. The full terms ship inside the kit.

What you are licensed to do

This package is licensed to the purchasing workspace, one time, for the exact version purchased. Everyone in that workspace may run it in production without limits: unlimited inference, on any hardware or cloud the workspace controls, for any lawful commercial purpose. You may modify the configuration and tooling for your own use.

Your outputs are yours

Everything the model produces for you belongs to you. You may use, sell, and build products on the model's outputs without restriction or royalty. Serving the model to your own customers as part of your product is use, not redistribution, and is fully allowed.

What you may not do

You may not redistribute, resell, sublicense, rent, publish, or otherwise make the package or its artifacts available to any third party. That covers the serving configuration, scripts, kit archive, and benchmark receipts, whole or in part, modified or not. The upstream model weights are not included in this kit. You fetch them separately from Hugging Face at the pinned revision, and they remain governed by the model's own license, not this package license. The license belongs to the purchasing workspace and cannot be transferred separately from it.

The base model keeps its own license

The underlying model remains under its upstream open-source license, which is included in this kit with attribution and a statement of RunInfra's modifications. Nothing in this license restricts rights the upstream license grants you for the ORIGINAL model; the restrictions above apply to RunInfra's optimized package.

Version-pinned, as measured

You purchased this exact version, proven against the exact engine version named in the receipt. It stays downloadable to your workspace and never expires. The benchmark receipt describes measurements taken at verification time on the named hardware; RunInfra does not promise future updates to this version, and later package versions are separate purchases.

Breach ends the license

If the workspace redistributes the package or its artifacts, this license terminates for that workspace. Sections about your outputs survive termination for outputs already produced.

Base model license

MIT, as published by the model author on Hugging Face. The full license text ships inside the kit.

Company

Who you are buying from

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.
Research
News
Contact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
Request profile
Stated in full on the Overview tab
Concurrency
1
Proof verified
Aug 1, 2026
Published
Aug 1, 2026
Speculative decoding verifies every draft token against the target model before emission, so served outputs are the engine's own outputs by construction.
No task-level score is claimed in this version.