RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Notes from the team.

Announcements, product releases, and research notes from the RunInfra team.

AllResearch noteAnnouncementEngineering
RunInfra
Research note
01LLM inference
02Quantization
03Speculative decoding
Article map03 signals / 0ZQ59AU
August 3, 2026

Lossless Inference

RunInfra
August 2, 2026

The fastest way to serve DeepSeek V4 Flash

Horizontal bar chart ranking five inference cost levers from accelerator choice at 2.3x to reasoning token volume at 14.3x.
July 31, 2026

$0.09 and $290.12: What Actually Moves Your Inference Bill

RunInfra
July 30, 2026

Serving Kimi K3 on vLLM was hard. Here is what we measured.

RunInfra
Engineering
01vLLM
02SGLang
03TensorRT-LLM
Article map03 signals / 03C5AMR
June 20, 2026

vLLM vs SGLang vs TensorRT-LLM: a reproducible benchmark

Deploy your first optimized model, measured before you ship

Describe the goal. RunInfra builds and optimizes the stack.

Start BuildingView Pricing
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy