Notes from the team.
Announcements, product releases, and research notes from the RunInfra team.
RunInfra is now inference for agents
With software alone, one B200 beats the LPU and gets close to Cerebras
Lossless Inference
The fastest way to serve DeepSeek V4 Flash

$0.09 and $290.12: What Actually Moves Your Inference Bill
Serving Kimi K3 on vLLM was hard. Here is what we measured.
vLLM vs SGLang vs TensorRT-LLM: a serving benchmark
Start building on open models
One API key for the OpenAI and Anthropic SDKs.
TLS in transit, AES-256 at rest
Workspace isolation
No training on your data
SOC 2 Type II