Skip to content
Backed by Y Combinator

Open models, built for agents

Fast responses. Lower-cost cached context. One API key for the OpenAI and Anthropic SDKs.

Inference speed, cost, and cache hits

Nemotron 3.5 Lightning 30B, 1 of 4

Output speed against cached input price: Nemotron 3.5 Lightning 30B on RunInfra at 541 tokens per second and $0.01 per 1M cached input tokens. Model only. 2 other listed providers of the same model.Price increases to the right; speed increases upward. Axes rescale to the current model's measurements. The green region is below the other listings' median price and above their median speed. The muted region is the opposite. The dotted line connects non-dominated points from the other listings only, not measurements between providers. RunInfra uses a different measurement method. The ring is RunInfra's 92.2% cache-hit rate. Last 24 hours. Use arrow keys to move between points, Enter or Space to inspect, and Escape to dismiss.Higher speed (tok/s)$0$0.015$0.03$0.0450150300450

Lower cached input price (USD / 1M tokens)

Your session stays warm

Session-aware routing keeps repeat calls close to cached context.

Five calls of one agent session, every one routed to the replica that already holds the sessionOne agent session makes 5 calls to replica 2 of 4 replicas. Call 1 is cold. Calls 2 through 5 hit the warm cache on the same replica.your agent, one sessionreplicascall 1coldcall 2hitcall 3hitcall 4hitcall 5hitreplica 1replica 2warmreplica 3replica 4
Without routing, a repeat lands warm one time in four.
  • One session, one replica.

    Repeat calls prefer the replica holding their context.

  • No hint needed.

    Automatic from the conversation or workspace. Send a session ID to keep calls together.

  • Reported, billed as cached.

    Cache hits appear in usage and are billed at the cached rate. Hits stay best effort.

What agents actually pay

Per 1M input tokens, at each model's measured cache hit rate.

Price per 1M input tokens, log scale

$0.01$0.02$0.05$0.10$0.20
  1. Effective input rate $0.031 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

    GLM 5.3 Flash

    List price $0.11Effective input $0.031
    3.5x

    lower than list at 99.0% cache hits

  2. Effective input rate $0.031 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

    DeepSeek V4.1 Flash

    List price $0.14Effective input $0.031
    4.5x

    lower than list at 99.4% cache hits

  3. Effective input rate $0.014 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

    Nemotron 3.5 Lightning 30B

    List price $0.05Effective input $0.014
    3.5x

    lower than list at 92.2% cache hits

  4. Effective input rate $0.016 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

    Qwen3.8 27B

    List price $0.10Effective input $0.016
    6.2x

    lower than list at 94.2% cache hits

  5. Effective input rate $0.012 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

    Ornith 1.5 35B

    List price $0.10Effective input $0.012
    8.3x

    lower than list at 98.5% cache hits

Compared with list: effective prices round up, multipliers round down. Request stream usage for cached tokens; your usage page shows billed figures. Cache hits are best effort. Cache measurement method.

The model library

Prices, capabilities, cache hits and speed. Open a model for limits and availability.

5 models

FiltersAll models
Capabilities
Input $0.11 per 1M input tokens. Cached input $0.03 per 1M cached input tokens. Output $0.45 per 1M output tokens. 99.0% cache hit, last 24 hours. 254.1 output tok/s, model only, before the gateway. Capabilities: Tool calling, JSON mode, Streaming, Image input. Effective input rate $0.031 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

GLM 5.3 Flash

Trending

Model capabilities: Tool calling / JSON mode / Streaming / Image input

$0.11
per 1M input tokens
$0.03
per 1M cached input tokens
$0.45
per 1M output tokens
99.0%
cache hit, last 24 hours
254.1
output tok/s, model only, before the gateway

Effective input $0.031 per 1M at that hit rate Effective input rate $0.031 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Input $0.14 per 1M input tokens. Cached input $0.03 per 1M cached input tokens. Output $0.58 per 1M output tokens. 99.4% cache hit, last 24 hours. 378 output tok/s, model only, before the gateway. Capabilities: Tool calling, JSON mode, Streaming. Effective input rate $0.031 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

DeepSeek V4.1 Flash

Trending

Model capabilities: Tool calling / JSON mode / Streaming

$0.14
per 1M input tokens
$0.03
per 1M cached input tokens
$0.58
per 1M output tokens
99.4%
cache hit, last 24 hours
378
output tok/s, model only, before the gateway

Effective input $0.031 per 1M at that hit rate Effective input rate $0.031 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Input $0.05 per 1M input tokens. Cached input $0.01 per 1M cached input tokens. Output $0.15 per 1M output tokens. 92.2% cache hit, last 24 hours. 540.9 output tok/s, model only, before the gateway. Capabilities: Tool calling, JSON mode, Streaming. Effective input rate $0.014 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Nemotron 3.5 Lightning 30B

Model capabilities: Tool calling / JSON mode / Streaming

$0.05
per 1M input tokens
$0.01
per 1M cached input tokens
$0.15
per 1M output tokens
92.2%
cache hit, last 24 hours
540.9
output tok/s, model only, before the gateway

Effective input $0.014 per 1M at that hit rate Effective input rate $0.014 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Input $0.10 per 1M input tokens. Cached input $0.01 per 1M cached input tokens. Output $0.40 per 1M output tokens. 94.2% cache hit, last 24 hours. 140 output tok/s, end to end, on the buyer path. Capabilities: Tool calling, JSON mode, Streaming, Image input. Effective input rate $0.016 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Qwen3.8 27B

Model capabilities: Tool calling / JSON mode / Streaming / Image input

$0.10
per 1M input tokens
$0.01
per 1M cached input tokens
$0.40
per 1M output tokens
94.2%
cache hit, last 24 hours
140
output tok/s, end to end, on the buyer path

Effective input $0.016 per 1M at that hit rate Effective input rate $0.016 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Input $0.10 per 1M input tokens. Cached input $0.01 per 1M cached input tokens. Output $0.40 per 1M output tokens. 98.5% cache hit, last 24 hours. 216 output tok/s, model only, before the gateway. Capabilities: Tool calling, JSON mode, Streaming, Image input. Effective input rate $0.012 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Ornith 1.5 35B

Model capabilities: Tool calling / JSON mode / Streaming / Image input

$0.10
per 1M input tokens
$0.01
per 1M cached input tokens
$0.40
per 1M output tokens
98.5%
cache hit, last 24 hours
216
output tok/s, model only, before the gateway

Effective input $0.012 per 1M at that hit rate Effective input rate $0.012 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours. Cache hits are best effort; the method is published at runinfra.ai/methodology.

Coding plans from $10 a month

Up to 99.4% cache hit rate, last 24 hours, so your plan goes further.

Connect your coding agents

The CLI finds them. You pick the model.

  • Claude Code
  • Codex
  • OpenCode
  • Cline
  • Droid
  • Grok Build

Let your coding agent set it up

  1. 1Paste the prompt into your coding agent
  2. 2Approve the sign-in in your browser
  3. 3Say yes to the setup plan

Start building on open models

One API key for the OpenAI and Anthropic SDKs.

Looking for another model?