Throughput
Tokens per second expresses how quickly a serving system produces tokens over an interval.
Measurement methodology
Version 2, updated Sep 20, 2026. How we define, date and limit every published number.
We update this method when our measurement rules change. Measurement dates change only when the underlying evidence changes.
Every result names the metric it measures. All definitions.
Tokens per second expresses how quickly a serving system produces tokens over an interval.
Time to first token measures elapsed time from request arrival until the first generated token becomes available.
The fiftieth and ninety-ninth percentiles mark different positions in an ordered latency distribution.
Quantization quality recovery is the measured retention of task performance after a model moves to lower precision.
Full definition: Comparison ratios are computed at render time from both absolutes and are never stored.
Full definition: Cache hit rate is the highest qualifying account's cached-input share on paid settled traffic: agent sessions of five or more turns, at least 10 million billed input tokens per account, excluding the SDK canary. Provider-reported cached counts are preferred, billing-derived otherwise; rates are floored to one decimal, published only at 70% or above, and revalidated hourly over 24 hours (7 days only where labeled). A labeled flagship rate can be borrowed only when a model's own gates decline to measure; the published figure is not a typical rate or a guarantee. After a failed refresh, the last cached rate can remain visible.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs