Most published inference numbers are single-stream. One request at a time, best case, and often without stating the concurrency at all. That tells you what a model can do. It doesn't tell you what happens when your app has users.
So I ran the same prompt set against thirteen open models on Hugging Face Inference Providers at three levels of load: 1, 25 and 100 concurrent requests per model. Same harness, same prompts, three runs, sequential. 2,340 requests.
The leaderboard is on the first tab of this Space. Every row reads across the three levels.
Method
Putting this first, because it decides the numbers.
One provider pinned per model. The same model id can be served by up to ten providers at different quantisation, hardware and price. An unpinned benchmark measures the router, not the model. Every row names the provider it used.
No client-side retries. The OpenAI SDK retries by default, which quietly deletes your 429s. A 429 or a 5xx is recorded as an error.
TTFT measured on a streamed response, not derived from total time. Throughput is measured over the generation window after the first token, because a whole-request rate flatters models with slow queueing.
Inter-token latency measured separately. The gap between streamed content chunks, reported as a p95. High TTFT with healthy ITL means requests are queueing. Healthy TTFT with rising ITL means generation is compute-bound. Reporting only one of them hides which problem you have.
Goodput is the headline metric. A request counts only if it meets every
threshold at once: first token under 2s, finished under 15s, ITL p95 under 80ms, format score at or
above 0.8, no error. Those thresholds describe a responsive chat application. They're an editorial
choice, not a law, which is why they're published in slo.json rather than buried.
Quality is deterministic format checking, not a judge model. Does the JSON parse with the required keys and types. Does the code compile and define the requested function. Does the summary respect the stated sentence limit. Graded partial credit: valid JSON wrapped in prose scores 0.8, code that parses but defines the wrong function scores 0.4. An optional LLM-as-judge rating is recorded in its own column and never blended in.
Outliers kept. Cold starts and queueing spikes are part of what a caller experiences.
Reasoning models handled explicitly. Hybrid models stream chain-of-thought in a separate field. That still stops the TTFT clock, but it isn't the answer, so it's counted apart and never scored. Without this, Qwen3 scores zero on everything, which is a harness bug and not a finding.
MoE models list active parameters separately. gpt-oss-120b is 117B total and 5.1B active. Charting it against dense models by total size produces a wrong chart.
Results
Seven of thirteen models never failed a request at any level
| Model | Provider | Goodput 1 / 25 / 100 | p95 TTFT 1 / 25 / 100 |
|---|---|---|---|
| Qwen3-Coder-30B-A3B | scaleway | 100 / 100 / 100 | 0.55 / 0.55 / 0.49 |
| Llama-3.1-8B | nscale | 95 / 95 / 95 | 1.01 / 0.77 / 0.77 |
| Qwen3-8B | nscale | 95 / 95 / 95 | 1.02 / 0.80 / 1.01 |
| gpt-oss-20b | groq | 93.3 / 90 / 91.7 | 0.56 / 0.41 / 0.47 |
| Qwen2.5-7B | featherless-ai | 91.7 / 85 / 91.7 | 0.96 / 0.77 / 1.01 |
| gpt-oss-120b | groq | 85 / 85 / 85 | 0.41 / 0.41 / 0.41 |
| Qwen2.5-Coder-7B | nscale | 85 / 85 / 85 | 0.89 / 1.03 / 0.78 |
Every error in the dataset appeared under load
At one concurrent user, all thirteen models returned a clean sweep: 780 requests, no failures. Four models start failing at 25. Six at 100.
If you only ever benchmark single-stream, your data contains no failure signal at all.
The model that looks healthy and isn't
DeepSeek-R1 on Novita scores 13.3%, 5.0% and 1.7% goodput across the three levels.
Over that same range its p95 first token never exceeds 2.2s, and it returns zero errors until 100 concurrent users. Latency and error rate both call it fine.
What fails is everything between those two metrics. Inter-token latency sits around 300ms p95 against an 80ms threshold, so the text visibly stutters. Output balloons from 363 to 866 tokens against a 53-token prompt, pushing most requests past the 15s total-time ceiling. Format score sits near 0.5.
This is the clearest argument I have for publishing goodput rather than a latency number. The failure is real, it's monotonic across load, and no conventional single metric registers it.
Cost runs backwards from quality
| Model | Provider | $/1M output | Goodput at 100 |
|---|---|---|---|
| Qwen2.5-Coder-7B | nscale | 0.03 | 85 |
| Llama-3.1-8B | nscale | 0.06 | 95 |
| Qwen3-8B | nscale | 0.18 | 95 |
| gpt-oss-20b | groq | 0.50 | 91.7 |
| gpt-oss-120b | groq | 0.75 | 85 |
| Qwen3-Coder-30B-A3B | scaleway | 0.912 | 100 |
| DeepSeek-R1 | novita | 2.50 | 1.7 |
The most expensive model measured has the worst goodput at every level. The cheapest holds 85% throughout, at a hundredth of the price.
Rates come from the Hugging Face router's own catalogue, which is what these runs were billed. They can differ from a provider's public pricing page: Groq lists gpt-oss-20b at $0.30 per million output tokens, while the router billed $0.50. Six models are served by a provider that bills per subscription rather than per token, so they have no per-1M rate and their cost cell is blank rather than estimated.
Throughput tracks the provider, not the model size
Groq runs gpt-oss-20b at 570–650 tok/s. Nscale sits around 190 for the 7–8B models. Featherless is around 40 for everything it serves. A 117B model on Groq is roughly twelve times faster than a 0.6B model on Featherless.
How much these numbers move
I ran the 1-user level twice, two hours apart, with an identical harness against the same endpoints. The two runs disagree by more than rounding.
| First run | Second run | |
|---|---|---|
| Models clearing 90% goodput | 4 | 6 |
| DeepSeek-R1 goodput | 1.7% | 13.3% |
| TTFT spread at 1 / 25 / 100 | 10x / 38x / 27x | 8x / 40x / 48x |
The spread moved in opposite directions at the top end between the two sets.
Shared inference endpoints are noisy neighbours by construction. You're measuring a service whose load you don't control, and each per-model figure here is 60 requests. Treat any single number as one sample.
What survived both sets, and is therefore worth acting on: which providers hold steady under load, that errors appear only under concurrency, and that DeepSeek-R1 fails the SLO at every level. A three-point difference in one model's goodput did not survive, and shouldn't be quoted.
I'd rather publish this than have someone else discover it by re-running me.
What this does not measure
Everything here is provider-side, from a datacenter, with a fixed prompt set. It describes how the provider behaves, not how your application will.
Your network path isn't in it. Your prompt distribution isn't in it. Your real traffic shape isn't in it, and neither are your rate limits. Percentiles over sixty requests are thin.
If you want numbers that describe your own system, run the harness against your own endpoint.
Reproduce it
Everything needed is in this Space:
- leaderboard-c1.csv, -c25.csv, -c100.csv — one row per model per level
- results/results.csv — one row per request
- prompts.json — the 20 prompts and their checks
- quality.py — the scoring code, so you can disagree with it
- slo.json — the thresholds that decide goodput
- model_shortlist.csv — the models and their pinned providers
- job.py and RUN.md — the runner, and how to run it on HF Jobs
Disagree with the SLO thresholds and the goodput column changes. That's the point of publishing them.
The tool is InferGauge, source-available under BSL 1.1: you can read it, run it internally and in production, and gate CI with it. You can't resell it as a hosted service. It converts to Apache-2.0 in 2030. infergauge.com