Latency, throughput, answer quality and error rates for free and open models on Hugging Face Inference Providers — measured at one concurrent user, then twenty-five, then a hundred.
Get InferGauge →
pip install infergauge
Measured with infergauge.com
Read the full write-up → Method, results, and how much the numbers move between runs
Seven of thirteen models never failed a single request — not at one caller, not at a hundred. Their goodput barely moves either: Qwen3-Coder-30B holds 100% at all three levels, Llama-3.1-8B and Qwen3-8B hold 95%, and gpt-oss-120b holds 85% with a p95 first token of 0.41s at both ends of the range.
Every failure in this dataset appeared under load. At one concurrent user all thirteen models returned a clean sweep — 780 requests, no errors. Four models start failing at twenty-five, six at a hundred. The gap between fastest and slowest first token widens from 8x to 48x across the same range.
DeepSeek-R1 declines steadily and nothing but goodput registers it — 13.3%, 5.0%, 1.7%. Its p95 first token never exceeds 2.2s and it returns no errors at all until a hundred callers. What fails is the part no single metric reports: roughly 300ms between tokens at every level, output ballooning from 363 to 866 tokens against a 53-token prompt, and a format score around 0.5.
The table above is the latest run. Every weekly run is also archived under its own date and never overwritten — the point of running weekly is the history, and a file rewritten every Monday has none in it. Open a date to see how that week came out.
Two runs of the same procedure, 13 models × 20 prompts × 3 repeats = 780 requests each. One at a single request in flight, one at 25 concurrent requests per model. The raw per-request results are published here so the table can be re-derived rather than taken on trust.
It is InferGauge's own Free-tier concurrency cap, so anyone on the free plan can reproduce this exactly. Requests run in parallel against one model at a time, so contention lands on that endpoint rather than being spread across unrelated providers.
The same thirteen models were benchmarked at one concurrent user twice on 25 September, two hours apart, with an identical harness against the same endpoints. The two runs disagree by more than rounding:
Shared inference endpoints are noisy neighbours by construction: you are measuring a service whose load you do not control. Treat any single figure here as one sample, not a constant. The findings that survive both runs — which providers hold steady, that errors appear only under load, that DeepSeek-R1 fails the SLO at every level — are the ones worth acting on. A 3-point difference in one model's goodput is not.
This is also why the raw per-request CSV is published. Percentiles over sixty requests are thin, and the underlying distribution is more informative than the summary.
Goodput is the share of requests that met every threshold below on the same request. It is deliberately harsher than any single column: a model can post a good median TTFT and still fail most requests once errors, stalls and malformed output are counted together.
These thresholds describe a responsive chat-style application. They are an editorial choice,
not a law, which is why they are published rather than buried — they live in
slo.json alongside the harness, and changing them changes the goodput column.
TTFT and ITL answer different questions. A TTFT that climbs under load while ITL stays flat means requests are queueing. A flat TTFT with ITL rising means the serving engine is compute-bound. Reporting only one of them hides which problem you have.
Deterministic format checks declared per prompt: does the JSON parse with the required keys and types, does the code compile and define the requested function, does the summary respect the stated sentence limit, does the answer carry the correct value. Scored 0–1 per prompt with graded partial credit — valid JSON wrapped in prose scores 0.8, code that parses but defines the wrong function scores 0.4.
Quality is computed over successful requests only. Where a model's error rate is high, its quality score rests on very few responses and should be read as provisional — TinyLlama's 100-user score rests on the 13 requests of 60 that survived, and DeepSeek-R1's 25-user score on 4.
No LLM-as-judge rating was collected in these runs. Where one is, it goes in its own column and is never blended into the deterministic score.
Tokens per second is measured over the generation window, after the first token. Where a provider does not stream its reasoning channel as deltas, that window can collapse to a sliver and the rate inflates absurdly — one raw measurement read 3778 tok/s. When the window is under a fifth of the request, the whole request is used instead. This makes throughput conservative on very short responses, which is the right direction to be wrong in.
Total and active parameters are listed separately. gpt-oss-120b is 117B total but
5.1B active; Qwen3-Coder-30B-A3B is 30B total and 3.3B active. Plotting MoE models
against dense models by total size produces a chart that is simply wrong.
Hybrid reasoning models stream chain-of-thought separately from the answer. That still stops the time-to-first-token clock — it is a token arriving — but it is counted apart and never scored. Qwen3 models run with reasoning disabled and the gpt-oss pair at low reasoning effort, both stated per model in the shortlist. DeepSeek-R1 runs unmodified, which is why it is slow and why its quality score is low: it answers correctly and ignores the stated format constraints.
Your own endpoint, your network path, your prompt distribution, your rate limits, or quality degradation on your own infrastructure. Those decide whether your application holds up, and they can only be measured from where your application runs.
Know exactly how your AI performs under load.
InferGauge is a local-first load testing and benchmarking tool for LLM applications. Every test
is one YAML file, the dashboard is live at localhost:8710 while the test runs, and
prompts, keys and results never leave your machine.
pip install infergauge
infergauge init -y
infergauge run
Two numbers tell you what is actually wrong, and most load testers report neither. TTFT is the wait before the first token — what a user feels, and it includes queueing. ITL is the gap between streamed tokens during generation.
High TTFT + healthy ITL → queueing. The model generates fine once it starts; requests are waiting in line. Fix with capacity or parallelism, not a smaller model.
Healthy TTFT + rising ITL → compute-bound. Generation itself is saturating. Fix with hardware, quantisation, or a smaller model.
Alongside those: goodput, the share of requests meeting every SLO at once, and the saturation knee — the concurrency where latency inflects. The table on this page is a two-point version of that curve. Run it against your own endpoint and you get the whole curve, live.
infergauge run --baseline latest --max-regression-pct 15
Exit code 1 if any metric worsens by more than 15% against the previous run. Performance regressions stop being something you discover in production.
InferGauge is source-available under the Business Source License 1.1: read and audit the source, use it internally and in production, gate your CI with it. What is not permitted is reselling it as a hosted service. It converts to Apache-2.0 in 2030. This Space's own code is MIT.