What happens when 100 people call at once

Latency, throughput, answer quality and error rates for free and open models on Hugging Face Inference Providers — measured at one concurrent user, then twenty-five, then a hundred.

Get InferGauge  → pip install infergauge Measured with infergauge.com

Read the full write-up  → Method, results, and how much the numbers move between runs

The short version

Seven of thirteen models never failed a single request — not at one caller, not at a hundred. Their goodput barely moves either: Qwen3-Coder-30B holds 100% at all three levels, Llama-3.1-8B and Qwen3-8B hold 95%, and gpt-oss-120b holds 85% with a p95 first token of 0.41s at both ends of the range.

Every failure in this dataset appeared under load. At one concurrent user all thirteen models returned a clean sweep — 780 requests, no errors. Four models start failing at twenty-five, six at a hundred. The gap between fastest and slowest first token widens from 8x to 48x across the same range.

DeepSeek-R1 declines steadily and nothing but goodput registers it — 13.3%, 5.0%, 1.7%. Its p95 first token never exceeds 2.2s and it returns no errors at all until a hundred callers. What fails is the part no single metric reports: roughly 300ms between tokens at every level, output ballooning from 363 to 866 tokens against a 53-token prompt, and a format score around 0.5.

How to read these numbers. Everything here is measured provider-side, from a datacenter, with one provider pinned per model. It describes how the provider behaves — not how your application will. Your network path, your prompts and your real traffic are not in this picture. For numbers that describe your own system, run InferGauge against your own endpoint.