# Running the benchmark on Hugging Face

Everything runs on HF infrastructure. The job pulls this harness out of the Space, calls
Inference Providers, and pushes `leaderboard.csv` back to the Space root — so the public
leaderboard updates itself.

## The one thing you need

A Hugging Face token with **Write** scope and Inference Providers enabled —
<https://huggingface.co/settings/tokens>. Write is required because the run commits
results back into this Space.

Everything below works on a **free account**. The compute is free; the only thing consumed
is your Inference Providers allowance.

---

## Option A — Colab notebook (free, no terminal)

`run_benchmark_colab.ipynb` in this repo. This is the recommended route.

1. Open <https://colab.research.google.com> → **File → Upload notebook** → pick that file.
2. Run cell 1 (installs), cell 2 (paste the token into the hidden prompt), cell 3 (run).
3. Leave `SMOKE` ticked for the first run — 15 requests, about a minute.
4. When it looks right, untick `SMOKE` and run cell 3 again for the full 780.

Colab's CPU runtime is free and this needs no GPU. The token lives only in the runtime's
memory; **Runtime → Disconnect and delete runtime** when you're done.

### Headless: submit an HF Job from the notebook, then close the laptop

Once the account has a credit balance, the same job can run on HF's own hardware instead
of in the Colab runtime — so nothing depends on the browser staying open. Paste this into
a cell, run it, and walk away:

```python
from huggingface_hub import run_job
import getpass

TOKEN = getpass.getpass("HF token (Write scope): ").strip()
SCRIPT = ("https://huggingface.co/spaces/Infergauge-load/"
          "llm-performance-and-load/raw/main/job.py")

job = run_job(
    image="ghcr.io/astral-sh/uv:python3.12-bookworm",   # must contain uv
    command=["uv", "run", SCRIPT],
    env={"IG_SMOKE": "0", "IG_REPEATS": "3", "IG_CONCURRENCY": "25"},
    secrets={"HF_TOKEN": TOKEN},
    token=TOKEN,
)
print("running:", job.url)
```

### The container timeout is real, and single-stream nearly hit it

HF Jobs kills a container after a fixed wall-clock ceiling — around 30 minutes on cpu-basic. The
concurrency-1 run on 25 Sep took **35 minutes** for 780 requests and was killed. It survived only
because `job.py` had already finished pushing; the job page still reads *Status: Error — Job
timeout* even though the data is complete and correct.

Do not rely on that margin. Pass a longer timeout when submitting a single-stream run:

```python
job = run_job(..., timeout="2h")
```

The load levels are not affected — 25 and 100 both finish in about 8 minutes, because the requests
go out in parallel. Only concurrency 1 runs long enough to matter.

If a run is killed *before* the push, nothing reaches the Space and the whole run is lost: results
are written only at the end. That is a known gap in the harness.

**The image must contain `uv`.** `python:3.12` does not, and the job dies in seconds with
`exec: "uv": executable file not found in $PATH`. `job.py` declares its dependencies in a
PEP 723 header, which is what `uv run` reads.

Check the job URL about 30 seconds after submitting. A container that cannot start fails
immediately; a healthy one shows `Installed N packages`, then `InferGauge benchmark job`,
then the per-request lines.

---

## Option B — HF Jobs from the CLI (needs a credit balance)

Runs on HF infrastructure with a built-in cron scheduler, but Jobs is pay-as-you-go:
<https://huggingface.co/settings/jobs> shows *"Add Credits to start your own Jobs"* until
a balance exists. CPU Basic is **$0.01/hour**, so a run costs a fraction of a cent — but
it is not zero, and a free account can't use it.

There is also **no web UI for Jobs**; it is CLI only.

```bash
curl -LSf https://hf.co/cli/install.sh | bash
hf auth login          # paste the token when prompted
```

**Smoke test first.** 3 models × 5 prompts ≈ 15 requests, about a minute:

```bash
hf jobs uv run --with-secrets HF_TOKEN \
  https://huggingface.co/spaces/Infergauge-load/llm-performance-and-load/raw/main/job.py
```

**Full run** — 13 models × 20 prompts × 3 repeats = 780 requests:

```bash
hf jobs uv run --with-secrets HF_TOKEN -e IG_SMOKE=0 -e IG_REPEATS=3 \
  https://huggingface.co/spaces/Infergauge-load/llm-performance-and-load/raw/main/job.py
```

**Monthly refresh**, 07:13 UTC on the 1st:

```bash
hf jobs scheduled uv run "13 7 1 * *" --with-secrets HF_TOKEN -e IG_SMOKE=0 \
  https://huggingface.co/spaces/Infergauge-load/llm-performance-and-load/raw/main/job.py
```

Watch it: `hf jobs ps`, `hf jobs logs <id>`, `hf jobs inspect <id>`.

Jobs can also be created from Python rather than the CLI — `run_job` and
`create_scheduled_job` in `huggingface_hub` — but they still bill against the same credit
balance, so this is not a way around Option B's cost.

---

## Option C — GitHub Actions (free, scheduled)

Public GitHub repos get unlimited Actions minutes with no card. This is the free
equivalent of Option B's cron schedule: put the harness in a repo, add the token as a
repository secret, and let the workflow run monthly and push `leaderboard.csv` back here.

Worth setting up once real numbers exist and the leaderboard needs keeping current — not
before.

## Environment variables

| Variable | Default | What it does |
|---|---|---|
| `HF_TOKEN` | — | Required, **write** scope. Used for inference *and* for pushing results |
| `IG_SMOKE` | `1` | `1` = 3 models × 5 prompts. `0` = the full shortlist |
| `IG_REPEATS` | `3` | Repeats per prompt on a full run |
| `IG_JUDGE` | *(blank)* | Optional judge model id. Costs extra tokens; scored in its own column |
| `IG_CONCURRENCY` | `1` | Parallel requests per model. `1` = single-stream; `25`/`100` = load |
| `IG_MAX_ERROR_RATE` | `0.5` | Refuse to publish above this failure share. Raise it for a load run |
| `IG_FORCE_PUSH` | *(blank)* | `1` publishes anyway, past the guard |
| `IG_REPO_ID` | `Infergauge-load/llm-performance-and-load` | Where to read the harness and push results |

## What it writes back

| Path in the Space | What |
|---|---|
| `leaderboard.csv` | Repo root. This is what `index.html` renders — always the most recent run |
| `leaderboard-c<N>.csv` | Per-concurrency copy, so a new run never destroys the previous level |
| `results/results.csv` | One row per request |
| `results/report.md` | Summary with charts |
| `results/charts/` | PNGs, light and dark |

## Before you publish a run

1. Fill `pricing.json` with each model's published rate at its pinned provider, and set
   `_checked` to the date. Models left blank get a blank cost cell — never a guess.
2. Smoke run. Confirm no unexpected errors.
3. Full run.
4. Read `report.md`, especially the error table. A model at 100% errors usually means the
   provider dropped it from the router, not that the model is bad. Take it off the
   shortlist or note it; don't publish it as a zero.
5. Update the measured date on the Space.

## Files

| File | What |
|---|---|
| `run_benchmark_colab.ipynb` | **Start here.** Free Colab runner, no terminal |
| `job.py` | The runner itself. Fetches the harness, runs it, pushes results back |
| `benchmark.py` | Runner, aggregation, charts, report |
| `quality.py` | Deterministic format scoring, plus the optional judge |
| `prompts.json` | 20 prompts across 5 task families, each with a check |
| `pricing.json` | Per-1M rates per model at its pinned provider — fill this in |
| `slo.json` | Published thresholds that decide the goodput column. Optional; defaults apply if absent |
| `model_shortlist.csv` | 13 models with size, licence, providers |
| `hf_client.py` | OpenAI-compatible client for the Inference Providers router |
