{
  "version": "1.0",
  "created": "2026-09-24",
  "notes": "20 prompts across 5 task families. Each carries a deterministic check so quality can be scored without a judge model; judge scoring is optional on top. Keep prompts stable across runs — changing them invalidates comparison with previous leaderboards.",
  "prompts": [
    {
      "id": "qa-01",
      "task": "short_qa",
      "prompt": "In one sentence, what does time to first token measure?",
      "max_tokens": 80,
      "check": {"type": "contains_any", "values": ["first token", "first output", "initial token"], "max_words": 45}
    },
    {
      "id": "qa-02",
      "task": "short_qa",
      "prompt": "What is the capital of Australia? Answer with the city name only.",
      "max_tokens": 16,
      "check": {"type": "contains_any", "values": ["Canberra"], "max_words": 5}
    },
    {
      "id": "qa-03",
      "task": "short_qa",
      "prompt": "Name the HTTP status code returned when a client is rate limited. Number only.",
      "max_tokens": 16,
      "check": {"type": "contains_any", "values": ["429"], "max_words": 5}
    },
    {
      "id": "qa-04",
      "task": "short_qa",
      "prompt": "In one sentence, explain the difference between p95 latency and average latency.",
      "max_tokens": 80,
      "check": {"type": "contains_any", "values": ["95", "percentile"], "max_words": 60}
    },

    {
      "id": "sum-01",
      "task": "summarization",
      "prompt": "Summarise in exactly two sentences:\n\nLoad testing an LLM endpoint differs from load testing a conventional web API in three ways. Responses stream, so a single latency number hides the wait before the first token and the gaps between tokens after it. Cost scales with tokens consumed rather than with requests served, so a slow expensive request and a fast cheap one look identical on a request-count graph. And answer quality can degrade under concurrency even while latency stays within limits, because providers shed work by truncating or routing to smaller models.",
      "max_tokens": 200,
      "check": {"type": "sentence_count", "value": 2, "tolerance": 1}
    },
    {
      "id": "sum-02",
      "task": "summarization",
      "prompt": "Compress the following into a single sentence under 25 words:\n\nThe saturation knee is the concurrency level at which p95 latency stops rising roughly linearly with added users and begins to inflect sharply upward. Below the knee, the system has headroom and additional users are absorbed. Above it, each additional user costs disproportionately more latency, and queueing dominates. Locating the knee answers the capacity planning question directly.",
      "max_tokens": 120,
      "check": {"type": "max_words", "value": 32}
    },
    {
      "id": "sum-03",
      "task": "summarization",
      "prompt": "Extract the three numbered findings from this test report as a bulleted list, nothing else:\n\nRun 47 completed with 2000 users over ten minutes. First, p95 latency crossed the six second SLA at approximately 380 concurrent users. Second, goodput fell to 71 percent at peak, driven almost entirely by TTFT breaches rather than errors. Third, projected monthly cost at the measured rate is 4,180 US dollars, assuming sustained peak traffic.",
      "max_tokens": 200,
      "check": {"type": "line_count_min", "value": 3}
    },
    {
      "id": "sum-04",
      "task": "summarization",
      "prompt": "Rewrite this for a non-technical executive in two sentences, no jargon:\n\nHigh TTFT combined with healthy inter-token latency is the queueing signature: the model generates at full speed once it starts, but requests are waiting in line for a worker. The fix is capacity or parallelism, not a smaller model.",
      "max_tokens": 160,
      "check": {"type": "sentence_count", "value": 2, "tolerance": 1}
    },

    {
      "id": "code-01",
      "task": "coding",
      "prompt": "Write a Python function `p95(values: list[float]) -> float` returning the 95th percentile using the nearest-rank method. Return only code in a single fenced block, no explanation.",
      "max_tokens": 300,
      "check": {"type": "python_compiles", "must_define": "p95"}
    },
    {
      "id": "code-02",
      "task": "coding",
      "prompt": "Write a Python function `tokens_per_second(output_tokens: int, total_s: float, ttft_s: float) -> float` that measures throughput over the generation window only (after the first token), returning 0.0 if that window is not positive. Code only, one fenced block.",
      "max_tokens": 300,
      "check": {"type": "python_compiles", "must_define": "tokens_per_second"}
    },
    {
      "id": "code-03",
      "task": "coding",
      "prompt": "Write a Python function `parse_sla(text: str) -> dict` that parses lines of the form `p95_latency_ms: 8000` into a dict with int values, ignoring blank lines and lines starting with #. Code only, one fenced block.",
      "max_tokens": 350,
      "check": {"type": "python_compiles", "must_define": "parse_sla"}
    },
    {
      "id": "code-04",
      "task": "coding",
      "prompt": "Write a bash one-liner that finds every .yaml file under the current directory modified in the last 7 days and prints its path and size. Return only the command in a fenced block.",
      "max_tokens": 150,
      "check": {"type": "contains_any", "values": ["find"]}
    },

    {
      "id": "reason-01",
      "task": "reasoning",
      "prompt": "A request waits 4.0 seconds before the first token, then streams 200 tokens with an average gap of 20ms. What is the total latency in seconds, and is the bottleneck queueing or generation? Answer in two lines: the number, then the diagnosis.",
      "max_tokens": 160,
      "check": {"type": "contains_any", "values": ["8", "8.0", "queue"]}
    },
    {
      "id": "reason-02",
      "task": "reasoning",
      "prompt": "Endpoint A: TTFT p95 0.4s, ITL p95 95ms. Endpoint B: TTFT p95 3.8s, ITL p95 18ms. For a chat UI where users read as text appears, which feels faster, and why? Two sentences.",
      "max_tokens": 160,
      "check": {"type": "contains_any", "values": ["A", "first token"]}
    },
    {
      "id": "reason-03",
      "task": "reasoning",
      "prompt": "A service meets its p95 latency SLA and its error-rate SLA, but goodput is 62%. Explain how that is possible in two sentences.",
      "max_tokens": 160,
      "check": {"type": "contains_any", "values": ["every", "all", "another", "other", "simultaneous"]}
    },
    {
      "id": "reason-04",
      "task": "reasoning",
      "prompt": "Cost is $0.0042 per request at 50,000 requests per day. What is the projected 30-day spend? Show the arithmetic in one line, then the total.",
      "max_tokens": 160,
      "check": {"type": "contains_any", "values": ["6300", "6,300"]}
    },

    {
      "id": "json-01",
      "task": "json",
      "prompt": "Return ONLY valid JSON, no prose and no code fence, matching exactly this shape: {\"model\": string, \"ttft_ms\": number, \"healthy\": boolean}. Use model \"demo\", ttft_ms 320, healthy true.",
      "max_tokens": 120,
      "check": {"type": "json_schema", "required_keys": ["model", "ttft_ms", "healthy"], "types": {"model": "str", "ttft_ms": "num", "healthy": "bool"}}
    },
    {
      "id": "json-02",
      "task": "json",
      "prompt": "Return ONLY valid JSON, no prose. An object with key \"metrics\" holding an array of exactly three objects, each with keys \"name\" (string) and \"value\" (number). Use latency 4.2, ttft 0.9, goodput 93.",
      "max_tokens": 200,
      "check": {"type": "json_schema", "required_keys": ["metrics"], "array_key": "metrics", "array_len": 3}
    },
    {
      "id": "json-03",
      "task": "json",
      "prompt": "Convert to JSON, output only the JSON object: p95 latency is 4200 ms, error rate is 0.4 percent, the run passed. Use keys p95_latency_ms, error_rate_pct, passed.",
      "max_tokens": 150,
      "check": {"type": "json_schema", "required_keys": ["p95_latency_ms", "error_rate_pct", "passed"], "types": {"p95_latency_ms": "num", "error_rate_pct": "num", "passed": "bool"}}
    },
    {
      "id": "json-04",
      "task": "json",
      "prompt": "Return ONLY a JSON array of the three test types named here, as lowercase strings, no prose: Load, Stress, Spike.",
      "max_tokens": 100,
      "check": {"type": "json_array", "length": 3, "lowercase": true}
    }
  ]
}
