Skip to content

Live eval runner over typed loaders

Kind: reference.

Parent: #98, #95. Depends on the stock llama.cpp path (#97).

This runner drives a tiny balanced slice of landed loaders through GenerationPort (grammar-JSON floor). It reports:

Counter Meaning
attempted generate calls made
schema_valid Calls that returned a value validating against the task schema
gold_match Schema-valid outputs equal to loader gold

Banking77 uses reports_unauthorized Noul agreement vs the six-intent proxy (fraud → true). BoolQ uses exact match on the answer Noul (no / yes).

This measures structure + accuracy on gold. It is not calibration, ECE, or probability quality (#50).

Supported loaders

Dataset Default limit Metric label
banking77 4 (balanced fraud / not_fraud) noul_agreement
boolq 4 (balanced no / yes) exact_match

CI and offline unit tests use vendored fixtures under tests/fixtures/. Live runs may download public Hub slices when fixture text is not injected.

Commands

Offline wiring (default CI):

uv run pytest evals/tests/unit/test_eval_runner.py -m unit -q

Opt-in live (router + model required):

export TYPEVET_LLAMA__DEFAULT_MODEL='<your-gemma-4-model-id>'
TYPEVET_LLAMA__TIMEOUT=600 \
  uv run python -m typevet_evals.cli.eval_runner --dataset boolq --limit 2

Proof runs that must not silently skip (nonzero on missing config, model, workload, or incomplete schema-valid completion; wrong gold labels stay in metrics only):

uv run python -m typevet_evals.cli.eval_runner --require-live --dataset boolq --limit 2

Run both loaders in one invocation:

uv run python -m typevet_evals.cli.eval_runner --dataset boolq --dataset banking77 --limit 2

Pytest live marker (BoolQ smoke):

uv run pytest evals/tests/live/test_eval_runner_live.py -m live -q

When TYPEVET_LLAMA__DEFAULT_MODEL is unset or the router is down, the CLI prints skip: … to stderr and exits 0; live pytest tests skip.

Release evidence runs that must not silently skip live collection:

export TYPEVET_REQUIRE_LIVE=1
uv run pytest evals/tests/live/test_eval_runner_live.py -m live -q

With TYPEVET_REQUIRE_LIVE set to a truthy value (1, true, yes, on), the same missing router or model conditions fail the test instead of skipping. Default behaviour is unchanged. The eval CLI equivalent remains --require-live.

Library API

from typevet_evals.runner.core import run_eval_tasks
from typevet_evals.runner.datasets import load_eval_tasks

tasks = load_eval_tasks("boolq", limit=2, boolq_jsonl_text=open("…").read())
report = run_eval_tasks(port, tasks, model="your-model-id")

port is any GenerationPort (FakeGenerationAdapter offline, LlamaCppGenerationAdapter live).