Live eval runner over typed loaders¶
Kind: reference.
Parent: #98, #95. Depends on the stock llama.cpp path (#97).
This runner drives a tiny balanced slice of landed loaders through
GenerationPort (grammar-JSON floor). It reports:
| Counter | Meaning |
|---|---|
attempted |
generate calls made |
schema_valid |
Calls that returned a value validating against the task schema |
gold_match |
Schema-valid outputs equal to loader gold |
Banking77 uses reports_unauthorized Noul agreement vs the six-intent proxy
(fraud → true). BoolQ uses exact match on the answer Noul
(no / yes).
This measures structure + accuracy on gold. It is not calibration, ECE, or probability quality (#50).
Supported loaders¶
| Dataset | Default limit | Metric label |
|---|---|---|
banking77 |
4 (balanced fraud / not_fraud) | noul_agreement |
boolq |
4 (balanced no / yes) | exact_match |
CI and offline unit tests use vendored fixtures under tests/fixtures/.
Live runs may download public Hub slices when fixture text is not injected.
Commands¶
Offline wiring (default CI):
Opt-in live (router + model required):
export TYPEVET_LLAMA__DEFAULT_MODEL='<your-gemma-4-model-id>'
TYPEVET_LLAMA__TIMEOUT=600 \
uv run python -m typevet_evals.cli.eval_runner --dataset boolq --limit 2
Proof runs that must not silently skip (nonzero on missing config, model, workload, or incomplete schema-valid completion; wrong gold labels stay in metrics only):
Run both loaders in one invocation:
Pytest live marker (BoolQ smoke):
When TYPEVET_LLAMA__DEFAULT_MODEL is unset or the router is down, the CLI
prints skip: … to stderr and exits 0; live pytest tests skip.
Release evidence runs that must not silently skip live collection:
With TYPEVET_REQUIRE_LIVE set to a truthy value (1, true, yes,
on), the same missing router or model conditions fail the test instead
of skipping. Default behaviour is unchanged. The eval CLI equivalent remains
--require-live.
Library API¶
from typevet_evals.runner.core import run_eval_tasks
from typevet_evals.runner.datasets import load_eval_tasks
tasks = load_eval_tasks("boolq", limit=2, boolq_jsonl_text=open("…").read())
report = run_eval_tasks(port, tasks, model="your-model-id")
port is any GenerationPort (FakeGenerationAdapter offline,
LlamaCppGenerationAdapter live).