Run a small live judgment eval¶
Kind: how-to.
Run one bounded live check after the router and Gemma 4 weights work. Pick
TPJEP eight (in-repo pytest) or read the frozen finvet-6 receipt on
#133. Both use
ScoringJudgmentAdapter and local candidate scoring.
Prerequisites¶
- Complete Run Gemma 4 on llama.cpp.
- Set
TYPEVET_LLAMA__BASE_URLif the router is nothttp://127.0.0.1:8090. - Pin model id
gemma-4-31b-24gib-kv11-decoderunless your catalog differs.
Live pytest skips when the router is down. A skip is not a pass.
Option A — TPJEP eight-task smoke¶
This path exercises eight vendored JevBench rows through JudgmentPort. Gold
expected values never enter model inputs.
Offline proof first:
uv run pytest evals/tests/unit/test_eval_tpjep_loader.py evals/tests/contract/test_tpjep_runner_offline.py -q
Live run (slow first load):
Artifacts land under scratchpad/tpjep/ (gitignored). Details:
TPJEP v0 eight-task runner.
Option B — frozen finvet-6 receipt (read-only)¶
typevet does not ship a finvet-6 pytest yet. The supervisor recorded a post-#152 live receipt on #133.
| Field | Value |
|---|---|
| Code tip | 3f25851 |
| Binding fix | #152 249f168 |
| Model | gemma-4-31b-24gib-kv11-decoder |
| Template | Degraded ChatML + Control→label mapping |
| Broad fraud agreement | 6/6 (was 3/6 pre-fix on the same six rows) |
| Semantic controls | pos_unauthorized and neg_ordinary both passed |
| Elapsed | 29.849 s for 18 workload fields + 6 control fields |
Interpretation stays honest: binding repair changed decisions away from
all-duplicate_charge. That evidence does not mean calibration or close
133. Full table:¶
After the run¶
- Compare
n_answeredandn_prob_validin the TPJEP summary JSON. - Do not report ECE or “calibrated” from these smokes alone.
- File or extend an issue when you need a new pin or a larger sample.