Judgment live receipts¶
Kind: reference.
This page lists measured small live judgment runs. Each row links primary evidence on GitHub. Do not copy numbers into product claims without the same pins.
finvet-derived six-row receipt (post-#152)¶
Source comment: #133 post-#152 frozen finvet-6 live receipt.
| Pin | Value |
|---|---|
| Commit tip | 3f25851 |
| Binding fix | #152 249f168 — Control→label in field instructions |
| Model | gemma-4-31b-24gib-kv11-decoder |
| Server build | b11176-f805c57a2 (as recorded on #133) |
| Template | Degraded ChatML + Control→label mapping |
| Sample | 6 finvet-style rows + 2 predeclared semantic controls |
| Elapsed | 29.849 s (24 scored fields total) |
Outcome summary¶
| Metric | Pre-fix baseline (#133 exploratory) | Post-#152 |
|---|---|---|
| Broad fraud agreement | 3/6 | 6/6 |
All Choice = duplicate_charge |
yes | no |
| Semantic controls | not reported on baseline row | both passed |
Semantic controls (post-#152)¶
| Control | passed | Modal choice / noul (short) |
|---|---|---|
| pos_unauthorized | True | unauthorized_transaction, noul ≈ 0.993 |
| neg_ordinary | True | not_fraud, noul ≈ 0.010 |
Per-row Choice, unauthorized noul, urgency, and broad flags are in the #133
comment table. Full probability maps were stored in
scratchpad/finvet6/post152_receipt.json (gitignored).
What this receipt does not prove¶
- Calibration, ECE, or production quality (#133 remains open).
- Native Gemma chat template parity (#129).
- Generalization beyond the frozen six messages.
TPJEP eight-task live smoke¶
| Pin | Value |
|---|---|
| Test | evals/tests/live/test_tpjep_smoke_live.py |
| Fixture | tests/fixtures/tpjep/eight_task_smoke.jsonl |
| Default model | gemma-4-31b-24gib-kv11-decoder |
| Command | uv run pytest evals/tests/live/test_tpjep_smoke_live.py -m live -q |
| Artifacts | scratchpad/tpjep/live_summary.json, live_attempts.jsonl |
Receipt rules live in tests/fixtures/tpjep/live_acceptance.py. See
TPJEP v0 eight-task runner.