Native typed judgments in typevet¶
Kind: explanation.
Status: draft.
typevet ships native System One–shaped judgment: Noul, Choice, and
Score questions, normalized answers, and probability maps from candidate
scoring. The product runs on local llama.cpp with Gemma 4 weights.
typevet owns this implementation. TypeLLM
and judgevet are research and
protocol references, not dependencies you must install to use typevet.
Grammar-JSON generation remains the transport floor for some paths. Typed judgment is the spine judgevet-shaped callers need later. See TypeLLM, Jev and judgevet for the family map.
What you get today¶
| Capability | Status |
|---|---|
JudgmentPort + offline fakes |
Yes — tutorial and contract fixtures |
ScoringJudgmentAdapter over CandidateScoringPort |
Yes — offline and llama.cpp scoring |
| TPJEP v0 eight-task smoke | Yes — offline runner + opt-in live pytest |
| finvet-derived six-message exploratory receipt | Documented on #133; not a shipped CLI |
| Public calibration or ECE headline | No — see limitations |
Where to start¶
- First typed judgment offline
— wire
ScriptedScoringFakeandScoringJudgmentAdapterwith no model. - Run a small live judgment eval — TPJEP eight or read the frozen finvet-6 receipt.
- Judgment live receipts — measured numbers and pins, with links to issue comments.
Structured JSON from GenerationPort is a sibling path. See
Call typevet from Python.
Limitations¶
Valid structure is not calibration. Schema-valid or probability-valid outputs do not prove task accuracy, ECE, or production readiness. Say which eval counted gold only when the run used a loader with explicit gold-match or broad-agreement rules.
Template honesty. On llama.cpp, the judgment session classifies the
served template through /apply-template. By default it accepts only native
Gemma 4 turns and fails before scoring on any other family
(Gemma 4 page).
The classifier reads the markers in the rendered output, not the template
source. The recorded llama.cpp receipt used a --chat-template-file override
(#233). Thus
native_gemma4_turn does not prove that the GGUF's own template ran. A
ScoringJudgmentAdapter built with neither a served family nor a framing
uses a degraded ChatML prefix for text only. With image input, it fails
instead
(judgment_scoring.py).
On vLLM, the server applies
its own chat template, and typevet does not classify it
(support matrix).
Receipts must name the template class. Do not treat a degraded template as
silent equivalence to hosted Jev.
Small samples. The finvet-derived six-row receipt on #133 is exploratory. It does not close the broader #133 quality study.
Binding and labels. Control-token scoring requires model-facing options to
match scored tokens (#152,
commit 249f168).
Pre-fix all-duplicate_charge failures were a demonstrated binding mismatch,
not proof that post-fix runs are calibrated.
No TypeLLM compatibility promise. typevet reimplements the open-weight decision ideas. It does not guarantee byte-for-byte TypeLLM or SGLang parity.