Skip to content

Native typed judgments in typevet

Kind: explanation.

Status: draft.

typevet ships native System One–shaped judgment: Noul, Choice, and Score questions, normalized answers, and probability maps from candidate scoring. The product runs on local llama.cpp with Gemma 4 weights. typevet owns this implementation. TypeLLM and judgevet are research and protocol references, not dependencies you must install to use typevet.

Grammar-JSON generation remains the transport floor for some paths. Typed judgment is the spine judgevet-shaped callers need later. See TypeLLM, Jev and judgevet for the family map.

What you get today

Capability Status
JudgmentPort + offline fakes Yes — tutorial and contract fixtures
ScoringJudgmentAdapter over CandidateScoringPort Yes — offline and llama.cpp scoring
TPJEP v0 eight-task smoke Yes — offline runner + opt-in live pytest
finvet-derived six-message exploratory receipt Documented on #133; not a shipped CLI
Public calibration or ECE headline No — see limitations

Where to start

  1. First typed judgment offline — wire ScriptedScoringFake and ScoringJudgmentAdapter with no model.
  2. Run a small live judgment eval — TPJEP eight or read the frozen finvet-6 receipt.
  3. Judgment live receipts — measured numbers and pins, with links to issue comments.

Structured JSON from GenerationPort is a sibling path. See Call typevet from Python.

Limitations

Valid structure is not calibration. Schema-valid or probability-valid outputs do not prove task accuracy, ECE, or production readiness. Say which eval counted gold only when the run used a loader with explicit gold-match or broad-agreement rules.

Template honesty. On llama.cpp, the judgment session classifies the served template through /apply-template. By default it accepts only native Gemma 4 turns and fails before scoring on any other family (Gemma 4 page). The classifier reads the markers in the rendered output, not the template source. The recorded llama.cpp receipt used a --chat-template-file override (#233). Thus native_gemma4_turn does not prove that the GGUF's own template ran. A ScoringJudgmentAdapter built with neither a served family nor a framing uses a degraded ChatML prefix for text only. With image input, it fails instead (judgment_scoring.py). On vLLM, the server applies its own chat template, and typevet does not classify it (support matrix). Receipts must name the template class. Do not treat a degraded template as silent equivalence to hosted Jev.

Small samples. The finvet-derived six-row receipt on #133 is exploratory. It does not close the broader #133 quality study.

Binding and labels. Control-token scoring requires model-facing options to match scored tokens (#152, commit 249f168). Pre-fix all-duplicate_charge failures were a demonstrated binding mismatch, not proof that post-fix runs are calibrated.

No TypeLLM compatibility promise. typevet reimplements the open-weight decision ideas. It does not guarantee byte-for-byte TypeLLM or SGLang parity.