How typevet works with Gemma 4¶
Kind: explanation.
This page is for a platform engineer who has not seen typevet. It explains what typevet does with Gemma 4, how the llama.cpp and vLLM backends differ, and what the receipts prove. It links to other pages for setup steps and architecture detail.
The problem typevet solves¶
A model that answers in free text gives the caller a string to parse. The caller does not know how sure the model was. typevet asks a typed question and returns a typed answer with a probability.
typevet uses three question types from the judgevet vocabulary:
- Noul: a yes or no question. The answer is the probability of
true. - Choice: pick one label from a list. The answer is the selected label, its confidence and a probability for each label.
- Score: pick a level on a rubric of two or more levels. The answer is the expected level, a confidence and a probability for each level.
A fourth path is schema-bound generation. The caller sends a prompt and a JSON Schema. typevet returns an object that passes that schema, or it raises an error. This path has no probability.
The judgment port (JudgmentPort) carries the
three question types. The generation port
(GenerationPort and AsyncGenerationPort) carries schema-bound generation.
See native typed judgments for the product scope.
How a typed judgment becomes a score¶
typevet does not let the model write an answer. It reads the model's next-token distribution at the answer position, before any sampling. It then keeps only the tokens that stand for the allowed answers.
The flowchart shows how one typed question goes from judge() to a typed
answer.
flowchart TD
J["Caller calls judge(state, questions, model)"] --> D["Question becomes a Decision with ordered labels"]
D --> C["Each label gets a digit control: 0, 1, 2, ..."]
C --> F["Field block lists each control with its label"]
F --> P["State and field block become one prefix; framing or served template sets turn markers"]
P --> R["One PRE_SAMPLING scoring request per question"]
R --> L["Server returns candidate logprobs"]
L --> V{"Exact token ids, in order, finite?"}
V -- no --> E["Call fails closed with an error"]
V -- yes --> S["Softmax turns the logprobs into probabilities"]
S --> A["Noul: P(true). Choice: top option. Score: expected level"]
The steps in the code:
ScoringJudgmentAdapter.judgeinadapters/outbound/judgment_scoring.pyreceives the call.normalize_questionandjudgment_original_labelsindomain/judgment_normalize.pymake the Decision. Noul gives(false, true). Choice gives the criteria keys. Score gives0ton-1.control_binding_pairsandbind_control_candidatesbind the controls. The tokenizer must encode each digit as exactly one token id.render_field_instructionsindomain/field_instructions.pywrites the field block. A Choice line reads0 → billing: Money. A Noul or Score line readsControl 0 → false. The last line asks for exactly one control string._field_prefixand_compose_prefixmake one prefix that ends at the answer. An injected framing, such asChatContentFramingfor vLLM, composes it. Otherwiseadapters/outbound/gemma/scoring_prefix.pyuses the served template family.execute_categorical_decisionindomain/decision_execute.pysends the request toCandidateScoringPort.score_candidates. The llama.cpp and vLLM adapters implement it.validate_result_against_requestindomain/candidate_scoring_validate.pychecks the result._softmaxand_greedy_indexcompute the probabilities at temperature 1 by default.answer_from_executionbuilds the typed answer.
The probabilities are conditional on the listed options. Mass that the model puts on other tokens does not show in the answer. The #207 study below shows why that matters.
The scoring port (CandidateScoringPort) is the
only part that talks to a server. The steps before and after it are pure
domain code. Library-first architecture
explains the layers.
How generation with a schema works¶
Four generation adapters send one chat completion each:
LlamaCppGenerationAdapter, AsyncLlamaCppGenerationAdapter,
VllmGenerationAdapter and AsyncVllmGenerationAdapter. All four follow the
same three steps.
- Check the schema before the request.
check_request_schemachecks the schema against the JSON Schema Draft 2020-12 meta-schema. A malformed schema raisesValueError, and no request is sent. - Ask the server to constrain the output. llama.cpp gets a nested
response_format.json_schema, which it turns into a grammar. vLLM getsstructured_outputs: {"json": <schema>}. - Validate the reply on the client.
validated_valueparses the content, rejects non-finite numbers and runsjsonschema.validate. A failure raisesSchemaValidationError.
The client check is not a formality. On one llama.cpp pin, the grammar did not
enforce number bounds or multipleOf (#129). The
client check is the guard for number bounds. The multipleOf call returned
empty content, and the adapter raises GenerationError for empty content
before any schema check.
What is specific to Gemma 4¶
Native turn template. Gemma 4 marks turns with <|turn> and <turn|>. On
llama.cpp, the judgment session renders one message through /apply-template.
It then classifies the served template as native Gemma 4, native Gemma 3,
degraded ChatML or unsupported. The llama.cpp judgment session accepts only a
native Gemma 3 or Gemma 4 family. It also requires a model that declares image
input. typevet then writes the prefix in that
family's turn markers itself, because /completion takes raw text. On vLLM,
the server applies the served chat template. typevet sends the prefix as plain
user content.
Thinking off. Gemma 4 can start a thinking section before it answers. typevet reads the score at the first answer token, so no thinking section may open there. The #207 research found the no-thinking prefill boundary to be the main driver of off-menu mass. Each backend turns thinking off in a different way:
- llama.cpp scoring adds the no-thinking prefill
<|channel>thought\n<channel|>after the model turn header. The answer token comes directly after it. - vLLM scoring and vLLM generation send
chat_template_kwargs: {"enable_thinking": false}. - llama.cpp generation sends the same
enable_thinking: falsevalue.
Digit controls and the one-token limit. The model answers with a digit,
not with the label text. One next-token read then covers the whole answer.
On the Gemma 4 tokenizer, "0" to "9" are
single tokens and "10" is two. A native question with more than 10 options
therefore fails before any scoring call, with the message
native Choice supports N options on this tokenizer; got M (#234). The
rendered digits and the scored token ids come from one function,
control_binding_pairs, so they cannot drift apart.
Image input. A judgment can take images. Both backends support image-conditioned scoring. On vLLM, generation can also take images. llama.cpp generation refuses images before any request. See Gemma 4 multimodal judgments for the detail.
The two backends side by side¶
| Aspect | llama.cpp | vLLM |
|---|---|---|
| Use | Local development and small receipts | Hosting and throughput |
| Weights in the receipts | Quantized GGUF files | BF16 google/gemma-4-31B-it at one revision |
| Scoring call | POST /completion with the raw prefix, n_predict 0, n_probs 262144 |
POST /v1/chat/completions with max_tokens 1, logprob_token_ids, return_tokens_as_token_ids |
| Scoring result | completion_probabilities[0].top_logprobs |
choices[0].logprobs.content[0].top_logprobs |
| Turn markers | typevet writes them (native Gemma 4 turn and prefill) | The server applies the chat template |
| Control tokenizer | Router /tokenize |
Server /tokenize with add_special_tokens false |
| Generation constraint | response_format.json_schema, which becomes a grammar |
structured_outputs with the JSON Schema |
| Parallel requests | No concurrency setting | AsyncVllmGenerationAdapter with a max_concurrency limit |
| Candidate limit per call | No adapter limit | At most 128 logprob_token_ids |
| Setup page | Run Gemma 4 on llama.cpp | Serve typevet on vLLM |
Three ports make the backends interchangeable:
CandidateScoringPorthas one method,score_candidates.LlamaCppCandidateScoringAdapterandVllmCandidateScoringAdapterboth implement it.ScoringJudgmentAdaptertakes either one.GenerationPortandAsyncGenerationPorthave onegeneratemethod. The four generation adapters implement them.ModelFramingPortcomposes the prefix. The vLLM session injectsChatContentFraming, which adds no turn markers. The llama.cpp session passes the served template family instead.
TYPEVET_BACKEND selects llama_cpp (the default) or vllm for the
composition root. The caller code that calls judge or generate does not
change.
What the receipts prove¶
Each claim below comes from one issue comment. Read the receipt for the full pin and its limits.
- llama.cpp CORD smoke (#203):
one live run of the CORD combined arm on llama.cpp build
b11223-4da633776with a native Gemma 4 turn. It passed 5 of 5 checks, and the acceptance command exited 0. The sample is 6 to 18 claims per check. It is smoke evidence, not calibration. - vLLM acceptance (#170):
one run on
vllm/vllm-openai:v0.30.0with BF16google/gemma-4-31B-iton one H100 80 GB. The generation, image scoring and CORD sets passed their pre-registered gates. It used 88 model calls. Two sets were record-only: reversed label order gave 0 flips, and 4 parallel generation calls gave 0 errors. - llama.cpp grammar enforcement (#129):
one run used build
b11243-fc07d781eand a Gemma 4 31B QAT Q4_0 GGUF. The grammar enforced stringenum, integer bounds and the JSON object root. It did not enforce number bounds ormultipleOf. typevet's client validation catches a number outside its bounds. ThemultipleOfcall returned empty content, which the adapter raises asGenerationError. - Off-menu mass fix (#207): one 6-option Choice prompt put 0.9105 of the next-token mass on the word "Control", not on a digit. After Choice options dropped the word "Control", the off-menu mass on that prompt was 1.99e-7. The winning label did not change. This was one prompt on llama.cpp only.
- H100 throughput (#236): Attempt 1 made no model call. One question had 13 options, and native Choice supports 10 on the Gemma 4 tokenizer (#234). Attempt 2 measured public datasets on one H100, and Performance on one H100 gives its numbers, calibration and limits.
What is not claimed¶
- Other models. Every receipt uses Gemma 4 31B. No other model is tested.
- Other quantizations. Each receipt covers its own weights file. #129
used a Q4_0 file. #203 ran the local alias
gemma-4-31b-kv9-q4km-mmand did not record its file type. That alias loads a Q2_K file today (#233). vLLM used BF16. The #170 receipt does not attribute any difference to the backend, because the weights differ. - Calibration quality. A valid structure and a valid probability do not prove accuracy or calibration. The receipts are small samples. See native typed judgments.