Skip to content

Gemma 4 multimodal judgments

Kind: explanation.

This page is for a platform engineer who must decide whether to send images with a typed question. It explains how typevet carries an image to Gemma 4 on llama.cpp and on vLLM. It also states what the receipts prove and where the evidence stops. For steps, follow the linked how-to pages.

What an image-conditioned typed judgment is

A typed judgment asks named Noul, Choice or Score questions about a state and returns typed answers with probabilities. An image-conditioned judgment adds one or more images to that call. Every scored field sees the same images, in the same order.

The image helps when the answer is in the pixels and not in the text. Examples are a receipt that supports or contradicts an expense claim, or a screenshot that shows which site is open. When the text alone settles the question, an image adds tokens and latency and no new evidence.

The answer is still a probability map over fixed labels. The model does not write free text that you must parse.

How images enter typevet

The domain type is ImageInput. It holds encoded bytes and a mime type. It accepts image/png, image/jpeg and image/webp. It refuses empty bytes and every other mime type with ScoringValidationError.

The caller passes images through the keyword-only media argument of JudgmentPort.judge. None or an empty tuple is the text path. The judgment adapter puts one MEDIA_MARKER (<__media__>) per image in front of the rendered state, one marker per line. The marker order is the image order.

Two request types enforce a count check. CandidateScoringRequest and GenerationRequest each count the markers in the prefix or prompt. When that count differs from the number of images, the request fails before any network call. This check stops an image from being sent without a place in the prompt.

How each backend receives images

The two backends take the same domain request and send it on different wires.

Topic llama.cpp vLLM
Session factory open_gemma_native_vision_judgment open_vllm_judgment
Endpoint POST /completion POST /v1/chat/completions
Image payload Raw base64 in prompt.multimodal_data, beside prompt_string One image_url block per image, with a data: URI
Media marker Replaced with the router marker from GET /props Split out. Each marker becomes one image block
Vision check /props must report vision, or the factory raises ValueError No probe. The server must accept images
Template typevet composes the native Gemma 4 turn itself The server applies its chat template
Thinking Prefix ends with the no-thinking channel prefill chat_template_kwargs sets enable_thinking to false
Scores n_probs over the full vocabulary, 262144 entries logprob_token_ids, at most 128 candidates
Prompt cache cache_prompt is false on every request No typevet setting
Image count limit None in typevet Tested server flag --limit-mm-per-prompt {"image":2}
Model pin gemma-4-31b-kv9-q4km-mm, Q4 GGUF with its projector google/gemma-4-31B-it, BF16, vLLM 0.30.0

On llama.cpp, the model must load with its multimodal projector (--mmproj). The router reports modalities.vision and a media_marker through GET /props. The router randomizes that marker for each server instance. The adapter caches the marker for each model id and substitutes it for MEDIA_MARKER.

The nested prompt object is not a style choice. A top-level multimodal_data beside a string prompt returns HTTP 200 and drops the image. The image-conditioned live smoke records that finding and the wire shape.

On vLLM, the server applies the chat template, so typevet sends plain content without turn markers. The factory makes no HTTP call before the first judge call. See Serve typevet on vLLM for the server flags and the settings.

The typed-generation path differs. The vLLM generation adapter sends images as content blocks. The llama.cpp generation adapters refuse a request with images before any HTTP call.

What is specific to Gemma 4

Gemma 4 uses a native turn format: <|turn>user, the user text, <turn|>, then <|turn>model. On llama.cpp, typevet composes this prefix itself. The media markers sit inside the user turn, before the state and the field instructions. After the model header, the prefix adds the no-thinking prefill <|channel>thought\n<channel|>. This matches /apply-template output with thinking turned off.

The llama.cpp factory checks the served template before it builds a port. It renders one turn through POST /apply-template and classifies the result. By default it requires native_gemma4_turn. A ChatML render, a Gemma 3 render or a mixed render raises ValueError before the first judge call.

The judgment adapter also fails closed. With a native Gemma 3 or Gemma 4 family, it wraps every prefix in that turn, with or without images. An imaged prefix and an image-omitted prefix then differ only in the media markers. Without a native family, a text-only request falls back to a degraded ChatML prefix. A request with images and no native family raises JudgmentValidationError before any scoring request. typevet never composes a ChatML prefix around an image.

On vLLM, typevet does not classify the template. The server owns it, and the tested pin is the official Gemma 4 checkpoint.

What the receipts prove

Each figure below comes from one recorded run. None is a calibration claim.

Image-conditioned live smoke, llama.cpp, Gemma 3. On gemma-3-4b-it-q4km-mm, the smoke asked one colour Choice question. The smoke how-to names that model id. It answered red, green and blue correctly, and the omitted answer differed from the imaged answers. Prompt tokens went from 103 without an image to 362 with an image, a gap of 259 (#142 acceptance). This run predates cache_prompt: false and native Gemma 3 turns.

PSAI vision smoke, llama.cpp, Gemma 3. Five public computer-use screenshots ran present, omitted and swapped controls. Paired ordering held on 5 of 5 rows. The image added 259 tokens, and no omitted row counted as a hit (#154 acceptance). This run also predates cache_prompt: false and native Gemma 3 turns. Run the PSAI vision smoke names the model id, gemma-3-4b-it-q4km-mm.

CORD receipt images, llama.cpp, Gemma 4. The CORD set holds six public receipts and 18 synthetic claims. On gemma-4-31b-kv9-q4km-mm with template native_gemma4_turn and build b11223-4da633776, the run used 43 requests. All five semantic checks passed, and the acceptance command exited 0 (#203 receipt). See Run the CORD expense smoke.

Check Limit Measured n
Answerable accuracy floor 0.67 1 12
Contradicted recall floor 0.5 1 6
False-clear rate ceiling 0.25 0 12
Insufficient abstention rate floor 0.5 1 6
Largest label share ceiling 0.8 0.3333 18

The text-only arm of that run had accuracy 0.667 and contradicted recall 0. The image drove the combined result. The CORD how-to records that one Gemma 4 image costs 245 prompt tokens on the CORD smoke router.

vLLM acceptance, Gemma 4. One run on the vLLM pin passed every pre-registered gate (#170 receipt).

Set Calls Result
PSAI image present 4 4 of 4 correct. The image added about 265 tokens
PSAI image swapped 4 4 of 4 passed the swap gate
PSAI image omitted 4 Not counted toward any gate. 1 of 4 matched gold by chance (omitted_credited: 1)
PSAI text only 4 4 of 4 correct
CORD 43 Accuracy 1.0, contradicted recall 1.0, false-clear 0, abstention 1.0, label share 0.33
CORD with labels in reverse order 18 Same five checks pass, with 0 label flips

The swapped control is the strongest signal here. The text stays the same, and only the image changes.

The two backends ran different weights: Q4 GGUF on llama.cpp and BF16 on vLLM. The receipts do not compare backends, and a difference is not a backend effect.

Limits

  • Image count. typevet sets no cap on images per request. The llama.cpp scoring adapter sends every image it gets. The tested vLLM server allows two images per prompt. typevet does not check that limit before the request.
  • Image size. ImageInput checks the mime type and non-empty bytes only. An 8 MiB payload passes. Pixel limits on the served projector are not characterized (#204).
  • Long-running service. No test covers long-running behaviour beyond the adapter lifetime rules (#204). The llama.cpp marker changes when the router reloads the model. Build a new session after a reload. A stale marker fails tokenization with HTTP 400. The native vision how-to describes session ownership.
  • One model pin per backend. Evidence covers gemma-4-31b-kv9-q4km-mm on llama.cpp and google/gemma-4-31B-it at one revision on vLLM 0.30.0. Other models, quantizations and versions are not tested.
  • Default model id. TYPEVET_LLAMA__MULTIMODAL_MODEL defaults to the Gemma 3 id gemma-3-4b-it-q4km-mm (settings). The native vision factory requires the Gemma 4 turn by default. Set that variable, or pass model=, to the Gemma 4 id. Otherwise, a router that serves the Gemma 3 id makes the factory raise ValueError when it opens the session.
  • No general OCR claim. The receipts cover six CORD receipts and a few screenshots. They say nothing about text extraction quality on other documents.
  • Small samples. Each set is one run with n of 18 or fewer per check. Treat the numbers as smoke evidence, not as accuracy you can expect in production.

The release support matrix lists each runtime limit with the test that proves it.