Connect Gemma 4 native vision judgment¶
Kind: how-to.
Use one public runtime factory to compose LlamaCppCandidateScoringAdapter and
ScoringJudgmentAdapter for Gemma 4 native-turn vision. The factory lives in
typevet.runtime and does not import typevet_evals.
Prerequisites¶
- llama.cpp router with a Gemma 4 multimodal model id and
--mmprojloaded. - Native template render from
POST /apply-template(no ChatML markers). TYPEVET_LLAMA__*variables set for your router (see Configuration).
Open a session¶
from typevet.adapters.inbound.settings import load_llama_settings
from typevet.runtime import open_gemma_native_vision_judgment
settings = load_llama_settings()
with open_gemma_native_vision_judgment(settings=settings) as session:
response = session.port.judge(state, questions, session.model)
Pass session.model to judge. The factory pins that id on
session.port; any other model argument raises JudgmentValidationError
before tokenization or scoring.
Defaults:
- Model id:
settings.multimodal_model(TYPEVET_LLAMA__MULTIMODAL_MODEL). - Timeout:
settings.timeout(TYPEVET_LLAMA__TIMEOUT, default 300s). - Template: must classify as
NATIVE_GEMMA4_TURNwhenrequire_gemma4=True.
Unsupported routers raise ValueError before the first judge call when:
/propsreports text-only modalities, or/apply-templateis not a supported native Gemma turn family.
Probe without holding a port¶
from typevet.runtime import probe_gemma_native_vision_support
meta = probe_gemma_native_vision_support(settings=load_llama_settings())
print(meta["served"], meta["vision"])
Lifecycle notes¶
- The context manager owns the scoring adapter lifecycle (
close()on exit). - Pass an existing
httpx.Clientonly in tests viahttp_client=. - Consumer live matrices may pass
tokenize_contentandscoring_port_wrapperhooks for dispatch ledgers without duplicating probe wiring. - For long-running services, prefer one session per request or explicit client ownership documented in your composition root.
See also Run the image-conditioned live smoke for lower-level router checks.