Skip to content

Runtime package

Kind: reference. This page is generated from the docstrings of typevet.runtime. The API index lists the other packages.

typevet.runtime

Thin orchestration facades over domain and ports (#148).

Examples:

from typevet.runtime import decide_categorical, judge_with_scoring
See Also

Attributes:

Name Type Description
decide_categorical callable

M1 categorical decision via scoring port.

judge_with_scoring callable

One-shot scoring-backed judgment helper.

ScoringJudgmentAdapter type

Sync JudgmentPort over CandidateScoringPort.

compose_scoring_prefix callable

Degraded ChatML scoring prefix composition.

GemmaNativeVisionSession type

Live session from open_gemma_native_vision_judgment.

open_gemma_native_vision_judgment callable

Gemma native vision factory context manager.

probe_gemma_native_vision_support callable

Router probe without a long-lived port.

VllmJudgmentSession type

Session from open_vllm_judgment.

open_vllm_judgment callable

vLLM chat-completions judgment context manager.

vllm_tokenize callable

vLLM /tokenize hook factory.

Classes

ScoringJudgmentAdapter

ScoringJudgmentAdapter(
    scoring_port: CandidateScoringPort,
    *,
    tokenize_content: Callable[[str], Sequence[int]],
    temperature: float = 1.0,
    served_template: ServedTemplateClass | None = None,
    pinned_model: str | None = None,
    framing: ModelFramingPort | None = None,
)

JudgmentPort implementation using injected candidate scoring.

Examples:

adapter = ScoringJudgmentAdapter(scoring_port, tokenize_content=tokenize)
adapter.judge("state", {"q": Noul()}, "model-id")

Attributes:

Name Type Description
_port CandidateScoringPort

Injected scorer; not closed by this adapter.

_tokenize Callable[[str], Sequence[int]]

Control-string tokenizer hook.

_temperature float

Softmax temperature forwarded to execute.

_served_template ServedTemplateClass | None

Served family for every prefix; None when unknown.

_pinned_model str | None

When set, judge rejects other model ids before tokenization or scoring IO.

_framing ModelFramingPort | None

When set, composes every prefix in place of the served-template choice.

framing and served_template are mutually exclusive: a framing composes every prefix, so a served family would not be read.

Raises:

Type Description
ValueError

Both framing and served_template are set.

Methods:
judge
judge(
    state: str | dict[str, Any] | list[Any],
    questions: Mapping[str, Question | Mapping[str, Any]],
    model: str,
    *,
    media: tuple[ImageInput, ...] | None = None,
) -> JudgmentResponse

Validate all questions, score sequentially, return typed answers.

Builds a scoring prefix whose field block maps each ordinal control string to the public answer label before calling the scorer. A native served family wraps every field prefix whether or not media is empty. When media is non-empty, every prefix carries one MEDIA_MARKER per image and every scoring request carries the same image tuple. Without a native family, text-only prefixes use ChatML. An injected framing composes every prefix instead.

Parameters:

Name Type Description Default
state str | dict[str, Any] | list[Any]

Content under evaluation (text or JSON-serializable value).

required
questions Mapping[str, Question | Mapping[str, Any]]

Named native questions.

required
model str

Backend model id forwarded to the scorer.

required
media tuple[ImageInput, ...] | None

Images to condition every scored field on, in order.

None

Returns:

Type Description
JudgmentResponse

JudgmentResponse with one typed answer per question id.

Raises:

Type Description
JudgmentValidationError

Invalid or mismatched model, wire shape, question payload, unsupported served template, media without a native Gemma 3 or Gemma 4 served template, or a framing prefix with the wrong media-marker count, before any scoring IO.

GemmaNativeVisionSession dataclass

GemmaNativeVisionSession(
    port: JudgmentPort,
    client: Client,
    model: str,
    served: ServedTemplateClass,
    capability: MediaCapability,
)

Live router session with a configured JudgmentPort.

Attributes:

Name Type Description
port JudgmentPort

Scoring-backed judgment for native media turns.

client Client

Shared HTTP client for the router base URL.

model str

Model id probed for template and vision capability.

served ServedTemplateClass

Classified native template family.

capability MediaCapability

Router media capability for model.

Examples:

from typevet.adapters.outbound.llama_cpp.gemma_native_vision_factory import (
    GemmaNativeVisionSession,
)

assert GemmaNativeVisionSession.__dataclass_fields__

VllmJudgmentSession dataclass

VllmJudgmentSession(port: JudgmentPort, client: Client, model: str)

Open vLLM session with a configured JudgmentPort.

Attributes:

Name Type Description
port JudgmentPort

Scoring-backed judgment pinned to model.

client Client

HTTP client the session sends requests on.

model str

Served model name that port accepts.

Examples:

from typevet.adapters.outbound.vllm.judgment_factory import (
    VllmJudgmentSession,
)

assert VllmJudgmentSession.__dataclass_fields__

Functions:

decide_categorical

decide_categorical(
    *,
    scoring_port: CandidateScoringPort,
    model: str,
    candidates: tuple[CandidateTokenSpec, ...],
    field: Decision | Mapping[str, Any],
    context: str = "",
    inject_prefix: bool = False,
    temperature: float = 1.0,
    **legacy: Any,
) -> CategoricalExecutionResult

Score categorical candidates and return the greedy choice plus distribution.

Validates the ask before any scoring IO. Compiles schema when given; accepts a pre-compiled Decision instead. Injected scoring_port values are never closed by this function (construct and own LlamaCppCandidateScoringAdapter at the call site when needed).

By default the library composes the scoring prefix from context plus rendered field instructions. Set inject_prefix=True and pass prefix= to score an exact caller-owned prefix (context does not alter it).

Parameters:

Name Type Description Default
scoring_port CandidateScoringPort

Offline fake or llama.cpp adapter implementing scoring.

required
model str

Backend model id forwarded to the scorer.

required
candidates tuple[CandidateTokenSpec, ...]

Single-token specs aligned with the decision choices.

required
field Decision | Mapping[str, Any]

Compiled Decision or JSON Schema with one categorical property.

required
context str

User or task text for the judgment (native path).

''
inject_prefix bool

When true, require prefix= and score it verbatim.

False
temperature float

Softmax temperature for the executor (default 1.0).

1.0

Other Parameters:

Name Type Description
legacy Any

prefix for inject mode; prompt as deprecated context alias. Other keys raise DecisionExecutionError.

Returns:

Type Description
CategoricalExecutionResult

Greedy value, full probability table, and raw logprobs from the executor.

Raises:

Type Description
DecisionExecutionError

Invalid ask, unsupported M1 shape, permutations other than 1, or unknown kwargs.

SchemaError

Schema cannot be compiled.

ScoringValidationError

Propagated from the scoring port or executor.

judge_with_scoring

judge_with_scoring(
    state: str | dict[str, Any] | list[Any],
    questions: Mapping[str, Question | Mapping[str, Any]],
    model: str,
    *,
    scoring_port: CandidateScoringPort,
    tokenize_content: Callable[[str], Sequence[int]],
    media: tuple[ImageInput, ...] | None = None,
    framing: ModelFramingPort | None = None,
    **settings: Unpack[ScoringAdapterSettings],
) -> JudgmentResponse

One-shot judgment via ScoringJudgmentAdapter.

Non-empty media needs served_template=NATIVE_GEMMA3_TURN or NATIVE_GEMMA4_TURN; an omitted, unknown or unsupported family fails closed before any scoring IO. A framing composes every prefix instead; None keeps the served-family path.

Parameters:

Name Type Description Default
state str | dict[str, Any] | list[Any]

Content under evaluation.

required
questions Mapping[str, Question | Mapping[str, Any]]

Named native questions.

required
model str

Backend model id.

required
scoring_port CandidateScoringPort

Injected candidate scorer.

required
tokenize_content Callable[[str], Sequence[int]]

Control-string tokenizer hook.

required
media tuple[ImageInput, ...] | None

Images to condition every scored field on, in order.

None
framing ModelFramingPort | None

Model framing that composes every prefix, or None.

None

Other Parameters:

Name Type Description
temperature float

Softmax temperature for execute.

served_template ServedTemplateClass | None

Served family for media prefixes; None means unknown.

Returns:

Type Description
JudgmentResponse

JudgmentResponse from a fresh adapter instance.

Raises:

Type Description
ValueError

Both framing and served_template are set.

open_gemma_native_vision_judgment

open_gemma_native_vision_judgment(
    *,
    settings: GemmaVisionSettings,
    model: str | None = None,
    require_gemma4: bool = True,
    n_vocab: int = DEFAULT_N_VOCAB,
    http_client: Client | None = None,
    tokenize_content: Callable[[str], Sequence[int]] | None = None,
    scoring_port_wrapper: Callable[[CandidateScoringPort], CandidateScoringPort]
    | None = None,
) -> Iterator[GemmaNativeVisionSession]

Open a judgment port for Gemma native-turn vision on a llama.cpp router.

Probes /apply-template and media capability before the first judge call. Unsupported template families or text-only models raise ValueError before scoring dispatch.

Parameters:

Name Type Description Default
settings GemmaVisionSettings

Router connection options from the composition root.

required
model str | None

Model id; defaults to settings.multimodal_model.

None
require_gemma4 bool

When true, require NATIVE_GEMMA4_TURN.

True
n_vocab int

n_probs budget for Gemma 4 class vocab scoring.

DEFAULT_N_VOCAB
http_client Client | None

Optional pre-built client (for tests); not closed on exit.

None
tokenize_content Callable[[str], Sequence[int]] | None

Optional tokenizer hook; defaults to router /tokenize.

None
scoring_port_wrapper Callable[[CandidateScoringPort], CandidateScoringPort] | None

Optional wrapper applied before JudgmentPort wiring.

None

Yields:

Type Description
GemmaNativeVisionSession

A session holding the configured JudgmentPort and probe metadata.

Raises:

Type Description
ValueError

When vision is unavailable or the template is unsupported.

probe_gemma_native_vision_support

probe_gemma_native_vision_support(
    *,
    settings: GemmaVisionSettings,
    model: str | None = None,
    require_gemma4: bool = True,
    http_client: Client | None = None,
) -> dict[str, Any]

Return probe metadata without constructing a long-lived judgment port.

Parameters:

Name Type Description Default
settings GemmaVisionSettings

Router connection options.

required
model str | None

Model id; defaults to settings.multimodal_model.

None
require_gemma4 bool

When true, require NATIVE_GEMMA4_TURN.

True
http_client Client | None

Optional pre-built client (for tests).

None

Returns:

Type Description
dict[str, Any]

Mapping with ok, model, served, and vision keys.

Raises:

Type Description
ValueError

When the router rejects the probe (same rules as open).

compose_scoring_prefix

compose_scoring_prefix(*, context: str, field_block: str) -> str

Compose degraded ChatML text ending at the assistant answer boundary.

Parameters:

Name Type Description Default
context str

Caller-owned user or task text for the judgment.

required
field_block str

Rendered field instructions from render_field_instructions.

required

Returns:

Type Description
str

Prefix string ending with CHATML_ASSISTANT_HEADER.

open_vllm_judgment

open_vllm_judgment(
    *,
    client: Client,
    model: str,
    base_url: str | None = None,
    tokenize_content: Callable[[str], Sequence[int]] | None = None,
    scoring_port_wrapper: Callable[[CandidateScoringPort], CandidateScoringPort]
    | None = None,
) -> Iterator[VllmJudgmentSession]

Open a judgment port for a vLLM chat-completions server.

The factory makes no HTTP call before the first judge call. The port rejects other model ids before tokenization or scoring IO. The caller owns client; the factory does not close it.

Parameters:

Name Type Description Default
client Client

HTTP client for the vLLM server, for example one that carries an Authorization header.

required
model str

Served model name; the port is pinned to it.

required
base_url str | None

Server root for scoring and /tokenize. Defaults to the client base_url.

None
tokenize_content Callable[[str], Sequence[int]] | None

Optional tokenizer hook; defaults to vllm_tokenize.

None
scoring_port_wrapper Callable[[CandidateScoringPort], CandidateScoringPort] | None

Optional wrapper applied before JudgmentPort wiring.

None

Yields:

Type Description
VllmJudgmentSession

A session holding the configured JudgmentPort.

Raises:

Type Description
ValueError

When no server root is known.

vllm_tokenize

vllm_tokenize(
    client: Client, model: str, *, base_url: str | None = None
) -> Callable[[str], tuple[int, ...]]

Return a hook that tokenizes text through vLLM /tokenize.

Parameters:

Name Type Description Default
client Client

HTTP client for the vLLM server.

required
model str

Served model name sent in each request.

required
base_url str | None

Server root. Defaults to the client base_url.

None

Returns:

Type Description
Callable[[str], tuple[int, ...]]

A function that POSTs ``{"model", "prompt", "add_special_tokens":

Callable[[str], tuple[int, ...]]

false}and returns the replytokens`` as a tuple.

Raises:

Type Description
ValueError

When no server root is known.