Runtime package¶
Kind: reference. This page is generated from the docstrings of
typevet.runtime. The API index lists the other packages.
typevet.runtime
¶
Thin orchestration facades over domain and ports (#148).
Examples:
See Also
- typevet.runtime.categorical: Categorical decision facade
- typevet.runtime.judgment: Scoring-backed judgment facade
Attributes:
| Name | Type | Description |
|---|---|---|
decide_categorical |
callable
|
M1 categorical decision via scoring port. |
judge_with_scoring |
callable
|
One-shot scoring-backed judgment helper. |
ScoringJudgmentAdapter |
type
|
Sync |
compose_scoring_prefix |
callable
|
Degraded ChatML scoring prefix composition. |
GemmaNativeVisionSession |
type
|
Live session from |
open_gemma_native_vision_judgment |
callable
|
Gemma native vision factory context manager. |
probe_gemma_native_vision_support |
callable
|
Router probe without a long-lived port. |
VllmJudgmentSession |
type
|
Session from |
open_vllm_judgment |
callable
|
vLLM chat-completions judgment context manager. |
vllm_tokenize |
callable
|
vLLM |
Classes¶
ScoringJudgmentAdapter
¶
ScoringJudgmentAdapter(
scoring_port: CandidateScoringPort,
*,
tokenize_content: Callable[[str], Sequence[int]],
temperature: float = 1.0,
served_template: ServedTemplateClass | None = None,
pinned_model: str | None = None,
framing: ModelFramingPort | None = None,
)
JudgmentPort implementation using injected candidate scoring.
Examples:
adapter = ScoringJudgmentAdapter(scoring_port, tokenize_content=tokenize)
adapter.judge("state", {"q": Noul()}, "model-id")
Attributes:
| Name | Type | Description |
|---|---|---|
_port |
CandidateScoringPort
|
Injected scorer; not closed by this adapter. |
_tokenize |
Callable[[str], Sequence[int]]
|
Control-string tokenizer hook. |
_temperature |
float
|
Softmax temperature forwarded to execute. |
_served_template |
ServedTemplateClass | None
|
Served family for
every prefix; |
_pinned_model |
str | None
|
When set, |
_framing |
ModelFramingPort | None
|
When set, composes every prefix in place of the served-template choice. |
framing and served_template are mutually exclusive: a framing
composes every prefix, so a served family would not be read.
Raises:
| Type | Description |
|---|---|
ValueError
|
Both |
Methods:¶
judge
¶
judge(
state: str | dict[str, Any] | list[Any],
questions: Mapping[str, Question | Mapping[str, Any]],
model: str,
*,
media: tuple[ImageInput, ...] | None = None,
) -> JudgmentResponse
Validate all questions, score sequentially, return typed answers.
Builds a scoring prefix whose field block maps each ordinal control
string to the public answer label before calling the scorer. A native
served family wraps every field prefix whether or not media is
empty. When media is non-empty, every prefix carries one
MEDIA_MARKER per image and every scoring request carries the same
image tuple. Without a native family, text-only prefixes use ChatML.
An injected framing composes every prefix instead.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
str | dict[str, Any] | list[Any]
|
Content under evaluation (text or JSON-serializable value). |
required |
questions
|
Mapping[str, Question | Mapping[str, Any]]
|
Named native questions. |
required |
model
|
str
|
Backend model id forwarded to the scorer. |
required |
media
|
tuple[ImageInput, ...] | None
|
Images to condition every scored field on, in order. |
None
|
Returns:
| Type | Description |
|---|---|
JudgmentResponse
|
|
Raises:
| Type | Description |
|---|---|
JudgmentValidationError
|
Invalid or mismatched model, wire shape, question payload, unsupported served template, media without a native Gemma 3 or Gemma 4 served template, or a framing prefix with the wrong media-marker count, before any scoring IO. |
GemmaNativeVisionSession
dataclass
¶
GemmaNativeVisionSession(
port: JudgmentPort,
client: Client,
model: str,
served: ServedTemplateClass,
capability: MediaCapability,
)
Live router session with a configured JudgmentPort.
Attributes:
| Name | Type | Description |
|---|---|---|
port |
JudgmentPort
|
Scoring-backed judgment for native media turns. |
client |
Client
|
Shared HTTP client for the router base URL. |
model |
str
|
Model id probed for template and vision capability. |
served |
ServedTemplateClass
|
Classified native template family. |
capability |
MediaCapability
|
Router media capability for |
Examples:
from typevet.adapters.outbound.llama_cpp.gemma_native_vision_factory import (
GemmaNativeVisionSession,
)
assert GemmaNativeVisionSession.__dataclass_fields__
VllmJudgmentSession
dataclass
¶
VllmJudgmentSession(port: JudgmentPort, client: Client, model: str)
Open vLLM session with a configured JudgmentPort.
Attributes:
| Name | Type | Description |
|---|---|---|
port |
JudgmentPort
|
Scoring-backed judgment pinned to |
client |
Client
|
HTTP client the session sends requests on. |
model |
str
|
Served model name that |
Examples:
from typevet.adapters.outbound.vllm.judgment_factory import (
VllmJudgmentSession,
)
assert VllmJudgmentSession.__dataclass_fields__
Functions:¶
decide_categorical
¶
decide_categorical(
*,
scoring_port: CandidateScoringPort,
model: str,
candidates: tuple[CandidateTokenSpec, ...],
field: Decision | Mapping[str, Any],
context: str = "",
inject_prefix: bool = False,
temperature: float = 1.0,
**legacy: Any,
) -> CategoricalExecutionResult
Score categorical candidates and return the greedy choice plus distribution.
Validates the ask before any scoring IO. Compiles schema when given;
accepts a pre-compiled Decision instead. Injected scoring_port values
are never closed by this function (construct and own
LlamaCppCandidateScoringAdapter at the call site when needed).
By default the library composes the scoring prefix from context plus
rendered field instructions. Set inject_prefix=True and pass prefix=
to score an exact caller-owned prefix (context does not alter it).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scoring_port
|
CandidateScoringPort
|
Offline fake or llama.cpp adapter implementing scoring. |
required |
model
|
str
|
Backend model id forwarded to the scorer. |
required |
candidates
|
tuple[CandidateTokenSpec, ...]
|
Single-token specs aligned with the decision choices. |
required |
field
|
Decision | Mapping[str, Any]
|
Compiled |
required |
context
|
str
|
User or task text for the judgment (native path). |
''
|
inject_prefix
|
bool
|
When true, require |
False
|
temperature
|
float
|
Softmax temperature for the executor (default 1.0). |
1.0
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
legacy |
Any
|
|
Returns:
| Type | Description |
|---|---|
CategoricalExecutionResult
|
Greedy value, full probability table, and raw logprobs from the executor. |
Raises:
| Type | Description |
|---|---|
DecisionExecutionError
|
Invalid ask, unsupported M1 shape, permutations
other than |
SchemaError
|
Schema cannot be compiled. |
ScoringValidationError
|
Propagated from the scoring port or executor. |
judge_with_scoring
¶
judge_with_scoring(
state: str | dict[str, Any] | list[Any],
questions: Mapping[str, Question | Mapping[str, Any]],
model: str,
*,
scoring_port: CandidateScoringPort,
tokenize_content: Callable[[str], Sequence[int]],
media: tuple[ImageInput, ...] | None = None,
framing: ModelFramingPort | None = None,
**settings: Unpack[ScoringAdapterSettings],
) -> JudgmentResponse
One-shot judgment via ScoringJudgmentAdapter.
Non-empty media needs served_template=NATIVE_GEMMA3_TURN or
NATIVE_GEMMA4_TURN; an omitted, unknown or unsupported family fails
closed before any scoring IO. A framing composes every prefix
instead; None keeps the served-family path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
str | dict[str, Any] | list[Any]
|
Content under evaluation. |
required |
questions
|
Mapping[str, Question | Mapping[str, Any]]
|
Named native questions. |
required |
model
|
str
|
Backend model id. |
required |
scoring_port
|
CandidateScoringPort
|
Injected candidate scorer. |
required |
tokenize_content
|
Callable[[str], Sequence[int]]
|
Control-string tokenizer hook. |
required |
media
|
tuple[ImageInput, ...] | None
|
Images to condition every scored field on, in order. |
None
|
framing
|
ModelFramingPort | None
|
Model framing that composes every prefix, or |
None
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
temperature |
float
|
Softmax temperature for execute. |
served_template |
ServedTemplateClass | None
|
Served family for media
prefixes; |
Returns:
| Type | Description |
|---|---|
JudgmentResponse
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
Both |
open_gemma_native_vision_judgment
¶
open_gemma_native_vision_judgment(
*,
settings: GemmaVisionSettings,
model: str | None = None,
require_gemma4: bool = True,
n_vocab: int = DEFAULT_N_VOCAB,
http_client: Client | None = None,
tokenize_content: Callable[[str], Sequence[int]] | None = None,
scoring_port_wrapper: Callable[[CandidateScoringPort], CandidateScoringPort]
| None = None,
) -> Iterator[GemmaNativeVisionSession]
Open a judgment port for Gemma native-turn vision on a llama.cpp router.
Probes /apply-template and media capability before the first judge
call. Unsupported template families or text-only models raise ValueError
before scoring dispatch.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
GemmaVisionSettings
|
Router connection options from the composition root. |
required |
model
|
str | None
|
Model id; defaults to |
None
|
require_gemma4
|
bool
|
When true, require |
True
|
n_vocab
|
int
|
|
DEFAULT_N_VOCAB
|
http_client
|
Client | None
|
Optional pre-built client (for tests); not closed on exit. |
None
|
tokenize_content
|
Callable[[str], Sequence[int]] | None
|
Optional tokenizer hook; defaults to router |
None
|
scoring_port_wrapper
|
Callable[[CandidateScoringPort], CandidateScoringPort] | None
|
Optional wrapper applied before |
None
|
Yields:
| Type | Description |
|---|---|
GemmaNativeVisionSession
|
A session holding the configured |
Raises:
| Type | Description |
|---|---|
ValueError
|
When vision is unavailable or the template is unsupported. |
probe_gemma_native_vision_support
¶
probe_gemma_native_vision_support(
*,
settings: GemmaVisionSettings,
model: str | None = None,
require_gemma4: bool = True,
http_client: Client | None = None,
) -> dict[str, Any]
Return probe metadata without constructing a long-lived judgment port.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
GemmaVisionSettings
|
Router connection options. |
required |
model
|
str | None
|
Model id; defaults to |
None
|
require_gemma4
|
bool
|
When true, require |
True
|
http_client
|
Client | None
|
Optional pre-built client (for tests). |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Mapping with |
Raises:
| Type | Description |
|---|---|
ValueError
|
When the router rejects the probe (same rules as open). |
compose_scoring_prefix
¶
Compose degraded ChatML text ending at the assistant answer boundary.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
context
|
str
|
Caller-owned user or task text for the judgment. |
required |
field_block
|
str
|
Rendered field instructions from |
required |
Returns:
| Type | Description |
|---|---|
str
|
Prefix string ending with |
open_vllm_judgment
¶
open_vllm_judgment(
*,
client: Client,
model: str,
base_url: str | None = None,
tokenize_content: Callable[[str], Sequence[int]] | None = None,
scoring_port_wrapper: Callable[[CandidateScoringPort], CandidateScoringPort]
| None = None,
) -> Iterator[VllmJudgmentSession]
Open a judgment port for a vLLM chat-completions server.
The factory makes no HTTP call before the first judge call. The port
rejects other model ids before tokenization or scoring IO. The caller
owns client; the factory does not close it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Client
|
HTTP client for the vLLM server, for example one that carries
an |
required |
model
|
str
|
Served model name; the port is pinned to it. |
required |
base_url
|
str | None
|
Server root for scoring and |
None
|
tokenize_content
|
Callable[[str], Sequence[int]] | None
|
Optional tokenizer hook; defaults to |
None
|
scoring_port_wrapper
|
Callable[[CandidateScoringPort], CandidateScoringPort] | None
|
Optional wrapper applied before |
None
|
Yields:
| Type | Description |
|---|---|
VllmJudgmentSession
|
A session holding the configured |
Raises:
| Type | Description |
|---|---|
ValueError
|
When no server root is known. |
vllm_tokenize
¶
vllm_tokenize(
client: Client, model: str, *, base_url: str | None = None
) -> Callable[[str], tuple[int, ...]]
Return a hook that tokenizes text through vLLM /tokenize.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Client
|
HTTP client for the vLLM server. |
required |
model
|
str
|
Served model name sent in each request. |
required |
base_url
|
str | None
|
Server root. Defaults to the client |
None
|
Returns:
| Type | Description |
|---|---|
Callable[[str], tuple[int, ...]]
|
A function that POSTs ``{"model", "prompt", "add_special_tokens": |
Callable[[str], tuple[int, ...]]
|
false} |
Raises:
| Type | Description |
|---|---|
ValueError
|
When no server root is known. |