Error reference¶
Kind: reference. This page maps domain failures, adapter behavior and retry
boundaries for typed generation. Sister shape: judgevet
docs/reference/errors.md.
Status: current code. A future error-hierarchy ADR (#43) may refine names or grouping; until then, treat the types below as authoritative.
Parent: #29.
Domain vs adapter errors¶
Domain errors describe typed generation and schema compilation inside
typevet.domain. They do not perform HTTP. Callers use them to classify bad
input, unsupported schema shapes and invalid model output.
Adapter errors are domain exception types raised from outbound adapters
(FakeGenerationAdapter, the llama.cpp adapters and the vLLM adapters).
Adapters translate httpx transport failures into TransportError, llama.cpp
and vLLM HTTP error statuses into BackendHttpError, other parse or shape
failures into GenerationError, and run jsonschema validation into
SchemaValidationError. The candidate scoring adapters raise
ScoringValidationError and ScoringUnsupportedCapabilityError, which
subclass GenerationError through ScoringError. They do not define a
parallel adapter-specific hierarchy.
Schema compilation uses SchemaError (ValueError subclass) from
typevet.domain when a JSON Schema mapping cannot be compiled. That happens
before any generation port call. It is not a GenerationError and is never
raised from GenerationPort.generate.
The generation port documents failure
modes every implementation may raise: TransportError, BackendHttpError,
other GenerationError parse failures, and fail-fast SchemaValidationError.
Generation failures¶
The package root typevet exports only the first four types.
typevet.domain.errors exports all of them:
| Type | Parent | Meaning |
|---|---|---|
GenerationError |
Exception |
A generation call failed before a valid GenerationResult existed (parse, shape, or empty fake). |
TransportError |
GenerationError |
The HTTP client failed before a usable response (status_code and body_snippet are None). |
BackendHttpError |
GenerationError |
llama.cpp or vLLM returned HTTP status 400 or above; carries status_code and truncated body_snippet (500 chars max). |
SchemaValidationError |
GenerationError |
Parsed output failed JSON Schema validation. Optional payload holds the rejected value. |
GenerationUnsupportedCapabilityError |
GenerationError |
The adapter cannot send a part of the request, such as images. Raised before any HTTP call. Import from typevet.domain or typevet.domain.errors. |
ScoringError |
GenerationError |
A candidate scoring call failed before a valid result existed. |
ScoringValidationError |
ScoringError |
Scores failed the coverage or finiteness rules. |
ScoringUnsupportedCapabilityError |
ScoringError |
The backend cannot honor the requested score stage or capability. |
SchemaValidationError stores the message in standard exception args. When
set, payload is the parsed object or mapping that failed validation (see
domain/errors.py).
Request validation (not generation errors)¶
GenerationRequest rejects bad asks in __post_init__ with ordinary Python
errors, not GenerationError:
| Condition | Type |
|---|---|
Blank prompt or model |
ValueError |
schema is not a mapping |
TypeError |
schema.type is set and not "object" |
ValueError |
A media item is not an ImageInput |
TypeError |
Count of MEDIA_MARKER in prompt differs from len(media) |
ValueError |
schema fails the JSON Schema Draft 2020-12 meta-schema, or json.dumps cannot encode it (adapter check) |
ValueError (schema is not a valid JSON Schema:) |
The last row comes from check_request_schema in
chat_completion.py.
The llama.cpp and vLLM generation adapters, sync and async, call it before the
request, so no request is sent. The check runs once for each distinct schema.
Fix the request; do not treat these as retryable generation failures.
Event loop misuse (not a generation error)¶
An async vLLM adapter that owns its HTTP client serves only the first event
loop that calls generate. This applies to the adapter from
async_vllm_generation_adapter and to AsyncVllmGenerationAdapter built with
client=None. A call on a different loop raises a plain RuntimeError
before any request:
| Condition | Type | Message |
|---|---|---|
generate runs on a different event loop from the first call |
RuntimeError |
build one adapter per event loop |
This RuntimeError is not a GenerationError. A caller that catches only
GenerationError does not catch it. The failure is permanent for that
adapter, so do not retry it. Build one adapter for each event loop. An adapter
with an injected client does not do this check.
Schema compilation failures¶
Import SchemaError from typevet.domain (not the package root):
| Type | Parent | Meaning |
|---|---|---|
SchemaError |
ValueError |
The JSON Schema mapping is outside the supported compile subset (compile_json_schema, Decision, dependency layers). |
compile_json_schema and helpers in
decision_compile.py raise
SchemaError with a human-readable message for unsupported keywords, bad
enums, dependency cycles and similar compile-time rules.
For some unsupported field types without a finite enum, compilation raises
NotImplementedError instead of SchemaError. That signals a missing feature,
not a malformed document the caller can correct by editing one field.
llama.cpp adapter mapping¶
LlamaCppGenerationAdapter
POSTs to v1/chat/completions with response_format json_schema. HTTP status
400 and above are treated as adapter failure (constant _HTTP_ERROR_STATUS).
| Condition | Raised type | Typical message prefix |
|---|---|---|
Request has media (images) |
GenerationUnsupportedCapabilityError |
llama.cpp generation adapters do not send images (no HTTP call) |
httpx.HTTPError on POST |
TransportError |
llama.cpp request failed: |
| HTTP status ≥ 400 | BackendHttpError |
llama.cpp HTTP {status}: (body_snippet truncated to 500 chars) |
| Response body is not JSON | GenerationError |
llama.cpp returned non-JSON HTTP body |
Missing or empty choices[0].message.content |
GenerationError |
llama.cpp response missing… or empty message content |
| Message content is not valid JSON | GenerationError |
model content was not valid JSON |
| Parsed JSON root is not an object | SchemaValidationError |
model JSON root must be an object (payload set) |
Parsed object holds a non-finite number (NaN, Infinity) |
SchemaValidationError |
structured output contains a non-finite number (payload set) |
jsonschema.validate fails on parsed object |
SchemaValidationError |
output failed schema: (payload set) |
This table describes local mapping only. It does not assert which HTTP statuses a given llama.cpp build returns for every failure mode. See Run Gemma 4 on llama.cpp for the live path.
vLLM adapter mapping¶
VllmGenerationAdapter
and AsyncVllmGenerationAdapter POST to v1/chat/completions with
structured_outputs json. Each call makes one POST and no retry.
vllm/http_mapping.py
maps the HTTP outcome. HTTP status 400 and above is an adapter failure
(HTTP_ERROR_STATUS in adapters/outbound/http_errors.py).
| Condition | Raised type | Typical message prefix |
|---|---|---|
Request schema is not a valid JSON Schema |
ValueError |
schema is not a valid JSON Schema: (no HTTP call) |
httpx.HTTPError on POST, including an httpx timeout |
TransportError |
vLLM request failed: |
| HTTP status ≥ 400 | BackendHttpError |
vLLM HTTP {status}: (body_snippet truncated to 500 chars) |
| Response body is not JSON | GenerationError |
vLLM returned non-JSON HTTP body |
Missing or empty choices[0].message.content |
GenerationError |
vLLM response missing… or vLLM returned empty message content |
| Message content is not valid JSON | GenerationError |
model content was not valid JSON |
| Parsed JSON root is not an object | SchemaValidationError |
model JSON root must be an object (payload set) |
Parsed object holds a non-finite number (NaN, Infinity) |
SchemaValidationError |
structured output contains a non-finite number (payload set) |
jsonschema.validate fails on parsed object |
SchemaValidationError |
output failed schema: (payload set) |
| Async adapter with an owned client runs on a second event loop | RuntimeError |
build one adapter per event loop (no HTTP call) |
A generation request with images is not refused. The adapter sends each image
as an image_url block.
VllmCandidateScoringAdapter
uses the same HTTP mapping for transport errors, status 400 and above, and a
body that is not JSON. It then reads the logprobs:
| Condition | Raised type | Typical message prefix |
|---|---|---|
More than 128 candidates, a stage other than PRE_SAMPLING, or a candidate with more than one token id |
ScoringUnsupportedCapabilityError |
vLLM accepts at most… or VllmCandidateScoringAdapter supports… (no HTTP call) |
| Response root is not an object | GenerationError |
vLLM chat completions response root must be an object |
Missing choices[0].logprobs.content[0].top_logprobs |
GenerationError |
vLLM response missing choices[0].logprobs.content[0].top_logprobs |
top_logprobs is not a list |
GenerationError |
vLLM top_logprobs must be a list |
An entry has no token or logprob, or its logprob is not a real number |
GenerationError |
vLLM top_logprobs entry must be… or vLLM top_logprobs logprob … is not a real number |
An entry token is not in token_id:N form |
GenerationError |
vLLM top_logprobs token … is not in token_id:N form |
| Two entries have the same token id | ScoringValidationError |
vLLM top_logprobs holds a duplicate entry |
top_logprobs has no score for a requested candidate |
ScoringValidationError |
missing scores for requested candidates: |
A candidate logprob is not finite, or is positive above 1e-6 |
ScoringValidationError |
non-finite logprob for candidate… or positive logprob for candidate… |
The /tokenize call of the vLLM judgment factory uses the same HTTP mapping.
A reply without a list of integer tokens raises GenerationError
(vLLM /tokenize response must hold a list of integer tokens).
This table describes local mapping only. It does not assert which HTTP statuses a given vLLM build returns for every failure mode.
Fake adapter mapping¶
FakeGenerationAdapter validates
a fixed or callable mapping with the same jsonschema path as llama.cpp:
| Condition | Raised type |
|---|---|
Constructor given no value, responder, or fail |
ValueError |
fail= configured |
Raises that exception as-is (tests use GenerationError) |
| No value source at generate time | GenerationError |
| Value fails schema | SchemaValidationError (fake output failed schema:) |
Retry and exception boundaries¶
typevet does not expose a retryable flag or a built-in retry loop on
generation adapters. Each generate call performs at most one HTTP round trip
(llama.cpp or vLLM) or one validation pass (fake).
Use these boundaries when a caller adds retries:
| Failure | Retry at generation layer? | Notes |
|---|---|---|
SchemaValidationError |
No (fail-fast) | Same prompt and schema may repeat the same invalid output; fix schema, prompt, or model. Adapter docstring: fail-fast on schema mismatch. |
TransportError, BackendHttpError (5xx) |
Optional caller policy | Not implemented in-repo; a supervisor may retry with backoff outside the adapter. |
BackendHttpError (4xx) |
Usually no | Router config, model id, or request the backend rejects. |
GenerationError (bad JSON shape on 2xx) |
Usually no | Non-recoverable response shape from the model or router. |
GenerationUnsupportedCapabilityError |
No | Use an adapter that supports the request, or remove the images. |
SchemaError, NotImplementedError |
No | Fix or narrow the schema before calling generation. |
GenerationRequest ValueError / TypeError |
No | Fix the request object. |
RuntimeError (build one adapter per event loop) |
No | Build one adapter for each event loop. |
except GenerationError catches SchemaValidationError because it subclasses
GenerationError. Use except SchemaValidationError when validation failures
need distinct handling (for example logging payload).
Each backend has its own mapping module:
llama_cpp/http_mapping.py
and
vllm/http_mapping.py.
Both use the status and snippet limits in
http_errors.py.
Adapters map httpx.HTTPError to TransportError and HTTP status ≥ 400 to
BackendHttpError. Other library or application errors propagate unless the
caller handles them.
The following example runs offline and prints nothing when its assertions pass. It demonstrates types only; it makes no network call.
from typevet import BackendHttpError, GenerationError, SchemaValidationError
from typevet.domain import SchemaError, compile_json_schema
assert issubclass(SchemaValidationError, GenerationError)
assert issubclass(BackendHttpError, GenerationError)
err = SchemaValidationError("missing key", payload={"x": 1})
assert err.payload == {"x": 1}
http_err = BackendHttpError(
"llama.cpp HTTP 500: oops", status_code=500, body_snippet="oops"
)
assert http_err.status_code == 500
assert http_err.body_snippet == "oops"
try:
compile_json_schema({"type": "array"})
except SchemaError:
pass
else:
raise AssertionError("expected SchemaError")
assert not issubclass(SchemaError, GenerationError)
Public surface¶
| Symbol | Import from |
|---|---|
GenerationError, TransportError, BackendHttpError, SchemaValidationError |
typevet or typevet.domain.errors |
SchemaError, compile_json_schema, Decision |
typevet.domain |
GenerationPort |
typevet.ports.generation |
ScoringError, ScoringValidationError, ScoringUnsupportedCapabilityError |
typevet.domain or typevet.domain.errors |
LlamaCppGenerationAdapter, FakeGenerationAdapter |
typevet.adapters.outbound |
VllmGenerationAdapter, AsyncVllmGenerationAdapter, VllmCandidateScoringAdapter |
typevet.adapters.outbound |
Do not document exception types that are not raised by the current tree. When #43 lands, revise this page to match the accepted hierarchy.