Skip to content

Configuration

Kind: reference. Environment variables for composition roots (CLI, MCP, live harness). Library callers pass explicit adapter constructor arguments instead.

Parent: #46. Log settings were added in #30.

Composition root vs library

Layer Reads environment Module
Composition root yes typevet.adapters.inbound.settings
Composition root yes typevet.adapters.inbound.backend_settings
Outbound adapter no typevet.adapters.outbound.llama_cpp
Diagnostics yes (stderr only) typevet.adapters.diagnostics.settings

Importing typevet does not read the environment. Call load_llama_settings or configure_from_environ at process startup in the inbound layer.

llama.cpp router

LlamaSettings holds connection options. load_llama_settings reads the mapping below. llama_cpp_adapter passes the values into LlamaCppGenerationAdapter without the adapter touching os.environ.

Environment name Field Type Default Notes
TYPEVET_LLAMA__BASE_URL base_url URL string http://127.0.0.1:8090 Trailing slash stripped
TYPEVET_LLAMA__TIMEOUT timeout float, seconds 300 Must be positive
TYPEVET_LLAMA__DEFAULT_MODEL default_model string or empty none Router model id for live pytest (any Gemma 4 GGUF alias); live skips when empty
TYPEVET_LLAMA__MULTIMODAL_MODEL multimodal_model string gemma-3-4b-it-q4km-mm Router model id for the image-conditioned live smoke; the id must declare image input
TYPEVET_LLAMA_URL base_url URL string (same) Legacy alias when nested name unset
TYPEVET_GEMMA_MODEL default_model string (same) Legacy alias when nested name unset

Nested names win when both nested and legacy names are set.

Future CLI hookup

There is no shipped CLI yet (pyproject.toml optional extra cli only adds Typer). When a CLI lands, it should:

  1. Call load_llama_settings() once at startup (and configure_from_environ() for stderr diagnostics).
  2. Build llama_cpp_adapter(settings) and pass the port into command handlers.
  3. Close the adapter on process exit.

Until then, scripts and live tests act as the composition root using the same helpers.

Backend selection

generation_adapter reads TYPEVET_BACKEND and builds one generation adapter.

Environment name Values Default Notes
TYPEVET_BACKEND llama_cpp, vllm llama_cpp Other values raise ValueError

llama_cpp builds llama_cpp_adapter(load_llama_settings()). vllm builds a VllmGenerationAdapter on the vllm_http_client client. Closing that adapter closes its client.

vLLM server

VllmSettings holds connection options. load_vllm_settings reads the mapping below.

Environment name Field Type Default Notes
TYPEVET_VLLM__BASE_URL base_url URL string none Required; trailing slash stripped
TYPEVET_VLLM__MODEL model string none Required; served model name
TYPEVET_VLLM__TIMEOUT timeout float, seconds 300 Must be positive
TYPEVET_VLLM__API_KEY api_key string or empty none Sent as Authorization: Bearer <key>
TYPEVET_VLLM__MAX_CONCURRENCY max_concurrency integer 1 Must be a positive integer; POST limit for one AsyncVllmGenerationAdapter
TYPEVET_VLLM__USER_AGENT user_agent string or empty none Sent as User-Agent only when set; otherwise the httpx default

The key does not appear in repr(VllmSettings). The sync and async clients are built the same way with or without a key. Only the Authorization header differs. So HTTPS_PROXY and the other proxy variables apply in both cases. The key must be ASCII. The adapters from generation_adapter and async_vllm_generation_adapter mask the key in errors. Each adapter checks each generation error, its attributes and its cause for the raw or JSON-escaped key. The attributes include strings inside a parsed payload, for example the payload of a SchemaValidationError. On a match, the adapter raises the same error type again. The new error shows *** for the key and has no cause or context. A server that echoes the Authorization header therefore cannot put the key into a BackendHttpError message. Successful results are not changed. Error messages name the variable, not its value. An invalid TYPEVET_VLLM__TIMEOUT or TYPEVET_VLLM__MAX_CONCURRENCY error has no cause, so a traceback does not show the value.

generation_adapter builds the sync adapter and does not read max_concurrency. async_vllm_generation_adapter builds an AsyncVllmGenerationAdapter from the TYPEVET_VLLM__* variables and does not read TYPEVET_BACKEND. Its httpx.AsyncClient has the same base URL, timeout, headers and proxy behaviour as the sync client, and max_concurrency sets its POST limit. Closing the adapter closes its client. The limit applies to one adapter only. Two adapters do not share it. With the default of 1, the adapter sends one request at a time. The adapter keeps one limit for each event loop. The httpx.AsyncClient of this adapter binds to the first event loop that uses it. The adapter records the first running loop that calls generate. A call on a different loop, for example from a second asyncio.run, raises RuntimeError("build one adapter per event loop") before any request. This error is not a GenerationError. Build one adapter for each event loop, for example inside each asyncio.run call. An AsyncVllmGenerationAdapter built with client=None does the same check. An adapter with an injected client does not.

vLLM live acceptance run

The opt-in test evals/tests/live/test_vllm_acceptance_live.py reads the variables above and these three. It skips unless TYPEVET_REQUIRE_LIVE is truthy. When it is truthy and a required variable is missing, the test fails before any network call.

Environment name Default Notes
TYPEVET_REQUIRE_LIVE none Set to 1 to run the paid run
TYPEVET_VLLM_RECEIPT none Required; the file must not exist and the nearest existing parent directory must be writable
TYPEVET_VLLM_POD_NOTES unknown Free text for the receipt, such as GPU, flags and Hugging Face revision; never put the key here

Diagnostic logging

See Diagnostic events for TYPEVET_LOG__FORMAT, TYPEVET_LOG__LEVEL, and TYPEVET_LOG__LOG_PROMPTS.