Serve typevet on vLLM¶
Kind: how-to.
Status: draft.
Use this page to run typevet generation and judgment against a vLLM OpenAI-compatible server. The steps follow the one tested pin from the #170 live receipt.
Supported versions¶
typevet supports one tested pin. Other versions are not tested.
| Item | Tested value |
|---|---|
| Server image | vllm/vllm-openai:v0.30.0 (/version returns 0.30.0) |
| Model | google/gemma-4-31B-it |
| Model revision | 842da3794eaa0b77d5f08bae87a17459d91ff475 |
| Weights | BF16, no quantization |
| GPU | One H100 with 80 GB of memory |
| Served model name | gemma-4-31b-it |
The tested server flags:
--max-model-len 8192
--gpu-memory-utilization 0.95
--max-num-seqs 4
--limit-mm-per-prompt {"image":2}
--logprobs-mode raw_logprobs
--served-model-name gemma-4-31b-it
The server requires an API key. The server reads it from VLLM_API_KEY.
Start the server¶
-
Optional: get a Hugging Face token. The Gemma 4 weights are not gated. The tested run supplied a token.
-
Start the stock image with about 150 GB of container disk. This example uses
docker run. The tested run used the image's default entrypoint with the same server arguments. The--gpus,--ipcand-poptions are standard Docker options that the tested run did not use.
export VLLM_API_KEY='<your-key>'
export HF_TOKEN='<your-hugging-face-token>'
docker run --gpus all --ipc=host -p 8000:8000 \
-e VLLM_API_KEY -e HF_TOKEN \
vllm/vllm-openai:v0.30.0 \
--model google/gemma-4-31B-it \
--revision 842da3794eaa0b77d5f08bae87a17459d91ff475 \
--max-model-len 8192 \
--gpu-memory-utilization 0.95 \
--max-num-seqs 4 \
--limit-mm-per-prompt '{"image":2}' \
--logprobs-mode raw_logprobs \
--served-model-name gemma-4-31b-it
- Wait for the model to load. In the tested run,
/v1/modelsanswered about 7 minutes after the server started.
Check the server¶
- Check the server version:
Expect {"version":"0.30.0"}.
- Check the served model name:
curl -s -H "Authorization: Bearer $VLLM_API_KEY" \
http://127.0.0.1:8000/v1/models | jq '.data[].id'
Expect "gemma-4-31b-it".
Configure typevet¶
Set the backend and the server settings in the environment:
export TYPEVET_BACKEND=vllm
export TYPEVET_VLLM__BASE_URL=http://127.0.0.1:8000
export TYPEVET_VLLM__MODEL=gemma-4-31b-it
export TYPEVET_VLLM__API_KEY="$VLLM_API_KEY"
export TYPEVET_VLLM__TIMEOUT=300
Configuration lists each variable, its default and its validation rule. These rules apply most often:
TYPEVET_BACKENDacceptsllama_cpporvllm. An empty value selectsllama_cpp.TYPEVET_VLLM__BASE_URLandTYPEVET_VLLM__MODELare required.TYPEVET_VLLM__MODELmust equal the--served-model-namevalue.TYPEVET_VLLM__TIMEOUTis in seconds. The default is 300 seconds.TYPEVET_VLLM__API_KEYmust be ASCII. An empty value sends no key.TYPEVET_VLLM__MAX_CONCURRENCYsets the POST limit for oneAsyncVllmGenerationAdapter. The default is 1.TYPEVET_VLLM__USER_AGENTreplaces the httpx defaultUser-Agentheader.
generation_adapter builds the sync adapter. That adapter does not read
TYPEVET_VLLM__MAX_CONCURRENCY. To send parallel requests, use
async_vllm_generation_adapter. It builds an AsyncVllmGenerationAdapter
with the same variables and uses TYPEVET_VLLM__MAX_CONCURRENCY as its limit.
Its HTTP client binds to the first event loop that uses it. Build one adapter
for each event loop, for example inside each asyncio.run call. A call on a
different loop raises RuntimeError("build one adapter per event loop")
before any request. A retry with the same adapter fails again.
Some proxies block requests that carry a library User-Agent header. In that
case, set TYPEVET_VLLM__USER_AGENT to a value that the proxy accepts. The
tested run used curl/8.9.1.
Key and network behaviour¶
- The client sends the key as
Authorization: Bearer <key>. repr(VllmSettings)does not show the key.- Error messages name the variable, never its value.
- The adapters from
generation_adapterandasync_vllm_generation_adaptermask the key in errors. The port fromopen_judgmentalso masks it. Each error shows***in place of the raw or JSON-escaped key. This includes a parsed payload. The error has no cause or context. - An
AsyncVllmGenerationAdapterthat you build yourself does not mask the key. - typevet sets no TLS or proxy options. The httpx defaults apply, so the
client verifies certificates and reads
HTTPS_PROXYand the other proxy variables. - The client is the same with or without a key. Only the
Authorizationheader differs.
Use an HTTPS URL or a local address. A plain HTTP URL sends the key without encryption.
Run one typed generation call¶
from typevet.adapters.inbound import generate, generation_adapter, load_vllm_settings
schema = {
"type": "object",
"properties": {"sentiment": {"type": "string", "enum": ["pos", "neg"]}},
"required": ["sentiment"],
"additionalProperties": False,
}
settings = load_vllm_settings()
with generation_adapter() as port:
result = generate(
port,
prompt="Classify the review: The blender works well.",
schema=schema,
model=settings.model,
)
print(result.value)
The result value matches the schema. The adapter raises a GenerationError
subclass when the server rejects the request.
Run one judgment call¶
from typevet.adapters.inbound.backend_settings import open_judgment
from typevet.domain import Choice
questions = {
"sentiment": Choice(
criteria={"pos": "Positive review", "neg": "Negative review"},
instructions="Classify the review.",
),
}
with open_judgment() as session:
response = session.port.judge(
"The blender broke after one day.",
questions,
session.model,
)
answer = response.choices["sentiment"]
print(answer.choice, answer.probabilities)
Pass session.model to judge. The session pins that name from
TYPEVET_VLLM__MODEL.
Run the live acceptance test¶
The live acceptance test runs five sets once and writes one JSON receipt.
Each run costs GPU time. The test skips unless TYPEVET_REQUIRE_LIVE is set.
TYPEVET_REQUIRE_LIVE=1 \
TYPEVET_BACKEND=vllm \
TYPEVET_VLLM__BASE_URL=http://127.0.0.1:8000 \
TYPEVET_VLLM__MODEL=gemma-4-31b-it \
TYPEVET_VLLM__API_KEY="$VLLM_API_KEY" \
TYPEVET_VLLM_RECEIPT=scratchpad/vllm/new-receipt.json \
TYPEVET_VLLM_POD_NOTES='H100 80 GB, vLLM v0.30.0, revision 842da37' \
uv run pytest evals/tests/live/test_vllm_acceptance_live.py -m live -q -s
TYPEVET_VLLM_RECEIPTmust name a new file. The test fails when the file already exists.- The nearest existing parent directory must be writable.
- Do not put the key in
TYPEVET_VLLM_POD_NOTES. - A missing or invalid variable fails the test before any network call.
The test prints the receipt path and its sha256 digest.
Limits¶
- The evidence is one pin and one run.
- Other models are not tested.
- Other vLLM versions are not tested.
- Quantized weights are not tested.
- Mixed scoring batches are not tested. See vllm-project/vllm#51789.
- The
/metricsKV-cache metric names are not tested. The tested run did not find a KV-cache usage metric. - The results do not compare backends. The llama.cpp baseline used a different quantization.