Use a vLLM server behind an API gateway¶
Kind: how-to.
Status: draft.
Use this page when an API gateway is in front of your vLLM server. Apigee, Kong and Azure API Management are examples. The gateway can use a path prefix, its own key header and extra headers. typevet sends them on each vLLM call. Issue #331 holds the contract.
No live gateway call proves these steps. Contract tests on httpx.MockTransport
prove them.
Before you start¶
- Start a vLLM server. See Serve typevet on vLLM.
- Get the gateway URL, the key header name and each required extra header.
- Use an HTTPS gateway URL.
Set the gateway variables¶
-
Set the backend and the gateway URL with its path prefix.
-
Set the gateway key.
-
Set the header that carries the key. This example uses
X-API-Key. -
Set the scheme before the key. An empty value sends the key with no scheme.
-
Set the extra headers as one JSON object of string values.
-
Optional: set a request-id header. typevet then sends a new UUID4 hex value in it on each request.
Check what typevet sends¶
With the variables above, each vLLM call goes to the gateway prefix:
| Call | URL |
|---|---|
| Generation, sync and async | https://gw.example.com/vllm/v1/chat/completions |
| Judgment scoring | https://gw.example.com/vllm/v1/chat/completions |
| Judgment tokenization | https://gw.example.com/vllm/tokenize |
A trailing slash on TYPEVET_VLLM__BASE_URL gives the same URLs.
Each request carries these headers:
X-API-Key: <gateway key>, and noAuthorizationheader.X-Tenant: acme.X-Request-Id: <32 hex characters>, when you set step 6.
These entry points build their clients from the same settings:
generation_adapter, async_vllm_generation_adapter and open_judgment.
The judgevet bridge and the served-template probe use open_judgment.
Thus they send the same headers.
Set no gateway variable to keep the direct behaviour.
typevet then sends Authorization: Bearer <key> as before.
Header rules¶
typevet checks the headers when it builds VllmSettings.
A failed check raises ValueError.
The message names the variable and never shows a value.
| Rule | Limit |
|---|---|
| Header name | An HTTP token of at most 128 bytes |
| Header value | ASCII characters 0x20 to 0x7E only, at most 2,048 bytes |
| Number of extra headers | At most 32 |
| Names and values together | At most 8,192 bytes |
typevet refuses these names in TYPEVET_VLLM__HEADERS, in any letter case:
- The hop-by-hop headers:
Connection,Keep-Alive,Proxy-Authenticate,Proxy-Authorization,TE,Trailer,Transfer-EncodingandUpgrade. HostandContent-Length.Content-Type,Accept,Accept-EncodingandUser-Agent. typevet sets them. OnlyTYPEVET_VLLM__USER_AGENTchanges the agent string.- The header in
TYPEVET_VLLM__AUTH_HEADER. Set the key inTYPEVET_VLLM__API_KEYinstead.
typevet sends each value as written.
It does not expand $NAME and does not run !command.
Handle gateway errors¶
- A gateway 429 raises
BackendHttpErrorwithstatus_code429. Your code decides whether to retry. typevet adds no retry. - Read
retry_after_secondsfor the wait thatRetry-Afterasks for. It isNonewhen the gateway sends no valid value. - Read
rate_limitfor thex-ratelimit-*andratelimit-*headers. It never holds the auth header or aTYPEVET_VLLM__HEADERSname. - An HTML error body from the gateway is not in the error.
body_snippetis empty. - typevet follows no redirect.
A 3xx status raises
BackendHttpErrorwith thatstatus_code. The error does not show theLocationheader. - Each error from these entry points masks the key and each value in
TYPEVET_VLLM__HEADERSwith***. The masked error has no cause or context.
Security lists the limits of this masking.
Retry hints gives the Retry-After rules.
Correlate a gateway log line¶
Do this step when a gateway log line must match a typevet result or error. Issue #356 holds the contract.
-
Set
TYPEVET_VLLM__REQUEST_ID_HEADERas in step 6. Configure the gateway to log that header. -
Read the id of each judgment question from the response.
from typevet.adapters.inbound.backend_settings import open_judgment with open_judgment() as session: response = session.port.judge("state", questions, session.model) for name, request_id in response.request_ids.items(): print(name, request_id)The id is the value in the scoring request of that question. When a scoring wrapper sends more than one request for a question, the last id is kept.
-
Read
request_idon aBackendHttpErroror aTransportError. It is the id of the request that failed. This works forgeneration_adapter,async_vllm_generation_adapterandopen_judgment. -
Search the gateway log for that id.
Without TYPEVET_VLLM__REQUEST_ID_HEADER, request_ids is empty and request_id is None.
A successful generation result does not carry the id.
The /tokenize requests of a judgment are not in request_ids.
typevet never logs the id. The id is not a secret, so typevet does not mask it.