Performance on one H100¶
Kind: reference.
This page lists the typed-judgment throughput that typevet measured on vLLM. The source of every number is the #236 attempt-2 live receipt. The pod cost comes from the #236 supervisor pod record. The data is from one run on one pod with one pin.
Latency at a glance¶
Each Banking77 record makes two scoring calls: a Noul and a Choice.
| In flight | Client time per record, p50 / p95 | Server time per request, p95 | Records/s |
|---|---|---|---|
| 1 | 0.24 s / 0.29 s | 0.3 s or less | 3.96 |
| 8 | 0.31 s / 0.38 s | 0.3 s or less | 25.1 |
| 32 | 0.82 s / 1.01 s | 0.5 s or less | 37.3 |
| 64 | 1.59 s / 1.88 s | 1.0 s or less | 39.6 |
One record takes less than one second up to 32 requests in flight. At 64 in flight, each record waits longer, but total throughput is 10 times the rate at 1. Client time includes the RunPod proxy from the operator machine. Server times are histogram bucket upper bounds.
Pins¶
| Item | Value |
|---|---|
| Server image | vllm/vllm-openai:v0.30.0 (/version returns 0.30.0) |
| Model | google/gemma-4-31B-it, served as gemma-4-31b-it |
| Model revision | 842da3794eaa0b77d5f08bae87a17459d91ff475 (short 842da37), per the contract |
| Weights | BF16, no quantization |
| GPU | One H100 80GB HBM3 SXM, RunPod secure cloud, data center US-MO-1 |
| Server flags | --max-model-len 8192 --gpu-memory-utilization 0.95 --max-num-seqs 64 --enable-prefix-caching --logprobs-mode raw_logprobs --served-model-name gemma-4-31b-it |
| API key | Required |
| typevet revision | a0cc4b9 |
| Client | The operator machine, through the RunPod proxy, User-Agent: curl/8.9.1 |
| Harness | evals/tests/live/test_public_throughput_live.py (under tests/live/ at a0cc4b9) |
| Run date | 2026-09-29 |
/v1/models does not show a model revision.
The revision comes from the pre-registered contract.
Method¶
The harness sends one judge call per record through the sync judgment port.
A thread pool sets the number of records in flight.
This page calls that number the concurrency level.
Each question in a record uses one scoring call.
| Measure | Definition |
|---|---|
| Wall-clock | Last record end minus first record start, on the client clock |
| Records/s | Answered records divided by wall-clock seconds |
| Scoring calls/s | Scoring calls divided by wall-clock seconds |
| Client latency | Seconds per sent record, on the client, proxy included |
| Server e2e | vLLM /metrics request latency histogram, per scoring call |
| Errors | Records without an answer |
Server percentiles are histogram bucket upper bounds. The smallest bucket edge is 0.3 s. Thus "≤0.3" means 0.3 s or less, with no more detail.
Banking77 balanced, 480 records¶
The set holds all 240 fraud rows and 240 other rows of the Banking77 test split.
Each record asks the reports_unauthorized Noul and the fraud_type Choice with 6 options.
Each record uses two scoring calls.
The mean prompt is 253.7 tokens per record, and the longest is 337 tokens.
| Level | Records | Wall-clock (s) | Records/s | Scoring calls/s | Client p50 / p95 / p99 (s) | Server e2e p95 (s, bucket bound) | Errors |
|---|---|---|---|---|---|---|---|
| 1 | 480 | 121.36 | 3.96 | 7.91 | 0.243 / 0.295 / 0.396 | ≤0.3 | 0 |
| 8 | 480 | 19.11 | 25.12 | 50.24 | 0.309 / 0.376 / 0.538 | ≤0.3 | 0 |
| 32 | 480 | 12.89 | 37.25 | 74.50 | 0.825 / 1.014 / 1.310 | ≤0.5 | 0 |
| 64 | 480 | 12.11 | 39.63 | 79.26 | 1.588 / 1.883 / 2.028 | ≤1.0 | 0 |
xychart-beta
title "Banking77-480: records/s by concurrency level"
x-axis "Concurrency level" ["1", "8", "32", "64"]
y-axis "Records/s" 0 --> 45
bar [3.96, 25.12, 37.25, 39.63]
Level 64 gives 10.0 times the records/s of level 1. From level 32 to level 64, records/s increases by 6%. Over the same step, the client p50 latency is about two times larger. The four levels together took 167.0 s.
DIFrauD SMS, 500 records, level 64¶
The set holds 500 rows of the DIFrauD SMS test split, drawn with seed 0.
Each record asks the is_scam Noul and uses one scoring call.
The mean prompt is 89.0 tokens per record, and the longest is 150 tokens.
| Level | Records | Wall-clock (s) | Records/s | Scoring calls/s | Client p50 / p95 / p99 (s) | Server e2e p95 (s, bucket bound) | Errors |
|---|---|---|---|---|---|---|---|
| 64 | 500 | 4.88 | 102.46 | 102.46 | 0.592 / 0.880 / 1.122 | ≤0.8 | 0 |
Calibration against finvet¶
The contract compares the expected calibration error (ECE) of each set with a finvet Jev baseline. ECE uses 10 equal-width probability bins. The pass threshold adds 0.03 to the finvet value.
| Set | Answered | ECE | finvet Jev ECE | Threshold | Base rate (typevet / finvet) | Result |
|---|---|---|---|---|---|---|
| Banking77-480 | 480 of 480 at each level | 0.0892 to 0.0905 | 0.17 | 0.20 | 0.500 / 0.50 | Pass |
| DIFrauD SMS-500 | 500 of 500 | 0.1578 | 0.07 | 0.10 | 0.184 / 0.196 | Fail, no parity |
The Banking77-480 ECE per level is 0.0902 (1), 0.0892 (8), 0.0905 (32) and 0.0905 (64).
DIFrauD fails parity because the model is overconfident on "scam". 166 records have a scam probability of 0.9 or more. Only 55.4% of those 166 records are scam. About 74 records that are not scam get a high "scam" probability.
The typevet rows differ from the finvet rows.
finvet samples rows differently from the typevet hash seed.
The finvet baselines have two decimals and came from the Jev service jev-1.13.0.
The Jev backend and hardware are unknown.
Cold start and cost¶
| Item | Value |
|---|---|
| Pod created | 20:41:39Z |
/v1/models ready |
20:48:25Z |
| Cold start | 6 min 46 s |
| Run | 20:49:00Z to 20:51:54Z |
| Pod price | $3.49 per hour |
| Pod cost, creation to deletion | About $0.75, from creation at 20:41:39Z to deletion at 20:54:27Z |
| Compute cost per 1,000 Banking77 records at level 64 | About $0.024 |
The cost per 1,000 records is arithmetic, not a measured value. It is 1,000 records at 39.63 records/s, priced at $3.49 per hour. It excludes cold start, idle time and the proxy.
Not measured¶
- Full Banking77 test split, 3,080 records. The runner stopped before the first scoring call. The harness shares one tokenizer-call cap of 256 across the whole run. Only 144 calls were left for 3,080 new texts, so the runner stopped before the first scoring call. This is a harness limit, not a limit of typevet or vLLM. No throughput, latency or ECE exists for this set.
- The finvet collections workload. One question has 13 options. Native Choice supports 10 options on the Gemma 4 tokenizer. This run did not measure that workload.
- GPU memory. The vLLM
/metricsendpoint does not show it.
Limits¶
- Public datasets only. No partner data was measured.
- The data is from one run, one pod and one model pin.
- The texts are short public texts of 89 to 254 mean prompt tokens per record. The throughput does not transfer to longer prompts.
- Client latency includes the RunPod proxy from the operator machine to US-MO-1.
- Server percentiles are histogram bucket upper bounds. Below 0.3 s they give no detail.
- The page makes no claim about other GPUs, models, precisions or vLLM versions.
- Calibration here is ECE on two sets. It is not a general quality claim.
Dataset attribution¶
- Banking77 by PolyAI (Casanueva et al., 2020), licensed under CC BY 4.0.
- DIFrauD (
difraud/difraudon Hugging Face), licensed under MIT.
No record text is in the receipt or on this page.