Image judgment throughput on one H100¶
Kind: explanation.
This page is for an enterprise reader who asks how many image judgments one GPU can give. It explains one measured run of typevet image judgments on one H100. It says what raised the rate, what did not, and whether the answers changed under load. It also says what the run cannot tell you.
Every number on this page comes from the committed throughput receipts. Setup facts that the receipts do not record come from the #336 run plan and the #336 result comment. The run is one run per level, on one model, one GPU and one pin.
The short answer¶
On this setup, 16 requests in flight gave the highest rate. The full sets ran at 4.4 to 5.3 judgments per second. That is 23 to 42 times the rate of the 0.4.0 runs. Those runs sent one request at a time over a longer network path.
At 32 in flight, the rate fell and each judgment waited longer. The verdicts for faces and checks did not change under load. For signatures, 1 to 3 near-tied verdicts moved under batching. These results are for this setup only.
What was measured¶
One judgment is one item: one typed judgment that asks all its questions. Each set asks different questions about different images:
| Set | Images per judgment | Questions per judgment | Full set |
|---|---|---|---|
| Faces (LFW View 2) | 2 | 3 | 200 pairs |
| Checks (synthetic) | 1 | 4 | 140 cases |
| Signatures (CEDAR) | 2 | 3 | 180 pairs |
The face page, the check page and the signature page explain each set and its questions.
typevet sends each question as a separate request. The request carries the images again. A face judgment is thus three requests, each with both images.
Setup¶
| Item | Value |
|---|---|
| GPU | One H100 SXM, RunPod secure cloud, data center US-MO-1 |
| Server image | vllm/vllm-openai:v0.30.0 (receipts record build 0.30.0) |
| Model | google/gemma-4-31B-it, BF16 |
| Model revision | 842da3794eaa0b77d5f08bae87a17459d91ff475 (receipt pin) |
| Server flags | --enable-prefix-caching --max-num-seqs 32 --max-num-batched-tokens 16384 --gpu-memory-utilization 0.95 --max-model-len 8192 --limit-mm-per-prompt {"image":2} --logprobs-mode raw_logprobs |
| Client | Operator machine in Arizona, direct TCP port to the pod |
| Network baseline | Median 0.073 s for 20 GET /v1/models calls |
| Run date | 2026-09-30 |
The receipts do not record the server flags. The flags come from the #336 run
plan. The result comment omits --logprobs-mode raw_logprobs.
Full sets at concurrency 16 and 32¶
Concurrency is the number of judgments in flight at one time. Latency is the client time for one judgment, network included.
| Set | Concurrency | Wall time | Judgments/s | Images/s | Latency p50 / p95 |
|---|---|---|---|---|---|
| Faces | 16 | 45.0 s | 4.45 | 8.89 | 3.62 s / 3.81 s |
| Faces | 32 | 58.6 s | 3.42 | 6.83 | 9.30 s / 10.19 s |
| Checks | 16 | 26.4 s | 5.31 | 5.31 | 2.92 s / 3.42 s |
| Checks | 32 | 28.3 s | 4.94 | 4.94 | 6.52 s / 7.31 s |
| Signatures | 16 | 40.6 s | 4.44 | 8.87 | 3.53 s / 4.05 s |
| Signatures | 32 | 54.0 s | 3.33 | 6.67 | 9.92 s / 11.61 s |
No judgment failed or was discarded in any run.
The concurrency sweep¶
The sweep used the same subset at each level. The subset is 60 face pairs, 63 check cases and 60 signature pairs.
| Set | 1 | 4 | 16 | 32 |
|---|---|---|---|---|
| Faces, judgments/s | 1.14 | 3.01 | 4.78 | 3.30 |
| Checks, judgments/s | 0.88 | 2.80 | 5.47 | 5.01 |
| Signatures, judgments/s | 1.09 | 3.11 | 4.49 | 3.28 |
| Faces, p50 latency | 0.84 s | 1.25 s | 3.19 s | 8.55 s |
| Checks, p50 latency | 1.13 s | 1.35 s | 2.73 s | 6.11 s |
| Signatures, p50 latency | 0.91 s | 1.27 s | 3.41 s | 8.50 s |
The rate rose from 1 to 16 in flight for every set. At 32 it fell for every set. The full-set runs show the same fall.
What raised throughput¶
Three changes raised the rate against the 0.4.0 runs:
- Concurrency. On the same pod, 16 in flight gave four to six and a half times the rate of 1.
- Prefix caching. The server reused cached prompt tokens across the separate question requests. At concurrency 16, 0.55 of prompt tokens came from the cache.
- Client placement. The client used a direct TCP port to a US pod. The 0.4.0 runs used the RunPod proxy to a pod in AP-IN-2.
The receipts do not split the gain between these three changes. The sweep isolates concurrency only. No run turned prefix caching off or moved the client back.
What did not: concurrency 32¶
At 32 in flight, every set ran slower than at 16. Latency more than doubled. The server counters show two changes at 32. For faces and signatures, requests waited in the server queue. For checks, the queue time stayed near zero. The prefix-cache hit rate fell for all three sets. The receipts do not show why the hit rate fell.
| Counter, full set | Faces c16 | Faces c32 | Checks c16 | Checks c32 | Signatures c16 | Signatures c32 |
|---|---|---|---|---|---|---|
| Requests | 600 | 600 | 560 | 560 | 540 | 540 |
| Prefix-cache hit rate | 0.550 | 0.041 | 0.550 | 0.382 | 0.544 | 0.066 |
| Multimodal cache hits | 921 of 1,200 | 1,200 of 1,200 | 483 of 560 | 560 of 560 | 888 of 1,080 | 1,080 of 1,080 |
| Preemptions | 0 | 0 | 0 | 0 | 0 | 0 |
| Queue time, sum | 0.018 s | 69.9 s | 0.012 s | 0.018 s | 0.011 s | 61.9 s |
At concurrency 16, the faces run used 217,600 cached tokens of 395,600 prompt tokens. The queue time is the server sum over all requests. Near zero means requests did not wait in the server queue.
Comparison with the 0.4.0 runs¶
The 0.4.0 receipts are single-stream runs of the same model on vLLM. The check and signature receipts pin the same revision. The face receipt records no revision. These runs sent one request at a time through the RunPod proxy to AP-IN-2. Prefix caching was off.
| Set | 0.4.0 judgments/min | Concurrency 16 judgments/min | Ratio |
|---|---|---|---|
| Faces | 11.4 | 267 | 23.4 |
| Checks | 7.6 | 319 | 41.7 |
| Signatures | 9.2 | 266 | 29.0 |
The ratio combines concurrency, prefix caching and the shorter network path. Do not read it as the gain from any one of them.
Answers under load¶
A higher rate is of no use if the answers change. The receipts let us compare answers at each level.
Within this session. Faces and checks gave the same verdicts at every
level. The face Noul probabilities were identical at every level. The check
payee_matches probabilities were identical too. The check amounts_match
probabilities moved by small amounts between levels. The largest move was
0.00004 in the subset and 0.00006 in the full set.
For signatures, the Noul probabilities were identical at every level. Some
Choice verdicts moved:
- In the subset, 1 of 60 verdicts at 16 and 2 of 60 at 32 differ from concurrency 1.
- In the full set, 3 of 180 verdicts differ between 16 and 32.
All moved verdicts are on genuine-versus-skilled-forgery pairs. The Choice
options for these pairs are near-tied. The leading option has a probability of
about 0.65 or less in each receipt checked. Batching moved these pairs between
same_writer and skilled_forgery_suspected.
Against the 0.4.0 receipts. Verdicts are equal for 199 of 200 faces, 140 of 140 checks and 176 to 179 of 180 signatures. The face difference exists already at concurrency 1. Three of the four signature differences also exist at concurrency 1. These differences do not come from concurrency. The receipts do not isolate the cause. The fourth signature difference is a near-tied pair that moved at concurrency 16.
At concurrency 32, two of the three signature pairs that differ at
concurrency 1 return to the 0.4.0 verdict. These pairs are original_27_19
and original_12_2. At the same level, the prefix-cache hit rate falls.
The run plan also asked for Noul probabilities within 0.001 of 0.4.0. That
did not hold everywhere. In the full sets, 7 face pairs and 10 signature pairs
differ by more. The largest face difference is 0.36, on a pair whose verdict
did not change. One check value differs by 0.007.
The #332 design (item 3) said that a mismatch makes the run a bug. The #336 run plan said that a mismatch is reported, not averaged away. This page reports it for that reason.
Image reuse and an open possibility¶
Each question is a separate request that sends the images again. The server caches do the image reuse. At concurrency 16, most multimodal cache lookups hit.
A future code change could ask all questions of a judgment in one request. That is an open possibility only. No receipt measures it.
The #342 design looked at that possibility and closed it for now. One prompt with several scoring positions changes what the model conditions on. So its probabilities would need new calibration evidence. Also, llama.cpp gives no prompt-side logprobs. One shared prefix with batched continuations works on llama.cpp only. The one small change left is about the questions of one judgment. It sends them at the same time instead of one after the other. That change belongs to #121. It needs a new concurrency 32 receipt before any claim.
Cost¶
The pod ran from 19:13:28Z to about 19:31:50Z, about 18.4 minutes. At $3.49 per hour, that is about $1.07. The pod took 6.7 minutes to become ready. That start-up time is part of the cost.
The same pod also ran the #329 Gemma held-out text receipt. The $1.07 is thus not the cost of the image runs alone.
On this setup, 1,000 judgments at concurrency 16 in the full sets take this much GPU time:
| Set | Seconds per 1,000 judgments | GPU cost per 1,000 judgments |
|---|---|---|
| Faces | 225 s | $0.22 |
| Checks | 188 s | $0.18 |
| Signatures | 225 s | $0.22 |
These figures exclude start-up time and idle time. They use the $3.49 hourly price from the run plan.
What you can take from this¶
- On this setup, about 16 judgments in flight gave the highest rate of the levels tried.
- On this setup, one H100 gave about 4.4 to 5.3 image judgments per second.
- On this setup, faces and checks gave the same verdicts at every level tried.
- Near-tied
Choiceanswers can move under batching. Check them if a decision depends on them.
What you cannot take from this¶
- No general vLLM or GPU claim. The run used one model, one GPU, one pin and one flag set.
- No variance. Each level ran once. The run has no repeats and no confidence interval. Run-to-run spread was measured on the llama.cpp checks slice only (#357). Its answers were identical, but its throughput was not compared.
- No other levels. The sweep tried 1, 4, 16 and 32 only. The best level can lie between them.
- No split of the gain. The receipts do not separate concurrency, caching and network.
- No single-request latency claim. Concurrency hides round-trip time. The client latency includes the network path from Arizona.
- No other image sizes or prompts. Other images, questions or token counts can give other rates.
- No quality claim. Throughput says nothing about whether the answers are correct. The set pages report accuracy.
Receipts¶
| Set | Sweep receipts | Full-set receipts | 0.4.0 receipt |
|---|---|---|---|
| Faces | evals/fixtures/lfw/receipts/face_match_vllm_throughput_c{1,4,16,32}.json |
face_match_vllm_throughput_full_c{16,32}.json |
face_match_vllm_receipt.json |
| Checks | evals/fixtures/checks/receipts/check_match_vllm_throughput_c{1,4,16,32}.json |
check_match_vllm_throughput_full_c{16,32}.json |
check_match_vllm_receipt.json |
| Signatures | evals/fixtures/cedar/receipts/signature_match_vllm_throughput_c{1,4,16,32}.json |
signature_match_vllm_throughput_full_c{16,32}.json |
signature_match_vllm.json |
The design is on #332, with the client placement amendment. The text throughput run is on the performance page.