Check images against a synthetic register¶
Kind: explanation.
This page is for an engineer who reads the check-match receipts. It explains what typevet measured when it compared one check image with one register row. It names the data and the pins, and it shows where the evidence stops. It is not a how-to. The synthetic checks reference names the modules, the metric rules and the receipt fields.
Every number on this page comes from the two committed receipts. Setup facts that the receipts do not pin come from the #316 run comments. The probabilities are model confidence. The checks are generated, so the results say nothing about real checks.
What was measured¶
Each case is one typed judgment. The call carries one check image and the register row as text: check number, date, payee and amount. It asks four questions:
| Question id | Type | Answer |
|---|---|---|
payee_matches |
Noul |
Probability that the check names the register payee |
amounts_match |
Noul |
Probability that both check amounts equal the register amount |
verdict |
Choice |
consistent, payee_mismatch, amount_mismatch, date_mismatch, unsigned or cannot_tell |
legibility |
Score |
0 to 4, how clearly the check text can be read |
The data is generated by typevet's own code with Pillow. The slice holds 20 register rows. Each row gives 7 variants, so the slice holds 140 cases, drawn with seed 0:
| Variant | What changes | Accepted verdict |
|---|---|---|
clean |
Nothing | consistent |
payee_changed |
The payee | payee_mismatch |
written_amount_changed |
The written amount only; the number amount equals the register | amount_mismatch |
both_amounts_changed |
Both amounts | amount_mismatch |
wrong_date |
The date | date_mismatch |
unsigned |
No signature strokes | unsigned |
low_legibility |
Clean content with a Gaussian blur | consistent or cannot_tell |
The design is in the #303 design comment. Amendment A1 sets the labels above. For amendment A2, the owner reviewed a contact sheet of all 140 renders. The owner signed off the labels before the live runs. A model check does not count as that human check.
The two runs¶
Each backend ran the slice once for this table. Later llama.cpp repeats are
in Limits of the data. Both runs used the
same renders: the slice SHA-256 starts 64a9621a2d55 in both receipts.
| Topic | llama.cpp | vLLM |
|---|---|---|
| Receipt | check_match_llama_cpp_receipt.json | check_match_vllm_receipt.json |
| Server | Local build b11277-eae11d221 |
Stock vLLM 0.30.0, one H100 on RunPod |
| Model | Alias gemma-4-31b-kv9-q4km-mm: Gemma 4 31B, Q2_K GGUF (see below) |
google/gemma-4-31B-it, revision 842da379, BF16 |
| Context | n_ctx 4096 tokens |
max_model_len 8192 tokens |
| Cases | 140, no backend failure | 140, no backend failure |
verdict accuracy |
0.821 (115 of 140) | 0.929 (130 of 140) |
| False-clear rate | 0.03 (3 of 100) | 0.00 (0 of 100) |
payee_matches ROC-AUC, ECE |
1.000, 0.000 | 1.000, 0.000 |
amounts_match ROC-AUC, ECE |
0.984, 0.027 | 1.000, 0.007 |
cannot_tell chosen |
0 of 140 | 0 of 140 |
| Noul–Choice agreement | 0.986 (138 of 140) | 0.993 (139 of 140) |
Mean legibility score, sharp renders |
4.00 | 4.00 |
Mean legibility score, blurred renders |
3.13 | 2.86 |
Mean legibility level, blurred renders |
3.00 | 2.90 |
| Mean latency | 5.44 s per case | 7.85 s per case, through a remote proxy |
The legibility score is the expected value over the five level
probabilities. The level is the most probable level.
The server, proxy and GPU facts come from the #316 run comments, not from the
receipts. The receipts pin only the llama.cpp alias, not the weights file.
The alias name says q4km, but it does not describe the decoder weights. The
#316 correction reports
model_ftype "Q2_K - Medium" from the server /props endpoint.
The two backends ran different weights: a Q2_K GGUF on llama.cpp and BF16
on vLLM. Do not credit any gap to the backend. The weights changed as well,
and one run each cannot separate the two causes.
Verdicts by variant¶
Each variant has 20 cases on each backend.
| Variant | llama.cpp verdicts | vLLM verdicts |
|---|---|---|
clean |
15 consistent, 5 unsigned |
19 consistent, 1 unsigned |
payee_changed |
20 payee_mismatch |
20 payee_mismatch |
written_amount_changed |
17 amount_mismatch, 3 consistent |
20 amount_mismatch |
both_amounts_changed |
20 amount_mismatch |
20 amount_mismatch |
wrong_date |
20 date_mismatch |
20 date_mismatch |
unsigned |
20 unsigned |
20 unsigned |
low_legibility |
16 unsigned, 3 consistent, 1 amount_mismatch |
11 consistent, 9 unsigned |
What the answers show¶
On sharp renders, both runs caught every changed payee, every changed date
and every change to both amounts. The hardest sharp variant was a written
amount that disagrees with the number amount. llama.cpp answered consistent
on 3 of those 20 checks. vLLM caught all 20.
The model never chose cannot_tell. Blur did not make it abstain. Instead,
blur moved the answer to unsigned: 16 of 20 blurred checks on llama.cpp and
9 of 20 on vLLM. The blurred signature strokes are the likely cause, but the
receipts cannot show why the model chose an answer.
Some clean checks also got unsigned: 5 of 20 on llama.cpp and 1 of 20 on
vLLM. Every render has a VOID mark that crosses the signature line, as the
design requires. That mark can hide or look like the strokes. The unsigned
variant scored 20 of 20 on both backends, but a model that says unsigned too
often also scores well there. Read the unsigned score together with the
clean and blurred rows.
What the probability means¶
Each Noul value is model confidence. It is not a calibrated match
percentage. An amounts_match value of 0.99 does not mean that 99 of 100 such
checks show the register amount.
The ROC-AUC and ECE values are high and low on this slice. The payee separation was complete on both backends. That result holds for these 140 generated renders only. It does not show that the probability is calibrated on other layouts, other fonts, scanned paper or other pins.
The false-clear rate and its limits¶
A false clear is a mismatch case answered consistent. The rate counts 100
cases: the five variants whose accepted verdict is not consistent. The
blurred variant never counts, because consistent is a correct answer there.
- llama.cpp: 3 of 100. All three are
written_amount_changedcases (r07,r09andr17). - vLLM: 0 of 100.
Zero in 100 is not proof of a zero rate. With no event in 100 cases, the rule of three puts the 95 % upper bound near 3 in 100. The count also mixes five variants with different difficulty. One run per backend gives no variance estimate.
Noul–Choice agreement and its limits¶
The agreement check asks whether the two Noul answers and the verdict say
the same thing. A Noul below 0.5 says "mismatch". The
reference states the full rule.
Agreement was 138 of 140 on llama.cpp and 139 of 140 on vLLM. Every disagreement was on a blurred render.
Agreement measures consistency, not truth. The three llama.cpp false clears
all agreed: the amounts_match value was 0.77, 0.86 and 1.00, and the
verdict was consistent. Both answers were wrong in the same direction. A
high agreement rate does not show that the answers are correct.
Layered use: a legibility gate before the field judgments¶
A real pipeline could run in layers. A typed quality check decides first whether the image is usable. Then the field judgments run. A person reviews each held image. The receipts suggest a first shape for that gate, but they do not measure it.
The legibility level was 4 on all 120 sharp renders on both backends. On the
20 blurred renders, the levels spread lower:
| Backend | Blurred levels | Gate "level 3 or lower: hold" | Blurred unsigned answers held |
|---|---|---|---|
| llama.cpp | level 2: 2, level 3: 16, level 4: 2 | Holds 18 of 20 blurred, 0 of 120 sharp | 14 of 16 |
| vLLM | level 1: 1, level 2: 5, level 3: 9, level 4: 5 | Holds 15 of 20 blurred, 0 of 120 sharp | 7 of 9 |
On these renders, the gate would remove most blurred unsigned errors before
the field judgments run. The gate also has a cost. It holds blurred checks that
the model answered correctly. On vLLM it holds 8 of the 11 correct consistent
answers. On llama.cpp it holds all 3 of 3. A person must then review those
checks too. It would not stop the three llama.cpp false clears.
Those checks were sharp, at level 4, and a legibility gate does not look at the
amounts.
The threshold was chosen after seeing this data. No code in typevet applies such a gate. #344 tested it once on new renders.
One test on a fresh seed¶
The rule was fixed before the run: level 3 or lower holds, level 4 passes. The owner signed off the seed-1 contact sheets. Then the same 20 rows times 7 variants ran once on local llama.cpp, from seed 1, with no rendering change. The verdict rule was also fixed first. The hypothesis holds when the gate holds 15 or more of the 20 blurred renders. It must also hold 2 or fewer of the 120 sharp ones.
| Measure | Seed 0 (#316) | Seed 1 (#344) |
|---|---|---|
| Blurred renders held | 18 of 20 | 17 of 20 |
| Sharp renders held | 0 of 120 | 0 of 120 |
| Accuracy without the gate | 0.821 | 0.864 |
| Accuracy of the passed cases | 0.918 | 0.976 |
The hypothesis holds on seed 1. Every sharp render scored level 4 again. The
gate has the same cost as before: it holds blurred checks the model would have
answered correctly. Two seeds on one backend do not calibrate the level. The
receipt is
evals/fixtures/checks/receipts/check_match_llama_cpp_seed1_receipt.json.
Option order of the verdict¶
Issue #105 asks if the
order of the six verdict options moves the answer. The study scores 21 seed-1
cases, rows r00 to r02, on llama.cpp. It lists the options in six orders from
a balanced Latin square. Each option takes each position one time. Ordering 0
is the order of the main run. Each order is one scoring request, so the study
sends 126 requests. The mean verdict is the arithmetic mean of the
probabilities, as in TypeLLM permutations="auto". The rule was set before
the run. A position spread of 0.05 or less means no position bias to fix. A
larger spread with no loss of accuracy gives an opt-in option. A larger spread
with a loss of accuracy is inconclusive. The receipt is
evals/fixtures/checks/receipts/check_match_orderings_llama_cpp_seed1.json.
The first run, on 21 cases (rows r00 to r02), was inconclusive. The position spread was 0.057, just above the limit. The mean verdict kept 17 of 21 cases right against 18 for the single order.
The full run on all 140 seed-1 cases (840 requests) settled it. The spread
was 0.033, with a 95% bootstrap interval of 0.020 to 0.049, all below the
limit. The mean verdict kept 118 cases right against 121 for the single
order; the interval of that difference is −0.050 to 0.000. Averaging never
helped a case. The 20 cases whose winner moved with the order (12 clean,
8 blurred) switched between consistent and unsigned. This is the near-tie
the main runs show on clean renders. The answer digits follow position, so the
spread also holds any bias toward a digit. typevet does not average
orderings. The receipt is
evals/fixtures/checks/receipts/check_match_orderings_llama_cpp_seed1_full.json;
the result is on #105.
What this is not¶
This is not a fraud-detection control and not a counterfeit-detection control. The evaluation makes no claim about MICR lines. Do not use these results to pay or refuse a check.
Generated checks are not evidence about real checks. Each render is not
negotiable by construction. The routing number fails the ABA check digit, the
account number prints as zeros, and a SPECIMEN mark crosses the face. The
bank name says that the bank is not real, and the signatures are seeded
synthetic strokes.
Limits of the data¶
- Generated layout. One layout, the bundled Pillow font and one blur kind. Real checks vary in paper, print, handwriting and scan quality.
- Small slice. 20 register rows, 20 cases per variant.
- Run-to-run spread, llama.cpp only. Issue #357 ran the seed-0
slice five more times on llama.cpp, with equal code paths. All five gave
the same 140 verdicts. The largest change of any recorded probability was 0.
Accuracy (0.821), false-clear rate (0.03) and each
NoulECE had a range of 0. Under the pre-registered rule, the one-run numbers are stable to the reported precision. The pooled 95% bootstrap interval of accuracy is 0.794 to 0.849. Its 700 rows repeat the same 140 cases, so it does not cover new cases. The first receipt,check_match_llama_cpp_receipt.json, has other code-path digests, so it is outside the statistics. Its answers are identical too. The vLLM run has no repeats. Wall times are reported, not compared. Other jobs shared the CPU during the repeats (see the evidence note on #357). The receipts are repeats 1, 2, 3, 4 and 5. - Different weights.
Q2_Kon llama.cpp and BF16 on vLLM. The receipts cannot separate a weights effect from a backend effect. They also pin only the llama.cpp alias, not the weights file. - A documented quantization, one run. Issue
#358 ran the seed-0
slice once on Google's QAT
Q4_0pair (google/gemma-4-31B-it-qat-q4_0-gguf, revision59dde245, aliasgemma-4-31b-qat-q4_0-mm), with the same server flags and build as the pin. Accuracy was 0.807 (113 of 140) against 0.821, and the false-clear rate 0.02 against 0.03. Thepayee_matchesECE stayed near 0, and theamounts_matchECE moved from 0.027 to 0.018. The run-to-run spread is 0, so the differences come from the weights. The accuracy sits inside the pooled interval above, so one run cannot rank the two quantizations. The receipt ischeck_match_llama_cpp_qat_receipt.json. - Labels. The owner checked the labels on a contact sheet once, for seed 0.
Latency and cost¶
One item is one case: one typed judgment that asks all four questions. The
receipts give wall_seconds and, per case, latency_seconds and
input_tokens.
| Backend | Items | Wall time | Mean seconds per item | Items per minute | Mean input tokens |
|---|---|---|---|---|---|
| llama.cpp | 140 | 762 s (12.7 min) | 5.44 | 11.0 | 1,598 |
| vLLM | 140 | 1,100 s (18.3 min) | 7.85 | 7.6 | 1,918 |
Read these numbers with their limits:
- One request at a time. Each run sent one case, then waited for the answer. This is not a throughput test. The H100 throughput reference is the #236 receipt on the performance page.
- Network time is included on vLLM. The vLLM requests went through the RunPod proxy to a pod in data center AP-IN-2. The check run and the signature run shared that pod at the same time.
- Input tokens differ. vLLM counted more input tokens per case than llama.cpp: 1,918 against 1,598. The cause was not checked.
The H100 pod cost about $2.24 in total (#316 pod record). That cost is shared with the signature run (#319) and the wording check (#309). It is not the cost of this run alone.
Pins¶
| Pin | llama.cpp | vLLM |
|---|---|---|
| Generator seed | 0 | 0 |
| Register rows, variants per row | 20, 7 | 20, 7 |
| Pillow | 12.3.0 | 12.3.0 |
| Slice SHA-256 (first 12) | 64a9621a2d55 |
64a9621a2d55 |
| Code baseline commit | 38cada3 |
ce63ddd |
| Wall time | 762 s | 1100 s |