Skip to content

Check images against a synthetic register

Kind: explanation.

This page is for an engineer who reads the check-match receipts. It explains what typevet measured when it compared one check image with one register row. It names the data and the pins, and it shows where the evidence stops. It is not a how-to. The synthetic checks reference names the modules, the metric rules and the receipt fields.

Every number on this page comes from the two committed receipts. Setup facts that the receipts do not pin come from the #316 run comments. The probabilities are model confidence. The checks are generated, so the results say nothing about real checks.

What was measured

Each case is one typed judgment. The call carries one check image and the register row as text: check number, date, payee and amount. It asks four questions:

Question id Type Answer
payee_matches Noul Probability that the check names the register payee
amounts_match Noul Probability that both check amounts equal the register amount
verdict Choice consistent, payee_mismatch, amount_mismatch, date_mismatch, unsigned or cannot_tell
legibility Score 0 to 4, how clearly the check text can be read

The data is generated by typevet's own code with Pillow. The slice holds 20 register rows. Each row gives 7 variants, so the slice holds 140 cases, drawn with seed 0:

Variant What changes Accepted verdict
clean Nothing consistent
payee_changed The payee payee_mismatch
written_amount_changed The written amount only; the number amount equals the register amount_mismatch
both_amounts_changed Both amounts amount_mismatch
wrong_date The date date_mismatch
unsigned No signature strokes unsigned
low_legibility Clean content with a Gaussian blur consistent or cannot_tell

The design is in the #303 design comment. Amendment A1 sets the labels above. For amendment A2, the owner reviewed a contact sheet of all 140 renders. The owner signed off the labels before the live runs. A model check does not count as that human check.

The two runs

Each backend ran the slice once for this table. Later llama.cpp repeats are in Limits of the data. Both runs used the same renders: the slice SHA-256 starts 64a9621a2d55 in both receipts.

Topic llama.cpp vLLM
Receipt check_match_llama_cpp_receipt.json check_match_vllm_receipt.json
Server Local build b11277-eae11d221 Stock vLLM 0.30.0, one H100 on RunPod
Model Alias gemma-4-31b-kv9-q4km-mm: Gemma 4 31B, Q2_K GGUF (see below) google/gemma-4-31B-it, revision 842da379, BF16
Context n_ctx 4096 tokens max_model_len 8192 tokens
Cases 140, no backend failure 140, no backend failure
verdict accuracy 0.821 (115 of 140) 0.929 (130 of 140)
False-clear rate 0.03 (3 of 100) 0.00 (0 of 100)
payee_matches ROC-AUC, ECE 1.000, 0.000 1.000, 0.000
amounts_match ROC-AUC, ECE 0.984, 0.027 1.000, 0.007
cannot_tell chosen 0 of 140 0 of 140
Noul–Choice agreement 0.986 (138 of 140) 0.993 (139 of 140)
Mean legibility score, sharp renders 4.00 4.00
Mean legibility score, blurred renders 3.13 2.86
Mean legibility level, blurred renders 3.00 2.90
Mean latency 5.44 s per case 7.85 s per case, through a remote proxy

The legibility score is the expected value over the five level probabilities. The level is the most probable level.

The server, proxy and GPU facts come from the #316 run comments, not from the receipts. The receipts pin only the llama.cpp alias, not the weights file. The alias name says q4km, but it does not describe the decoder weights. The #316 correction reports model_ftype "Q2_K - Medium" from the server /props endpoint.

The two backends ran different weights: a Q2_K GGUF on llama.cpp and BF16 on vLLM. Do not credit any gap to the backend. The weights changed as well, and one run each cannot separate the two causes.

Verdicts by variant

Each variant has 20 cases on each backend.

Variant llama.cpp verdicts vLLM verdicts
clean 15 consistent, 5 unsigned 19 consistent, 1 unsigned
payee_changed 20 payee_mismatch 20 payee_mismatch
written_amount_changed 17 amount_mismatch, 3 consistent 20 amount_mismatch
both_amounts_changed 20 amount_mismatch 20 amount_mismatch
wrong_date 20 date_mismatch 20 date_mismatch
unsigned 20 unsigned 20 unsigned
low_legibility 16 unsigned, 3 consistent, 1 amount_mismatch 11 consistent, 9 unsigned

What the answers show

On sharp renders, both runs caught every changed payee, every changed date and every change to both amounts. The hardest sharp variant was a written amount that disagrees with the number amount. llama.cpp answered consistent on 3 of those 20 checks. vLLM caught all 20.

The model never chose cannot_tell. Blur did not make it abstain. Instead, blur moved the answer to unsigned: 16 of 20 blurred checks on llama.cpp and 9 of 20 on vLLM. The blurred signature strokes are the likely cause, but the receipts cannot show why the model chose an answer.

Some clean checks also got unsigned: 5 of 20 on llama.cpp and 1 of 20 on vLLM. Every render has a VOID mark that crosses the signature line, as the design requires. That mark can hide or look like the strokes. The unsigned variant scored 20 of 20 on both backends, but a model that says unsigned too often also scores well there. Read the unsigned score together with the clean and blurred rows.

What the probability means

Each Noul value is model confidence. It is not a calibrated match percentage. An amounts_match value of 0.99 does not mean that 99 of 100 such checks show the register amount.

The ROC-AUC and ECE values are high and low on this slice. The payee separation was complete on both backends. That result holds for these 140 generated renders only. It does not show that the probability is calibrated on other layouts, other fonts, scanned paper or other pins.

The false-clear rate and its limits

A false clear is a mismatch case answered consistent. The rate counts 100 cases: the five variants whose accepted verdict is not consistent. The blurred variant never counts, because consistent is a correct answer there.

  • llama.cpp: 3 of 100. All three are written_amount_changed cases (r07, r09 and r17).
  • vLLM: 0 of 100.

Zero in 100 is not proof of a zero rate. With no event in 100 cases, the rule of three puts the 95 % upper bound near 3 in 100. The count also mixes five variants with different difficulty. One run per backend gives no variance estimate.

Noul–Choice agreement and its limits

The agreement check asks whether the two Noul answers and the verdict say the same thing. A Noul below 0.5 says "mismatch". The reference states the full rule.

Agreement was 138 of 140 on llama.cpp and 139 of 140 on vLLM. Every disagreement was on a blurred render.

Agreement measures consistency, not truth. The three llama.cpp false clears all agreed: the amounts_match value was 0.77, 0.86 and 1.00, and the verdict was consistent. Both answers were wrong in the same direction. A high agreement rate does not show that the answers are correct.

Layered use: a legibility gate before the field judgments

A real pipeline could run in layers. A typed quality check decides first whether the image is usable. Then the field judgments run. A person reviews each held image. The receipts suggest a first shape for that gate, but they do not measure it.

The legibility level was 4 on all 120 sharp renders on both backends. On the 20 blurred renders, the levels spread lower:

Backend Blurred levels Gate "level 3 or lower: hold" Blurred unsigned answers held
llama.cpp level 2: 2, level 3: 16, level 4: 2 Holds 18 of 20 blurred, 0 of 120 sharp 14 of 16
vLLM level 1: 1, level 2: 5, level 3: 9, level 4: 5 Holds 15 of 20 blurred, 0 of 120 sharp 7 of 9

On these renders, the gate would remove most blurred unsigned errors before the field judgments run. The gate also has a cost. It holds blurred checks that the model answered correctly. On vLLM it holds 8 of the 11 correct consistent answers. On llama.cpp it holds all 3 of 3. A person must then review those checks too. It would not stop the three llama.cpp false clears. Those checks were sharp, at level 4, and a legibility gate does not look at the amounts.

The threshold was chosen after seeing this data. No code in typevet applies such a gate. #344 tested it once on new renders.

One test on a fresh seed

The rule was fixed before the run: level 3 or lower holds, level 4 passes. The owner signed off the seed-1 contact sheets. Then the same 20 rows times 7 variants ran once on local llama.cpp, from seed 1, with no rendering change. The verdict rule was also fixed first. The hypothesis holds when the gate holds 15 or more of the 20 blurred renders. It must also hold 2 or fewer of the 120 sharp ones.

Measure Seed 0 (#316) Seed 1 (#344)
Blurred renders held 18 of 20 17 of 20
Sharp renders held 0 of 120 0 of 120
Accuracy without the gate 0.821 0.864
Accuracy of the passed cases 0.918 0.976

The hypothesis holds on seed 1. Every sharp render scored level 4 again. The gate has the same cost as before: it holds blurred checks the model would have answered correctly. Two seeds on one backend do not calibrate the level. The receipt is evals/fixtures/checks/receipts/check_match_llama_cpp_seed1_receipt.json.

Option order of the verdict

Issue #105 asks if the order of the six verdict options moves the answer. The study scores 21 seed-1 cases, rows r00 to r02, on llama.cpp. It lists the options in six orders from a balanced Latin square. Each option takes each position one time. Ordering 0 is the order of the main run. Each order is one scoring request, so the study sends 126 requests. The mean verdict is the arithmetic mean of the probabilities, as in TypeLLM permutations="auto". The rule was set before the run. A position spread of 0.05 or less means no position bias to fix. A larger spread with no loss of accuracy gives an opt-in option. A larger spread with a loss of accuracy is inconclusive. The receipt is evals/fixtures/checks/receipts/check_match_orderings_llama_cpp_seed1.json.

The first run, on 21 cases (rows r00 to r02), was inconclusive. The position spread was 0.057, just above the limit. The mean verdict kept 17 of 21 cases right against 18 for the single order.

The full run on all 140 seed-1 cases (840 requests) settled it. The spread was 0.033, with a 95% bootstrap interval of 0.020 to 0.049, all below the limit. The mean verdict kept 118 cases right against 121 for the single order; the interval of that difference is −0.050 to 0.000. Averaging never helped a case. The 20 cases whose winner moved with the order (12 clean, 8 blurred) switched between consistent and unsigned. This is the near-tie the main runs show on clean renders. The answer digits follow position, so the spread also holds any bias toward a digit. typevet does not average orderings. The receipt is evals/fixtures/checks/receipts/check_match_orderings_llama_cpp_seed1_full.json; the result is on #105.

What this is not

This is not a fraud-detection control and not a counterfeit-detection control. The evaluation makes no claim about MICR lines. Do not use these results to pay or refuse a check.

Generated checks are not evidence about real checks. Each render is not negotiable by construction. The routing number fails the ABA check digit, the account number prints as zeros, and a SPECIMEN mark crosses the face. The bank name says that the bank is not real, and the signatures are seeded synthetic strokes.

Limits of the data

  • Generated layout. One layout, the bundled Pillow font and one blur kind. Real checks vary in paper, print, handwriting and scan quality.
  • Small slice. 20 register rows, 20 cases per variant.
  • Run-to-run spread, llama.cpp only. Issue #357 ran the seed-0 slice five more times on llama.cpp, with equal code paths. All five gave the same 140 verdicts. The largest change of any recorded probability was 0. Accuracy (0.821), false-clear rate (0.03) and each Noul ECE had a range of 0. Under the pre-registered rule, the one-run numbers are stable to the reported precision. The pooled 95% bootstrap interval of accuracy is 0.794 to 0.849. Its 700 rows repeat the same 140 cases, so it does not cover new cases. The first receipt, check_match_llama_cpp_receipt.json, has other code-path digests, so it is outside the statistics. Its answers are identical too. The vLLM run has no repeats. Wall times are reported, not compared. Other jobs shared the CPU during the repeats (see the evidence note on #357). The receipts are repeats 1, 2, 3, 4 and 5.
  • Different weights. Q2_K on llama.cpp and BF16 on vLLM. The receipts cannot separate a weights effect from a backend effect. They also pin only the llama.cpp alias, not the weights file.
  • A documented quantization, one run. Issue #358 ran the seed-0 slice once on Google's QAT Q4_0 pair (google/gemma-4-31B-it-qat-q4_0-gguf, revision 59dde245, alias gemma-4-31b-qat-q4_0-mm), with the same server flags and build as the pin. Accuracy was 0.807 (113 of 140) against 0.821, and the false-clear rate 0.02 against 0.03. The payee_matches ECE stayed near 0, and the amounts_match ECE moved from 0.027 to 0.018. The run-to-run spread is 0, so the differences come from the weights. The accuracy sits inside the pooled interval above, so one run cannot rank the two quantizations. The receipt is check_match_llama_cpp_qat_receipt.json.
  • Labels. The owner checked the labels on a contact sheet once, for seed 0.

Latency and cost

One item is one case: one typed judgment that asks all four questions. The receipts give wall_seconds and, per case, latency_seconds and input_tokens.

Backend Items Wall time Mean seconds per item Items per minute Mean input tokens
llama.cpp 140 762 s (12.7 min) 5.44 11.0 1,598
vLLM 140 1,100 s (18.3 min) 7.85 7.6 1,918

Read these numbers with their limits:

  • One request at a time. Each run sent one case, then waited for the answer. This is not a throughput test. The H100 throughput reference is the #236 receipt on the performance page.
  • Network time is included on vLLM. The vLLM requests went through the RunPod proxy to a pod in data center AP-IN-2. The check run and the signature run shared that pod at the same time.
  • Input tokens differ. vLLM counted more input tokens per case than llama.cpp: 1,918 against 1,598. The cause was not checked.

The H100 pod cost about $2.24 in total (#316 pod record). That cost is shared with the signature run (#319) and the wording check (#309). It is not the cost of this run alone.

Pins

Pin llama.cpp vLLM
Generator seed 0 0
Register rows, variants per row 20, 7 20, 7
Pillow 12.3.0 12.3.0
Slice SHA-256 (first 12) 64a9621a2d55 64a9621a2d55
Code baseline commit 38cada3 ce63ddd
Wall time 762 s 1100 s