Two-image signature comparison¶
Kind: explanation.
This page is for an engineer who reads the signature-match receipts. It explains what typevet measured with two signature images per call. It names the data and pins, and it shows where the evidence stops. It is not a how-to. The CEDAR loader reference names the files, the cache and the request shape.
The owner approved CEDAR for one experiment: how typed two-image judgments behave. There is no customer use or support.
What was measured¶
Each pair is one typed judgment. The call carries both images, in order. Image 1 is a genuine signature and image 2 is the questioned signature. The call asks three questions:
| Question id | Type | Answer |
|---|---|---|
same_writer |
Noul |
Probability that the same person wrote both signatures |
verdict |
Choice |
same_writer, different_writer, skilled_forgery_suspected or cannot_tell |
image_quality |
Score |
0 to 4, how clearly image 2 shows the signature |
The data is the CEDAR offline signature set: 55 writers, each with 24 genuine signatures and 24 skilled forgeries. The loader fetches the archive from CEDAR and checks a pinned SHA-256 digest. The slice holds 180 pairs, 60 of each kind, drawn with seed 0:
| Pair kind | Image 2 |
|---|---|
genuine_genuine |
Another genuine signature by the same writer |
genuine_skilled |
A skilled forgery of the same writer's signature |
genuine_random |
A genuine signature by another writer |
A "random forgery" is a genuine signature by a different writer. Nobody tried to copy the reference. The design is in the #304 design comment.
The two runs¶
Each backend ran the slice once. There are no repeats.
| Topic | llama.cpp | vLLM |
|---|---|---|
| Receipt | signature_match_llama_cpp.json | signature_match_vllm.json |
| Server | Local build b11277-eae11d221 |
Stock vLLM v0.30.0, one H100 GPU on RunPod |
| Model | Alias gemma-4-31b-kv9-q4km-mm: Gemma 4 31B, Q2_K GGUF |
google/gemma-4-31B-it, revision 842da379, BF16 |
| Context | n_ctx 4096 tokens |
max_model_len 8192 tokens |
| Pairs | 180, no backend failure | 180, no backend failure |
Same-writer accuracy (verdict) |
0.744 | 0.761 |
ROC-AUC of same_writer, all pairs |
0.926 | 0.913 |
| ROC-AUC, genuine against skilled | 0.853 | 0.827 |
| ROC-AUC, genuine against random | 1.000 | 1.000 |
| ECE (10 bins) | 0.319 | 0.307 |
Skilled false accept, same_writer ≥ 0.5 |
0.983 (59 of 60) | 0.900 (54 of 60) |
Skilled false accept, verdict same_writer |
0.650 (39 of 60) | 0.650 (39 of 60) |
Noul–Choice agreement |
0.85 (153 of 180) | 0.90 (162 of 180) |
cannot_tell rate |
0.0 | 0.0 |
| Mean latency | 4.31 s per pair | 6.53 s per pair |
Same-writer accuracy counts skilled_forgery_suspected and different_writer
both as "different writer". A skilled forgery is a different-writer pair.
The llama.cpp receipt pins the model alias, not the file type or a GGUF hash.
The q4km in the alias does not describe the decoder weights. The local router
reported the file type Q2_K - Medium for this alias
(#319 correction).
The verdicts by pair kind:
| Pair kind | Backend | same_writer |
skilled_forgery_suspected |
different_writer |
cannot_tell |
|---|---|---|---|---|---|
genuine_genuine |
llama.cpp | 53 | 5 | 2 | 0 |
genuine_genuine |
vLLM | 56 | 4 | 0 | 0 |
genuine_skilled |
llama.cpp | 39 | 16 | 5 | 0 |
genuine_skilled |
vLLM | 39 | 14 | 7 | 0 |
genuine_random |
llama.cpp | 0 | 0 | 60 | 0 |
genuine_random |
vLLM | 0 | 0 | 60 | 0 |
The two backends ran different weights: a Q2_K GGUF on llama.cpp and BF16
on vLLM. Do not credit a gap between the columns to the backend. The weights
changed as well, and one run each cannot separate the two causes.
What the answers show¶
On this slice, every random forgery was rejected. All 120 random pairs, 60 per backend, got
the verdict different_writer. Every one had a same_writer probability below
0.01. The ROC-AUC against random pairs is 1.000 on both backends.
Skilled forgeries are mostly accepted. On both backends, 39 of 60 skilled
forgeries got the verdict same_writer. The Noul accepted even more of them:
59 of 60 on llama.cpp and 54 of 60 on vLLM.
The Choice catches about a third of skilled forgeries. It rejected 21 of 60 on
each backend. The Noul almost never does: at the 0.5 threshold it rejected 1
of 60 on llama.cpp and 6 of 60 on vLLM. Every skilled forgery that the Noul
rejected, the Choice rejected too.
The two answers disagree most on skilled forgeries. On llama.cpp, 20 of the 27
disagreements are skilled pairs. On vLLM, 15 of the 18 are. In those pairs the
verdict says "different writer" while the same_writer probability stays
high. Of the 21 skilled pairs that the Choice rejected, 16 on llama.cpp and
14 on vLLM still had a probability of 0.9 or more.
The reliability bins show why the ECE is high. On llama.cpp, the top bin
(0.9 to 1.0) held 113 pairs with a mean probability of 0.999. Only 51% of them
were same-writer pairs: 58 genuine pairs and 55 skilled forgeries. On vLLM, the
top bin held 112 pairs with a mean of 0.999, and 53% were same-writer pairs.
The Noul gives a skilled forgery the same near-certain value as a genuine
pair.
The model never chose cannot_tell in verdict. On this slice, the option to
abstain was never used, so the slice says nothing about when the model
abstains.
The image_quality score does not separate the kinds. Every genuine and
skilled image 2 got level 4 on both backends. Among random pairs, one image got
level 0 on llama.cpp, and two images got level 3 on vLLM.
What the Noul alone would miss¶
A caller that reads only the same_writer probability and accepts at 0.5 would
accept 90% to 98% of skilled forgeries on this slice. A higher threshold does
not help much. At 0.9, the Noul still accepted 55 of 60 skilled forgeries on
llama.cpp and 53 of 60 on vLLM.
One reading: the probability behaves like an answer to "do these look like the
same hand?". A skilled forger tries to make that answer yes. The Choice has an
explicit skilled_forgery_suspected option. That option gives the model a
place to say that a copy is a forgery. On skilled forgeries it used that option 16 and 14 times. It
is still wrong on 39 of 60 skilled forgeries. This reading is a hypothesis. The
runs did not test it.
The Noul ranks some skilled forgeries below genuine pairs, with a ROC-AUC of
0.853 and 0.827. But its values for both kinds sit near 1.0, so no single
threshold separates them. At a threshold, the Noul separates different
writers and not much else. Do not read one number as a verdict on a signature.
Read both answers. Treat disagreement between them as a sign that the pair is
hard, not as a detector.
What the probability means¶
The same_writer probability is model confidence. It is not a match
percentage and not a forensic score. A value of 0.999 does not mean that 999 of
1000 such pairs share a writer. On this slice, about half of such pairs did not.
The ECE values of 0.319 and 0.307 hold for 180 CEDAR pairs only. They do not show how the probability behaves on other signatures, other writers or other pins.
What this is not¶
This is not a fraud control and not a document examination. The evaluation has no trained examiner, no original document, no pen-pressure or stroke-order data and no chain of custody. Do not use these results to accept or refuse a signature, a cheque, a contract or a person.
Limits of the data¶
- Real writers. CEDAR signatures come from real people, not from a generator. The results describe these 55 writers and their forgers only.
- Writer population and scans. This page did not verify the writers' languages, scripts, ages or regions from CEDAR documentation. It did not verify the scanner or the collection period either. Treat these as unknown. Do not assume the results transfer to other scripts or other scan conditions.
- Skilled forgers. This page did not verify who made the forgeries or how much practice they had. The forgery skill level is unknown.
- Possible memorization. CEDAR is public. Its images may be in the model's training data. This was not tested.
- One slice, one run. Each backend ran 180 pairs once. The slice gives no variance estimate. Five writers appear twice in each kind. Run-to-run spread was measured on the checks slice only. Five llama.cpp repeats there were identical (#357).
Latency and cost¶
One item is one pair: one typed judgment that asks all its questions. The
receipts give wall_seconds and, per pair, latency_seconds and
input_tokens.
| Backend | Items | Wall time | Mean seconds per item | Items per minute | Mean input tokens |
|---|---|---|---|---|---|
| llama.cpp | 180 | 776 s (12.9 min) | 4.31 | 13.9 | 1,063 |
| vLLM | 180 | 1,176 s (19.6 min) | 6.53 | 9.2 | 2,045 |
Read these numbers with their limits:
- One request at a time. Each run sent one pair, then waited for the answer. This is not a throughput test. The H100 throughput reference is the #236 receipt on the performance page.
- Network time is included on vLLM. The vLLM requests went through the RunPod proxy to a pod in data center AP-IN-2. The signature run and the check run shared that pod at the same time.
- Input tokens differ. vLLM counted more input tokens per pair than llama.cpp: 2,045 against 1,063. The cause was not checked.
The H100 pod cost about $2.24 in total (#316 pod record). That cost is shared with the check run (#316) and the wording check (#309). It is not the cost of this run alone.
Data terms and storage¶
CEDAR publishes no licence. The CEDAR page links the archive under "Published Data Sets" and needs no sign-in. It states no terms of use. typevet uses the data for research only.
The repository stores no signature bytes. Receipts hold pair ids and typed answers only. The loader fetches the archive from CEDAR at run time into a cache outside the repository. It refuses a file whose SHA-256 differs from the pinned value. The loader reference states the policy for fixtures and tests.
Signatures and biometric law¶
Illinois BIPA is the Illinois Biometric Information Privacy Act. Its definition of "biometric identifier" (740 ILCS 14/10) excludes "writing samples" and "written signatures".
GDPR can still apply. Article 4(14) and Recital 51 set the test. An image is biometric data when specific technical means process it to identify a person. A same-writer judgment can be such processing. Article 9 can then apply to the signature image.
This page is not legal advice. It states why the repository keeps no signature images, not whether a use is lawful.