Two-image face matching¶
Kind: explanation.
This page is for an engineer who reads the face-match receipts. It explains what typevet measured with two face images per call. It names the data and pins, and it shows where the evidence stops. It is not a how-to. The LFW loader reference names the files, the cache and the request shape.
What was measured¶
Each pair is one typed judgment. The call carries both images, in order, and asks three questions:
| Question id | Type | Answer |
|---|---|---|
same_person |
Noul |
Probability that the two faces are the same person |
verdict |
Choice |
same_person, different_person or cannot_tell |
face_visibility |
Score |
0 to 4, how clearly image 2 shows the face |
The data is LFW View 2 (pairs.txt, 10 folds). The loader fetches the archive
from the scikit-learn figshare mirror and checks a pinned SHA-256 digest. The
slice holds 200 pairs: 100 same-person pairs and 100 different-person pairs,
20 per fold, drawn with seed 0. The design is in the
#292 design comment.
The two runs¶
Each backend ran the slice once. There are no repeats.
| Topic | llama.cpp | vLLM |
|---|---|---|
| Receipt | face_match_llama_cpp_receipt.json | face_match_vllm_receipt.json |
| Server | Local build b11243-fc07d781e |
Stock vLLM v0.30.0, one H100 80 GB on RunPod |
| Model | Alias gemma-4-31b-kv9-q4km-mm: Gemma 4 31B, Q2_K GGUF |
google/gemma-4-31B-it, revision 842da379, BF16 |
| Context | n_ctx 4096 tokens |
max_model_len 8192 tokens |
| Pairs | 200 | 200 |
verdict accuracy |
0.955 (95 of 100 same, 96 of 100 different) | 0.975 (96 of 100 same, 99 of 100 different) |
| ROC-AUC | 0.9956 | 0.9993 |
| ECE | 0.0617 | 0.0236 |
cannot_tell rate |
0.0 | 0.0 |
face_visibility, same pairs |
level 3: 5, level 4: 95 | level 3: 1, level 4: 99 |
face_visibility, different pairs |
level 3: 9, level 4: 91 | level 3: 3, level 4: 97 |
| Mean latency | 4.18 s per pair | 5.27 s per pair, through a remote proxy |
Two earlier llama.cpp attempts stopped on keep-alive disconnects (#305). The #305 fix now retries such a close once, so later runs keep keep-alive on. The vLLM pod cost about 1.69 USD, shared with #231.
The two backends ran different weights: a Q2_K GGUF on llama.cpp and BF16 on
vLLM. Do not credit the accuracy gap to the backend. The weights changed as
well, and one run each cannot separate the two causes.
What the answers show¶
The probabilities cluster at the ends. On llama.cpp, 84 pairs had a
same-person probability below 0.1 (mean 0.0012), and 107 pairs had one of 0.9
or more (mean 0.996). Of those 107 pairs, 8 were different-person pairs.
The five same-person pairs that llama.cpp missed got the verdict
different_person with a verdict probability from 0.59 to 0.88.
Their same-person probability was 0.86 to 0.99.
The two answers disagree on those pairs.
The model never chose cannot_tell in verdict. On this slice, the option to
abstain was never used, so the slice says nothing about when the model
abstains.
The face_visibility score does not separate the classes. Almost every image
2 got level 4 in both classes. LFW images are mostly news photographs of public figures, so the faces are
almost always visible. The field needs harder images before it can show
anything.
What the probability means¶
The same-person probability is model confidence. It is not a calibrated match percentage. A value of 0.996 does not mean that 996 of 1000 such pairs show the same person.
The ECE values are low on this slice. That result holds for 200 LFW pairs only. It does not show that the probability is calibrated on other images, other people or other pins.
What this is not¶
This is not an identity-verification control and not a KYC control. The evaluation has no ID document, no liveness check and no spoof check. Do not use these results to admit or refuse a person.
Limits of the data¶
- Population. LFW over-represents light-skinned, male, Western public figures photographed by news media. The results do not transfer to other populations.
- Image kind. The results do not transfer to selfies, ID-document photos or poor lighting.
- Possible memorization. The people in LFW are public figures. They may be in the model's training data. The model could recognize a known face instead of comparing two faces. This was not tested.
- One slice, one run. Each backend ran 200 pairs once. The slice gives no variance estimate. Run-to-run spread was measured on the checks slice only. Five llama.cpp repeats there were identical (#357).
Face images are biometric data¶
A face image is biometric data under laws such as Illinois BIPA and GDPR Article 9. BIPA is the Illinois Biometric Information Privacy Act.
The repository stores no face bytes. Receipts hold pair ids and typed answers only. The loader fetches the images at run time into a cache outside the repository. LFW has no formal licence. The photographers keep the image copyright, and the dataset is for research use. The loader reference states the policy for fixtures and tests.