Post-hoc calibration of Noul probabilities¶
Kind: explanation.
This page is for an engineer who reads the typevet receipts. It explains one offline study from #343. The study asks one question. Can a small map, fitted after the model answers, fix the poor calibration in the committed receipts? If it cannot, the owner must consider fine-tuning.
The study made no model call. It read receipts that typevet had already
committed. The fitters are plain Python in typevet_evals.calibration. The
full result is in
post_hoc_receipt.json.
Why calibration matters here¶
A Noul answer gives a probability. Calibration asks whether that probability matches the observed frequency. Some receipts show a large gap. The DIFrauD seed wording on Gemma has an ECE of 0.17 to 0.19 on three of its five receipts. The CEDAR signature runs have an ECE of about 0.31.
A calibration map does not change the model. It changes only the number that the caller reads. It is cheap to fit and cheap to apply.
The plan, fixed before any fit¶
The #343 contract set the plan before any fit. The study did not change it after the results.
Series. Each series is one probability column of one receipt. There are 16 series:
- DIFrauD, five held-out receipts, columns
seedandevolved(158 rows). - Generated checks, llama.cpp and vLLM (140 rows). The probability is the largest verdict probability. The label is whether the verdict was correct.
- CEDAR signatures, llama.cpp and vLLM (180 pairs).
- LFW faces, llama.cpp and vLLM (200 pairs).
Split. A row goes to the calibration half when
int(sha256(id).hexdigest(), 16) is even. Otherwise it goes to the
evaluation half. The fitters see only the calibration half. All metrics use
the evaluation half.
Methods. Each fitter clips probabilities to [1e-6, 1 - 1e-6] first.
| Method | Map | Fit |
|---|---|---|
| Temperature | sigmoid(logit(p) / T) |
Golden-section search on log T in [-3, 3] |
| Platt | sigmoid(a * logit(p) + b) |
Newton steps with step halving |
| Isotonic | Non-decreasing steps | Pool adjacent violators; value of the nearest lower knot |
Metrics. ECE uses 10 equal-width bins. It reuses the face-match ECE function. The study also reports the Brier score and the accuracy at 0.5. A bootstrap of 1,000 resamples gives a 95% interval of the ECE change. The interval is context only.
Rule. A method meets the rule on a series when three things are true:
- ECE drops by 0.03 or more.
- The Brier score drops.
- Accuracy drops by 0.01 or less.
A series with an evaluation-half ECE below 0.05 before the fit does not count toward the decision.
Decision. The pre-registration named six series with an ECE of 0.10 or more. Two more series pass that bar under the same ECE convention (the #309 llama.cpp seed wording at 0.113 and the llama.cpp checks at 0.165); they were not named, and adding them does not change the outcome, because isotonic leaves both above 0.05. Post-hoc calibration is sufficient when one method meets the rule on 4 or more of the six named series. That method must also leave the evaluation-half ECE below 0.05 on each of those 4.
Result on the six decision series¶
Each cell is before → after on the evaluation half. "Rule" is the pre-registered rule. "Counts" means the rule is met and the ECE after is below 0.05.
| Series | n cal / eval | Method | ECE | Brier | Accuracy | Rule | Counts |
|---|---|---|---|---|---|---|---|
| DIFrauD #252 Gemma llama.cpp, seed | 68 / 90 | Temperature | 0.198 → 0.200 | 0.173 → 0.148 | 0.811 → 0.811 | no | no |
| Platt | 0.198 → 0.071 | 0.173 → 0.074 | 0.811 → 0.889 | yes | no | ||
| Isotonic | 0.198 → 0.057 | 0.173 → 0.048 | 0.811 → 0.956 | yes | no | ||
| DIFrauD #252 Gemma vLLM, seed | 68 / 90 | Temperature | 0.214 → 0.202 | 0.209 → 0.153 | 0.778 → 0.778 | no | no |
| Platt | 0.214 → 0.020 | 0.209 → 0.083 | 0.778 → 0.889 | yes | yes | ||
| Isotonic | 0.214 → 0.020 | 0.209 → 0.082 | 0.778 → 0.889 | yes | yes | ||
| DIFrauD #309 vLLM, seed | 68 / 90 | Temperature | 0.215 → 0.202 | 0.210 → 0.154 | 0.778 → 0.778 | no | no |
| Platt | 0.215 → 0.024 | 0.210 → 0.083 | 0.778 → 0.889 | yes | yes | ||
| Isotonic | 0.215 → 0.012 | 0.210 → 0.081 | 0.778 → 0.889 | yes | yes | ||
| DIFrauD #309 vLLM, evolved | 68 / 90 | Temperature | 0.187 → 0.188 | 0.184 → 0.139 | 0.811 → 0.811 | no | no |
| Platt | 0.187 → 0.020 | 0.184 → 0.083 | 0.811 → 0.889 | yes | yes | ||
| Isotonic | 0.187 → 0.008 | 0.184 → 0.087 | 0.811 → 0.878 | yes | yes | ||
| CEDAR signatures, llama.cpp | 91 / 89 | Temperature | 0.306 → 0.236 | 0.302 → 0.202 | 0.674 → 0.674 | yes | no |
| Platt | 0.306 → 0.126 | 0.302 → 0.115 | 0.674 → 0.843 | yes | no | ||
| Isotonic | 0.306 → 0.035 | 0.302 → 0.096 | 0.674 → 0.876 | yes | yes | ||
| CEDAR signatures, vLLM | 91 / 89 | Temperature | 0.282 → 0.219 | 0.280 → 0.192 | 0.719 → 0.719 | yes | no |
| Platt | 0.282 → 0.099 | 0.280 → 0.128 | 0.719 → 0.831 | yes | no | ||
| Isotonic | 0.282 → 0.036 | 0.280 → 0.117 | 0.719 → 0.831 | yes | yes |
Isotonic counts on 5 of the 6 series. Platt counts on 3. Temperature counts on none. So by the pre-registered rule, post-hoc calibration is sufficient. The study does not recommend a fine-tuning issue.
The other ten series¶
These series are reported but are not part of the decision. Each cell is the evaluation-half ECE; the ECE before the fit is in the second column.
| Series | ECE before | Temperature | Platt | Isotonic |
|---|---|---|---|---|
| DIFrauD #252 Gemma llama.cpp, evolved | 0.033 (excluded) | 0.052 | 0.046 | 0.017 |
| DIFrauD #252 Gemma vLLM, evolved | 0.034 (excluded) | 0.047 | 0.047 | 0.021 |
| DIFrauD #252 Jev, seed | 0.103 | 0.052 (rule met) | 0.089 | 0.089 |
| DIFrauD #252 Jev, evolved | 0.076 | 0.029 (rule met) | 0.022 (rule met) | 0.033 |
| DIFrauD #309 llama.cpp, seed | 0.141 | 0.134 | 0.057 (rule met) | 0.061 (rule met) |
| DIFrauD #309 llama.cpp, evolved | 0.107 | 0.110 | 0.067 (rule met) | 0.076 |
| Checks, llama.cpp | 0.189 | 0.129 (rule met) | 0.110 (rule met) | 0.059 (rule met) |
| Checks, vLLM | 0.093 | 0.069 | 0.085 | 0.087 |
| LFW faces, llama.cpp | 0.080 | 0.074 | 0.051 | 0.054 |
| LFW faces, vLLM | 0.028 (excluded) | 0.030 | 0.017 | 0.014 |
Why temperature fails¶
Temperature scaling has one parameter. It can only pull probabilities toward 0.5 or push them away. It cannot move the point where the map crosses 0.5.
The Gemma DIFrauD seed runs need that move. Their fitted temperatures are about 4 to 6.4, but ECE barely changes. Platt adds an intercept, and isotonic has no fixed shape. Both move the crossing point.
This is also why accuracy rises after Platt and isotonic. A map that moves the crossing point changes which rows read positive at 0.5.
Transfer between tasks¶
The study also fitted a map on one task and applied it to another. It used DIFrauD #309 llama.cpp seed and CEDAR signatures on llama.cpp. No method met the rule in either direction. The signature map made DIFrauD worse: ECE rose from 0.141 to between 0.237 and 0.325.
So each task and backend needs its own map. One shared map does not work.
Use a calibration map¶
#352 lets a caller
apply a fitted map at judgment time. The map is one JSON file per task and
question. Its schema id is typevet.calibration_map/1.
load_calibration_map(path, sha256=...) reads the file. It refuses the file
when the sha256 of its bytes differs from the caller's digest.
CalibratedJudgment(inner, maps, task_id=..., backend=...) wraps a judgment
port. The caller declares the task and the backend once. The wrapper refuses
a map for another task or backend. It refuses a response from another model.
Maps do not transfer, so these checks fail closed.
The answer carries the calibrated Noul. JudgmentResponse.calibration
records the raw value, the calibrated value, the method and the map digest.
Each refusal is a CalibrationMapError with no values in its message.
A map with a top-level "levels": n field is a Score map. It is one pooled
map for all the level probabilities. The wrapper maps each level and rescales
the levels to sum to 1. Then it sets the score to the expected level and the
confidence to the largest level probability. For a Score answer, the record
holds the raw and calibrated score and both sets of level probabilities. No
Score map has a held-out result: #343 fitted no Score series.
The recorded after metrics describe the pooled fit before the rescale, not
the rescaled levels a caller reads. A pooled map cannot correct a bias at one
level.
A map for a Choice question is refused before the inner call. A map whose
kind or level count does not match the question is also refused before the
call. This applies to typed questions and to wire dictionaries. A wire Score
without a criteria list is checked only after the call, by the answer check.
The model check runs for every map on every response. A mapped question that the call
does not ask is skipped. The map stores the receipt digest in lower case. A
map whose evaluation failed the rule is recorded, not refused.
The typevet_evals.calibration_artifact module writes the map. It fits one
method on the calibration half and measures the evaluation half. It records
the sha256 of the source receipt. It returns the sha256 of the bytes it wrote,
and the caller pins that digest. The steps are in
Calibrate a task with your own receipts.
Limits¶
- Small halves. Each evaluation half has 70 to 98 rows. A 10-bin ECE on so few rows is noisy. The bootstrap intervals are wide.
- One split. The study used one fixed split. A different split can give different numbers.
- Isotonic gives hard values. Some isotonic steps are exactly 0 or 1. Such a value claims certainty. Log loss is infinite when such a row is wrong.
- Boundary fits. The Jev evolved temperature fit reached the search
bound,
T = exp(-3). The Jev seed Platt fit gave a very large slope. On those Jev calibration halves, the probabilities almost separate the labels. - Same distribution. Each map was fitted and tested on rows from one receipt. A map can fail when the data or the wording changes.
- No live check. No model call used a fitted map. This page says nothing about a deployed map.
Receipts¶
post_hoc_receipt.json: every series, method, fitted parameter, metric and interval, the transfer rows, the decision and the sha256 of each source receipt.- The source receipts are under
evals/fixtures/difraud/receipts/,evals/fixtures/checks/receipts/,evals/fixtures/cedar/receipts/andevals/fixtures/lfw/receipts/.