Calibration review, second quarter
Forecasting. — 2026-07-06. — 2 tracings. — Provisional record.
Ninety-one resolved questions. Overconfident in the eighty-to-ninety band, and the reason is not what I assumed.
Ninety-one questions resolved this quarter. The reliability curve is close to the diagonal below seventy per cent and drifts above it after that: when I said eighty-five I was right about seventy-eight per cent of the time.
The comfortable story is that this is ordinary overconfidence and the fix is to shade high estimates downward. I tried that on last year’s data and it made the Brier score worse, because the shading also flattened the cases where I was right to be sure.
The less comfortable story, which the question-level breakdown supports, is that the miss is concentrated almost entirely in questions where I had a stake in the answer. Strip those out and the curve straightens. Keep only those and it bends twice as hard.
What I am changing
Not the numbers. The workflow. Any question where I would be pleased by one of the outcomes now gets written down as a separate class, forecast last, and reviewed against the others rather than against the diagonal. If the split holds for another two quarters it is a real effect and not a slice of noise I went looking for.
Tracings
Drawn from this record.
- N102On being wrong in publicNotes
The review only works because the numbers were public before they resolved. A private forecast is a memory, and memory edits itself.
- F102Forecasting my own projectsForecasting
Traced from
Records that point here.