Calibration record
Every number below is computed weekly from live resolutions only, with bootstrap confidence intervals (1000 resamples). When an interval crosses zero we say so; nothing here is a claim until it separates.
snapshot 2026-09-06 05:05 UTC · 143 resolved · 267 scored forecasts · 543 shadow scores · 54 abstained
Brier by verdict class
the core experiment: does the verdict class carry information about error?| Verdict | n | Brier | Bootstrap CI |
|---|---|---|---|
| CAUTION | 33 | 0.234 | 95% CI [0.175, 0.299] |
| CONSTRAINED | 234 | 0.275 | 95% CI [0.246, 0.306] |
ABSTAIN rows never publish a probability, so they cannot appear here; their value is measured below, through the shadow.
Value of abstention
an ungated shadow model forecasts every question; we compare its error on the questions we abstained fromshadow on published set
0.246
n=333
shadow on abstained set
0.243
n=210
abstention value (diff)
-0.003
95% CI [-0.040, 0.034]
Positive means the questions our gate refused were genuinely harder, even for an ungated model. The interval currently crosses zero: direction only, not yet a claim.
Layer vs shadow on the published set
same questions, same evidence pool; the only difference is the decision layerRisk-coverage
verdict classes are the natural thresholds of a selective forecaster| Publish up to | Coverage | Risk (Brier) | n |
|---|---|---|---|
| CONSTRAINED | 72.9% | 0.275 | 234 |
| CAUTION | 83.2% | 0.270 | 267 |
naive AURC 0.229 · with few points this is descriptive, not comparative