Calibration record
early period · n<30Every number below is computed weekly from live resolutions only, with bootstrap confidence intervals (1000 resamples). When an interval crosses zero we say so; nothing here is a claim until it separates.
snapshot 2026-07-27 15:24 UTC · 27 resolved · 45 scored forecasts · 84 shadow scores · 6 abstained
Brier by verdict class
the core experiment: does the verdict class carry information about error?| Verdict | n | Brier | Bootstrap CI |
|---|---|---|---|
| CONSTRAINED | 45 | 0.236 | 95% CI [0.176, 0.296] |
ABSTAIN rows never publish a probability, so they cannot appear here; their value is measured below, through the shadow.
Value of abstention
an ungated shadow model forecasts every question; we compare its error on the questions we abstained fromshadow on published set
0.248
n=63
shadow on abstained set
0.285
n=21
abstention value (diff)
+0.037
95% CI [-0.088, 0.167]
Positive means the questions our gate refused were genuinely harder, even for an ungated model. The interval currently crosses zero: direction only, not yet a claim.
Layer vs shadow on the published set
same questions, same evidence pool; the only difference is the decision layerRisk-coverage
verdict classes are the natural thresholds of a selective forecaster| Publish up to | Coverage | Risk (Brier) | n |
|---|---|---|---|
| CONSTRAINED | 88.2% | 0.236 | 45 |
| CAUTION | 88.2% | 0.236 | 45 |
naive AURC 0.208 · with few points this is descriptive, not comparative