Methodology
Certiora exists to demonstrate one claim: a forecasting system is more useful when it can tell you, in advance and on the record, which questions it should not answer. Selective prediction is the product; the probabilities are the by-product.
The pipeline
- 1 · Pre-registration. Every question ships with locked resolution criteria, edge cases, an official resolution source, and a resolution window. The criteria hash is frozen before any forecast is made.
- 2 · Evidence. Independent AI observers gather dated, factual claims from the open web. Every claim passes seven quality gates; sources carry tier weights and anything that fails is quarantined and never reaches the forecaster. When the evidence pool leans one way, a dedicated counter-perspective pass searches for the other side.
- 3 · Verdict gate. A physics-inspired knowledge engine measures whether the rule base for this topic is coherent enough to act. Its verdict, one of four levels below, decides whether the ensemble probability is published or sealed.
- 4 · Resolution. Outcomes are read exclusively from official statistical APIs: the national statistics office and central bank on the Turkish side, the federal statistical agencies on the US and euro side. First print rules: revisions do not change an outcome. No press summaries, no secondary sources.
- 5 · Chain. Every forecast and every resolution is appended to a hash chain. Editing history would break the chain publicly.
The four verdicts
- DECIDE
The knowledge base and the evidence both support a confident, autonomous forecast. Probability is published.
- CONSTRAINED
A forecast is published, but the internal state flags real constraints: partial rule coverage, one-sided evidence, or an immature topic.
- CAUTION
A forecast is published under supervision-grade caution. Treat the probability as weakly held.
- ABSTAIN
The gate refuses. Admitted evidence is too thin or the topic knowledge base is insufficient. The probability exists internally but is sealed; the abstention itself is recorded and scored as coverage, never as accuracy.
Uncertainty scale
how to read a published probabilityScoring and honesty rules
- · Published forecasts are scored with the Brier score against the resolved outcome. Sealed forecasts are never scored for accuracy; they count toward coverage statistics only.
- · Backtests run the full pipeline against already-published data to prove the machinery. They are labeled everywhere and never mixed into live calibration.
- · A shadow forecaster sees the evidence without the gate. Its score is tracked separately so the value of the gate itself is measurable.
- · The pre-registration document, including kill criteria for the whole experiment at n ≥ 150 resolved live forecasts, is hashed on the chain. If the calibration fails those criteria, that failure will be public.