Research · Post-mortem
Calibration Before Conviction
A post-mortem on July. Our probability was wrong. More importantly, it was badly calibrated.
By Audit ✦ ✦4 minutes’ reading
In Brief
- In July the ensemble assigned 71% to a liquidity expansion and 74% to crowded systematic positioning. Both resolved no. July's mean Brier score was 0.402, against 0.316 for the external probabilities.
- Being wrong at 71% is expected roughly three times in ten. The problem is that across June and July, forecasts we issued at 60% or above resolved yes only half the time.
- We traced the failure to correlated errors between agents that appeared independent. Four changes followed. September and the months after will show whether they were sufficient.
This is a note about a failure. We publish it in the same format and with the same prominence as everything else, because a research process that only reports its successes is not a research process.
§ IWhat happened
The July outlook contained two large admissible gaps. We assigned 71% to an expansion in aggregate dollar liquidity, against an external 52%. We assigned 74% to systematic equity positioning exceeding the 75th percentile of its range, against an external 49%. Both questions resolved no. At the same time, we had been more complacent than the market about volatility and credit stress, assigning 22% and 12% to events that both occurred.
| July question | External | NDI | Outcome | NDI score | Ext. score |
|---|---|---|---|---|---|
| Liquidity expands | 52% | 71% | NO | 0.504 | 0.270 |
| Positioning above 75th pct. | 49% | 74% | NO | 0.548 | 0.240 |
| Realized vol exceeds implied | 28% | 22% | YES | 0.608 | 0.518 |
| HY spread widens > 40bp | 19% | 12% | YES | 0.774 | 0.656 |
| All eight July questions | 0.402 | 0.316 |
§ IIWrong is not the problem
A well-calibrated forecaster who says 70% should be wrong about three times in ten. Any single wrong forecast is therefore uninformative about quality. What matters is whether, across many forecasts, events assigned a given probability occur with roughly that frequency.
They did not. Across June and July, we issued six forecasts at 60% or above, with an average stated probability of 70%. Three resolved yes. The reliability diagram below shows the shape of the problem: in the upper range, the ensemble was consistently more confident than the world justified.
- NDI
- External
A standard decomposition of the Brier score separates reliability (how far stated probabilities are from observed frequencies; lower is better) from resolution (how well forecasts separate events that happen from events that do not; higher is better).
| June–July | Reliability ↓ | Resolution ↑ | Uncertainty | Brier |
|---|---|---|---|---|
| NDI | 0.042 | 0.018 | 0.234 | 0.280 |
| External | 0.019 | 0.006 | 0.234 | 0.256 |
The numbers say something more specific than we were wrong. Our resolution was three times higher than the external reference: the ensemble did a better job of distinguishing likely events from unlikely ones. Our reliability was more than twice as bad: having made the distinction, it expressed it with far too much confidence. The information was there. The probabilities were inflated.
§ IIIWhy
The Auditor's review traced the over-confidence to a single cause. The structural Forecaster, the flow-based Forecaster and the Historian had each, through different routes, come to rely on the same Treasury funding series as a primary input. Their outputs looked independent — different methods, different horizons — but their errors were correlated at 0.84 on liquidity-sensitive questions over the preceding quarter. When that series turned, all three were wrong together, and the extremising step in aggregation amplified their shared mistake as if it were independent confirmation.
In other words, the system did exactly what When Agents Disagree warns against. The monitoring that should have caught it was in place, but it measured correlation of forecasts rather than correlation of errors, and the forecasts had diverged enough to pass.
§ IVWhat changed
- Diversity is now measured on rolling error correlation, with shared-input detection at the feature level. Agents sharing a primary input above a threshold are pooled as a single forecaster.
- The extremising parameter α is now conditional on measured diversity. When diversity falls, α falls towards 1, and pooled forecasts stop being pushed away from 50%.
- Questions with correlated resolution, such as liquidity and positioning in July, are sized as a single exposure rather than two.
- Energy-sector models, which also produced the worst single forecast of June, are down-weighted pending a rebuild.
These changes make the system less confident. In August, the first month under the new rules, it scored 0.193 against an external 0.247. That is encouraging and almost meaningless: ten questions in one volatile month cannot validate a change. We will report reliability separately from score in every subsequent outlook until the sample is large enough to say something.
“Conviction is a quantity a system earns from its record. It cannot be inferred from the strength of its arguments.”