Skip to content
AI Committee8 min read

How to Read a Forecast Calibration Curve (Ours Says We Are Overconfident)

BlofinX Quantitative Research TeamUpdated 2026-08-25
Key Takeaways
  • Calibration asks a different question from accuracy: when a model says 70%, does it happen about 70% of the time?
  • Measured on 61 forecasts to 25 August 2026, BlofinX's stated confidence averaged 68.2% while the outcome rate was 34.4% — roughly 34 points of overconfidence. The live figure is on the track record page and supersedes this one.
  • The curve is inverted, not merely shifted: accuracy FALLS as stated confidence rises, and of the 6 forecasts stated at 80-90% confidence, none resolved as called.
  • That inversion is the finding worth knowing. A high confidence score, on this evidence, was not a reason to size up.
  • Applying a calibration layer reduced out-of-sample expected calibration error from 0.51 to 0.20, which is why displayed confidence is now adjusted before you see it.

There are two separate questions you can ask of any forecaster. The first is how often they are right. The second, which almost nobody asks, is whether their confidence means anything — when they say they are 80% sure, does it happen 80% of the time? That second question is calibration, and our own answer to it is unflattering enough that we think publishing it is the most useful thing on this site.

Expectancy calculatorTurn a claimed win rate into the payoff it would need to break even, before you trust the confidence.

Accuracy and calibration are not the same thing

A forecaster who says '55% confident' on every call and is right 55% of the time is perfectly calibrated and barely useful. A forecaster who is right 80% of the time but claims 99% confidence is highly accurate and badly calibrated — and the second failure is the one that hurts you, because confidence is what you size a position on.

How to read the diagram

A reliability diagram plots stated confidence on one axis and measured outcome rate on the other, with forecasts grouped into confidence buckets. Perfect calibration is the 45-degree line. Points below the line mean the forecaster claimed more certainty than the outcomes justified; points above mean they were underconfident. What matters is not any single point but the shape: a curve that stays below the line, and especially one that falls as confidence rises, is telling you the confidence score is not carrying information.

What ours looks like, as at 25 August 2026

Across 61 forecasts carrying a stated confidence, the mean claimed confidence was 68.2% and the measured outcome rate was 34.4%. That is an overconfidence gap of about 34 percentage points, an expected calibration error of 0.41 and a Brier score of 0.40. Broken into buckets, the pattern is worse than a flat offset.

Stated confidence   Forecasts   Actually resolved as called
40-50%                     1                         0.0%
50-60%                    11                        54.5%
60-70%                    29                        41.4%
70-80%                    14                        21.4%
80-90%                     6                         0.0%

// Source: GET /api/predictions/calibration, 25 Aug 2026.
// Confidence and outcome moved in OPPOSITE directions.

The inversion is the part that matters

Our most confident forecasts were our worst. Of the six issued at 80-90% stated confidence, none resolved as called, while the least confident bucket was the only one at parity. On this sample, a high confidence score was not a reason to size up — if anything it was a mild warning. We have no comfortable explanation for that, and the sample is small enough that some of it is noise, but it is what the data says and we would rather print it than quietly drop the bucket.

What we changed because of it

Rather than continue displaying raw model confidence, we fit a calibration layer and validated it out of sample. That reduced out-of-sample expected calibration error from 0.51 to 0.20 and the Brier score from 0.42 to 0.18. The confidence figure shown in the terminal is the adjusted one, which is why it is usually lower and less dramatic than the underlying model's own output.

How to use this when judging anyone else

Ask any forecaster for their reliability diagram. Not their accuracy — their calibration. Almost no one in this industry publishes one, and the reason is visible above: it is the chart most likely to embarrass you. A service that will not show you how its confidence has performed is asking you to size positions on a number that has never been checked.

Summary

Calibration measures whether a forecaster's confidence means anything. Ours, measured across 61 forecasts to 25 August 2026, ran about 34 points above outcomes, and the highest-confidence bucket performed worst of all. We publish the curve, we adjust what we display because of it, and we would suggest asking any competing service for the same chart.

Frequently Asked Questions

What is a calibration curve?

A chart comparing stated confidence against measured outcomes. If forecasts made at 70% confidence come true about 70% of the time, the model is calibrated. Deviation below the diagonal means overconfidence.

What does an expected calibration error of 0.34 mean?

It is the average gap between claimed confidence and observed outcome rate, weighted by how many forecasts fall in each bucket. 0.34 is large — it means stated confidence was, on average, about 41 percentage points away from what happened.

Why would a model's most confident predictions be its worst?

On a 61-forecast sample some of this is noise, and the highest bucket holds only 6 forecasts. A plausible mechanism is that agreement among correlated inputs produces high confidence without producing independent evidence. We report the pattern rather than explaining it away.

Is the confidence shown in the terminal the raw model output?

No. It is calibrated before display, which cut out-of-sample calibration error from 0.51 to 0.20. The raw figure was materially overconfident and we do not think it should be what a member sizes a position on.

Related guides

Test Quantitative AI Signals Live

Access 14-Agent AI Consensus, real-time L2 orderbook spoofing radar, and quarter-Kelly position sizing capped at 2% of capital.