Every prediction is written down before kickoff and graded afterwards. The first prediction stands — nothing is revised once a result is known.
Is it honest?
Accuracy can be gamed by only predicting easy matches. Calibration cannot: to be calibrated the model has to be right about its own uncertainty, so its 35% calls must fail about 65% of the time. Each row groups predictions by how confident the model was, and compares what it claimed against what happened.
- 30-40%75 callssaid36.8%landed42.7%
- 40-50%85 callssaid44.2%landed41.2%
- 50-60%56 callssaid54.7%landed57.1%
- 60-70%28 callssaid65%landed78.6%
- 70-80%6 callssaid74.5%landed100%
Grey is what the model claimed, blue is what happened. Landing at or slightly above the claim means it is honest and a little underconfident — the safe direction to be wrong. Buckets with few calls will move a lot; treat small samples as noise.
By competition
Open a competition for every prediction it has made, graded match by match.