The experiment
Hypothesis: a public rating model with no inside information produces match and title forecasts at least as accurate as a naive baseline, and close to the betting market. Forecasts are frozen before the first ball.
Lock status
Not locked yet. Ratings, fixtures and the knockout format are verified. Forecasts lock after a final rating refresh following the December international window, and before 7 January 2027.
Scoring
- Each group match: Brier score of the locked win/draw/loss forecast. 0 is perfect, 2 is confidently wrong.
- Baseline: one-third each outcome, which always scores 0.667.
- Title: log loss of the locked model and market title probabilities for the eventual champion.
- Claim strength: 36 group matches is a small sample. A gap below about 0.03 in mean Brier is noise, not signal.
0
Matches scored
–
Model mean Brier
–
Baseline mean Brier
Backtest: 2019 and 2023 Asian Cups
The same model, with ratings frozen the day before each tournament, scored on 72 group matches and 30 knockout ties. Lower is better.
| Forecast | Group Brier | Group log loss | Knockout Brier |
|---|---|---|---|
| This model (settings in use) | 0.457 | 0.782 | 0.234 |
| Best fit to both tournaments | 0.458 | 0.784 | 0.215 |
| Coin flip / one third each | 0.667 | 1.099 | 0.250 |
Favourite calibration (group matches)
| Given 35-50% | predicted 44% | won 64% | n=11 |
| Given 50-65% | predicted 57% | won 50% | n=24 |
| Given 65-80% | predicted 73% | won 70% | n=20 |
| Given 80-100% | predicted 89% | won 94% | n=17 |
What it shows
- Far better than guessing on group matches. Knockouts are close to a coin flip, as expected for tight ties.
- Settings tuned on one tournament did worse on the other’s group matches, so the defaults stay (goal rate 1.28, host +100, uncertainty 60).
- No sign that favourites are overrated. Qatar won both editions without being a top-five pre-tournament rating.
- Ratings are rebuilt from open results data and mapped to the eloratings.net scale (fit error 28 points).