命中率不适合评估概率预测。本文解释多分类 Brier 评分与有序概率评分 RPS、校准为何比准确率更重要,并公开线上基线:1,889 场样本外比赛 Brier 0.2195、RPS 0.1114。
研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。
Ask how accurate a football forecast is and someone will answer with a hit-rate: "we got 54% of winners right". This number is close to meaningless. In a typical league season, favourites win around 45–50% of matches, so a coin that always picks the favourite scores similarly — and a model that confidently picks upsets will be punished even when its probabilities are honest. Evaluating a probability forecast requires scoring rules designed for probabilities. This article explains the two we use, and publishes our live baselines.
The Brier score is the mean squared error of a probability forecast. For a football match with outcomes (home, draw, away), the multi-class form sums the squared error across all three:
BS = (p_home − o_home)² + (p_draw − o_draw)² + (p_away − o_away)²
where o is 1 for the observed outcome and 0 otherwise. The score ranges from 0 (perfect) to 2 (worst), and 2/3 is what you get by predicting the empirical base rates blindly. Two properties make it the workhorse: it is a proper scoring rule — the best strategy is to state your true belief — and it decomposes into calibration and refinement terms, which matters next.
Football outcomes are ordered: home win > draw > away win in terms of home advantage. The RPS exploits this ordering by comparing cumulative probability against cumulative outcome. Getting home-vs-away right while shuffling the draw is penalised less than confusing the extremes. Because of this ordering sensitivity, RPS has become the de facto standard in the academic football-forecasting literature, which is why we report it alongside Brier.
A model is calibrated when, among all matches to which it assigns 60% home-win probability, roughly 60% of home teams actually win. Calibration is checkable directly with reliability curves, and it is the property that makes probabilities trustworthy downstream. A model can have an unimpressive hit-rate and still be excellent if it is sharp and calibrated; a model with a flashy hit-rate but poor calibration will mislead anyone who uses its numbers as probabilities. This is why our gate process refuses to ship parameter changes that improve hit-rate while degrading calibration.
We maintain a rolling out-of-sample evaluation of the club model (bivariate Poisson with Dixon-Coles correction, blended with a gradient-boosting estimate of goal expectancy). Current production baselines over the most recent 1,889 evaluated matches:
| Model path | Brier | RPS | Hit-rate | n |
|---|---|---|---|---|
| Pure Dixon-Coles | 0.2227 | 0.1134 | 51.24% | 1,889 |
| DC + GBM blend (production) | 0.2195 | 0.1114 | 51.46% | 1,889 |
For context, published academic benchmarks for football 1X2 forecasting typically land around RPS ≈ 0.19–0.21 on harder multi-league panels with different evaluation windows, so direct comparison requires care; the point of publishing these numbers is not a leaderboard claim but a fixed, dated baseline that every future model change on this site must beat across a time-based holdout before it ships.
A model can be sharp but badly calibrated: it may rank matches correctly while systematically overstating favourites. This is why we report calibration alongside Brier and RPS rather than trusting a single number. In one memorable experiment, a temperature-scaling adjustment improved calibration on our calibration window by five points of hit-rate — and reversed into a measurable regression on the untouched test window. The fix was to discard it. Calibration computed on data the model has influenced is calibration theatre; only held-out calibration counts.
Reliability diagrams are the fastest diagnostic: bucket predictions by stated probability and compare against observed frequencies. A monotone S-curve hugging the diagonal is what health looks like. What you want to watch for is the tail behaviour — models are usually most overconfident exactly where their predictions are most extreme, and those are the predictions that matter most.
Football outcomes are ordinal: home win, draw, away win — adjacent categories share structure that one-hot encoding throws away. The Ranked Probability Score integrates the cumulative distribution, so predicting 1 when the match ends in a draw is penalised less than predicting 2. Brier treats every miss identically; RPS acknowledges that some misses are nearer misses. We track both, plus multi-class log loss for market comparisons, and we require improvement across the family — not a cherry-picked metric — before any model change ships. A candidate that wins on one score while regressing on the others is a candidate we have learned to distrust.
Proper scoring rules reward honest distributions, but they cannot distinguish lucky calibration from real calibration in samples of a few hundred matches — the confidence intervals on our per-league splits are wide enough that we no longer draw conclusions below roughly 200 matches per bucket. Scoring rules also say nothing about economic value: a model can improve Brier and still lose to the closing line, as our 6,478-match study demonstrated. Metrics measure distance from truth; only the market measures whether that distance is monetisable. We publish both and let readers weigh them.