fbsa.app FBSA 研究
研究笔记 / 方法论

如何为足球概率模型打分:Brier、RPS 与校准——附我们自己的数字

FBSA 研究 · 评估框架 · 9 分钟阅读

命中率不适合评估概率预测。本文解释多分类 Brier 评分与有序概率评分 RPS、校准为何比准确率更重要,并公开线上基线:1,889 场样本外比赛 Brier 0.2195、RPS 0.1114。

研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。

Ask how accurate a football forecast is and someone will answer with a hit-rate: "we got 54% of winners right". This number is close to meaningless. In a typical league season, favourites win around 45–50% of matches, so a coin that always picks the favourite scores similarly — and a model that confidently picks upsets will be punished even when its probabilities are honest. Evaluating a probability forecast requires scoring rules designed for probabilities. This article explains the two we use, and publishes our live baselines.

The multi-class Brier score

The Brier score is the mean squared error of a probability forecast. For a football match with outcomes (home, draw, away), the multi-class form sums the squared error across all three:

BS = (p_home − o_home)² + (p_draw − o_draw)² + (p_away − o_away)²

where o is 1 for the observed outcome and 0 otherwise. The score ranges from 0 (perfect) to 2 (worst), and 2/3 is what you get by predicting the empirical base rates blindly. Two properties make it the workhorse: it is a proper scoring rule — the best strategy is to state your true belief — and it decomposes into calibration and refinement terms, which matters next.

The ranked probability score (RPS)

Football outcomes are ordered: home win > draw > away win in terms of home advantage. The RPS exploits this ordering by comparing cumulative probability against cumulative outcome. Getting home-vs-away right while shuffling the draw is penalised less than confusing the extremes. Because of this ordering sensitivity, RPS has become the de facto standard in the academic football-forecasting literature, which is why we report it alongside Brier.

Calibration beats accuracy

A model is calibrated when, among all matches to which it assigns 60% home-win probability, roughly 60% of home teams actually win. Calibration is checkable directly with reliability curves, and it is the property that makes probabilities trustworthy downstream. A model can have an unimpressive hit-rate and still be excellent if it is sharp and calibrated; a model with a flashy hit-rate but poor calibration will mislead anyone who uses its numbers as probabilities. This is why our gate process refuses to ship parameter changes that improve hit-rate while degrading calibration.

Our live baselines (published, not promised)

We maintain a rolling out-of-sample evaluation of the club model (bivariate Poisson with Dixon-Coles correction, blended with a gradient-boosting estimate of goal expectancy). Current production baselines over the most recent 1,889 evaluated matches:

Model pathBrierRPSHit-raten
Pure Dixon-Coles0.22270.113451.24%1,889
DC + GBM blend (production)0.21950.111451.46%1,889

For context, published academic benchmarks for football 1X2 forecasting typically land around RPS ≈ 0.19–0.21 on harder multi-league panels with different evaluation windows, so direct comparison requires care; the point of publishing these numbers is not a leaderboard claim but a fixed, dated baseline that every future model change on this site must beat across a time-based holdout before it ships.

Practical checklist

If you are evaluating any football probability model — ours or anyone else's — demand these four things: (1) proper scoring rules (Brier and/or RPS), never bare hit-rate; (2) a time-based out-of-sample split, never random k-fold, which leaks information across the season timeline; (3) a calibration check, not just a summary score; (4) fixed baselines with sample sizes, so improvements can be verified rather than asserted.

Calibration is a separate axis from sharpness

A model can be sharp but badly calibrated: it may rank matches correctly while systematically overstating favourites. This is why we report calibration alongside Brier and RPS rather than trusting a single number. In one memorable experiment, a temperature-scaling adjustment improved calibration on our calibration window by five points of hit-rate — and reversed into a measurable regression on the untouched test window. The fix was to discard it. Calibration computed on data the model has influenced is calibration theatre; only held-out calibration counts.

Reliability diagrams are the fastest diagnostic: bucket predictions by stated probability and compare against observed frequencies. A monotone S-curve hugging the diagonal is what health looks like. What you want to watch for is the tail behaviour — models are usually most overconfident exactly where their predictions are most extreme, and those are the predictions that matter most.

Why RPS for football specifically

Football outcomes are ordinal: home win, draw, away win — adjacent categories share structure that one-hot encoding throws away. The Ranked Probability Score integrates the cumulative distribution, so predicting 1 when the match ends in a draw is penalised less than predicting 2. Brier treats every miss identically; RPS acknowledges that some misses are nearer misses. We track both, plus multi-class log loss for market comparisons, and we require improvement across the family — not a cherry-picked metric — before any model change ships. A candidate that wins on one score while regressing on the others is a candidate we have learned to distrust.

Limitations

Proper scoring rules reward honest distributions, but they cannot distinguish lucky calibration from real calibration in samples of a few hundred matches — the confidence intervals on our per-league splits are wide enough that we no longer draw conclusions below roughly 200 matches per bucket. Scoring rules also say nothing about economic value: a model can improve Brier and still lose to the closing line, as our 6,478-match study demonstrated. Metrics measure distance from truth; only the market measures whether that distance is monetisable. We publish both and let readers weigh them.