独立泊松模型系统性低估低比分平局。本文拆解 Dixon-Coles 的 τ 修正项如何对比分矩阵对角线重新分配概率,以及我们如何在防止时间泄漏的前提下拟合与监控该参数。
研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。
The default way to model a football scoreline is to give each team a Poisson goal expectation — say 1.6 expected goals for the home side and 1.1 for the away side — multiply the probability mass functions, and read the full score matrix off the grid. It is elegant, interpretable, and wrong in a specific, well-documented place: the low-scoring corner of the matrix.
Fit an independent Poisson model to a few seasons of real results and compare predicted frequencies with observed ones. The pattern is remarkably stable across leagues: draws — specifically 0-0 and 1-1 — occur more often than independent Poisson predicts, while mid-range scores like 1-0 and 2-1 in both directions occur slightly less often. The independence assumption is what fails: the two teams' goal counts are not independent, because for long stretches of a match both sides respond to the same game state. When the score is level, particularly late, both teams' incentives shift in ways that suppress goals — and Poisson knows nothing about incentives.
In their paper Modelling Association Football Scores and Inefficiencies in the Football Betting Market, Dixon and Coles kept the Poisson framework but introduced a multiplicative correction — usually written τ(x,y) — applied only to the low-score cells (0-0, 1-0, 0-1, 1-1) of the matrix. The single parameter, often called rho, shifts probability mass between these cells: negative rho inflates the draw diagonal (more 0-0 and 1-1 than independence implies), positive rho does the opposite. Crucially, all other cells are left untouched, so the model's marginal goal distributions stay intact.
The same paper contributed a second idea that matters just as much in practice: exponential decay of historical results. A match played eight months ago carries less information about today's team strength than one played last week. The half-life of that decay is a hyperparameter — our own grid search over 90–360 days found one well-performing setting for our data, but the honest lesson was that the optimum is data-dependent and must be re-validated, not assumed.
Our club model is bivariate Poisson with the Dixon-Coles correction, re-estimated on a rolling window with exponential decay. Two implementation details are worth documenting because they are where most home-grown versions go wrong:
In our ablations, removing the correction degrades the draw column of the score matrix visibly and moves outcome-level Brier/RPS meaningfully — but the honest headline is that the correction is worth on the order of a fraction of a percent of Brier, while the choice of scoring rule, out-of-sample discipline and team-strength estimation each matter an order of magnitude more. The Dixon-Coles adjustment is the best-known part of the model, but it is a finishing touch, not the engine. Modellers who spend a week tuning rho and an hour on leakage control have their priorities inverted.
The Dixon-Coles adjustment modifies the joint probability of low-scoring outcomes — 0-0, 1-0, 0-1 and 1-1 — by a dependence parameter commonly called rho. With independent Poisson goals, those four scorelines are assumed unrelated; in real football they are not, because late-game state changes everything. A team leading 1-0 stops attacking; a team drawing 1-0 behind pushes forward and concedes on the counter. The result is systematically more draws and more 1-0s than independent Poisson predicts, and rho is the parameter that absorbs this.
The practical magnitude is easy to underestimate. In our production blend the correction is worth roughly one percentage point of Brier score on out-of-sample club matches — small on paper, but in a field where our entire Dixon-Coles × gradient-boosting blend moved the needle from 0.2227 to 0.2195, it is a meaningful share of everything we have gained from modelling beyond the baseline.
The original 1997 paper contributed a second idea that matters as much as rho: matches should not count equally. Dixon and Coles weighted historical games with an exponential decay, an early acknowledgment that team strength is a drifting quantity. Our implementation follows this with an exponentially decaying likelihood whose half-life we tuned by grid search — the surprising result, documented in a separate note, is that the optimal half-life was measured in years, not months, suggesting modern squad stability changes slower than intuition suggests.
Dixon-Coles is a correction, not a theory: rho is a global parameter that cannot distinguish a stalemate born of tactical caution from one born of two depleted attacks. Bayesian hierarchical models and dynamic rating approaches model the same phenomenon with more structure, at the cost of more assumptions to defend. We have kept DC because it is transparent, cheap to fit, and its failure modes are legible — when it drifts, we can see where. For a research site that publishes its own numbers, legibility is worth more than elegance.