统计目标模型可解释、样本高效;梯度提升灵活、能吃结构化特征。我们按 0.6/0.4 权重融合 Dixon-Coles 比分矩阵与 GBM 概率,在 1,889 场样本外比赛上将 Brier 从 0.2227 降到 0.2195。本文讲融合设计与防泄漏重放。
研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。
There are two classic ways to predict football outcomes. The statistical way: model goals directly with a Poisson-family distribution, get a full score matrix, and read 1X2 probabilities off it. The machine-learning way: throw structured features into a gradient-boosted tree model and let it learn the mapping to outcomes. Each has a clear weakness — the goal model is rigid about what it can express; the GBM is opaque and data-hungry. Our international-match pipeline blends the two, and this article documents how.
The goal model is a Dixon-Coles-corrected Poisson with exponentially decayed team strengths (the decay itself chosen by our out-of-sample gate — see the half-life article). It produces a full score matrix, which gives us 1X2 probabilities, over/under probabilities and score heatmaps from one coherent object. Its virtue is structure: goals are modeled as goals, so related markets stay mutually consistent, and it works even with sparse data — exactly the situation with national teams, which play a handful of matches a year.
The GBM consumes a deliberately small feature set: Elo-style ratings, last-10 form, goal difference, and rest between matches — features computed strictly from information available before each match. Its virtue is flexibility: it captures interactions and non-linearities the Poisson family cannot express. Its vice is variance: with limited matches, an unregularized GBM will happily memorize.
The final probabilities are a linear mix: p = W_DC · p_poisson + (1 − W_DC) · p_gbm. We selected W_DC on the calibration band only; 0.6 won — the goal model carries the majority, which matches the data regime (sparse matches favor structured models). The improvement on the holdout was real but modest: over 1,889 out-of-sample international matches, Brier fell from 0.2227 to 0.2195 and the ranked probability score from 0.1134 to 0.1114. Modest is fine. The gate exists precisely so that modest-but-true beats dramatic-but-false.
Three honest notes. First, the feature set is small because most proposed features died at the gate (injury data, rest-day fine tuning, congestion flags) — the blend's quality owes more to rejection than to inclusion. Second, the blend weight is one global number; we did attempt to make it situation-dependent and the holdout said no, consistent with everything else in this series. Third, the GBM's contribution is real but small; if we had to keep only one component for national-team data, it would be the goal model — the GBM earns its 40% in aggregate, not in any single match.
The result now powers our international dashboard. The bigger point for other modellers: hybridization is not exotic, it is bookkeeping — two model families, one honest evaluation loop, and a weight fitted once on a band the holdout never saw during tuning.
The Poisson-Dixon-Coles side brings structure: a parametric goal model whose behaviour is understood, whose parameters are interpretable, and whose weaknesses (independence assumptions, global corrections) are documented. The gradient-boosting side brings pattern memory: it finds interactions in engineered features — Elo differences, recent form trajectories, goal-difference trends, rest intervals — that no parametric form we tried captured. Alone, each leaves measurable accuracy on the table; alone, each fails in characteristic ways. The blend is an average of a structured model and a memorised one, weighted toward structure.
The weights were not chosen by taste. A grid search over the mixing parameter selected 0.60 toward the Dixon-Coles side, and the out-of-sample numbers that earned deployment were: Brier score 0.2227 → 0.2195 and ranked probability score 0.1134 → 0.1114 over 1,889 club matches — improvements replicating across both held-out windows, not just one. A stacking architecture evaluated in the same programme scored better still in experimentation, but had not cleared every gate at the time of writing; it remains a candidate rather than a claim.
The most instructive result in this programme was negative: transplanting the club blend's parameters to our national-team corpus made every metric worse. National-team football breaks the assumptions quietly — squads assemble a handful of times a year, rest asymmetries are extreme, tournament stakes distort incentives, and the club-derived form signals simply do not exist in the same form. The national-team system runs its own weights, with no temperature correction and a different blend ratio, and attempts to unify the two corpora under shared parameters failed our gates twice. A model is a claim about a data-generating process; when the process changes, the claim does not travel.
The blend is fitted and validated on recent seasons of top-flight European club football; its gains are smaller in league-years with regime breaks, and untested outside that domain. Gradient boosting is a black box by temperament — we audit it with feature importance and partial dependence, but audit is not explanation, and part of why the DC side retains the larger weight is exactly that its behaviour stays legible. Finally, mixing weights are a global choice; a per-match adaptive mixture is the obvious refinement, and it is on the roadmap gated, like everything else here, on held-out evidence.