fbsa.app FBSA 研究
研究笔记 / 负结果

为什么平局是最难预测的赛果

FBSA 研究 · 赛果研究 · 8 分钟阅读

平局约占足球比赛的四分之一,但我们尝试的每一种"平局专项"改进——平局标记特征、平局选择规则、比分矩阵对角线加权——都在保留样本上反转。本文解释平局难在何处,以及为什么接受校准的概率比追逐平局更诚实。

研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。

Of the three outcomes in a football match, the draw is the strange one. It is common — roughly a quarter of matches — it carries no directional information about which team was better, and it is the outcome most likely to make a good model look stupid in a single weekend. It is also the outcome that attracts the most "special handling" ideas. We tried several. All of them failed out-of-sample. Here is what we learned about why.

Draws are mostly absence of signal

A goal model produces expected goals for each side; the draw probability is the mass the score matrix puts on the diagonal. When two teams are evenly matched, or when both are off-form, the diagonal swells — but so does the probability of a one-goal either way. The draw lives in the dense center of the distribution, where small shifts in expected goals move win/draw/loss probabilities in complicated, partially offsetting ways. Unlike a strong favourite, which is a positive signature a model can latch onto, a likely draw is often just the absence of any strong signature. That is precisely the kind of event where features stop discriminating.

What we tried, and what happened

Draw indicator features. Match-level covariates designed to flag draw-prone fixtures: tightly-matched strength pairs, low combined expected goals, late-season low-stakes flags. On the holdout band, none improved proper scoring rules; several degraded calibration.

Draw selection rules. Logic that overrides or boosts the model's draw pick under certain conditions. This is the most seductive idea in the category — "the model is too timid on draws, let's push them." On the holdout, the pushed draws lost more often than the base model, because the conditions that select "likely draws" in-sample are exactly the conditions under which draws are most over-represented in noise.

Diagonal re-weighting. Ex extra probability mass on the score matrix diagonal (the Dixon-Coles rho, or a heavier-handed draw bonus). A globally fitted rho is part of our production model; but league-specific or sharpened diagonal corrections failed the gate — the per-league rho estimates were too noisy, and the improvements seen during fitting did not transfer.

The uncomfortable statistics of a 1-in-4 event

The deeper problem is sample arithmetic. If draws are ~25% of matches, and your best achievable discrimination on them is weak, then any draw-special rule is being fitted on the noisiest quarter of your data with the weakest signal. With a few thousand training matches, the sampling variance of draw-focused statistics dwarfs the real effects you are hoping to capture. The result is what we kept seeing: beautiful in-sample structure, reversal out-of-sample.

What we do instead. We let the model say "28%" and accept it. A calibrated draw probability, scored by Brier/RPS across thousands of matches, captures all the draw information the model genuinely has — no more, no less. Every attempt to squeeze extra draw signal out of our data has either been noise or duplicated what the calibrated probability already encodes. Until a feature survives the holdout gate, the honest position is that our draw knowledge is already in the number.

Draw-picking is where prediction content and prediction theatre diverge most sharply. A confident "this week's draws" column is entertaining; a calibrated probability is falsifiable. We chose falsifiable, and this article is the receipt.

The geometry of the problem

Under a Poisson-style goal model, the probability of a draw is not a free parameter — it is a derived quantity, the sum of the diagonal cells of the scoreline matrix. Its value is governed by how close the two teams' expected goals are and how low-scoring the match profile is. Two mismatched teams can produce a sensible draw probability almost by accident; two evenly matched teams in a cagey, low-scoring league produce a draw probability so high that any independent-Poisson calibration visibly strains. The draw, in other words, is where the model's assumptions about goal dependence are stressed hardest — which is precisely why the Dixon-Coles correction, aimed at the low-score cells, buys back so much of the missing probability mass.

Our negative results cluster here for a related reason: every draw-specific adjustment we have tried — flagging likely draws, sharpening the diagonal, re-weighting drawn outcomes in the likelihood — amounts to tuning the same handful of cells, and each attempt improved calibration in-sample while failing replication out-of-sample. The information that distinguishes an imminent 1-1 from an imminent 2-1 is mostly not in the historical scorelines the model sees.

What we have not tried

Two directions remain genuinely open. Match-state dynamics — how probability shifts when a draw becomes valuable to both teams, as in the final rounds of a league season — are invisible to our static pre-match features. And refereeing style, which plausibly influences the draw rate through game management, is absent from our data entirely. We note them here because a research site that lists only the ideas it tested is advertising half its notebook.

Limitations

Draw rates are league-specific and era-specific; a calibration learned across mixed competitions averages away effects that are real within any single league. Sample sizes compound the problem — draws are a minority outcome, so even thousands of matches yield a few hundred draws per league-season, wide enough intervals to hide moderate effects. And our evaluation rewards calibrated whole distributions, not draw-recovery in isolation, so a genuine draw improvement too small to move Brier or RPS across the threshold would look identical to noise. Some of what we filed as negative results may be real effects that our instruments cannot yet resolve.