fbsa.app FBSA 研究
研究笔记 / 负结果

主场优势:一个全局数字胜过六十个联赛补丁

FBSA 研究 · 负面结果 · 7 分钟阅读

直觉上,每个联赛的主场优势应该不同。但按联赛锐化主场参数、按主客分裂强度估计、引入休息天数差与欧战拥挤特征,全部未能通过样本外门禁。为什么"更细"不等于"更好"——这是方差与样本量的账。

研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。

Home advantage in football is real, large and remarkably durable: across our data, home teams win far more than they lose, in every league and every season. Once you have a goal model, the obvious refinement stares at you — surely each league's home edge is different? High-altitude venues, long travel distances, culture, officiating. So we tried to capture that structure several ways. All of them failed out-of-sample. The single global home parameter outperformed every finer version. This article is about why.

What we tried

Per-league home parameters. Fit a separate home-advantage term for each competition, so that (say) a famously fortress-like league could get a bigger coefficient. Home/away split strengths. Instead of one attack/defence rating per team, estimate separate ratings for home and away performance, letting the model learn who is "a different team at home." Rest-days difference. A feature for the gap in days since each side's previous match. Fixture congestion. Flags for midweek European fixtures and squad rotation pressure. Four plausible refinements, four holdout failures.

The arithmetic that kills fine-graining

A league season contains a few hundred home matches per club community, but the home advantage signal in any single match is dwarfed by match-to-match noise in goals. To estimate a league-specific home edge precisely enough to beat a global one, you need the true between-league differences to be both large and stable. They are neither: measured home advantages fluctuate by season about as much as they differ between leagues, which means most of what a per-league estimate captures is that league's last couple of seasons of luck.

Sixty league parameters fitted on noisy data produce sixty chances to overfit. The global parameter, fitted on everything, trades a small bias (all leagues share one number) for a large variance reduction (one stable estimate). With effects this noisy, the trade is not close — the global estimate wins, and our holdout confirmed it every time.

The shrinkage intuition. If you must have per-group parameters, the right tool is partial pooling: each league's estimate is shrunk toward the global mean in proportion to its sample noise. But note what that converges to when data is thin — roughly the global number everywhere. Our experiments suggest our data is exactly that case, which is why we kept the global parameter and logged the refinement attempts as negative results.

The general lesson

"Finer granularity" feels like insight, but every extra parameter is a loan taken out against your sample size. The question is never "could league A have a different home edge from league B?" (it could) but "can my data tell me reliably how much?" When the honest answer is no, the coarser model is not lazy — it is correct. This is the same force behind our failures on per-league draw corrections and per-league decay constants, and it will apply to whatever refinement you are currently considering: check the sample size per group before you fall in love with the structure.

Home advantage remains one of the strongest and most robust signals in our model — globally parameterized, re-validated continuously, and, as of every experiment we have run, unimprovable by making it fancier.

The bias-variance tradeoff, concretely

A global home-advantage parameter is estimated from every match in the corpus: low variance, but it forces one number onto environments that genuinely differ — altitude venues, travel-heavy leagues, cultures of crowd intensity. A per-league parameter adapts to each environment but is estimated from a fraction of the data, and when a league's sample wobbles, so does its parameter. Our experiments came down on the side of the global number, and the reason is the variance side of the tradeoff: the between-league differences in home advantage were smaller and less stable than the estimation noise introduced by trying to capture them. Every per-league variant — split home and away strengths, per-league sharpening, rest-day differentials, continental congestion adjustments — failed replication out-of-sample.

This is a recurring lesson in applied modelling: the theoretically finer model is only better if the data can afford it. Ours could not, and pretending otherwise showed up immediately as degraded test-window scores.

The natural experiment we did not have

The pandemic empty-stadium seasons are the cleanest study of home advantage ever conducted — dozens of leagues simultaneously removed crowds while keeping everything else constant. Published analyses of that period found home advantage shrank measurably but did not vanish, suggesting the crowd is a large but partial component. We mention this because it defines what our own parameters cannot: our single global number averages over eras that include that shock, and a richer model with an era term could in principle do better. We have chosen to let the decay-weighted likelihood absorb era effects slowly rather than model them explicitly — a defensible simplification, and an honest limitation.

Limitations

Our corpus is the top European leagues, where travel distances are modest and stadium quality relatively uniform; home advantage in leagues spanning continents or altitudes is plausibly more heterogeneous, and the global parameter would fit it worse. The result is also conditional on our evaluation horizon: per-league parameters might win with twice the data per league. And we treat home advantage as multiplicative on goal expectancy, whereas the true mechanism — refereeing bias, travel fatigue, crowd effects on officials — may not decompose that neatly. The global number is not the truth; it is the best-defended approximation our evidence supports.