fbsa.app FBSA 研究
研究笔记 / 方法论

球队强度的保鲜期:我们的指数衰减半衰期搜索

FBSA 研究 · 时间衰减 · 8 分钟阅读

历史比赛该按多快的速度贬值?我们在 90 到 360 天的常规网格里没有找到能通过样本外门禁的设置,最终胜出的是一个 8 年的半衰期。本文解释指数加权、时间三分切验证,以及为什么"网格最优"不等于"可上线"。

研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。

Every team-strength model has to answer an uncomfortable question: how much is a result from three years ago worth today? Fit equally on all history and your ratings move at the speed of continental drift. Weight only recent form and every September looks like a title race. The standard answer is exponential decay: each historical match carries weight w = 0.5^(age / h), where h is the half-life — the age at which a result counts half as much as a match played today.

The obvious search, and why it failed

The textbook half-life for football ratings is somewhere between a season and a year. So that is where we searched first: a grid from 90 to 360 days, evaluated on a rolling out-of-sample window with strict chronological replay. Not one setting in that range survived our validation gate. Some looked better on the fitting band; every one of them failed to hold up on the holdout. That is not evidence that exponential decay is wrong — it is evidence that a grid tuned on a narrow band of history will happily fit noise, and our gate is designed to catch exactly that.

The unexpected winner: eight years

When we widened the grid far beyond the conventional range, a clear outlier emerged: a half-life of roughly 2,920 days — about eight years. In other words, a match from eight years ago still carries half the weight of last week's. That sounds absurd until you think about what team strength actually is: a slow-moving composite of academy pipelines, club culture, financial muscle and fanbase. It does not halve every season. Recent form matters, but the low-frequency signal — which clubs are structurally strong — persists for years, and a long half-life lets the model accumulate it instead of constantly forgetting.

Crucially, this setting did not merely win the grid. It passed the same three-way time-split gate (fit / calibration / holdout) that every other candidate failed, with consistent improvement across the holdout band. It has been running in production since late September 2026.

Why "grid-optimal" does not mean "deployable"

Our deployment rule. Every model change must pass a chronological 70/10/20 split: parameters are fitted on 70% of history, hyperparameters are selected on a 10% calibration band, and only results on the final 20% holdout decide deployment. A candidate that wins on the calibration band but reverses on the holdout is rejected — no matter how good the story sounds. This rule has killed far more ideas than it has approved, and we consider that its main feature.

The trap. If you evaluate twenty hyperparameter settings and pick the best, you have fit twenty coin flips. The holdout band is the only place where the twenty-first, untouched flip can tell you something honest.

Practical takeaways

If you build your own ratings: use exponential decay, but expect your optimal half-life to depend on your data — league coverage, sample size and how your features aggregate all shift it. Treat anything inside the conventional 90–360-day window as a hypothesis, not a default. And be suspicious of any source that quotes a single "best" decay constant without describing how it was validated. The half-life is not a fact about football; it is a fact about your data, and it deserves the same out-of-sample skepticism as every other parameter.

Why eight years beat eight months

Our grid covered half-lives from 90 days to 2,920 days. The intuition going in was that football moves fast — transfers, managerial changes, tactical fashions — and a half-life of one to two seasons would win. The data disagreed emphatically. Every candidate in the 90-to-360-day band failed our validation gates; 2,920 days, roughly eight years, was the only setting that survived. The likely explanation is signal-to-noise: a short half-life lets a single hot streak or injury-free run dominate the strength estimate, and the variance this introduces costs more than the responsiveness gains. Team strength, it turns out, has a stable core that decays slowly, and most apparent form is noise around it.

This echoes a result familiar to Elo practitioners: the best-performing chess and football rating systems tend to be older and slower than their designers expect. Responsiveness feels right; stability scores better.

The exception that proves the rule

An eight-year half-life is not a universal constant, and we are careful not to oversell it. Our evaluation window covers an era of relatively stable squad structures in the top European leagues. A rule change that reshaped the game — a new competition format, a structural shock like the pandemic empty-stadium season, or a shift in transfer economics — would eventually make eight years too slow. The half-life is a parameter to be re-validated, not a truth; our pipeline re-checks it on every learning cycle precisely because the "right" answer can move.

Limitations

Half-life is entangled with every other component of the model: the same grid run on a different feature set, or on the national-team corpus where squads turn over faster and fixtures are sparser, need not agree with the club result. We also evaluate at the aggregate level — a half-life that is optimal on average may still be wrong for promoted teams, newly promoted managers, or clubs under new ownership. Fine-grained adaptation is a legitimate next step, but it multiplies parameters, and our discipline requires every added parameter to earn its keep on held-out data before it ships.