fbsa.app FBSA 研究
研究笔记 / 负结果

伤停数据给了我们什么——以及没给什么

FBSA 研究 · 负面结果 · 7 分钟阅读

我们尝试了四种伤停特征:缺阵人数、出场时间加权、球员层级加权、门将缺阵单独建模。在时间切分的保留带上,没有一个能稳定改善模型。本文公开完整实验设计与负结果。

研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。

If a team's best striker and first-choice goalkeeper are both out, everyone agrees that matters. Squad availability is the most intuitive "hidden signal" in football modelling, and the most common next step for anyone who has built a baseline goal model. So we built it — four ways. All four failed our out-of-sample gate. This article documents what we tried and why we now believe the honest conclusion is that, in our data, availability adds nothing our model can use.

What we tested

We evaluated four feature designs, each fed into the same rolling replay that produced our published baselines:

1. Absence count. The number of unavailable players per team, matched to the fixture date. The bluntest measure, and the most common one in public discussions.
2. Minutes-weighted absence. Share of recent minutes lost, so that a squad rotation player missing counts less than a 90-minute every-week starter.
3. Hierarchy-weighted absence. A manual importance tier per player (starter / rotation / squad), weighting absences by tier.
4. Goalkeeper absence. First-choice goalkeeper out as a separate binary feature, on the theory that keeper changes affect goals conceded through a different channel than outfield changes.

Why none of it survived

Three forces conspire against availability features, and our experiments ran into all of them.

Timing and leakage. Real lineups are published about an hour before kickoff. If you use lineup-accurate availability in a backtest, you are using information that is unavailable at decision time; if you use only pre-match injury reports, your data quality collapses — reports are incomplete, stale, and occasionally misleading. There is no clean middle ground, and both ends of the spectrum produce dishonest evaluations.

Redundancy. Much of what availability measures is already visible to a strength model: a team missing key players concedes more and scores less, which recent-results weighting already picks up one or two weeks later. The feature's marginal information, after conditioning on form, is small.

Signal-to-noise. Over our evaluation window, the effect sizes of even the well-designed variants (minutes-weighted, goalkeeper) were indistinguishable from zero on the holdout band. Point estimates wobbled around zero across seasons; nothing was stable enough to survive a three-way time-split gate.

Our conclusion, and its limits. We logged the negative result and marked all four variants as "do not retry" under the current data and pipeline. That is a statement about our data and our baseline — not a law of nature. A pipeline with fast, reliable lineup data captured at kickoff, or a model without strong recent-form features, could reach a different conclusion. What we are confident about is the method: any availability feature must be evaluated with decision-time information only, on a chronological holdout, before anyone believes it.

The broader lesson generalizes past injuries: the features that feel most insightful are usually the ones the market and the recent-results signal already price in. The gap between "obviously relevant" and "measurably additive" is where most feature engineering goes to die, and publishing the deaths is part of doing honest research.

Why the signal is weaker than it looks

Injuries feel like information the market cannot fully price: a key striker is out, therefore downgrade the attack. The reasoning is sound at the level of a single match, which is what makes the aggregate failure surprising. Three mechanisms reconcile the two. First, squad depth — top clubs replace a starter with a near-equivalent, and our rating already captures squad quality at the club level. Second, tactical adaptation: opponents and the affected team both adjust, damping the raw effect. Third, and most importantly, the market prices injuries in real time, and our match-level features already partially reflect the consequences through recent results.

There is also a data limitation that our experiments made concrete: publicly available injury sources record a binary fact — the player is unavailable — without minutes lost, replacement quality, or positional context. A hundred minutes lost from a squad player is not a hundred minutes lost from the talisman, but the binary feed treats them identically. Weighting schemes we tested (by minutes, by market value, by squad-role) were attempts to reconstruct that missing structure from noisy proxies, and none of them beat the simple absence count by a margin our gate would accept.

What we did not test

Our data covered player availability, not freshness or output. Load management — the accumulated fatigue that precedes soft-tissue injury — lives in physical-performance data we do not ingest. Likewise, we tested goalkeeper absence as a separate variant and found nothing, but with only a handful of elite goalkeepers moving between clubs in any window, the test had limited power. These remain open questions rather than closed ones, which is exactly why the negative result is filed with its caveats rather than buried.

Limitations

The study spans top-flight European club football over recent seasons; lower divisions, where a single key player can be a far larger share of team output, may behave differently. Injury reporting itself is noisy — clubs manage information strategically, and late team-news surprises arrive after our daily learning cycle. Finally, absence is endogenous in a subtle way: players are rested when matches matter less, so measured effects conflate importance of the fixture with importance of the player. Untangling that would require fixture-importance modelling that is on the roadmap but not yet built.