一支球队持续超出 xG,是能力还是运气?我们测算了 xG 幸运度(实际进球减 xG)的跨期持续性,相关系数只有 0.07——几乎没有可用的预测信号。本文讨论 xG 的正确用法与常见的过度解读。
研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。
Expected goals (xG) assigns each shot a probability of becoming a goal, based on location, situation and shot type. Summed over a match, it answers: "how many goals would an average finisher have scored from these chances?" The gap between actual goals and xG — call it finishing over- or under-performance — is one of the most discussed numbers in football analytics, and one of the most misused. A team scoring 10 more than its xG generates a wave of "they're clinical" commentary; a team 8 under generates "they need a striker." We wanted to know: does over-performance persist? And if not, can it still be used as a feature?
We computed each team's xG differential — actual goals minus xG, both for and against — over rolling periods and measured how strongly over-performance in one period predicts over-performance in the next, after accounting for the overlap in opposition and schedule. The result: a correlation of about 0.07. In practical terms: knowing that a team ran hot last month tells you almost nothing about whether it will run hot this month. Most of what looks like "clinical finishing" at monthly resolution is variance in shot outcomes — the same goals that went in last month rattle off the post this month.
This is consistent with what the finishing-skill literature has found for years: genuine, stable differences between finishers exist, but they are small, slow to reveal themselves, and drowned in sampling noise at team-season and shorter horizons. A 0.07 correlation is not "zero," but a signal this weak cannot carry a feature on its own.
Over-performance is informative about the past: it tells you results ran ahead of process. That has value in specific places — most notably in explaining why a team's league position might misrepresent its underlying level, and in adjusting how much weight "lucky" results deserve when estimating team strength. In our pipeline, xG earns its place as an input to strength estimation, blended with results, rather than as a direct predictive signal. The goals a team deserved tell you about its process; the goals it actually scored tell you about its results; strength estimation wants process, cleaned of as much result noise as possible.
There is also a structural caveat: xG models are themselves estimates, built from shot data whose quality varies across leagues and seasons. Two xG numbers from different providers can differ meaningfully for the same match. Any feature built on provider-specific xG inherits that provider's biases — another reason we treat it as one input among several rather than an oracle.
xG is one of the best public lenses on football, and simultaneously one of the most over-claimed. Its predictive magic is in describing chance creation, not in forecasting finishing. If a model you are evaluating leans on recent over-performance as a positive signal, you now know what we measured: that signal has almost no persistence, and what looks like edge is usually last month's coin flips wearing a lab coat.
We measured the persistence of finishing over-performance: correlation between a team's goals-minus-xG differential in one period and the next. The coefficient came out at 0.07 — statistically indistinguishable from zero for our purposes. Read carefully, this is a claim about teams, not players: at club level, over a window long enough for our strength estimates, finishing skill differences wash out almost entirely. A club that overperformed its xG in the autumn was not a club to expect continued overperformance in the spring.
It is not a claim that finishing skill does not exist. Elite forwards demonstrably repeat above-average conversion across seasons; the effect is real at the player level. But it is small relative to shot-volume and chance-quality effects, it does not survive aggregation to the club-season scale our model operates on, and — the practical point — it is already partially embedded in the xG totals themselves, since chance quality and finisher quality are correlated by recruitment.
Shooting-luck regression is a dead end as a feature, but xG remains among the most informative quantities we ingest — as a measure of underlying performance, not as a luck-adjuster. A team whose results outrun its xG is a team our results-based ratings will overrate for a while; xG tells us by how much, even if it cannot tell us when the correction lands. The distinction is between predicting luck (impossible, r=0.07) and recognising it (straightforward). Our pipelines use xG for the second job and explicitly decline the first.
xG models are themselves calibrated on historical shot distributions and inherit their biases — shot-quality models trained on top-league data misprice long-range attempts and set pieces, and different xG vendors disagree meaningfully on the same match. Our persistence test also used a specific windowing; at player-month resolution the skill signal would strengthen, at club-season resolution it would weaken further. Finally, tactical change can create genuine, persistent finishing regimes — a new system generating higher-quality chances looks like "luck" regressing slower than it should. The measurement is honest; its interpretation has boundaries.