拟合期 +5 个百分点的改善,到了保留带变成反向——这样的陷阱我们踩过很多次。本文公开我们的时间三分切门禁(70/10/20)、网格与测试带分离的纪律,以及一连串在门禁前倒下的候选特征。
研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。
Every modelling project has a graveyard: ideas that looked brilliant in fitting and died in evaluation. Most projects keep the graveyard private, which is why published feature lists read like victory parades. Ours is public. This article describes the gate we run every candidate through, and then lists — by name — the features that failed it, so that the next person tempted by the same idea can skip a few funerals.
Head-to-head history between two clubs feels like gold: some teams "always beat" others. Our H2H direction factor improved calibration-band metrics by about +5 points — a large, exciting effect. On the holdout it reversed. The in-sample signal was mostly a handful of long-running rivalries plus noise; once those rivalries changed character, the factor flipped sign. This is the canonical shape of an overfit feature: big in fitting, gone or negative in validation.
Temperature calibration for the international pipeline. A grid search found an "optimal" temperature of 1.2 on the calibration band; on the holdout, the best value was the identity (1.0). Temperature scaling was a pure negative contribution for that pipeline — the model was already calibrated.
Rest-days difference. The intuition (a team with three more rest days should have an edge) is sound; the measured effect in our data was indistinguishable from zero out-of-sample.
Per-league sharpening. Tuning home advantage, draw correction or decay per league multiplied the number of fitted parameters by sixty and the noise by more.
Squad availability features (four variants — covered in their own article) and xG over-performance as a standalone signal (persistence ≈ 0.07) died the same way.
Notice the pattern: every failed idea was plausible. Plausibility is what makes them dangerous — each one carries a story convincing enough to skip validation. But the arithmetic of repeated search is unforgiving: try twenty plausible features, and a couple will look great on any given band by luck alone. Without a untouched holdout, your project converges to a museum of coincidences.
The gate also changes what you work on. Ideas that can only produce tiny, fragile effects (per-league sharpening) become visibly worse bets than ideas that produce structural gains (better team-strength estimation, honest blending of model families, fixing data quality). Discipline is not a tax on creativity; it is a compass.
You do not need our exact percentages. You need three things: time-ordered splits (never random), a holdout you touch once, and a log of failures you actually consult before retrying an idea. If your evaluation cannot produce a negative result that changes your behavior, you do not have an evaluation — you have a highlight reel.
Every feature gate is a statistical test, and a battery of tests guarantees false positives. Run twenty candidate features against a fixed evaluation window and, at the five percent significance level, one will look genuine by pure chance. Our head-to-head direction factor is a case study: on the calibration window it added five points of hit-rate — the largest single-feature gain we had measured. On the test window the sign flipped. Nothing about the feature had changed; the calibration window had simply been lucky, and the gate did its job.
The defence is procedural rather than statistical: fixed evaluation windows that candidates never see, a strict rule that improvements must replicate out-of-sample, and a written log of every rejected candidate. The log matters more than it sounds — it converts silent selection bias into an auditable record, and it stops the same seductive feature from being re-tested until it finally passes by luck.
Of all the improvements our gate has rejected, the pattern that recurs most often is a calibration win that inverts. Calibration adjustments are fitted to the residuals of the calibration window, which means they are the easiest kind of change to overfit — the adjustment memorises window-specific mispricing rather than learning structure. Temperature scaling, draw-bias corrections and per-league sharpening have all followed the same arc: celebrated on calibration data, convicted on test data. We now treat any calibration-only gain with suspicion until it has been replicated twice.
The gate is conservative by design, which means it also rejects real improvements occasionally — the cost of avoiding false positives is paid in false negatives. Our time-based 70/10/20 split assumes the future resembles the past; a genuine regime change would be invisible to it until after the fact. And a single test window, however carefully held out, is still one draw from the distribution of futures. We mitigate rather than solve this: thresholds are deliberately strict, and any feature that passes is monitored in production, because passing the gate is evidence, not proof.