fbsa.app FBSA 研究
研究笔记 / 方法论

带保留门禁的特征挖掘:为什么大多数改进活不下来

FBSA 研究 · 实验纪律 · 9 分钟阅读

拟合期 +5 个百分点的改善,到了保留带变成反向——这样的陷阱我们踩过很多次。本文公开我们的时间三分切门禁(70/10/20)、网格与测试带分离的纪律,以及一连串在门禁前倒下的候选特征。

研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。

Every modelling project has a graveyard: ideas that looked brilliant in fitting and died in evaluation. Most projects keep the graveyard private, which is why published feature lists read like victory parades. Ours is public. This article describes the gate we run every candidate through, and then lists — by name — the features that failed it, so that the next person tempted by the same idea can skip a few funerals.

The gate

Chronological 70/10/20 split. All data is ordered by time. The earliest 70% fits parameters. The next 10% (calibration band) is where hyperparameters and feature candidates are tuned. The final 20% (holdout) is touched exactly once per candidate: its numbers decide deployment. No exceptions, no re-tuning after peeking.

One look per candidate. If you check the holdout repeatedly while iterating, it becomes a second training set and stops being honest. We log the holdout result once and move on.

Improvement must transfer. A candidate that wins the calibration band but reverses on the holdout is rejected and recorded. Grid-band improvement alone never ships.

A typical death: the H2H direction factor

Head-to-head history between two clubs feels like gold: some teams "always beat" others. Our H2H direction factor improved calibration-band metrics by about +5 points — a large, exciting effect. On the holdout it reversed. The in-sample signal was mostly a handful of long-running rivalries plus noise; once those rivalries changed character, the factor flipped sign. This is the canonical shape of an overfit feature: big in fitting, gone or negative in validation.

More corpses

Temperature calibration for the international pipeline. A grid search found an "optimal" temperature of 1.2 on the calibration band; on the holdout, the best value was the identity (1.0). Temperature scaling was a pure negative contribution for that pipeline — the model was already calibrated.
Rest-days difference. The intuition (a team with three more rest days should have an edge) is sound; the measured effect in our data was indistinguishable from zero out-of-sample.
Per-league sharpening. Tuning home advantage, draw correction or decay per league multiplied the number of fitted parameters by sixty and the noise by more.
Squad availability features (four variants — covered in their own article) and xG over-performance as a standalone signal (persistence ≈ 0.07) died the same way.

Why the gate matters more than the features

Notice the pattern: every failed idea was plausible. Plausibility is what makes them dangerous — each one carries a story convincing enough to skip validation. But the arithmetic of repeated search is unforgiving: try twenty plausible features, and a couple will look great on any given band by luck alone. Without a untouched holdout, your project converges to a museum of coincidences.

The gate also changes what you work on. Ideas that can only produce tiny, fragile effects (per-league sharpening) become visibly worse bets than ideas that produce structural gains (better team-strength estimation, honest blending of model families, fixing data quality). Discipline is not a tax on creativity; it is a compass.

For your own project

You do not need our exact percentages. You need three things: time-ordered splits (never random), a holdout you touch once, and a log of failures you actually consult before retrying an idea. If your evaluation cannot produce a negative result that changes your behavior, you do not have an evaluation — you have a highlight reel.

The multiple-comparisons trap

Every feature gate is a statistical test, and a battery of tests guarantees false positives. Run twenty candidate features against a fixed evaluation window and, at the five percent significance level, one will look genuine by pure chance. Our head-to-head direction factor is a case study: on the calibration window it added five points of hit-rate — the largest single-feature gain we had measured. On the test window the sign flipped. Nothing about the feature had changed; the calibration window had simply been lucky, and the gate did its job.

The defence is procedural rather than statistical: fixed evaluation windows that candidates never see, a strict rule that improvements must replicate out-of-sample, and a written log of every rejected candidate. The log matters more than it sounds — it converts silent selection bias into an auditable record, and it stops the same seductive feature from being re-tested until it finally passes by luck.

Why calibration gains are the most deceptive

Of all the improvements our gate has rejected, the pattern that recurs most often is a calibration win that inverts. Calibration adjustments are fitted to the residuals of the calibration window, which means they are the easiest kind of change to overfit — the adjustment memorises window-specific mispricing rather than learning structure. Temperature scaling, draw-bias corrections and per-league sharpening have all followed the same arc: celebrated on calibration data, convicted on test data. We now treat any calibration-only gain with suspicion until it has been replicated twice.

Limitations

The gate is conservative by design, which means it also rejects real improvements occasionally — the cost of avoiding false positives is paid in false negatives. Our time-based 70/10/20 split assumes the future resembles the past; a genuine regime change would be invisible to it until after the fact. And a single test window, however carefully held out, is still one draw from the distribution of futures. We mitigate rather than solve this: thresholds are deliberately strict, and any feature that passes is monitored in production, because passing the gate is evidence, not proof.