核心结论:在 6,478 场俱乐部比赛的滚动样本外检验中,模型在 log-loss 上输给收盘赔率(1X2:0.987 对 0.967;大小球 2.5:0.684 对 0.666),所有正 EV 分组的模拟 ROI 为 -3.4% 至 -12.6%,CLV 全面 ≤ 0。文末给出重启该方向研究的两条预注册门槛。
研究笔记正文当前以英文发布,目录页与摘要提供多语言版本。
Every football modelling project eventually faces the same question: if the model is any good, why doesn't it make money against the market? This article is our answer, published as a full negative result. We took our club-level probability model, played it against bookmaker closing odds on 6,478 matches in a rolling out-of-sample design, and it lost. Not narrowly — consistently, across every strategy bucket we tested.
Closing odds aggregate the money and information of the entire betting market minutes before kickoff. Decades of research on market efficiency suggest the closing line is the single most accurate public estimate of match probabilities available. If your model cannot beat the closing line on a proper scoring rule, any apparent "edge" at earlier timestamps is far more likely to be noise or stale information than genuine forecasting skill.
So the test design is simple and unforgiving: for each match, compare the model's 1X2 probability distribution to the distribution implied by closing odds, using log-loss (equivalently, cross-entropy). Lower is better. Then, separately, simulate a value-betting strategy: whenever the model probability implies positive expected value at the available price, place a unit. Track profit and closing line value (CLV) over time.
| Market | Model log-loss | Closing-implied log-loss | Verdict |
|---|---|---|---|
| 1X2 (match outcome) | 0.987 | 0.967 | Market wins |
| Over/under 2.5 goals | 0.684 | 0.666 | Market wins (narrowly) |
The 1X2 gap of 0.020 in log-loss is decisive over thousands of matches. The goals market is closer — our model's mean calibration there is genuinely decent — but "close to the market" and "better than the market" are different claims, and only one of them pays.
We then stratified all entries by model-implied expected value. If the model had any real edge, positive-EV buckets should show positive returns. They did not:
| Strategy bucket | Simulated ROI |
|---|---|
| All EV buckets (1X2) | −3.4% to −12.6% |
| CLV achieved | ≤ 0 across the board |
The CLV result is the more damning one. Beating the closing line on the bets you choose is the standard evidence of genuine picking skill; our selections never managed it even once at the aggregate level.
Two reasons. First, honesty is the product: this site exists to document how a probability model behaves under honest evaluation, and a negative result evaluated honestly is worth more than a cherry-picked hot streak. Second, it defines the roadmap. Our current candidates for genuine improvement are new information, not new parameters: odds movement data (being collected since September 2026) and shots-on-target integration, both of which add signal the closing line reacts to but our current feature set never sees.
Across 6,478 club matches evaluated with rolling out-of-sample predictions, our model scored a multi-class log loss of 0.987 on the 1X2 outcome, against 0.967 for the de-vigged closing odds — a gap of roughly two percent. On over/under 2.5 goals the picture was closer: 0.684 versus 0.666, with our mean calibration actually competitive. Every expected-value bucket we tested — from marginal (+2%) to aggressive (+8%) edges — produced negative simulated return, and bettors following the model would have systematically failed to beat the closing line (closing line value at or below zero throughout).
The bucket structure matters more than the headline. Losses concentrated exactly where an inefficient market would leave money on the table, which is to say: they didn't. If the model had found soft prices anywhere, the high-EV buckets would have outperformed the low-EV ones. They did not, and the monotonic relationship between model-implied edge and realised loss is itself evidence that the market already incorporates the information our features carry.
A common misreading of studies like this one is that football markets are unbeatable. The correct reading is narrower: over our sample period, at the odds sources we sampled, the specific information set we used was already priced. Closing-line efficiency is a moving target — it says the consensus of all participants a minute before kickoff is hard to beat, not that every informational corner is exhausted. Sharp bettors do find edges; they tend to live in lower-profile leagues, in timing (team news reaction speed), or in data the consensus has not yet absorbed, such as injury micro-structure or weather-adjusted tactical shifts.
For a research site, the honest conclusion is the one we published: within our scope, we found nothing. That conclusion has a shelf life. As new data sources arrive — shots on target tracking is next on our roadmap — the study will be rerun and the results published either way.
Three caveats bound these conclusions. First, odds were sampled from a limited set of sources, and closing-line efficiency varies by bookmaker; a sharper closing line implies more margin for error elsewhere, not less. Second, the sample is dominated by the top European leagues, where market scrutiny is most intense — results may not transfer to second divisions. Third, our simulation assumed flat staking and no in-play hedging; sophisticated execution could narrow, though in our view not close, the observed gap. None of these caveats, in our judgement, changes the direction of the finding.