How we trained our model

How we trained our model no shortcuts

Most prediction sites are a feeling wearing a percentage. Ours is the opposite: an engine that had to earn every number it publishes · against twelve years of reality, an adversarial test bench, and a market that punishes wishful thinking.

The raw material

  • 86,000+ matches across 21 leagues and 12 seasons (2014–2026), every file checksummed and source-stamped.
  • 21,500+ matches with true expected-goals data across the top five leagues · not estimates, measured shot quality.
  • 116,448 individual shots with pitch coordinates and 129,000+ player-match rows harvested for the next research wave.
  • Per match: result, shots, shots on target, corners, fouls, cards, referee, schedule load, rest days, promotion status · plus full odds boards and the sharpest closing lines in the market.

The training discipline

The engine is trained walk-forward only: it predicts every round using strictly what was knowable before kickoff, then moves one step forward · across ten full out-of-sample seasons. Random shuffles, peeking, or retro-fitting are structurally impossible: a dedicated leakage firewall is enforced by the test suite on every single change.

We battle-tested more than a dozen mathematical model families · count processes, dynamic ratings, shot-quality models, regularised learners and layered ensembles · and kept only what survived paired statistical judgement with multiple-testing correction. The losers are documented, not deleted: a rejected idea is evidence too.

The gauntlet every change must survive

  • 299 automated gates run on every change · calibration, leakage, contract and chaos tests.
  • Tens of thousands of Monte Carlo paths stress bankroll, variance and edge-overconfidence before anything ships.
  • Placebo controls: every "edge" must beat a deliberately random twin, or it dies.
  • Cluster-aware bootstrap + multiple-testing correction: correlated bets never masquerade as independent evidence.
  • AI agent fleets · differently-prompted adversarial reviewers · red-team the codebase and try to break the engine on purpose. Countless hours of it.

Why you can trust the numbers

Every published probability is calibrated against reality (our holdout calibration error sits at market level), stamped with the exact engine build that produced it, and written to an append-only archive the day it was made. We can replay any prediction we have ever published · and so can history. See the track record and The Honesty Ledger, where every settled call is judged in public.

Plain-language glossary

The terms our coach and prediction pages rely on, in plain English. Honesty includes not hiding behind jargon.

Expected goals (xG)
A measure of chance quality. It estimates how likely an average shot was to be scored from where and how it was taken, then adds those up. It shows whether a result was deserved or just lucky.
Calibration
Whether our percentages mean what they say. If we label many games ‘60% home win’ and home teams win about 60% of them, the model is calibrated. We publish our calibration error, and it is the one thing we can actually prove.
Closing line value (CLV)
Did you get a better price than the market’s final, sharpest price before kickoff? Beating that closing line over many bets is the only reliable sign of skill, because it is the hardest price to beat.
The vig (margin, de-vig)
The cut the bookmaker builds into every price. Add up the implied chances of all outcomes and they total more than 100%, and the extra is the margin. De-vigging strips it out to read the true probability. The margin is why most bettors lose over time.
Expected value (EV)
The average result of a bet if you could repeat it forever. Positive means the price is in your favour, negative means it is against you. Most public prices are slightly negative because of the margin.
No bet
Our default answer. When our calibrated probability does not clearly beat the price on offer, the disciplined call is to not bet. Saying no bet protects your bankroll, and it is often the most valuable thing we publish.
Base rate
How often something normally happens, measured over a large sample. A trend only means something compared with its base rate and a big enough sample, otherwise it is noise.
Regression to the mean
Extreme runs tend to fade back toward normal. A team far above or below its xG usually drifts back, so hot and cold streaks predict less than they feel like they should.
Variance
Short-term luck. Even a genuine edge loses over small samples purely by chance, which is why a handful of bets tells you almost nothing. It is why the coach needs at least 20 bets before it says anything.
Drawdown
The drop from the highest point your bankroll reached to a later low. Knowing your worst drawdown tells you how much swing you have already survived, and need to be able to survive.
Bankroll and staking
Your bankroll is the money set aside for betting. Staking is how much of it you risk per bet. A disciplined norm is 1 to 2% per bet at a flat size, while above 5% is over-exposure that variance can wipe out.
Fractional Kelly
A formula that sizes a bet to its edge. Because edges are usually overestimated, we use a fraction of it (a quarter) plus hard caps to stay safe. With no proven edge, a flat stake is safer.
Parlay (acca)
One bet combining several legs that must all win. It multiplies the bookmaker margin, not your edge, so without a real edge the expected value drops fast with each leg. Fewer legs, flat stakes.
Favourite-longshot bias
Bettors tend to overpay for long odds and underpay short ones, so longshots are usually the worse value. If your high-odds bets lose money over many tries, this is the likely leak.
Walk-forward testing
Testing the model the honest way: predict each round using only what was known before kickoff, then step one round forward, never peeking at the future. It is how we avoid fooling ourselves.
Loss-chasing and tilt
Raising your stakes to win back losses, or betting on emotion. It is the single most-cited marker of gambling harm, and the coach watches your own logged bets for it.
PGSI
The Problem Gambling Severity Index, a short, standard nine-question self-check about the last 12 months. It is a tool for reflection, not a diagnosis, and the coach includes it.

What we will not claim: that we beat the closing line. Nobody honest does, consistently, on public data. The engine’s edge is discipline · calibrated probabilities, an open No-bet log, and risk mathematics that keep you in the game. The exact recipes, weights and thresholds stay our trade secret.

Not financial advice. No model beats the sharp closing market on public data. The value here is honest probabilities, No-bet calls, and CLV history.
Updated: 2026-07-28 09:16:08 · build 3c83206 · calibration isotonic-1x2.v1
Glossary · 18+ · Please gamble responsibly. BeGambleAware / Gambling Therapy