Model Lab
This page replays filter, staking and calibration rules against the same earliest pre-event system records with settled outcomes. It does not use synthetic fixtures, and the variants are not separate trained production models. No money is attached; promotion requires the real-data gates below.
Variants
· 12 rule variants · same real pre-event recordsCumulative profit (staked units)
· Same fixtures · different rulesSide-by-side metrics
| Variant | Picks | Settled | W–L | Win% | Flat ROI | Staked ROI | +/− units | Avg edge | CLV | Brier |
|---|---|---|---|---|---|---|---|---|---|---|
Baseline (live) | 474 | 195 | 127–68 | 65.1% | -4.81% | -3.46% | -16.51u | -1.81pp | -1.81pp | 0.2112 |
Baseline · edge ≥ 3pp | 1 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.11u | +3.23pp | +3.23pp | 0.3120 |
Baseline · edge ≥ 7pp | 0 | 0 | 0–0 | — | — | — | +0.00u | +0.00pp | — | — |
Baseline · half stake | 474 | 195 | 127–68 | 65.1% | -4.81% | -3.46% | -8.25u | -1.81pp | -1.81pp | 0.2112 |
Baseline · ≥ 4 books | 418 | 195 | 127–68 | 65.1% | -4.81% | -3.46% | -16.51u | -1.73pp | -1.73pp | 0.2112 |
Edge ≥ 5pp | 0 | 0 | 0–0 | — | — | — | +0.00u | +0.00pp | — | — |
Market-shrunk | 1 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.00u | +2.26pp | +3.23pp | 0.3013 |
Market-shrunk · light | 4 | 4 | 1–3 | 25.0% | -58.25% | -58.25% | -2.33u | +2.41pp | +2.68pp | 0.3790 |
Favourites only | 336 | 166 | 116–50 | 69.9% | -0.71% | -0.71% | -1.18u | -2.09pp | -2.09pp | 0.2037 |
Value underdogs | 0 | 0 | 0–0 | — | — | — | +0.00u | +0.00pp | — | — |
Calibrated (Platt-lite) | 6 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.00u | +2.33pp | +0.98pp | 0.3023 |
Calibrated · strong shrink | 22 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.00u | +2.94pp | -0.06pp | 0.1877 |
Methodology
· How a variant gets promotedEvery variant is scored on the same auto-logged picks the live baseline model sees. Variants only differ in their filter rule (which picks they take) and staking rule. Outcomes (won / lost / void) come from the same reconciled results — no rewriting history.
- Baseline is whatever you see on the public track record today (recommended stake, no filter).
- Challenger variants paper-trade alongside. No money is attached.
- A challenger is promoted to baseline only when it beats the live model on staked ROI over ≥100 settled picks with positive CLV. Single-week noise doesn't move the needle.
- Brier score is computed on each variant's emitted probability (calibrated/shrunk variants use the transformed prob), so calibration improvements show up directly.
Research only. Past simulated performance is not indicative of future results.
Monte Carlo simulator
· Stress-test any model: ours or yoursRuns thousands of simulated betting seasons against the inputs below to show the full distribution of outcomes — not just the average. Use it to see how often a model bankrupts, what a realistic worst-case looks like, and whether the median path is actually profitable. Upload your own picks as CSV to bootstrap from real history instead of assumed parameters.
CSV with columns edge_pct, odds, stake (header optional, stake optional). When uploaded, the simulator samples real picks instead of using the edge% / odds inputs below. Staking rule from inputs is used when stake is 0 or missing.