Model Lab
Baseline vs challenger paper-trading. Every variant is evaluated on the exact same auto-logged fixtures as the live track record — only the filter rule (which picks to take) and staking rule change. No money is attached; this is how a quant desk evolves a model in public.
Variants
· 12 models · same fixturesCumulative profit (staked units)
· Same fixtures · different rulesSide-by-side metrics
| Variant | Picks | Settled | W–L | Win% | Flat ROI | Staked ROI | +/− units | Avg edge | CLV | Brier |
|---|---|---|---|---|---|---|---|---|---|---|
Baseline (live) | 324 | 195 | 127–68 | 65.1% | -4.81% | -3.46% | -16.51u | -1.61pp | -1.61pp | 0.2112 |
Baseline · edge ≥ 3pp | 1 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.11u | +3.23pp | +3.23pp | 0.3120 |
Baseline · edge ≥ 7pp | 0 | 0 | 0–0 | — | — | — | +0.00u | +0.00pp | — | — |
Baseline · half stake | 324 | 195 | 127–68 | 65.1% | -4.81% | -3.46% | -8.25u | -1.61pp | -1.61pp | 0.2112 |
Baseline · ≥ 4 books | 272 | 195 | 127–68 | 65.1% | -4.81% | -3.46% | -16.51u | -1.47pp | -1.47pp | 0.2112 |
Edge ≥ 5pp | 0 | 0 | 0–0 | — | — | — | +0.00u | +0.00pp | — | — |
Market-shrunk | 1 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.00u | +2.26pp | +3.23pp | 0.3013 |
Market-shrunk · light | 4 | 4 | 1–3 | 25.0% | -58.25% | -58.25% | -2.33u | +2.41pp | +2.68pp | 0.3790 |
Favourites only | 269 | 166 | 116–50 | 69.9% | -0.71% | -0.71% | -1.18u | -1.78pp | -1.78pp | 0.2037 |
Value underdogs | 0 | 0 | 0–0 | — | — | — | +0.00u | +0.00pp | — | — |
Calibrated (Platt-lite) | 2 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.00u | +2.36pp | +1.99pp | 0.3023 |
Calibrated · strong shrink | 4 | 1 | 0–1 | 0.0% | -100.00% | -100.00% | -1.00u | +2.72pp | +0.13pp | 0.1877 |
Methodology
· How a variant gets promotedEvery variant is scored on the same auto-logged picks the live baseline model sees. Variants only differ in their filter rule (which picks they take) and staking rule. Outcomes (won / lost / void) come from the same reconciled results — no rewriting history.
- Baseline is whatever you see on the public track record today (recommended stake, no filter).
- Challenger variants paper-trade alongside. No money is attached.
- A challenger is promoted to baseline only when it beats the live model on staked ROI over ≥100 settled picks with positive CLV. Single-week noise doesn't move the needle.
- Brier score is computed on each variant's emitted probability (calibrated/shrunk variants use the transformed prob), so calibration improvements show up directly.
Research only. Past simulated performance is not indicative of future results.
Monte Carlo simulator
· Stress-test any model: ours or yoursRuns thousands of simulated betting seasons against the inputs below to show the full distribution of outcomes — not just the average. Use it to see how often a model bankrupts, what a realistic worst-case looks like, and whether the median path is actually profitable. Upload your own picks as CSV to bootstrap from real history instead of assumed parameters.
CSV with columns edge_pct, odds, stake (header optional, stake optional). When uploaded, the simulator samples real picks instead of using the edge% / odds inputs below. Staking rule from inputs is used when stake is 0 or missing.