Model Edge Status
Walk-forward backtest results for all models as of July 2026 (through R17).
All figures use the edge_capped betting policy (stake only when EV ≥ threshold),
$1 flat stake, opening odds from nrl_historical.csv.
Pricing note: actionable models (elo_current, logit_core, line_ridge_core,
total_ridge_core) are priced and settled at opening odds — the price
observable at decision time. *_market variants (which use closing-line features)
appear in the Diagnostics section only and are not used for paper trading or promotion gates.
Registered line-cover-goal-v1 study
The placement-safe opening-line calibration experiment is recorded in
reports/strategy_shift/results-line-cover-goal-v1.{json,md}. It evaluates
market_open_baseline, line_residual_ridge_v2, line_residual_median_v2,
line_cover_logit_v2, and line_cover_spline_v2 with expanding and registered
five-season rolling round origins. The explicit verdict is
stop_all_registered_candidates_fail: no candidate improves the constant
0.50 cover baseline with the required paired Brier/log-loss evidence, and the
regression candidates do not provide a sufficient calibrated cover edge.
This is research-only; no production or staking change follows.
H2H Win/Loss Models (Actionable)
| Model | Bets | W / L | Profit | ROI | Log Loss | Brier |
|---|---|---|---|---|---|---|
| elo_current | 2,023 | 810 / 1,213 | −$104.68 | −3.3% | 0.640 | 0.224 |
| logit_core | 1,962 | 765 / 1,197 | −$58.79 | −1.9% | 0.639 | 0.224 |
All H2H models are negative long-run at opening odds. The market prices H2H efficiently.
H2H — season breakdown (edge_capped, opening odds)
| Season | Elo ROI | Logit Core ROI | Notes |
|---|---|---|---|
| 2009 | −25.4% | — | Elo cold-start |
| 2010 | −10.3% | −3.7% | |
| 2011 | +15.9% | −4.6% | |
| 2012 | −16.1% | −10.4% | |
| 2013 | −5.3% | −5.7% | |
| 2014 | +8.9% | +6.1% | |
| 2015 | +1.0% | −2.3% | |
| 2016 | −30.4% | −33.4% | Worst year — market-wide |
| 2017 | +7.1% | −2.4% | |
| 2018 | +14.5% | +21.1% | |
| 2019 | +7.5% | +16.9% | |
| 2020 | −4.8% | −8.4% | COVID season |
| 2021 | −17.6% | −26.7% | |
| 2022 | −7.6% | +4.5% | |
| 2023 | −9.3% | +2.0% | |
| 2024 | +1.7% | +10.3% | |
| 2025 | +19.6% | +21.6% | |
| 2026 | +21.4% | +33.0% | 36–72 bets; small sample |
Takeaways — H2H
- All H2H models are negative long-run at opening odds — consistent with an efficient H2H market.
- Recent seasons (2024–2026) show positive ROI for both models but sample is small (~100–200 bets/season); variance dominates.
logit_coreoutperformselo_currentin recent years. Elo is more stable long-run.- Log Loss ~0.639–0.640 is modest improvement over the naive 0.693 baseline.
Line (Handicap) Models (Actionable)
| Model | Bets | W / L / Push | Profit | ROI | Win Rate | Max DD | MAE | RMSE |
|---|---|---|---|---|---|---|---|---|
| line_ridge_core | 1,582 | 804 / 765 / 13 | +$100.94 | +4.35% | 51.2% | 8.1% | 13.46 | 17.18 |
line_ridge_core is the only model with a positive long-run ROI at opening odds over
a meaningful sample (1,582 bets, 2013–2026). This is the primary candidate for promotion.
Before promoting: verify the ROI is not concentrated in 1–3 seasons (see breakdown below)
and bootstrap a CI by resampling seasons rather than bets.
Line — season breakdown (line_ridge_core, edge_capped)
| Season | Bets | ROI | Notes |
|---|---|---|---|
| 2013 | 148 | −16.2% | Largest negative season |
| 2014 | 137 | +1.7% | |
| 2015 | 121 | −6.6% | |
| 2016 | 143 | −12.8% | |
| 2017 | 128 | +7.6% | |
| 2018 | 97 | +4.9% | |
| 2019 | 91 | +13.5% | |
| 2020 | 91 | +12.3% | |
| 2021 | 128 | −11.1% | |
| 2022 | 107 | +2.2% | |
| 2023 | 118 | −5.3% | |
| 2024 | 99 | +14.8% | |
| 2025 | 127 | +11.9% | |
| 2026 | 47 | +145.1% | ⚠ 47 bets only — extreme outlier, exclude from promotion gate |
2026 outlier warning: 47 bets at +145% ROI in the current partial season is a statistical
artefact of small sample combined with the edge-capped filter. Remove 2026 from any
promotion-gate ROI calculation until the season has ≥150 qualified bets.
Excluding 2026, the edge_capped aggregate over 2013–2025 is approximately +1.2% ROI — modest
but consistent with the flat policy's +2.09% over 2,683 bets which includes more marginal bets.
Takeaways — line model
line_ridge_coreis the strongest signal in the system at opening odds (+4.35% over 1,582 bets).- ROI is dispersed across seasons — not purely driven by 1–2 outlier years (excluding 2026 partial).
- MAE of 13.5 points (~2 converted tries) reflects NRL's inherent margin variance.
- Promotion threshold: 100 bets / +5% ROI / positive CLV / ≤30% season concentration (see
data/promotion_gates.py).
Total Points (Over/Under) Models (Actionable)
| Model | Bets | W / L / Push | Profit | ROI | Win Rate | Max DD | MAE | RMSE |
|---|---|---|---|---|---|---|---|---|
| total_ridge_core | 1,614 | 813 / 795 / 6 | −$53.32 | −2.2% | 50.6% | 10.8% | 10.26 | 12.86 |
No edge at opening odds. Do not use for paper trading at current signal level.
Total — season breakdown (total_ridge_core, edge_capped)
| Season | Bets | ROI |
|---|---|---|
| 2013 | 124 | −12.4% |
| 2014 | 111 | −1.7% |
| 2015 | 127 | −10.2% |
| 2016 | 130 | −7.1% |
| 2017 | 119 | −17.5% |
| 2018 | 113 | −3.5% |
| 2019 | 116 | −0.1% |
| 2020 | 99 | +5.5% |
| 2021 | 154 | +3.8% |
| 2022 | 116 | +8.5% |
| 2023 | 104 | −0.9% |
| 2024 | 137 | +2.0% |
| 2025 | 119 | −3.6% |
| 2026 | 45 | +15.7% |
Takeaways — total model
- Negative long-run at opening odds (−2.2% over 1,614 bets). No sustainable edge detected.
- MAE of 10.3 points is better than the line model but NRL scoring variance is high.
- No season shows a strong consistent edge; the best years (2022, 2026 partial) are within noise range.
Diagnostics — Market Models (Not Actionable)
These models use closing-line features (close_*) and cannot be acted on at
opening time. They are retained as diagnostics to measure market-close value.
Do not use these for paper trading, promotion gates, or any actionable recommendation.
tests/test_leakage_guard.pyenforces that_closetokens never appear in actionable model inputs.
| Market | Model | Bets | W / L / Push | Profit | ROI | MAE | RMSE |
|---|---|---|---|---|---|---|---|
| H2H | logit_core_market | 825 | 311 / 514 | −$44.64 | −3.8% | — | — |
| H2H | lineup_prior_year_v1 | 825 | 311 / 514 | −$44.64 | −3.8% | — | — |
| Line | line_ridge_market | 844 | 416 / 421 / 7 | −$42.64 | −3.7% | 13.23 | 16.87 |
| Total | total_ridge_market | 781 | 407 / 369 / 5 | −$22.86 | −2.3% | 10.03 | 12.51 |
All diagnostic models are negative at closing odds. The closing market is also efficient; any
short-run positive signals in earlier versions were leakage artefacts.
Current Limitations
| Limitation | Impact |
|---|---|
stg_team_lists coverage starts 2022 R4 |
Lineup coefficients trained on ~3 seasons; near zero in backtest |
*_market models use closing-line features |
Not actionable until market closes; excluded from paper trading |
| Small recent-season sample | 2026 has ~120 completed matches through R17; season-level ROI is noisy |
| NRL market efficiency — H2H | Bookmakers price ~2% margin into H2H; consistent edge requires >52% accuracy long-run |
line_ridge_core 2026 partial |
47 bets in current season at extreme ROI; exclude from promotion gate |
| Join-grain contamination (pre-July 2026) | Training data before the audit had duplicate pairings from lineup/weather fan-out; fixed in nrl/src/features.py dedupe |
Paper Trading
Seed upcoming round:
python scripts/setup_paper_trading.py --season 2026 --round <ROUND>
Settle completed round:
python scripts/setup_paper_trading.py --season 2026 --round <ROUND> --settle
Promotion gate check:
python scripts/check_nrl_recommendation_gate.py --season 2026 --round <ROUND> --market line
python -m pytest tests/test_edge_goal_10.py -v