Edge Roadmap — July 2026
Status: review-owned. Updated after each strategy-shift phase review.
North star: a strategy that clears the promotion gates out of sample —
≥100 settled paper trades, ROI > +5%, positive mean AND median CLV — placed at
opening prices. Everything on this roadmap either moves toward that or gets cut.
Where we stand (evidence, not vibes)
| Claim | Status | Evidence |
|---|---|---|
| The OPENING price is soft | Established | line_ridge_core +4.4% ROI exists only at open (CI −4.4/+19); open-priced models' edge signals confirmed by the close at 54–55% (CLA z = 2.7–3.4); close-featured diagnostics settle negative |
| The closing price is efficient | Established | Every model priced/settled at close is long-run negative |
| H2H market beatable with team-level features | Rejected so far | λ ≈ 0.02–0.10 across the bench — the market does ≥90% of the predicting |
| ss_strength_v1 > Elo | Established (small) | G1 pass, best raw ECE on bench; λ 0.055 — a better component, not an edge |
| Compound-Poisson scoring > ridge | Rejected (as built) | G3/G4 fail both runs; one iteration remains before retirement |
| Lineup/RAPM channel adds signal | SIGNAL CONFIRMED, GATE NOT CLEARED | 2026-07-10: two bugs fixed (date-key join + lowercase team-key mismatch). With real deltas: G7 PASS λ=0.141 (lineup carries market-beating information); G6 FAIL by 0.5 milli-log-loss (0.6266 vs 0.6261 ablation). CLA z=5.22 — market moves toward RAPM picks. See phase3-rapm-v3 / -decay-v2 / phase4-stats-prior-v2 reports. |
| Team-list timing edge (Tuesday window) | Untested hypothesis | The one mechanism consistent with everything above; needs live paper trades + CLV |
| Line-move direction predictable from Tuesday features | CONFIRMED + positive backtested CLV (2026-07-25) | line-move-v1b: full 66.8% vs market_form_only 59.0% vs open_line_only 53.5%. Lineup channel adds 7.74pp over market structure, paired z=4.18 (mean-reversion-only REJECTED — open line alone barely beats base). CLV leg: mean +1.47 line points, 95% CI [1.08,1.87] excludes zero, 8 seasons. Backtested, not live. See review-2026-07-25.md |
| Lineup channel moves the LINE (Law 3 for line market) | CONFIRMED (2026-07-25) | Market-only ablation isolates it: 7.74pp lift over market+form, paired z=4.18. Single-feature ablations were masked by multicollinearity. First gate-clearing confirmation of the lineup channel |
| Favourites underpriced at open (models confirm) | Established (2026-07-12) | CLA segmentation: backing the favourite = 62.8–65.0% agreement (z 5.5–6.9); backing dogs = 41–49% (anti-signal); alive in 2023+ era. Defines open_fav_v1. See docs/edge-review-2026-07-12.md |
open_fav_v1 backtest (both legs) |
CLA strong, ROI unproven — live CLV adjudicates | 2026-07-16: OF1 PASS (CLA 0.737 / 0.690 on 2023+, 307 eligible); ROI report-only: flat +1.2% CI [−7.2,+9.1], edge_capped −1.1% CI [−9.9,+7.9]. Filter carries the signal (G2/G6 fail individually). First 2 paper trades seeded R20. See results-open-fav-v1-roi + review-2026-07-16 |
| Underdog business-logic overlay | Tested and REJECTED (4 of 6 segments) | 2026-07-16 pre-registered diagnostic: no segment clears CLA ≥0.55 ∧ z ≥3 ∧ n ≥100 (best: away dog 0.637 z=2.62). origin_week + rain segments UNEVALUABLE (data gaps) — one re-run allowed after backfill, then closed |
| H2H↔line cross-market divergence | Tested and KILLED | Line stopped correcting toward H2H-implied margins after 2018 (59.5% pre-2018 → 37–43% since). Archived; revisit only with exchange data |
Instrument state: leakage guards test-enforced; CLA in every report;
bootstrap CIs require ≥5 seasons AND ≥100 bets; results contract (schema v3)
round-trips between sessions; UI surfaces models/strategy/ops/DQ/lineage.
Re-sequenced 2026-07-12 (docs/edge-review-2026-07-12.md is the detail)
Priority order now: N1 open-fav-v1 backtest verdict → N2 live evidence
quota every round (both strategies) → everything else. cpois RETIRED;
phase3-rapm-v4 demoted to idle rounds; H2H feature engineering FROZEN
pending a new information source; Betfair exchange data priced in September
(N4 — the only route to a materially better backtest).
Horizon 0 — superseded (kept for history: the week of R19)
- DONE ✓ Re-run RAPM with fixed join —
phase3-rapm-v3committed 2026-07-10.
Two bugs fixed: kickoff-date join and lowercase team-key mismatch. First real
deltas confirmed (abs sum 6731 across 2018–2026). Results: G7 PASS λ=0.141,
G6 FAIL 0.6266 vs 0.6261 (ablation). CLA z=5.22. - DONE ✓ Decay and stats-prior re-runs —
phase3-rapm-decay-v2and
phase4-stats-prior-v2. Same pattern: G7 PASS, G6 FAIL by 0.5 milli-log-loss. - Keep the live loop fed: R19 close snapshots per kickoff
(schedule_close_snapshots.py), CLV update post-round,rapm_timing_v1
paper trades at team-list drop. The Jobs page shows if any of this slips.
Horizon 1 — next 3–4 rounds
- Line market first. The only established soft spot is opening lines, and
the strongest model is line_ridge_core. Two cheap upgrades, gated on MAE +
open-line ROI CI: (a) feedss_strength_v1margin + variance into the
cover-probability (replace the pooled residual σ with the state-space
variance per match); (b) add the now-actually-populated lineup deltas as
features. Report as--phase line-v2. - CLA-gate every new feature family: if adding a family doesn't lift CLA
or log-loss vs its ablation control, it doesn't ship. No exceptions —
this is what kept the fake RAPM result out of production twice. - Accumulate
rapm_timing_v1evidence — target ≥40 trades by end of
H1 with CLV populated; interim read at 40 (direction only, no promotion).
Horizon 2 — by end of season 2026
- Promotion decision on the line channel: line_ridge(-v2) at open with
shrunk-Kelly staking — does the season's paper sample clear the gates?
That's the first candidate with a realistic shot. - RAPM verdict: with real deltas, either the lineup channel earns λ ≥ 0.10
somewhere (season-grain OR timing-grain) or the whole family is archived
with the evidence attached. - cpois final iteration (fed by ss strengths, negative-binomial tries) —
pass G3/G4 or retire. - 2027 pre-season: whatever survived gets a pre-registered live plan;
open-price capture automation (the Monday open snapshot is the single most
valuable data point we collect) hardened before R1.
Review 2026-07-11 — reading phase3-rapm-v3 correctly
The G6/G7 split is not a contradiction; the two gates measure different things
and the split tells us what to build next:
- λ 0.141 vs control 0.103 and CLA 0.583 (988 signals, z=5.22) vs 0.550
(733): the lineup delta makes the model disagree with the open MORE OFTEN
and be confirmed by the close MORE OFTEN — the strongest closing-line
agreement of any actionable model. Direction is right. This is information
the opening market does not have. - G6 fail by 0.0005 log-loss with slightly worse ECE (0.019 vs 0.016):
the delta's MAGNITUDE is noisy — directionally right, poorly scaled, so
probability quality doesn't improve even though betting-relevant signal does. - Consequence (two work items, pre-registered):
1. Magnitude calibration: fit the delta through a monotone/clipped
transform inside the walk-forward (e.g. per-season standardization +
winsorize at ±2σ before the logit). Gate: G6 vs the same control, plus
ECE must not degrade. Phase labelphase3-rapm-v4.
2. Timing channel unchanged as the decisive test — CLA 0.583 at
season-grain deltas is exactly the precondition the Tuesday-window
hypothesis needed.rapm_timing_v1trades are now the highest-value
evidence in the program. - Shrunk-Kelly ROIs on these models (39–48 bets) are small-sample noise and
correctly not significant — do not quote them. - UI note: model cards fixed this round (deterministic report pick by content
timestamp, cross-report metric fallback so no track ever renders blank,
market-grouped ordering H2H → Line → Total, primary-model fallback metrics).
Standing guardrails
- Nothing skips the results contract; failed gates get pushed.
- No gate edits in result commits; ROI never quoted without its CI.
- Closing data: diagnostic labels only, enforced by tests.
- The catalogue (/models), strategy page, jobs page, glossary, and this
roadmap update together — a roadmap the UI contradicts is worse than none.