Model changelog

Every change to the model and its rules, with the date, the evidence, and what it did. Reverted changes stay in the record. The live tallies for current findings are on Nudges; the method is on How it works.

Rule

Weather rule; pools; the sharp crowd

2026-09-18

What. Outdoors: 15+ mph wind shrinks The System toward 50 by a fifth, 50F or colder adds 2 to the favorite, over 75F shrinks by a tenth; forecasts from Open-Meteo. Survivor and confidence-pool optimizers. A track-record-weighted crowd scored as a forecaster.

Evidence. Weather on 1999–2022 outdoor games: wind 15+ favorites 62% vs 65% forecast (626), cold 69% vs 66% (1,517), heat 61% vs 63% (742); +9 a season, positive in each era.

Effect. Rule 10 of The System; new Pools page; new leaderboard row.

Reverted

Refit on 24 seasons: blend to 25% Elo, tight bump to 2, lock press and playoff weighting removed

2026-09-18

What. The System was refitted on 1999–2022 (FiveThirtyEight's QB-adjusted Elo vs nflverse closing spreads, 6,064 games) instead of our 2023–25 replays alone.

Evidence. Per season over 24: market 977, 25/75 blend 980, 50/50 blend 961, Elo 890. Tight-agreement favorites 69.2% vs 67.7% forecast; an 8-point bump cost 40 a season. Lock 379–71 (84%) vs 80%; pressing by 5 cost points over 450 locks. Playoffs: Elo 1,610, market 1,617 over 248 games.

Effect. The System now scores a hair above the market over the long run and is expected to, rather than 180 points above it on a fitted sample.

Rule

Game plan modes

2026-09-18

What. Protect (k 0.9), Straight, Chase (k 1.15) on The System's numbers, recommended from your gap to the leader.

Evidence. Stretching by k: 0.9 = 961/season, 1.0 = 980, 1.05 = 984, 1.15 = 977; weekly sd 54 / 60 / 62 / 68.

Effect. Chase is nearly free in expectation; Protect costs about a point a week.

Rule

The System defined

2026-09-18

What. Nine slider rules, each tested; the page lists what was rejected.

Evidence. Caps at 90/85/80, market-only on disagreements, and coin flips to 50 all cost points in both periods.

Effect. One accountable forecaster row on every leaderboard.

Model

Resting-starters lean targets locked teams

2026-09-18

What. From week 15, games involving a team locked into its seed or eliminated (from the simulation) are forecast 25% Elo / 75% market; other week 18 games 50/50.

Evidence. Week 18 was Elo's worst slice: Brier 0.246 vs the market's 0.203.

Effect. Replaces the blanket week 18 lean with a targeted one.

Reverted

Tight-agreement bump set to 8 (later cut to 2)

2026-09-17

What. The blend pushed 8 points toward the favorite when Elo and the market agreed within 4 points.

Evidence. On 2023–25: favorites 74% vs 66% forecast, +217 points. On 1999–2022 the surplus is 1.5 points and the +8 bump costs 40 a season.

Effect. Overfit to three seasons. Reduced to 2 the next day.

Model

Half offseason reversion

2026-09-17

What. Ratings revert halfway to 1505 between seasons instead of a third.

Evidence. Tuning on 2023–25 replays: +37 points over three seasons; teams Elo lagged the market on were regime changes.

Effect. Small; kept because the direction matches what the market sees early each season.

Model

Rookie quarterback priors

2026-09-17

What. Rookies start at a draft-slot prior (≈40 for the first pick, low 20s late, 15 undrafted) instead of zero.

Evidence. Fitted on 2021–25 rookie starters; league mean start ≈ 49. Daniels's week 2 adjustment moved from −25 to +55.

Effect. Three of the fifteen worst misses were rookie games.

Model

QB warm-up bug fixed

2026-09-17

What. A quarterback's first rolling value started 15 below a team value that had just jumped to the game's full value, so every new QB looked like a −280 backup for weeks.

Evidence. 141 of 855 replayed games had an adjustment pinned at −100, 37 on both sides. After the fix: 34, none on both sides. Elo's replay points rose from 2,815 to 2,927.

Effect. The largest single improvement to the model.

Rule

Lock and bold call of the week

2026-09-17

What. Two frozen picks per week: strongest agreement between Elo and the market, and their biggest disagreement on Elo's side.

Evidence. 2023–25: lock 59–6, bold 34–32. Later confirmed on 1999–2022: 379–71 and 217–233.

Effect. The lock replaced the bold call as the featured pick after the backtest.

Model

Markets as forecasters

2026-09-17

What. Sportsbooks, Kalshi and Polymarket stored per game and scored on every leaderboard on their closing numbers; the blend introduced.

Evidence. Market vs Elo over the replays: 3,483 vs 2,841 points.

Effect. The market became the benchmark, and the blend the site's own forecaster.

Model

FiveThirtyEight's Elo, seeded from their final ratings

2026-09-17

What. 1505 mean, K 20, home field 48, bye +25, playoff ×1.2, margin-of-victory multiplier, QB adjustment at 3.3 per VALUE point. Seeded from 538's end-of-2022 ratings and replayed through 2023–25.

Evidence. 538's published methodology.

Effect. The starting point.