Research
Everything the model believes was tested first, and this is the record: what was asked, on what data, how, what came out, and what changed because of it. The findings feed the Nudges page, where each one is tallied live this season, and the changelog, where every change is dated. The scripts are in the open-source repo; anyone can re-run them.
Ongoing
The sharp crowd
2026-09-18Question. Does weighting players' picks by their track record beat the raw crowd, or the market?
Data. This season's picks, weighted 1 + points/50.
Method. Scored as a forecaster on every leaderboard beside the raw crowd.
Result. Too early. Needs a few dozen players and a few weeks.
What changed. A row on the board; a Nudges entry.
The 2026 season as the test set
2026-09-18Question. Do the findings hold on games they were not found on?
Data. Every 2026 game as it finishes.
Method. Live tallies per finding on the Nudges page; the model scoreboard on Predictions and the recap; calibration on the Model page.
Result. Updated on every visit. After week 1: the closing spread led, The System second, Elo well back; tight-agreement favorites 5 of 6.
What changed. Anything that fails a full season gets revisited, and the changelog records it.
Studies
The 24-season backtest
2026-09-18Question. Which rules for combining Elo and the market hold up over decades, not just recent seasons?
Data. FiveThirtyEight's QB-adjusted Elo probabilities joined to nflverse closing spreads, 1999–2022: 6,064 decided games.
Method. Score every forecaster and rule variant as a player (rescaled Brier, playoffs doubled), per season, and compare by era.
Result. Market 977 points a season, 25% Elo / 75% market blend 980, 50/50 blend 961, Elo 890. Tight-agreement favorites won 69.2% vs 67.7% forecast; a 2-point bump helps, an 8-point bump costs 40 a season. Lock rule 379–71 (84%). Bold call 217–233 (48%). Playoffs: Elo and market level over 248 games.
What changed. The blend went to a quarter Elo with a 2-point nudge; the lock press and playoff Elo weighting fitted on three seasons were removed. Reproduce: npm run backtest:long
Which weekly pick rule wins
2026-09-17Question. Given one pick a week, which game should it be?
Data. Our 2023–25 replays (855 games), then confirmed on 1999–2022.
Method. Three rules scored week by week: biggest Elo–market disagreement, strongest agreement, shrunk relative edge; plus seven gap-based tiebreak variants.
Result. Strongest agreement 59–6 recently, 84% over 24 seasons, against an 80% forecast. Biggest disagreement about 48%, with Elo claiming 62%. Gap-based tiebreaks did not beat the simple rule (58–7, 56–9).
What changed. The lock of the week uses the agreement rule; the disagreement rule stays as the bold call, labeled as the contrarian view. Reproduce: npm run backtest
Tuning the Elo constants
2026-09-17Question. Are FiveThirtyEight's constants still the right ones?
Data. 2023–25 replays from their 2022 seed.
Method. Replay with each constant varied, then coordinate descent; score in prediction-game points.
Result. Turning the QB adjustment off costs 357 points over three seasons, so it matters. Half offseason reversion gains 37. Home field 35 (vs 48) and QB multiplier 4.5 (vs 3.3) gained 13 each, inside the noise.
What changed. Reversion set to one half. Home field and the QB multiplier left alone. Reproduce: npm run tune
The quarterback warm-up bug
2026-09-17Question. Why were so many games showing a −100 QB adjustment?
Data. 855 replayed games.
Method. Count games with the adjustment pinned at the cap; trace the first rolling value of each new quarterback.
Result. 141 games capped, 37 on both sides including Mahomes and Hurts in 2023 week 2. A new QB's first value started 15 below a team value that had just jumped to the game's full value.
What changed. New quarterbacks start at the game value when neither side has history, at a draft-slot prior when rookies, a bit below the team otherwise. Capped games fell to 34, none double. Elo gained 112 points over the replay.
Rookie quarterback priors
2026-09-17Question. What should a rookie start at?
Data. nflverse weekly QB stats and draft data, rookie starters 2021–25.
Method. Mean VALUE over each rookie's first eight starts, regressed on the log of draft pick.
Result. Prior ≈ 40 − 3.2·ln(pick): about 40 for the first pick, low 20s late, 15 undrafted; league mean start ≈ 49.
What changed. Rookies start at the prior. Jayden Daniels's week 2 adjustment moved from −25 to +55.
Where Elo fails
2026-09-18Question. What do Elo's worst misses have in common?
Data. 855 replayed games with closing lines.
Method. Slice Elo's Brier score and its 60 worst misses by disagreement size, week, kickoff slot, rest, division, and margin.
Result. In 55 of the 60 worst misses the market liked the winner more. 15+ point disagreements: Elo's side won 48% while Elo said 65%. Week 18 Brier 0.246 vs the market's 0.203. Monday night home teams won 49% against 58% forecast (63 games). Half the worst misses were decided by a field goal or less.
What changed. Week 18 and locked-team games lean on the market. Monday and Thursday effects are on watch on the Nudges page.
Weather
2026-09-18Question. Do wind, cold and heat move results beyond what the market already prices?
Data. 1999–2022 outdoor games with weather on record in nflverse: 4,499 games.
Method. Favorite win rate versus forecast by wind and temperature band; then rule variants scored on top of The System by era.
Result. Wind 15+ mph: favorites won 62% vs 65% forecast (626 games). 50°F or colder: 69% vs 66% (1,517). Over 75°F: 61% vs 63% (742). Combined rule +9 points a season, positive in every era. Scoring runs about 1.5 points under the total in wind.
What changed. Rule 10 of The System, with kickoff forecasts from Open-Meteo.
How hard to lean
2026-09-18Question. What happens to expected points and variance when you stretch or shrink every number?
Data. 1999–2022 with the corrected blend.
Method. Scale each probability's distance from 50 by k and score per season with weekly standard deviation.
Result. k 0.9: 961 a season, sd 54.5. k 1.0: 980, sd 59.8. k 1.05: 984, sd 62.5. k 1.15: 977, sd 67.7. Expectation barely moves between 0.95 and 1.15; swing moves a lot.
What changed. The game plan's Protect, Straight and Chase modes, recommended from your gap to the leader.
Rules that lost
2026-09-18Question. Do the intuitive slider rules help?
Data. Both periods.
Method. Score each rule on top of the blend.
Result. Capping confidence at 90 / 85 / 80 lost 16 / 60 / 148 points on the replays. Taking the market outright on big disagreements lost 51. Forcing coin flips to 50 lost 29. Pressing the lock by 5 lost points over 450 locks. Weighting Elo 75% in the playoffs lost over 248 playoff games.
What changed. None of them are in The System, and the page says so.
The bye bonus
2026-09-18Question. Is +25 Elo for a team off a bye earning its keep?
Data. 2023–25 replays, then 1999–2022 regular season with nflverse rest days (bye = 10+ days).
Method. Win rate of the team off a bye vs what Elo (with the bonus) and the market said.
Result. Recent: home off a bye 51% vs 61% said (80 games), which looked bad. Long run: home off a bye 58.8% vs 60.2% (502), away off a bye 44.4% vs 44.1% (545). Calibrated.
What changed. Nothing. The recent shortfall was a small-sample fluke, and the Nudges entry is closed.
Something you'd like tested? The data is all here and the scripts run in a minute. New results land on this page as they come.