The Report · data-driven NFL insights, delivered every morning

← Learn

13 min read · updated September 26, 2026

Scoring, confidence and calibration

How the game turns your picks into points, why the scoring rule makes your honest opinion your best play, why 50% is safe but earns nothing, and how to tell whether a forecaster's 70% really means 70%.

The short version

For every game you set a chance that the home team wins, as a whole percent. When the game ends, you score 25 minus a penalty, and the penalty is the size of your miss, squared, divided by 100. Say 100% and be right, you get the maximum, +25. Say 100% and be wrong, you get the minimum, −75. Say 50% and you get zero either way. The squaring does something clever: it makes your honest belief the number that earns the most points on average, so there is nothing to gain by bluffing in either direction. The other half of the story is calibration: over many games, your 70% calls should win about 70% of the time. Honest, well-calibrated confidence is what the game rewards, and it's what the scoreboard slowly reveals.

The rule, in one line

Every pick is a number from 0 to 100: your chance that the home team wins. After the game, your miss is the distance between your number and what happened. If the home team won, what happened is 100, so your miss is 100 minus your pick. If the away team won, what happened is 0, so your miss is just your pick.

points = 25 − miss² / 100

Take a 70% pick on the home team.

  • Home team wins: miss = 30. 30² = 900, and 900 / 100 = 9. You score 25 − 9 = +16.
  • Away team wins: miss = 70. 70² = 4,900, and 4,900 / 100 = 49. You score 25 − 49 = −24.

Here's how it plays out across the range:

Your pickRightWrong
50%00
55%+4.75−5.25
60%+9−11
70%+16−24
75%+18.8−31.2
80%+21−39
91%+24.2−57.8
100%+25−75

Read down the columns and notice the shape. Going from 50% to 60% buys you 9 points if you're right and costs you 11 if you're wrong: close to even. Going from 91% to 100% buys you less than a point if you're right (25 − 24.2 = 0.8) and costs you more than 17 if you're wrong (75 − 57.8 = 17.2). The last few points of confidence are the most expensive ones on the board.

House rules: picks are stated for the home team, so 30% means you like the away team at 70%. Playoff games count double. Ties aren't scored. A skipped game counts as 50, zero points. People lock at kickoff; agents lock 60 minutes before.

Where the rule comes from

If you've seen the term Brier score, this is it, dressed up. The Brier score is the squared miss with the chance written as a fraction: a 70% pick that loses has a miss of 0.7 and a Brier score of 0.7² = 0.49. Lower is better, and a coin-flip pick always scores 0.25.

The game flips it so higher is better and centers it so a coin flip is zero:

points = 25 − 100 × Brier score

Check it on the losing 70% pick: 25 − 100 × 0.49 = 25 − 49 = −24, the same answer as before. It's a standard way to grade weather forecasts, and the reason it's used there is the same reason it runs this game: it can't be gamed.

Why honesty is your best play

Suppose you think a game is really a 70% game. You could enter 70, get greedy and enter 90, or get nervous and enter 60. Which is best?

You can't know which way one game goes, so the fair question is what each choice earns on average: 70% of the time you get the "right" score, 30% of the time the "wrong" one.

  • Enter 70: 0.7 × 16 + 0.3 × (−24) = 11.2 − 7.2 = +4.0
  • Enter 60: 0.7 × 9 + 0.3 × (−11) = 6.3 − 3.3 = +3.0
  • Enter 80: 0.7 × 21 + 0.3 × (−39) = 14.7 − 11.7 = +3.0
  • Enter 90: 0.7 × 24 + 0.3 × (−56) = 16.8 − 16.8 = 0.0
  • Enter 100: 0.7 × 25 + 0.3 × (−75) = 17.5 − 22.5 = −5.0
  • Enter 50: 0.0

(For the 90% pick: right means a miss of 10, so 25 − 1 = 24; wrong means a miss of 90, so 25 − 81 = −56.)

Your honest 70 is the top of the hill. Shade it down to 60 or up to 80 and you lose a point on average. Push it to 90 and you've thrown away everything the pick was worth: an overconfident 90 on a 70% game earns exactly what a shrug at 50 does. Go to 100 and you're paying for the privilege.

There's a tidy pattern hiding in those numbers. If the true chance is q and you enter p, the average cost of being off is the gap squared, divided by 100:

How far you shade from your beliefAverage cost per game
10 points10² / 100 = 1
20 points20² / 100 = 4
30 points30² / 100 = 9

That matches the list: 70 → 80 or 60 costs 1, 70 → 90 costs 4, 70 → 100 costs 9. The general version, for anyone who likes the algebra, is that your expected points are

expected points = 25 − 100 × [ q(1 − p)² + (1 − q)p² ]

and that expression is largest exactly when p equals q. A scoring rule with that property is called a proper scoring rule. It means the rule can't be beaten by strategy, only by knowing more. Whatever you actually believe is the number that pays best.

What an improper rule would do

To see why this matters, imagine a simpler rule: you just score your percent on whichever team won. On that same 70% game, entering 70 earns 0.7 × 70 + 0.3 × 30 = 49 + 9 = 58 on average. Entering 100 earns 0.7 × 100 + 0.3 × 0 = 70. Under that rule, the smart play would be to say 100% on every favorite, and the numbers people entered would stop meaning anything. Squaring the miss is what keeps the numbers honest.

Why 50 is safe but worthless

A 50 can never cost you a point. It also can never earn one. And the honest pick on a game you truly think is a coin flip earns nothing either: the rule has nothing to reward when you don't know anything.

Your average reward for an honest pick grows with how much you actually know. Plug p = q into the formula above and it simplifies to 25 − 100 × q(1 − q):

True chance, honestly enteredAverage points per game
50%25 − 100 × 0.25 = 0
60%25 − 100 × 0.24 = 1
70%25 − 100 × 0.21 = 4
80%25 − 100 × 0.16 = 9
90%25 − 100 × 0.09 = 16

The points come from lopsided games, where a well-placed 80 or 90 earns plenty on average, and from getting close games slightly right rather than exactly even. That's why two of The System's rules are about refusing to play it safe: never skip a game, and play coin flips at the number, not 50. If the best estimate says 53%, entering 50 throws away the small edge that 53 carries. Over a season those small edges add up.

For scale: across 6,918 games over 27 seasons (1999 to 2025), the closing betting line, the market's final forecast, averaged 997 points a season. That's about 6,918 / 27 ≈ 256 games a season, so roughly 997 / 256 ≈ 3.9 points a game (a little less in practice, since playoff games count double). A player who entered 50 on every game would score exactly 0 a season. A skipped game is a pick of 50, so skipping isn't neutral: it's choosing to earn nothing where the forecast would have earned something.

Confidence: when to go big

If honesty is the rule, the question becomes how honest your confidence really is. Two easy mistakes pull in opposite directions.

Overconfidence feels like conviction. You've watched the team, you know the matchup, you type 85. The table shows why that's risky: from 70 to 85 a win gets you about 7 more points (an 85 that wins scores 25 − 2.25 = +22.75, next to +16), but a loss costs you about 23 more (an 85 that loses scores 25 − 72.25 = −47.25, next to −24). If the game was really 70%, you've paid 15² / 100 = 2.25 points a game on average for that feeling.

Timidity feels safe. You believe 75 and enter 60 just in case. That costs 15² / 100 = 2.25 a game too. The square doesn't care which direction you missed in.

The site tested this idea directly when building The System. Capping confidence, never going above 90, or 85, or 80, cost points every time it was tried. The lock of the week (the game where Elo and the market are most sure of the same team) went 434–76 over the 27 replayed seasons, which is 434 / 510 ≈ 85%, against a forecast of 79%. The locks were, if anything, a bit underconfident. Yet pressing those locks up by 5 extra points still cost points over the 510 of them, because a loss at a very high number is so expensive. The lesson isn't "be timid" or "be bold". It's "be accurate", and let the number be as big as the evidence says, no bigger.

Calibration: does your 70% mean 70%?

A forecaster is calibrated if the things it calls at 70% happen about 70% of the time, the things it calls at 60% happen about 60% of the time, and so on down the line. Calibration is how you check whether a forecaster's confidence is honest, including your own.

The way to check is to sort picks into bins and count. Here's a made-up example of what one player's season might look like:

Pick range (favorite)PicksFavorite wonWin rate
50–59%402255%
60–69%503264%
70–79%301757%
80–89%10990%

The first two rows look calibrated: the win rates land inside their bins. The 70s look overconfident: calls in the 70s won only 57%. The 80s went 9 for 10, which looks great but is too few games to mean much.

That last point matters. Even a perfectly calibrated forecaster bounces around. Twenty honest 70% calls should win 14 on average, but the typical swing is about √(20 × 0.7 × 0.3) = √4.2 ≈ 2 wins either way, so 12 or 16 wins is entirely normal. Calibration needs hundreds of games per bin before small differences mean anything.

Calibration in the site's own history

The 27-season backtest is full of calibration checks, and they're some of the most useful numbers on the site.

  • When Elo and the market disagree by 10 or more points, Elo's side won 50% of the time over 1,102 games, while Elo said 65% and the market said 48%. Elo was badly overconfident on exactly those games; the market was close.
  • When Elo and the market agree tightly (the same favorite at 55% or more, within 4 points of each other), favorites won 69.8% against a 67.6% forecast, over 2,607 games. A small underconfidence, which is why The System nudges those games up by 2.
  • The bold call of the week (the game where Elo disagrees most with the market) hit 48%, about 247–263, while Elo claimed 65%.
  • Wind of 15 mph or more, outdoors: favorites won 62% against a 65% forecast over 626 games, so The System pulls those numbers a fifth of the way toward 50.

Each became a rule, or a reason not to make one. That's the practical use of calibration: it tells you which direction to correct.

Calibrated isn't the same as good

A forecaster can be calibrated and still not very useful. One that entered the league's overall home win rate on every game would come true at exactly the rate it said, but it would never tell you which games are lopsided, so it would earn very little. Good forecasting needs numbers that mean what they say and numbers that separate the likely from the unlikely. The scoring rule measures both at once; calibration tables are the diagnosis.

The crowd, and the sharp crowd

The site also scores two forecasts built out of players' picks.

The crowd is the plain average of people's picks on a game. It's scored once three or more people have picked, and shown after kickoff. The bot crowd is the same thing for agents.

Averaging picks has a mathematical advantage under this scoring rule. Because the penalty is a square, the average pick always scores at least as well as the average of the individual scores. Try two players on the same game, one at 80% and one at 60%. The crowd says 70%.

Result80% pick60% pickTheir average scoreCrowd at 70%
Home wins+21+9+15+16
Away wins−39−11−25−24

Either way the crowd beats the two players' average by a point: their errors partly cancel before the square is taken.

The sharp crowd is a weighted average that leans on players who've been scoring well. Each person's weight is 1 plus their season points divided by 50, if their season points are positive; anyone at zero or below counts with a weight of 1. So a player on +100 for the season counts 1 + 100 / 50 = 3 times as much as a newcomer. Take the same two picks, with the 80% player on +100 and the 60% player at −20: (3 × 80 + 1 × 60) / 4 = 300 / 4 = 75%, against the plain crowd's 70%. Whether leaning on past scorers helps is itself a question the scoreboard answers over time, which is why both crowds are scored side by side.

What this means for playing

  • Enter what you believe. The rule is built so your real opinion is your best play. Any shading costs, on average, the gap squared over 100.
  • Don't hide at 50. It's the one pick guaranteed to earn nothing. A skipped game is a 50.
  • Size your confidence to your evidence. The top of the scale is expensive when you're wrong. A 100% pick that loses scores −75; it takes about five winning 70% calls to earn that back (5 × 16 = 80).
  • Judge yourself on many games. A good week or a bad one says little.
  • Remember what you're up against. The market's closing forecast is the hardest number to beat, and Elo trails it by more than 100 points a season (997 − 892 = 105). The System runs about even with the market, +14 a season with a standard error of 10, and clearly ahead of Elo. Beating either is hard; that's the game.

If you also bet, the same ideas apply to prices; the other articles here cover how. Bet only where it's legal, only if you're of age, and only within your means.

Where to see it on the site

A few terms

  • Miss: the distance between your pick and what happened (100 if the home team won, 0 if not).
  • Brier score: the squared miss, with chances written as fractions. Game points are 25 minus 100 times it.
  • Proper scoring rule: a rule where your honest belief earns the most points on average.
  • Calibration: whether calls made at a given percent come true about that often.
  • Overconfident / underconfident: calls that win less / more often than they claim.
  • Crowd: the average of people's picks on a game. The sharp crowd weights people by their season points.