Essay · Sports Analytics

Preseason: Simulating a Million College Pick'em Entries Across 10,000 Seasons

I am a data scientist. I am not a sports analytics person. Until this year, none of my models had anything to do with football.

My interest in sports analytics began with the movie Moneyball. I watched it more than a decade ago. At the time, I had a statistics degree and enough programming knowledge to be useful. I had not heard of data science. One of the things I loved about that movie was the idea that a room full of people who all agree can be wrong together, and that the disagreement is something you can test. A football contest with a million entrants turns out to be a very good place to look for that.

I've played ESPN's College Football Pick'em for years, the way a fan plays it. Ten games a week, pick the winners, rank how sure you are about each one. This year I wanted to see if I could build something that plays it better than I do. The goal: Beat the small group of friends I play against.

Winning it across all of ESPN is a separate matter. About a million people enter. Whoever finishes first has to forecast football games more accurately than Las Vegas does, then go against the grain on the handful of games where everybody else agreed, then be right about those. Vegas, to my knowledge, is the best forecaster in sports. Beating it is not a thing that is likely to happen. I built a model to try anyway.

This is the first essay in a series about that attempt. I'll share what worked, and what did not. This essay is the setup. Before you can say whether a model is any good, you need something to compare it against, and a fair way to do the comparing. That is what this one builds.

The game

ESPN's College Football Pick'em hands you ten games a week. You pick a winner in each one. Then you rank them: The game you are surest about gets 10 points, the next gets 9, on down to 1, and you use each number exactly once. A correct pick adds whatever confidence you put on it to your score. A wrong pick adds nothing. A perfect week is 55 points.

The scores add up all season. Nationally, one entry wins. Everybody else is playing for the group they joined, which in my case is a handful of friends. Beating them is goal one. The national leaderboard is the hard version, and it is the one worth measuring against.

Week 1 · ten gamesNext number: 10
at
at
at
at
at
at
at
at
at
at
Ten numbers to spend. A perfect week is 55 points.
The Week 1 card of the 2026 season. These are the ten games, and the first of them kicks off on September 5. Tap the team you think wins in the game you are surest about and it gets the 10. The three buttons at the bottom fill the card in different ways.

Three ways to fill it out

There is no standard way to do this. Some people pick underdogs. Some pick their favorite teams. Plenty of people just know things about particular programs and fill the card out from memory.

Three approaches are mechanical. Hand any one the same ten games and it fills the card out the same way every time, which is what makes them fair to measure a model against.

The first is the Crowd. ESPN shows you what share of everyone who filled out a card took each side, so the rule is to pick with the majority and give your highest confidence to the game where the majority is largest.

The second is Vegas. Take whichever team Las Vegas favors, and assign confidence by the size of the point spread. The biggest spread gets the 10, because that is the game Vegas is surest about.

The third is the AP poll, which is a panel of sportswriters voting every week. Take the higher ranked team. When neither team is ranked, fall back to where they finished last season, then to their record, then to their point differential, then to whoever is at home. Sort the ten games by how wide those gaps are. Call that card the Writers.

The Vegas rule is close to what I did for years, though I never checked whether anybody else plays that way. I would open a spreadsheet, this being before I learned Python, and type in the Vegas line for all ten games. Then I sorted on that column, biggest spread at the top. The favored team in each game was my pick, and its position in the sorted list was my confidence. Then I went back and moved a few games around based on nothing but how I felt about them. Georgia Tech got a bump, because that is where I got my master's degree. Auburn got a bump, because I like watching Auburn. An SEC team playing somebody from another conference usually got a bump too.

Run all three over the Week 1 card above and something strange happens. The Crowd and Vegas pick the same ten teams. Not most of them. All ten. A million casual entrants and the entire Las Vegas market do not disagree about a single game.

The Writers break with them three times, taking Tulane, California and Wyoming. Not one of those breaks involves a ranked team. In all three the ladder runs out of poll and falls back to last season, once to the final ranking and twice to the record.

What the Crowd and Vegas disagree about is confidence. The Crowd is surest about Cincinnati, which Vegas ranks eighth. Vegas is surest about LSU, which the Crowd ranks sixth. Identical picks, different orderings, and since the whole score comes from the confidence you assigned, the two cards end the week in different places.

gamethe crowdvegasthe writers
Clemson at LSU6LSU10LSU9LSU
Tulane at Duke7Duke9Duke7Tulane
Boston College at Cincinnati10Cincinnati8Cincinnati6Cincinnati
Baylor at Auburn8Auburn7Auburn1Auburn
Louisville at Ole Miss9Ole Miss6Ole Miss10Ole Miss
Wyoming at Colorado State3Colorado State5Colorado State3Wyoming
UNLV at Hawai'i4UNLV4UNLV2UNLV
Western Kentucky at Nevada5Western Kentucky3Western Kentucky5Western Kentucky
SMU at Florida State1SMU2SMU8SMU
UCLA at California2UCLA1UCLA4California
The same ten games, three ways. The pick sits under the confidence number it was given. The Crowd and Vegas take the same side in all ten games and differ on confidence. The Writers take the other team three times, and those are the ones in red. Red here means the Writers disagree. Nobody knows who is wrong yet, because the season starts on September 5.

Those are the three rules. There is a fourth card every week, the one my model produces, and the rest of this essay series is about that one. For now it is just the Machine, a fourth entry that has to be measured against the other three.

Vegas is very good, and Vegas is not right

Following Vegas is a strong strategy. Across five seasons of real Pick'em cards, taking the Vegas favorite every week and ranking by spread collects about 69 percent of the points available. If you've ever wondered why your office pool is harder to win than it looks, that is a good part of the answer. A lot of people are effectively running a similar strong strategy.

Strong is not the same as correct. Those five seasons hold 707 games, and Vegas lost 253 of them. In 11 of the 71 weeks your 10 was on a team that lost.

Here are two real weeks from the 2025 season. Both cards start filled out by the Vegas rule, and the buttons switch them to the Crowd and the Writers.

fill both weeks in like
2025, week 3ranked by the point spread
108Notre Dameover 16Texas A&M70%lost
9North Texasover Washington State67%
8Pittsburghover West Virginia65%lost
7Memphisover Troy61%
66Georgiaover 15Tennessee58%
511South Carolinaover Vanderbilt59%lost
4Minnesotaover California57%lost
3App Stateover Southern Miss57%lost
212Clemsonover Georgia Tech55%lost
1Tulaneover Duke55%
scored 23 of 55 · 6 of 10 picks lost
2025, week 13ranked by the point spread
106Oregonover 16USC76%
9Kennesaw Stateover Missouri State67%
8SMUover Louisville61%
720Tennesseeover Florida58%
615Georgia Techover Pittsburgh58%lost
511BYUover Cincinnati55%
4Florida Internationalover Jacksonville State58%
3Southern Missover South Alabama55%lost
2East Carolinaover UTSA54%lost
1TCUover 25Houston53%
scored 44 of 55 · 3 of 10 picks lost
Two Saturdays, played three ways. Left: six of ten picks lost, including the top slot, and the Vegas card scored 23 of 55. Right: A good week, 44 of 55. Each rule fills out both weeks the same way, so only the games are different. A number in front of a team is the AP poll ranking that week. Switch the rule at the top to see how the Crowd and the Writers handled the same two Saturdays.

Look at the top line on the left. Notre Dame, favored by a touchdown at home, roughly a 70 percent chance to win. If you used the Vegas strategy, then that is the game you were surest about, the one carrying your 10, and it lost to Texas A&M. The week scored 23 out of 55.

The card on the right is the one I remember, because I was at one of those games. Georgia Tech was ranked 15th and favored by two and a half points against an unranked Pittsburgh. Tech had Haynes King at QB, who was on fire all season, and had the momentum going in. A win that night would have sent us to the ACC championship game. Nobody in that stadium believed we were losing, but Vegas had the game close to a coin flip. I put my 10 on Georgia Tech. The Vegas card had it sixth. Tech lost, and my feelings cost me four points more than a stranger with a spreadsheet would have lost. The Writers agreed with me, for whatever that is worth. They put Georgia Tech in their second highest slot.

Kennesaw State, where I work as a researcher and am doing my PhD, is on that same card at 9, which is exactly where Vegas had them. They won. So on one Saturday in November both of my schools were on the same ten game ballot, and they split.

I doubt I am the only one doing that. A lot of those million entries belong to somebody's alma mater, and those people are bumping their team up the same way I did. It shows up in the crowd percentages, and later in this series I try to use it.

Simulating the season

Here is the question I could not answer by looking at a leaderboard: What score does it actually take to win this thing, and how often would a good strategy get there?

The trouble is that a season only happens once. 2022 played out how it played out, and there is no version where three close games go the other direction. So I built the contest in software and ran it over and over.

Given a season's weekly cards, the model gives every game a probability that the home team wins. Walk the schedule, flip a weighted coin at each game's own probability, and that is one alternate version of the season. Repeat ten thousand times.

Then you need opponents. Weak competition would make our own card look brilliant for no reason. So the million opponents were built two ways, one leaning toward the real ESPN pick percentages, the other spread between picking like the public and picking exactly like our model. Everything below was run under both.

Every one of the million cards is then scored against every one of the ten thousand seasons. That is ten billion card seasons, which sounds impossible on a desk machine and is not, because scoring a card is a dot product and ten billion dot products is one matrix multiplication. Before any of it runs, the fast version is checked against a slow and obviously correct loop and has to agree exactly, or the run stops. What comes out is an exact count of how many of the million entries landed on every possible score.

So far the whole thing is rigged in my favor. The seasons were drawn from my model's numbers. My card is those same numbers sorted, biggest confidence on the likeliest game. Grade that card against seasons built from the numbers it was sorted by, and it wins. My model could be terrible and it would still win.

Letting them grade each other

The way out is to stop letting one model be the scorekeeper.

Each of the four cards is built on a source that also produces probabilities. The market has them, sitting inside the point spreads. The Crowd has them, in the share of entries taking each side. The Writers have them, in the gaps between ranked teams. My model has them by construction. So instead of running the season once under my numbers, we ran it four times, once under each source, and scored all four cards each time.

Every card is now being marked by four different opinions, including its own. Marking is one multiplication repeated: The confidence a card spent on a pick, times that source's chance the pick wins. The Vegas rule put 6 points on Georgia Tech in November and Vegas gave Georgia Tech 58 percent, so under Vegas that game was worth 3.5. Sum the season, do it sixteen times, and you have a grid with the cards down the side and the sources across the top.

Four of those sixteen cells are a card being marked by the opinion it was built from, and those are the ones to ignore. My card spends its biggest number on the game my model likes most, so my model marking my card puts the biggest number on the biggest probability, and the second biggest on the second, all the way down. That is the highest total those ten numbers can make. It would still be the highest if my probabilities were nonsense. Everything worth reading is off the diagonal.

Every number in the grid is points, added up across five seasons of real cards. Those five seasons hold 71 weeks and 703 games, with 3,837 points sitting on the table in total. A card that got every single pick right would score 3,837. The Vegas card really scored 2643.

card ↓graded by →
the Machine
Vegas
the Crowd
the Writers
actually scored
the Machine
66%2,524
65%2,483
62%2,367
61%2,345
66%2,537
Vegas
63%2,406
68%2,616
62%2,377
60%2,301
69%2,643
the Crowd
62%2,377
64%2,465
65%2,492
59%2,277
65%2,491
the Writers
59%2,282
59%2,277
60%2,316
59%2,278
59%2,265
Points scored over 703 real games in 71 weeks, 2021 to 2025. Every point available across all of them adds up to 3,837.
Four cards, four opinions about what was likely, and one column of facts. Points out of 3,837, across 703 real card games from 2021 to 2025. The first four columns are projections: What each card would average if games kept landing at that source's probabilities. The last column is the points each card really scored. The underlined cell in each column is a source grading its own work. Drop the diagonal and watch the ranking change.

Three results, in the order they surprised me.

The first is how large the effect being corrected for actually is. Vegas scores 2616 under its own probabilities and averages 2361 under the other three. That 254 point swing is more than twice the gap between the best card and the worst once the diagonal is gone. Self grading is the largest single effect in the table.

The second is the Writers, who finish last in three of the four columns. Under the Writers' own probabilities, the Writers' own card does not win. The Machine does. The Writers card is the only one not sorted by its own source's probabilities. It sorts by gaps in poll position, and a five place gap in the poll is not a five place gap in the odds.

The third is what the real games say. The last column is the 703 games as they were really played. Off the diagonal, the Machine has the best average of the four. In the column that counts, it finishes second, 106 points behind Vegas.

The off diagonal answers whose card makes the best use of a given set of beliefs. The last column answers whose beliefs were closer to true. My card is a good sorting of my probabilities. Vegas simply had better probabilities.

What it takes to win

The numbers below all come from the 2022 season. A perfect 2022 is 731 points, which is what fourteen perfect weeks add up to.

points on the table
731
a perfect 2022, and the ceiling every entry in the field is measured against
the average entry
410 to 442
what a million simulated opponents scored, across every assumption we ran
what it took to win
551 to 570
the top score in a field of a million, in that same season
and the target moves
546 to 633
the winning score across all five seasons. There is no fixed number to aim at
The million opponents, with no strategy in them. Every number here describes the million opponents and the season they played, not any particular card. The ranges cover both field models and all four sets of probabilities.

Across all five seasons, under every set of assumptions we ran, the winning score landed between 546 and 633. The bar itself moves by 87 points depending on how the season fell and how good you assume the field to be.

So where does a good card land? A card played strictly by the Vegas rule, graded by anybody except Vegas, finishes somewhere between 78,865th and 229,577th out of a million. That is a wide range. It depends on which forecaster you believe and how good you assume the other million entries are.

Play that card every year and it reaches the top one percent between 3 and 12 times in a hundred seasons, the top hundred somewhere between once in 560 seasons and once in 3,300, and first place not once in ten thousand simulated tries.

The card is fine. First place goes to whichever entry got the luckiest. There are almost certainly other people running analytics on this contest, and some of them are likely better at it than I am. It still would not be enough. The gap between a very good card and a winning one is made of games nobody can call.

To beat the luckiest of a million, you have to pick games the field got wrong. So we went looking for a rule that tells you which ones. Lopsided games where the model is unsure, games the crowd agrees with us on most, over-hyped teams, programs the polls misjudge year after year. None of it survived testing. We'll dig into this in essay three.

One finding is worth keeping here, because it inverts the standard advice. Some common advice tells you to put your biggest number on the game where you break from the field. Putting your smallest was better, every time we looked. Deviating on purpose only adds randomness. It works when you know something the field does not, and that is the hard part.

Our own test is weaker on this than I would like. The simulated field holds a million different cards, and the real contest does not.1

You can try it yourself. Below is that same Georgia Tech and Pittsburgh week card, fixed in place, played over and over. One run could be a fluke, and so could ten. Run it hundreds of times and the law of large numbers kicks in, and the average across your simulated runs approaches the true average. It is the same reason the simulation upstream ran ten thousand seasons instead of one.

--
points out of 55
A season winner averages about 41 of 55 a week.
average so far--
04155
Press the button and the ten games get decided again.
The same card, over and over. These are the ten games from that Georgia Tech and Pittsburgh Saturday, with the exact picks Vegas would have made. Press the button and each game is decided again by its real probability, so a 70 percent favorite comes in about seven times out of ten. Nothing about the card changes. Only the results do.

The season is the experiment

That is the game. Ten games, ten numbers, a million people, and a scoring rule with exactly one winner. The Crowd has a strategy. The Writers have one. Vegas has the best of the three. None of them is likely to win, and I have now spent ten thousand simulated seasons establishing that my own model, played straight, is not likely to win either.

I am playing it anyway, every week, with the card posted before kickoff and the score posted after.

The next essay in this series is the Machine. What data goes into it, how the prediction pipeline was built, which model families were tried and how they were scored against each other, and the ideas that looked excellent and died under testing. The ones that failed get explained rather than skipped.

That piece is also the one that stays live. The weekly cards go there, all four of them, and so does the running scoreboard: The Machine against the Crowd against Vegas against the Writers, cumulative points, updated every week until the season ends.

How this was built

For anyone who wants to check the work or build something like it.

Code written in Python. Every row of data carries the time it became knowable, and a query for a given Saturday returns only what existed before kickoff, which is what keeps a backtest honest. Betting lines and crowd picks are allowed as something to measure against and never as an input to the model I built for my own picks all season.

The simulation ran on a single small desktop machine, an ASUS Ascent GX10.

I built this with Claude. Fable 5 helped me with architecture, planning, and the rulings on what was worth doing. Opus 5 wrote the code. The decisions, the errors, and the choice of what to publish are mine.

seasonweeksgamesperfectentriesseasons
2021141417811,000,00010,000
2022141367311,000,00010,000
2023141407701,000,00010,000
2024141387501,000,00010,000
2025151488051,000,00010,000
Field modelsTwo, run separately. One where every opponent leans toward the public, and one where opponents span a range from casual to sharp.
Truth sourcesFour. Every season and every field is run once under each strategy's probabilities, which is what the tournament matrix reads.
Outcome drawOne weighted coin flip per game, at the probability the truth source gives that game.
ScoringOne matrix multiplication. It is checked against a slow elementwise loop for exact equality before each run is allowed to start.
Card seasons scored400,000,000,000 across all forty cells.
RuntimeAbout forty seconds a cell.
ResumabilityOne file per cell. A cell already on disk is skipped, so a crash costs one cell at most.
The simulation, in full. Every cell is one season, under one assumption about the field, scored on one source's probabilities. All forty are on disk and can be regenerated from the same command.

Things the simulation does not do, which matter if you are reading the numbers closely:

  • Game outcomes are drawn independently of each other. Real Saturdays are not independent, and a weather system or a wave of upsets moves several games at once. There is a correlation setting in the code. Turning it up to a tenth moves the reported percentiles by two to four points, so it moves the number without changing the answer.
  • Every one of the million opponents is drawn separately, so the simulated field holds a million distinct cards. A real field does not. Large numbers of real entrants submit the same handful of obvious cards, and a block of identical cards behaves differently from a million different ones. We cannot tell how many, because ESPN publishes only the totals for each game.
  • How often a card finishes first is still not measurable. Even at ten thousand seasons, the count of outright wins for any single card is somewhere between zero and four, and no honest number comes out of that. It needs an estimator fitted to the tail rather than more brute force, and that is later work.

References

  1. Clair, B. and Letscher, D. (2007). Optimal Strategies for Sports Betting Pools. Operations Research 55(6), 1163 to 1177.Separates the share of entrants taking each side from the chance each side actually wins, and shows that in a large enough pool a card of favorites is differentiated from nobody. The reason the field has to be simulated at all.
  2. Levitt, S. D. (2004). Why are gambling markets organised so differently from financial markets? The Economic Journal 114(495), 223 to 246.Bookmakers forecast outcomes better than bettors do, and bettors lean toward favorites and toward home teams. Both halves of that are in the simulated field.
  3. Hardy, G. H., Littlewood, J. E. and Pólya, G. (1952). Inequalities, 2nd ed. Cambridge University Press, Theorem 368.The rearrangement inequality. Sorting your confidence the same way as the probabilities you drew the season from maximises the sum, which is the arithmetic behind the circularity problem and the reason the tournament matrix exists.
  4. Coleman, B. J., Gallo, A., Mason, P. M. and Steagall, J. W. (2010). Voter Bias in the Associated Press College Football Poll. Journal of Sports Economics 11(4).AP voters favor teams from their own state and conference and lean on simple signals like number of losses. The poll is a real opinion with human errors in it, which is what makes it a useful fourth judge.
  5. Coles, S. (2001). An Introduction to Statistical Modeling of Extreme Values. Springer, chapter 3.Return levels and the generalised extreme value distribution. The method for the one question this simulation still cannot answer, which is how often a given card finishes first.