The real AI benchmark isn’t MMLU. It’s which model produces the best NFL first-round mock the day before the draft, with the reporting landscape at its absolute noisiest, and with the actual scoreboard twenty-four hours away from lighting up in public.
As a New York Jets fan, the offseason is our Super Bowl, so I’ve become increasingly deep in the NFL Draft process over the past decade+. I’ve tried mocks using AI in previous years, but the models werent quite there. They are in 2026. So, I ran the bench on Claude Opus 4.7 and ChatGPT Pro 5.4 with one identical sentence:
“generate a predictive 2026 nfl draft for the entire first round write up your explanation for each pick. include trades.”
Here are the mocks:
Mock drafts are basically a 32-slot prediction problem where every slot depends on the slot before it, every team has private information you don’t have, and real money is on the line for everyone involved except you. Trade behavior cascades: one unexpected move at pick 3 shifts the next 15 selections. Teams actively run smokescreens in the pre-draft cycle (leaked interest, faked visits, selectively-placed rumors) because misdirecting the 31 competitors behind them is worth real draft capital. And the information asymmetry is enormous: team doctors see medicals you don’t, scouts see character interviews you don’t, and GMs hear what other GMs are offering in trade calls you’ll never be on.
The track record reflects this. In 2025, Daniel Jeremiah (widely regarded as the best mock drafter working) hit 19 of 32 first-round picks in exactly the right slot on his final mock, which was the best performance of the major analysts. Mel Kiper, the OG Draft analyst and king of voluptuous hair, nailed 4 exact slots in his 2025 final mock, his worst showing was 2023 when he got only 3 right in the entire round. No major analyst in 2025 cracked 20% on exact-slot accuracy. The better metric, “did you correctly identify the 32 players who went in round 1,” usually lands the top analysts in the 70-75% range, which is to say roughly 8 of the 32 names on a “really good” final mock aren’t even in round 1 the next night. That’s the baseline. Predicting round 1 correctly is hard for the humans whose entire job is predicting round 1 correctly, and hard in ways that get harder the closer you move to actual picks.
Which is exactly why it’s a good benchmark for AI models. It’s an evaluation where the ceiling is legibly bounded by how much real information you can get your hands on, how well you weigh it, and how much you’re willing to commit to specific falsifiable predictions. No model (and no human) gets 32 of 32. But the gap between 8 of 32 and 22 of 32 is the difference between a lazy synthesis of consensus and a genuinely informed read.
Mendoza to the Raiders at 1. Neither model hedged, neither model considered an alternative, and neither model should have. Indiana’s Heisman winner, 6’5” and 236, lands in a Vegas organization that spent the offseason building the landing pad around him: Kirk Cousins inked as the one-year bridge, Tyler Linderbaum signed in free agency to anchor the line, Klint Kubiak running the offense under Pete Carroll’s first full year in the building. The card has been in for weeks. If Mendoza somehow doesn’t go first on Thursday, every other pick in both mocks cascades sideways and nothing I’m about to write means anything.
From pick 2 onward, the two boards stop agreeing about almost anything.
Schefter reported on Monday that Francis Mauigoa, the best offensive lineman in the class on tape, has a herniated disc in his back. Asymptomatic at the moment of the rechecks at the NFL Combine in Indianapolis. But Schefter also reported that some front-office execs think Mauigoa will need the surgery “at some point either way,” and if it flares up in training camp, that’s a three-month window and potentially his whole rookie season.
Claude read this the way a cap-aware front office would read it. Mauigoa drops to Miami at 11. Not because 11 is a convenient landing slot but because the Dolphins have the right combination of variables to absorb the risk: Mauigoa played college ball at The U, trained there, lives in the city, and new GM Jon-Eric Sullivan came over from Green Bay where the entire organizational identity runs through the offensive line. Miami’s reportedly told their team doctors to clear him under their own standard. Eleven is where a team stops worrying about how far he could fall and starts worrying about whether they’ll still be able to take him.
ChatGPT sent Mauigoa to the Giants at 5. That pick doesn’t survive contact with the reporting. New York just acquired pick 10 from Cincinnati in the Dexter Lawrence trade, meaning they’d have taken him one round later for free if they’d simply waited, and they’d have spent the spare slot on one of the real needs (safety, where Caleb Downs is sitting right there). Drafting an injured player five slots higher than the market says you need to is the kind of double-payment that gets front-office executives fired.
Caleb Banks broke his fourth metatarsal the night before his on-field combine work. Ran a 5.04 the next morning on the broken bone, which is either the most impressive thing any prospect has done this cycle or the clearest possible preview of the stubbornness that keeps getting him hurt. Third foot injury on the same foot in twelve months. Surgery on March 9. His camp sent NFL teams a letter on Wednesday, twenty-four hours before the draft, saying he’s on pace to be fully cleared for football activities in early June.
Claude read the letter and kept Banks out of round 1. ChatGPT took him at 14 to Baltimore. A first-round pick who can’t step on a field until June is a pick whose rookie year starts in August, with OTAs and most of minicamp already gone, on a foot that has now broken three separate times. Jeremiah has Banks slotted 51st on his top-150. The distance between pick 14 and pick 51 is basically the distance between ChatGPT’s read and the consensus view of every draft analyst who isn’t ChatGPT.
Ty Simpson. One QB goes in the first round (Mendoza) and then there’s a gap, and Simpson is the name everyone’s trying to locate in that gap. Mocks have him anywhere from pick 9 to pick 65. Schefter’s Monday column was specific about why: the Cardinals and Jets are the two teams to watch, with Arizona the more interesting case because they pick 34th and would need to jump into round 1 to grab him. The reason to do that, the actual load-bearing reason, is the fifth-year option, which disappears for any quarterback taken after pick 32. For a GM evaluating an unproven three-quarter-season starter, another year of team control is real money, somewhere in the $25-30M range depending on the QB tier.
Claude’s mock threads this exactly. Arizona trades up from 34 to 26, sending pick 34 plus a 2026 third to Buffalo, grabs Simpson, gets the option. Buffalo has no second-rounder this year (they traded it to Chicago in the DJ Moore deal) so they’re specifically motivated to recoup Day 2 capital. The whole trade has an incentive structure on both sides that matches the Jimmy Johnson trade value chart to within a rounding error.
ChatGPT put Simpson at 16 to the Jets and called it a day. Which, to be fair, is also where Schefter floated him. The Jets have real interest. But they also have Geno Smith on the roster, Aaron Glenn is the head coach with his back against the wall in year one, and the reporting all week has framed the Jets’ round-1 needs as win-now picks rather than developmental QBs. The Jets also canceled their top-30 visit with David Bailey, which Tom Pelissero read as a near-certain signal that Bailey is the pick at 2. Both of those signals point away from Simpson going 16. Claude’s slide to 26 is the sharper read because it takes the option-cliff economics seriously instead of just matching Schefter’s headline.
Both models sent David Bailey top three, though they disagreed on which top-three slot. Claude had him going to the Jets at 2. ChatGPT slid him to Arizona at 3. The real answer depends on what Jets GM Darren Mougey told his draft room on Wednesday night, and that information is not in the public domain yet.
Here’s what the public reporting does tell us: DraftKings has Bailey at -145 to go 2nd, Arvell Reese at +110. Brugler thinks Reese is the best player in the entire class. Graziano’s late-week sources lean Reese. Pelissero thinks it would be a great surprise if the Jets don’t take Bailey. The market says Bailey; the evaluators lean Reese; the Jets canceled the visit with Bailey, which could mean anything from “we already know everything” to “we’re quietly off him.” Bailey led FBS in pressure rate at 20.2% and tied the national sack lead at 14.5. Reese is longer, more versatile, and has the kind of LB/EDGE positional flexibility that travels across defensive coordinator changes.
Claude hedged this pick at 70% Bailey. ChatGPT just took him outright. Whoever gets this one right is getting the most contested pick in the entire first round right, and neither model deserves credit for it unless Thursday night cooperates.
Claude modeled four trades, each of which has to be wrong in a specific and recognizable way to fail. Saints up from 8 to 3 to get Reese, sending 8 plus a 2026 second and a 2027 second, on the back of this reporting from nola.com: only 4 of 38 Saints draft trades under the Loomis/Ireland regime have been trade-downs, which means a trade-up from a top-10 slot is basically the modal Saints behavior. Eagles up from 23 to 15 for Kadyn Proctor, sending 23 plus a 2026 third, on the back of Lane Johnson turning 36 in five weeks and the rest of the Tier 1 tackle class being gone by the time pick 23 comes up. Cardinals up to 26 on the Simpson option-cliff economics already described. Chiefs to 29 via the McDuffie deal, which is already in the books.
Each of those trades names a specific compensation package, which means each one can be cleanly falsified by whatever actually happens Thursday. If the Saints trade up to a different slot, wrong. If the Eagles stay at 23, wrong. If the Cardinals hold at 34 and take Simpson in round 2, wrong. The falsifiability is the point.
ChatGPT listed six trades at the bottom of its mock, with no compensation attached to any of them: Lions up to 13 with Rams, Browns up to 20 with Cowboys, Chiefs up to 22 with Chargers, Dolphins up to 25 with Bears, Cowboys down to 24, Bears down to 30. Just arrows. No picks, no future capital, no explanation of why the receiving team would accept the move. These can be directionally right without being recognizably right, which means they’re hard to grade and harder to trust. Six vague trades in a bench that rewards specificity is weaker than four specific trades, even if the raw count looks better on a share card.
Nothing above is a score. Thursday night is the score. What I’ll be looking at Friday morning: the raw pick-accuracy count for each model, which projected trades actually happened with recognizable compensation, whose medical reads matched front-office behavior, whether Simpson landed where Claude said or where ChatGPT said or somewhere neither model predicted.
A few of the bets have clean asymmetric outcomes. Mauigoa going to Miami at 11 makes Claude look prescient; Mauigoa going to the Giants at 5 makes ChatGPT look like it heard something the rest of the draft community missed. Banks going in round 1 anywhere is a clear Claude miss, while Banks sliding to round 2 cashes for ChatGPT not at all (ChatGPT took him at 14) and hurts their accuracy count outright. Simpson is the single highest-variance bet in the class and whoever gets closest to his actual landing spot takes the biggest per-pick points swing.
Thanks for reading Dave's Quick Hits! This post is public so feel free to share it.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.