Welcome friends!
Happy March Madness! Otherwise known as peak season for those of us who love both sports and small investments with longer odds but (very) outsized return potential.
In the spirit of both the college hoops madness and the 10x more mad LLM arms race, Matt decided to turn his entries in the C2V pool into a foundational model showdown among the four current front runners -- ChatGPT, Claude, Gemini, and Grok.
This is purely for research purposes, of course. Definitely not an admission that Matt’s more recent Human-I brackets have been a thinly veiled cry for help. Just so we’re clear.
And yes, we know that DeepSeek is inarguably also a front-runner but, well, if you want to repeat this exercise there and send us the results, we’ll include them. Actually, don’t send the results, we’ll just take them down over the phone. But not our regular phones, we’ll grab a burner and text you the number. Or maybe just face to face. Preferably somewhere with no cameras or drone access.
Okay, enough preamble, let’s get to it!
Obviously, we still have three more rounds and a title game before we know how everyone did with their picks, and the points in this pool double each round, so it’s hard to handicap at this point (though we have shared the current rankings below).
That said, there is obviously a huge amount of randomness in this tournament every year for a variety of reasons, not least of which is that you have a bunch of mostly 18 – 22 year-old kids who have collectively played together for all of 30 games, suddenly thrust into a spotlight that’s many times brighter than anything they’ve previously experienced, playing in single-elimination games. Hence the “Madness” moniker.
So even with perfect data on everything from team performance metrics to public perception and historical tournament trends, and the capacity to model and weight every potential scenario based on evey potential factor, in near real time, someone who has never watched a college basketball game and picks teams based on mascots or alums they know from various schools will probably end up winning your pool anyway.
To us, this experiment is really less about the results and more about how we got there and how these models performed as a stand-in for a smart human who knows college hoops and has years of experience following this tournament and entering pools.
So here’s our review of the process and our experience working with our new college basketball analyst recruits.
This was relatively simple – I gave each model the pool size, the scoring system, and the explicit goal of winning the pool (as opposed to accurately picking each game).
That last instruction is an important distinction and every model referenced it in its response as being a key part of the rationale for their picks. We’ll see how well they implemented it, but at least we know they were paying attention (always a good sign with a new employee).
Top Tier
The winner here, in an absolute landslide, was Claude. This response could not have been better organized and easy to follow.
Most notably, young Claude:
Provided the rationale for each round and each material upset prediction alongside the actual picks (vs my having to jump around trying to match picks to reasoning).
Supplemented the round-by-round text with a full bracket graphic, which made it significantly quicker and easier to input those picks into the CBS app than with Gemini’s and Grok’s text-only outputs (Claude took maybe 20 - 30% as long as these other two).
Claude’s output also just looked incredibly professional vs the more sterile feel of the others. Not that the appearance ultimately matters as long as the content and readability are there, but there’s something to be said for a model that endulges dumb human things like vibe and feel that definitely don’t matter, but somehow also do.
Mid Tier
Gemini and Grok were fine in terms of readability but a little choppy in organization. Notably, the rationale for how each approached the pool overall, and various rounds and upsets was half at the top and half at the bottom, with the picks in the middle, which wasn’t ideal. But positive marks for both models for organizing the picks by region and round, for ease of input into the CBS brackets.
Should We Be Concerned Tier
In a legitimate shocker, given both its near-100% focus to-date on the consumer market and substantial first-mover advantage, ChatGPT came in dead last, and it wasn’t remotely close. No exaggeration, this was on the order of a first round 15 over 2 in terms of outcome vs expectations.
Start to finish, ChatGPT’s reponse was a hot mess. It bounced back and forth between commentary on this year’s tournament and historical trends and strategies so frequently that it was legitimately hard to distibguish between actual picks and illustrative examples; and the rationale for various approaches to overall strategy, regions, rounds, and highlighted matchups/upsets were so scattered across the response, it read like ChatGPT had had a stream of consciousness, brainstorming session in the conference room and then just texted a picture of the whiteboard at the end.
In addition to the lack of organizaton, there were also some major completeness issues that contributed to the difficulty of translating ChatGPT’s advice to a bracket (discussed more below), but overall, getting this bracket built took several times longer than the other three combined.
Same rankings here as in the prior category, but with a larger spread between tiers.
Mid Tier
Grok and Gemini remained roughly tied for second here, with neither exactly covering itself in glory, and the gap was considerably wider with Claude than in the prior category. The primary difference was the substantial demerits earned by both for making egregious mistakes on the easiest part of the test – the structure of the bracket itself.
Grok lowlights:
In the East region, it picked both Kansas and St. John’s to make the Sweet 16 even though they played each other in the prior round.
Same thing in the Midwest with Texas Tech and Alabama
In the Final Four, Grok had:
Game 1: Duke (East region) over Michigan (Midwest)
Game 2: Arizona (West) over Houston (South).
Problem is, neither matchup is actually possible since this year’s Final Four starts with the East region champ playing the winner of the South, and West playing Midwest
It did manage to stumble into a Duke vs Arizona final that is actually possible, though, so… better lucky than good?
Gemini nitpick:
Gemini ended its first response to my asking it to pick every game by asking me if I wanted it to pick every game (having only picked the Sweet 16 and beyond on the first pass). Not a huge deal, and Gemini remedied the issue on the second go around, but you do want your employees to make sure they fully understand the ask before heading back their desks.
Gemini lowlights:
Gemini’s Elite 8 included an East regional final of Duke over Alabama. The only problem with that — Alabama is in the Midwest region.
The “key upsets” section had a real doozy (this section’s big finale pick, no less):
The Cinderella: Utah State (9-seed) to the Sweet Sixteen. Pick them to beat (8) Villanova and then upset (1) Michigan/Florida in the 2nd round. This is the “pool-winning” move.
The issue here – neither Michigan nor Florida are in Utah State’s region. In fact, they’re not even in the same region as each other (and of course, they can’t be, because they’re both one-seeds). I suppose Utah State upsetting a team they won’t actually play really would be a stunner, though? Yeesh.
While I did enjoy Gemini’s liberal (and situation-appropriate) use of tournament analyst lingo all over its response (including in the quote above), having its cinderalla pick upset a second round opponent they definitely won’t see in the second round as a key part of its “pool-winning move” is a tough look. Like mugging to the crowd as your corner three is hitting the side of the backboard.
Also, that Utah State “pool-winner” was from Gemini’s initial response. By the time I had confirmed I wanted picks for all the games (roughly 30 seconds later), Utah State was already lying in a pile of pumpkins with a bunch of mice and a missing slipper, having lost to Arizona in round 2. On the bright side, Arizona is actually a team Utah State could play in the second round, so there’s that?
We’ll call this a tie again; and suggest to Grok and Gemini that while we love the team building efforts, maybe don’t hit the Wednesday $20 all-you-can-drink beer night at McSwiggin’s quite so hard next time.
This is Already Getting Hard to Watch Tier
As far as rough starts go… well, here’s ChatGPT’s initial response (from last Wednesday, the day before the tournament started):
When I pointed out that Selection Sunday had been 4 days prior, it responded with a “hey, you’re right, the bracket is available” mea cupla, but then it proceeded to really just not even bother doing the assignment.
Like Gemini, it did not make all the picks the first time around. Unlike Gemini, it did not ask if I wanted it to make all the picks. Also unlike Gemini, when I subsequently asked it to make all the picks, it still didn’t do it. After a third unsuccessful attempt, I just had to take what it did pick and then make the rest out of its straight-from-a-generic-gambling-podcast advice (pick two 12s over 5s, pick one 11 over 6, etc.).
I’m really not sure how else to say it — ChatGPT just streight-up half-assed the assignment. Obviously something that’s potentially in play for any human hire, but how is that even possible for an AI agent? Isn’t a big part of the pitch for AI agents that they’ll never show up to work disheveled and glassy-eyed, smelling like an unholy cocktail of Miller Light and bong water? I guess I should just be happy there’s no odor feature? This was legitimately jarring.
Management Fast Track Tier
I’ll tell you what, this Claude kid is really going places. Just clean, concise, and mistake-free in all respects. No notes.
Each model did helpfully list the places it went to collect relevant inputs (high marks for everyone there). and there were only two tiers.
Solid Research Tier
Three-way tie among Claude, Gemini, and ChatGPT. All used at least a half dozen sources, including those providing team performance metrics, betting odds, and crowd psychology.
Ken Pomeroy Tier
Grok did make vague references to other sources but it didn’t cite them and it’s rationale for almost every game for which it provided specific analysis referenced Ken Pom’s rankings.
Don’t get me wrong, Ken Pom is a legend in the college hoops metrics community for a reason. Not just his rankings and supporting metrics, but analysis on over- and under-seeded teams, lower seeds that are bad matcups for higher-seeded opponents and so on.
In fact, I’ve generally based most of my brackets on his rankings and commentary for the better part of the past two decades. So, he should absolutely be a part of any analysis, but if an LLM really is just giving me the same thing I’ve always done (albeit in considerably less time), it’s not exactly the most compelling sales pitch.
Last, but certainly not least, what do my new analysts’ brackets actually look like? Here are the rankings after two rounds (out of 80 entries):
Gemini: Current Points - T9; Max Possible Points - T28
Claude: Current Points - T13; Max Possible Points - T13
Grok: Current Points - T22; Max Possible Points - T6
ChatGPT: Current Points - T60; Max Possible Points - T34
As we noted in the lead, there is so much randomness in NCAA tourney outcomes, it’s really hard (and frankly, unfair) to judge a model on any one year’s outcomes, but we can at least get some idea of how well the models did on the game theory aspect of the ask (which each model did explicitly acknowledge focusing on —even (artificial) human train wreck, NepoBabyGPT got this one).
Some observations of note:
Each model specifically noted that big first round upset teams mostly end up losing their next game, so in a pool where point totals double each round, the reward for nailing those first round upsets mostly isn’t worth the risk of the higher seeds potentially going four or five rounds deep (and they picked accordingly). Good observation — sounds obvious, but people often miss this and start trying to differentiate their brackets too early.
Sweet 16 - Still nothing crazy here. Gemini and Claude went straight chalk (all 1 - 4 seeds), Grok had one 5-seed, and ChatGPT had two 5s and a 6:
Well, the zags didn’t work out great. Only ChatGPT’s Tennessee pick remains alive, while two Vandys and a Wisconsin were sent packing.
The Zags also didn’t work out great (see what I did there?), having been picked by all four models, along with Kansas and of course, Florida (though 98% of the world likely took that same hit). Three of the four models were also bit by the annual Virginia face plant.
Elite 8 - carryover from the prior round’s damage should be relatively minimal here, though everyone had Florida advancing (probably still amoing the broader 90%+), Gemini had Virginia, and Claude and ChatGPT both had Gonzaga moving on.
Final Four - this is where the game theory seems to have broken down a bit, with overlap on 15 of the 16 picks…
… or not, as that 16th pick was Gemini taking Virginia
Title Game - same here as everyone has Duke-Arizona other than Claude (Duke-Michigan)
Interestingly, one of the specific game theory points made by three of the four was that everyone was taking Duke, so Arizona was the higher EV pick (with Claude sticking with Duke).
In the future, are the LLMs going to have to start factoring in who the other LLMs are likeliest to pick? And if they know that the others are doing the same thing, do they go back the other way? We could have a fun AI-Wallace Shawn/Princess Bride thing brewing here.
Further to the Duke-Arizona duopoly, of the current top 21 in the C2V pool, there are 6 Arizonas and 6 Dukes.
Within that same group, it seems Michigan was your game theory winner with only 2 of the top 21, along with 4 Houstons. As for the other 3, though…
… well, this is the flipside of going contrarian. It’s a great strategy… right up to where Florida lays a giant egg in the second round.
Hopefully our assessment of our new hoops analyst hires won’t surprise anyone by now:
Gemini and Grok both seem like solid kids — a few sloppy mistakes to clean up, but both generally had a good first week and are certainly worth keeping around to see how they progress.
Claude is just pure CEO material. Sky’s the limit for this young go-getter.
And as for ChatGPT… well, in the immortal words of Judge Smails, the world needs ditch diggers too.
Seriously, though, I’m still reeling from this performance.
Maybe college hoops just isn’t ChatGPT’s thing, but at $168 billion of invested capital (and counting) and a $730 billion valuation, shouldn’t pretty much everything be its thing?
One may also say, who cares, it’s just a dumb basketball pool. True, but if it was this hard to interact with ChatGPT doing something that required exactly one incredibly simple and straightforward prompt, how are you going to fare working with ChatGPT on any of the millions of far more complex use cases it’s supposed to excel at?
Well, we can officially add forestry to the list of $200-billion-plus US industries where 95% of companies still ask field workers produce chain-of-custody, transaction, and other critical records with pen and paper, and where management teams must collect all of that information, manually input it into inventory, invoicing, and accounting systems, manually reconcile those systems with each other and with the records of suppliers/customers (all of whom generally run on equally archaic systems), monitor inventory movements and payments, and make real time decisions with incomplete, weeks-old data.
Enter our newest portfolio company, Waldo, the end-to-end operating system for the timber supply chain. Built by an exceptional leadership team that pairs an industry veteran (multi-generation logging executive), with a CEO and CTO who bring a combined 30+ years of SaaS startup experience, Waldo replaces paper trip tickets with real-time, audit-ready e-tickets that act as the system of record connecting landowners, loggers, truckers, and mills on one platform.
Each transfer of custody is logged and confirmed by the delivering and receiving parties via a simple mobile app interface, which flows through to the back office, automatically updating inventory and accounting records and generating invoices in real time.
The reduction in administrative time and the elimination of operational delays and disputes resulting from conflicting systems of record (caused by significant lags in time between field operations and back-office recordkeeping) provide immediate, quantifiable cost savings, as well as both qualitative and quantitative reductions in operational bottlenecks, dispute resolution and operational and compliance risk for each stakeholder in the chain.
Chris pushes back on the idea that all SaaS companies are suddenly in trouble. His view is more nuanced:
Some SaaS businesses are absolutely feeling pressure
Certain moats are getting thinner as software becomes easier to build
But outcomes still depend heavily on the vertical, buyer, and sales environment
In legacy industries, the real challenge is often change management, trust, and execution — not writing code
AI still depends on strong systems of record, clean data, permissions, governance, and integrations
That makes the right SaaS platforms more valuable, not less
The takeaway: SaaS is not dead — but the market is getting less forgiving of weak products and more rewarding for platforms that truly own workflow and data.
Armilla recently spoke with Eticas founder and CEO Gemma Galdon-Clavell about what it really takes to build trustworthy AI. The conversation explored why testing models in development is not enough, and why real-world auditing, measurable standards, and accountability matter once AI is live. It also highlighted the growing role of insurance in helping companies treat AI risk as a real business risk, not just a compliance exercise.
BriefCatch has introduced RealityCheck, a new authority verification tool designed to help lawyers and judges catch fabricated quotations, misstated holdings, unsupported legal arguments, and hallucinated case citations before documents are filed. Built into the drafting workflow, the tool combines citation validation with AI-assisted analysis to help legal teams improve accuracy, reduce risk, and strengthen the integrity of court filings.
Paladin and Practising Law Institute (PLI) have partnered to give law students access to skills-based training alongside real pro bono opportunities. The collaboration combines Paladin’s student-facing pro bono platform with PLI’s on-demand legal education resources, helping students build practical experience, strengthen core legal skills, and serve clients more confidently.
YMH Studios has partnered with Magellan AI to offer full attribution across audio, streaming, and video at no cost to advertisers and agencies running campaigns across its network. The move removes a common friction point in podcast ad deals by making multi-platform measurement a standard benefit, helping buyers track performance more easily across RSS, Spotify, and YouTube.
Bricks & Bytes highlights 50 construction tech startups gaining traction across AI, robotics, materials, modular construction, and field operations. Civ Robotics is featured in this piece that underscores how the strongest companies are solving real operational problems — from layout and procurement to low-carbon materials and jobsite automation — in an industry where measurable results matter more than hype.
Eden has two open roles
Gripp has one open role
Founding Sales Executive. You can contact Tracey directly at tracey@gripp.ag.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.