There is a specific moment when an AI handicapping model goes from being a useful tool to being a liability. It happens when the model generates a probability, converts that probability to a fair-odds price, and then — before you can act on it — the pool it’s playing in is too small to hold the bet without destroying the edge.
This is the pari-mutuel liquidity problem. It is the most important structural constraint in AI-assisted horse racing handicapping, it is almost never discussed, and the tools proliferating right now, such as EquinEdge, Horse Race Ready, a half-dozen gradient-boosted models on Substack and Discord, essentially never address it. They give you a win probability. They do not tell you the pool at which that probability becomes actionable.
This piece does.
The Daily Report is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
The Foundational Math
Pari-mutuel betting is not wagering against a bookmaker. It is wagering against every other person in the same pool. The house takes a percentage off the top — generally 10 to 30 percent depending on the track and wager type — and all remaining money is pooled together. Odds are calculated when the horses break from the gate: if 40 percent of the win pool sits on one horse, that horse returns roughly even money after takeout and breakage, regardless of how good its actual probability of winning is.
The implication for AI modeling is mechanical and unavoidable. When your model says a horse has a 40 percent win probability and the pool is pricing it at 30 percent — implying fair odds of roughly 3/2 against a tote price of 5/3 — you appear to have a meaningful overlay. But the moment you bet into that pool, you change the odds. In small pools, a single bet can crush your own price and erase the very edge you thought you found.
This is not a hypothetical concern. It is the mathematical ceiling of every probability-to-price conversion in racing.
Where the Handle Actually Lives — and Where It Doesn’t
To understand the severity of the liquidity problem, you need to understand the structure of racing handle in 2025 and 2026.
Total U.S. handle for 2025 fell to nearly $11.03 billion, down from more than $11.26 billion in 2024. But that headline number conceals a profound distribution problem. Handle is not evenly spread across tracks, race types, or even pool types.
When adjusted for inflation, since 2003 handle has fallen by 57.3% and there doesn’t appear to be anything on the horizon that will reverse the trend. The sport that peaked at $15.18 billion in nominal dollars in 2003 now generates $11 billion, and in real purchasing power, the collapse is even more severe. The average field size in 2024 was 7.45 starters per race, a thin field by any historical standard.
The distribution of that $11 billion matters enormously to the AI bettor. Kentucky Derby day alone generates hundreds of millions. The top 15 to 20 graded stakes programs: Breeders’ Cup, the Triple Crown, the big Saratoga and Santa Anita meets, account for a disproportionate share of total U.S. handle. The rest of the race card, every day, everywhere else, is running on thin pools that the academic and practitioner literature consistently identifies as structurally inefficient in ways that work against you.
Maiden races at small tracks. Tuesday afternoon claiming races at Parx or Thistledown. Off-the-turf reruns with four starters. These are the environments where AI models are most aggressively marketed to retail bettors and they are precisely the environments where pari-mutuel liquidity is weakest, pool impact is greatest, and the market is most susceptible to distortion.
The CAW Layer: The Other Pool Problem
Before we can discuss what retail bettors with AI models should do, we have to discuss the ecosystem those bettors are operating inside. Because the liquidity problem has another dimension that most retail handicappers don’t fully reckon with: the pools are not just thin. They are actively predatory.
Computer-assisted wagering is now estimated at $3 billion to $4 billion a year, accounting for some 30 percent or more of annual racetrack handle. CAW players, algorithmic syndicates operating with real-time data feeds, sub-second batch wagering capability, and preferential access, are the dominant force in most competitive pools.
CAW bets usually come seconds before a race begins and often result in dramatic odds changes that aren’t reflected in pari-mutuel pools until it’s too late for average bettors to do anything about it. At a recent Keeneland meet, a horse left the starting gate at 8-1 but won the race at 3-1. That’s a 63 percent odds drop — a $10 win bet expected to pay $80 paid only $30.
The rebate structure is the mechanism that makes this sustainable. Average takeout across all wagers nationally can be estimated at 20 percent. CAW platforms and tracks negotiate a much lower rate, estimated at 5 to 9 percent — meaning they retain at least 11 percent of their bet as a rebate. Bet $100,000 on multiple wagers on a race, lose $10,000, then get rebated $11,000, they are ending up $1,000 ahead regardless of outcome.
The October 2025 RICO class action filed by Hagens Berman against Churchill Downs, NYRA, and the Stronach Group alleging that nearly $4 billion in annual wagering is affected by these practices and that inside bettors enjoy “no-risk, no-loss wagering opportunities,” is the most dramatic expression of what retail handicappers have understood for years: the pools they’re betting into are not level playing fields.
This is the context for every AI model marketed to retail bettors. Your model generates a probability. A CAW syndicate with direct tote access and a 10 percent rebate is also generating a probability, a better-funded, faster-acting, more data-rich probability, and betting it into the same pool, seconds before post. The model-to-price conversion that looked profitable at 9:55 a.m. may be destroyed by 9:59 a.m.
The Pool Size Threshold:
Where Do Models Actually Work?
With that context established, we can address the central practical question: at what pool size does an AI model’s edge survive?
There is no single published threshold, but the existing literature and practitioner data allow us to construct a working framework.
The self-defeating bet calculation. In a win pool of $10,000 with a 20 percent takeout, the effective pool available for distribution is $8,000. A $500 bet on a horse represents 5 percent of the total pool. If your model gives that horse a 30 percent win probability — implying fair odds of about 7/3 — and the horse is currently at 3/1, your bet moves the pool distribution meaningfully. Five percent of additional money on a horse that represented, say, 20 percent of prior betting shifts the tote price noticeably. You are betting against yourself.
A $1,000 win bet is a drop in the bucket if you wager into a $1 million win pool — it represents just 0.1 percent of the pool and might not even knock a 20-1 shot down to 19-1. That math inverts catastrophically in a $10,000 pool. The same $1,000 bet is 10 percent of the pool and will move every horse in the field.
The practical threshold for meaningful pool impact resistance depends on bet size, but a $100 unit bettor targeting single-race win pools needs a minimum win pool of approximately $50,000 to ensure their bet represents less than 0.2 percent of the total, small enough to be price-neutral. At $50 unit, the threshold is around $25,000. Below those thresholds, pool impact is non-negligible and the model’s probability-to-price conversion is unreliable.
The favorite-longshot bias compounds the problem in small fields. The structural mispricing of racing markets — longshots chronically overbet, favorites underbet — is well-documented going back to Griffith’s 1949 analysis and confirmed in datasets covering millions of races. The rate of return on win bets declines as risk increases: betting randomly yields average returns of roughly -23 percent, while betting the favorite in every race yields returns closer to -15 percent, and betting 100-1 longshots produces returns of approximately -61 percent. NBER
In small fields — maidens with five or six starters, off-the-turf reruns, low-level claiming races — this bias is amplified because there are fewer price discovery participants. The favorite-longshot bias at North American racetracks has moderated over time as CAW players and sophisticated syndicates have driven money more efficiently into favorite pools. But this moderation has happened in large, liquid markets. Maiden races at regional tracks are still inefficiently priced and the inefficiency typically works against the retail bettor because the handful of informed bettors (trainers, connections) who know these horses push the prices most, while the recreational money that might create value on the favorite is absent.
The data problem in maiden races is categorical, not just quantitative. AI models trained on past performance data have a fundamental input failure in maiden races: first-time starters have no race history. Speed figures, pace profiles, class history — the core variables that give a model predictive power in open allowance and stakes racing — are simply absent. The model is forced to rely on workout data (inconsistently recorded and of marginal predictive value), breeding information (useful but noisy), and trainer/jockey statistics (useful but highly pool-dependent). Machine learning handicapping can uncover hidden patterns in racing data, including horse performance trends, jockey and trainer statistics, and track conditions — but it depends on data availability and quality. In maiden special weight races with multiple first-time starters, that data availability collapses.
The result is a model with high nominal confidence intervals and thin actual predictive validity operating in pools with insufficient liquidity to absorb the bet without moving the price.
The Daily Report is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
Where the AI Edge Actually Survives
This is not an argument against AI handicapping. It is an argument for precision about where that handicapping is and is not actionable.
The conditions under which an AI model’s edge survives pari-mutuel pool impact are relatively well-defined:
Win pool threshold: $100,000 minimum for meaningful retail bet sizing. Below this level, pool impact becomes a material factor for anyone betting more than $50 units. Above $500,000, pool impact concerns are largely irrelevant for retail bettors.
Race type: Stakes and graded stakes, not maiden or low-level claiming. The predictive variables that power AI models — speed figures, class, pace projections — are most reliable and most richly represented in horses with extensive race histories running at class levels where data is dense. EquinEdge’s top win-percentage horse wins 32.9 percent of the time — a meaningful edge above the breakeven win rate of roughly 20 percent for a 4-1 shot — but that figure reflects performance across all race types. In maiden and low-level claiming races, model accuracy drops substantially because input quality drops.
Track tier: Major ADW-accessible tracks with commingled pools. The pool sizes at Churchill Downs, Saratoga, Santa Anita, Gulfstream, and Keeneland are categorically different from those at Penn National, Delta Downs, or Mountaineer. Commingled international pools, which aggregate betting from North America, Australia, Hong Kong, and Europe, create the deepest, most efficient markets and the most robust price discovery. These are the environments where model-generated probabilities survive contact with the tote.
Exotic bets in deep pools, not exactas in shallow ones. The counterintuitive implication of the liquidity problem is that large-field graded stakes exactas and trifectas can offer better model expression than win bets in smaller fields. A deep exacta pool at a major track allows you to play your model’s confidence distribution across finishing combinations without pool impact concerns — and the combinatorial complexity of the exotic pools means there are more mispricings to find. In contrast, if you’ve bet a Pick 4 sequence in a pool with only $250 wagered, with a 20 percent takeout, the most you can win is $200 — and even a quartet of 50-1 winners would return less than four straight win tickets.
The CAW-excluded pools. There is an emerging natural experiment in the industry that AI retail bettors should be watching closely. After NYRA began cutting off CAW players from win bets with two minutes to post, a NYRA executive confirmed that win pool payouts were “substantially higher” compared to other pools where CAW remained active. Santa Anita implemented a similar restriction for parts of its 2025 meeting. These CAW-restricted win pools — whatever their other characteristics — are the closest thing to a fair test environment that exists in American racing. They are also the pools where your model’s probability-to-price conversion is most likely to be reflected in the final tote number.
A Practical Framework for the AI Bettor
Synthesizing the above, a rigorous retail AI handicapper in 2026 should operate with the following constraints:
Minimum pool filter: Set a hard floor. Serious practitioners use $50,000 to $100,000 minimum win pool as a prerequisite for any wager. Below this level, pass — regardless of what the model says.
Overlay buffer: Bet only when the tote meets or beats your fair odds by a buffer that covers estimation error and late money. A 10 percent overlay buffer is a practical starting point. In maiden races, increase this buffer substantially — to 20 to 25 percent — to account for the model’s degraded input quality.
Race type hierarchy: Grade your bet sizing by race type. Full unit in graded stakes and open allowance at major tracks. Half unit in non-graded stakes and restricted allowances. Quarter unit or avoid in maiden and low-level claiming. Zero in maiden races with multiple first-time starters unless the pool exceeds $200,000.
Exotic expression over win-pool dominance: In thin pools, the win pool is where your money moves the price most. The exacta and trifecta pools for the same race may be proportionally deeper because more combinations are in play. Use your model’s finishing distribution to structure exotic combinations rather than concentrating in the win pool.
Track-to-track pool monitoring: Not all $11 billion in annual handle flows equally. Before race day, check Equibase or TVG for pool data from prior days of the current meet. A track averaging $800,000 in win pool per race is a different betting environment than one averaging $80,000. This is not optional due diligence — it is the primary filter that determines whether your model’s edge is expressible.
The Industry Context:
Why This Gets Worse Before It Gets Better
The liquidity problem is structural and the structural forces are all moving in the wrong direction for retail bettors.
Handle was last up in 2021, but that was an outlier against pandemic-suppressed 2020 comparisons. Prior to that, the last handle increase was in 2018. The real handle decline since 2003 is close to 60 percent inflation-adjusted. Handle in the first quarter of 2025 was down 14.41 percent compared to the first quarter of 2024. These are not marginal fluctuations. The pool shrinkage is structural and secular.
Meanwhile, CAW players now account for as much as 30 to 40 percent of total pari-mutuel handle.
As retail bettors exit — driven by the late-odds drops, the difficulty of competing against algorithmic syndicates, the superior entertainment value of sports betting — the pools that remain are increasingly dominated by the very players your AI model is least equipped to compete with.
The AI handicapping market is growing at exactly the moment the market conditions for AI handicapping are deteriorating. More models, thinner pools, more sophisticated competition. The tools are getting better; the environment they operate in is getting worse.
That doesn’t mean the edge is gone. It means the edge is smaller, more concentrated, and available only in specific conditions that most retail bettors are not systematically filtering for.
The AI model is not the problem. Using it without a pool filter is.
The Daily Report is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.