“She was beautiful and seemingly quite intelligent, what with her pentameter search system. There wasn’t a reason in the world not to find her appealing.”
— Haruki Murakami, Hard-Boiled Wonderland and the End of the World
Right now, two big shifts are co-occurring in society, and I don’t think their intersection has been adequately explored. The first is the expansion of gambling in the form of prediction markets that allow you to bet onarbitrary events, not just a tightly restricted subset. The second, of course, is the rise of Large Language Models and the chatbots that are now woven into everyday life.
Together, they could trigger a dangerous and chaotic breakdown of society.
Everyone, at all times, is incentivized to affect future events to their benefit, and to avoid eventualities that will cause them harm. For example, if someone is being annoying, their peers start spending less time around them. Similarly, a worker who wants to be promoted will attempt to be more productive, or at least will strive to be perceived as such by their boss or superior.
The rise of prediction markets changes this calculus significantly because suddenly there’s a direct payoff to changing almost anything. In 2020, there was no incentive for anyone to tamper with weather stations, because what benefit could there be in fudging the daily high? That changed in 2026, when Polymarket created a market on what the daily high in Paris would be. Someone put down a small $120 bet that it would pass 22°C, even though the market gave that outcome less than a 1% chance. Minutes later, the temperature reading suddenly jumped above the line. The spike came from a single sensor. Experts later suggested the culprit may have been a hairdryer pointed at the official airport thermometer.
Generalized prediction markets alone are an enormous shift that society is utterly unprepared for. Our legal system is designed to protect us from the excesses of human nature and is specifically designed to punish crimes that people are likely to commit. There are so many bad actions a person could take that the law does not adequately punish simply because it is not in human nature to do those things. For example, car theft usually carries a stricter penalty than tampering with a weather sensor, and not because stealing a car does more harm than sabotaging government equipment. Changing the weather data for a whole city has the potential to cause catastrophic damage that far exceeds, in aggregate, the harm a single person suffers after losing their car. Of the two crimes, though, only stealing a car provides a clear personal benefit. Until now.
The arbitrariness of these markets is the reason the perverse incentives they create are so destructive. Prediction market fraud is just insurance fraud by another name, and various societies dealt with that successfully for hundreds of years. However, insurance is usually personalized to events in an individual’s life, like property damage and death, and not global happenings. There are exceptions, but they are often natural disasters or geopolitical events that are difficult to influence.
These limits keep the damage contained and make insurance fraud relatively easy to detect and legislate against. Instead of investigating everything, which is impractical, we can tell the police to investigate mysterious occurrences that are unlikely, controllable, and provide direct material benefit, like an insured barn mysteriously burning down. Forcing—making a bet pay off by causing the outcome yourself—was not impossible before, but the opportunities for it were rare and legible, so we could watch the few that mattered. To draw an analogy, most crime is committed by someone close to the victim, with an obvious motive, which makes it easier to solve. The random act, with no connection and no reason behind it, is the one that goes cold.
Other people have spelled out the danger here better than I can, so I won’t dwell further on these risks. If Polymarket or Kalshi truly becomes an everything prediction market, the results will be immediately catastrophic. To stop that, our legal system will have to adapt drastically, to a more punitive, observable version of current restrictions on insurance fraud and insider trading. Maybe some markets will become illegal. If safeguards aren’t instituted, the world will be filled with perverse incentives to manipulate real world events.
Imagine you live next to a powerful geyser that periodically jets out a stream of water hundreds of feet into the air. It’s strong enough to fling any large stone near its mouth high into the air. Peculiarly, you notice that all of them end up in two places— on top of a boulder, or in a pit. Both are right beside the geyser. It’s practically miraculous— aside from a few freak occurrences once or twice a year, all rocks left by the mouth of the geyser end up in those two places.
Furthermore, you notice that it’s about twice as likely to end up on the boulder, probably because of the precise dynamics of the water.
Hmm…
One day, you decide to make some extra cash, and you set up a colorful betting stand by the geyser and take a fee from each transaction. It’s easy money— you know the odds, and you’re bound to make money over time. Every night, you place the 30 pound Official Orange Geyser Stone in its specified location, and every morning, if it’s on the boulder, you pay out to the boulder people; if it’s in the pit, you pay out to the pit people. You have specialized construction equipment to get it back and do it the next day.
Except, after a few months of doing this, you notice something weird happening. For some reason the chance of the stone ending up in the pit has drastically increased— it’s happening two thirds of the time. Should you adjust the odds? Maybe something changed. Well, you try that, but then the odds change again!
What happened? Well, clearly there were people taking the stone overnight and dropping it in the pit. It’s an easy cash grab.
The interesting part is that the specifics of the situation are clearly relevant to how it reacted to the betting market, because dropping it in the pit is much easier than placing it on the boulder. The forcibility of the situation changed what was likely to occur. I am defining forcibility as the aggregate ability of individual actors or small groups to effect a change in outcome.
This is much different from fraud in normal incentive structures. There, something has to be both to your benefit and in your control for it to pose a threat. Here, anything with a betting market attached has a payoff, so the biggest threats are simply whatever’s easiest to control.
Think about what this means in the aggregate, over all the possible things that could happen. Whenever a small group has disproportionate power to force one specific outcome, we should expect that outcome to happen more often—regardless of whether the group has strong incentives to cause it.
Would the world be a better or worse place if this were true?
Let’s pause on another technology that’s spreading even faster and reaching even more people, the Large Language Model. They’re clearly useful: they can do everything from making presentations to discovering new math.
But LLMs are far more than productivity tools. They’re reshaping society. Education is reckoning with the implications of an always available answer key, a kind of omnipresent Chegg. Software engineers are writing less code. People are going to ChatGPT for advice, for therapy, even for love.
With this kind of power, it’s no surprise that LLMs can actively and intentionally change opinions. They’re at least as effective at changing people’s opinions as most people are. If you give them a person’s psychological profile, they can use that data to persuade them even more effectively. There is a robust body of work showing that language model generated messages can change our views. I don’t want to overstate these points, because often the effect is small, but it is evidently present.
Less studied is the magnitude of these effects over longer interactions. The prior research I cited tested the effects of a single LLM-generated message against either a human-written alternative or no message at all. But that’s not how most people actually use them. The average LLM conversation is probably about 5-10 messages. This is probably driven up by a few really long conversations, but people are not solely using it for one-off prompts.
The research we do have points to them being equally, if not more, persuasive over longer time scales. For instance, a study in 2025 found that GPT-4 was a more persuasive debate partner than most humans.
It’s hard to generalize these results across models. Each generation and family of model differs a lot, the most-used ones change constantly, and the studies are quite expensive to run. It is also hard to gauge how strong chatbot influence is over real conversations because the majority of studies focus on single-message interactions, where the effects are weakest. This setting is more comparable to reading a well-written op-ed than to having a conversation with a persuasive friend who knows you and your opinions well.
Nevertheless, there is enough information here to say they change the way we think. Do they also change the way we act?
“If they can change opinions,” you might be thinking, “of course they can change how we act!” Well, remember that these studies are really testing if a person clicks a different button on a survey before and after an interaction. Is this the same as changing who you’re gonna vote for in the 2026 election? It’s not a stretch to say these are just different questions, and treating them as the same is reasonable but still speculative. Research supports the idea that opinions have to change a lot to significantly impact actions.
Fortunately, some clever researchers have worked out how to test this more realistically. A recent study gave human subjects a small sum of money to donate and tested if language models could influence their choice of charity. They could, and outperformed humans who wrote a similar appeal. A 20 minute conversation with a chatbot was even more effective than a single message.
Even without a specific, human-provided objective, these models change how we act. Interacting with a sycophantic AI for just three weeks made people less likely to go to their friends and family for support. In a different study, three-quarters of people acted on personalized life advice from a chatbot, even though the study’s organizers did not ask them to do so. (Also notable is that following the chatbot’s advice didn’t improve their wellbeing.)
Beyond changing our minds, AI can also quietly limit our options by reducing the diversity of choices, even without tilting the average choice. To understand how this is possible, imagine a video game with a difficulty scale from 1-10 in which the average choice is a 5. By putting a warning at the top of the screen— “A difficulty of 5 is recommended for most players”— the spread decreases even if the average choice is still a 5. There are fewer unique playing experiences because more people are playing on difficulties close to 5.
Ample research shows that ChatGPT usage decreases diversity of choice. For instance, a 2025 study found that brainstormers using AI were paradoxically judged as more creative, even though their ideas were less diverse. In other words, individuals came up with out-of-the-box ideas, but they all came up with the same out-of-the-box ideas. In contrast, the group without ChatGPT produced a much wider and more diverse set of ideas in aggregate. Other studies have replicated this finding in other domains.
True, these studies don’t directly cover decision-making; perhaps people act differently when faced with real-world consequences. However, research on chatbot interaction indicates that AI can affect the tilt of both preferences and decisions. There’s little reason to think the same wouldn’t hold for the variety of our decisions.
As more of our decisions pass through chatbots, which will only grow as ChatGPT and Claude spread, their outcomes should grow less diverse.
Scenarios with fewer possible outcomes are naturally more predictable. But does that mean those outcomes are guaranteed to be more forcible? It’s hard to say. On an individual level, it does create forcible vulnerabilities that a bad actor can exploit. By randomly shrinking a wide open field of human choices down to a few predictable paths, the AI creates bottlenecks. Some of these will, by chance, turn out to be forcible.This is also where the two halves of the story merge. The market gives you a reason to force an outcome, and the widespread use of LLMs means you have a decent shot at actually forcing it.
Many large global events—like a nationwide election or a major policy shift—are ultimately just millions of these smaller, AI-influenced decisions stacked on top of each other. If individual human choices are becoming more predictable because we are all using the same types of AI to process information, then the behavior of the public as a whole becomes predictable as well.
This predictability is the exact mechanism that allows bad actors to force macro-events into reality. Throughout history, the easiest way to control a population has been to exploit their predictable reactions to information. When Adolf Hitler allegedly burned down the Reichstag parliament building, he did it because he knew exactly how the German public would react to a communist threat—they would panic and willingly hand him emergency powers. He forced a massive political outcome by controlling the input event. This was a fairly obvious power play. As chatbot usage becomes widespread, and human behavior becomes more predictable, there could be many such opportunities to identify where a small nudge can force a specific outcome.
You may scoff that these windows of opportunity would be very difficult to find. This may be true, but remember that this whole situation exists because of LLMs. As these models become more intelligent, perhaps they can identify the biases other LLMs will instill. Perhaps they will have a good sense of the biases they themselves impose. Future bad actors may not have to control the models; it might be possible to ask them.
LLM-induced predictability would be a thorn in society’s side even without Polymarket, but it’s survivable. Business leaders could use this type of power to influence public opinion and do real damage to society. Politicians could pull similar tricks. This is bad, but we managed to survive the initial rise of hyper-targeted social media and scandals like Cambridge Analytica—tools that ultimately helped keep political incumbency rates high even as overall public approval of government plummeted. We’re still here. I have no doubt we could survive— although not thrive— the further development of manipulatory tools.
The escalation lies in the ubiquity and scale of the incentives. By finding just a single prediction to manipulate, anyone can become fabulously wealthy. Prediction markets have announced again and again their intention to expand; look at Polymarket’s new ad, where they implore the viewer to ask a question, any question at all! The implicit pitch is that anything you can wonder about deserves its own market. The more questions, the more risk.
A bona fide market is a completely different beast than the Polymarkets and Kalshis of today. Right now, prediction markets are trivial and largely dominated by sports gambling. Even markets with societally salient outcomes, like congressional races, suffer from low trading volume. A hard-boiled prediction market, where you can wager on absolutely anything and spin up your own contracts at will, would have unimaginable consequences. To be clear, this is a theoretical exercise, and it only plays out if we actually get the hard-boiled version. A neutered market, where regulators ban half the contracts and police the rest like insider trading, sidesteps most issues I’ve raised. Yet commentators dream of a world where one can buy a latte by placing a bet against someone delivering it to you. This is inextricably tied to a world in which there are many, many forcible bets.
The market, obviously, will figure this out over time and you won’t be able to print money indefinitely as counterparties adjust their pricing and expectations. But at what cost? Reality itself has already measurably changed. As people identify these highly leverageable situations, they place bets they can manually force, and then they go force them. The market’s pricing adjusts retroactively, but the aggregate effect is that destructive, highly forced outcomes have fundamentally become more likely to happen.
What does a world look like when the most easily forced outcome becomes the most likely to happen? Probably very chaotic. It looks like a world where infrastructure fails because a short-seller needs a water treatment plant to go offline for an hour. It looks like a world where automated AI bots scrape localized prediction markets, match them against systemic vulnerabilities, and orchestrate hyper-targeted supply chain disruptions to get that sweet, sweet alpha. Breaking things is much easier than building them, after all. Many institutions are fragile enough already, and the extra predictability LLMs supply might be the nudge that finally pushes people to act. It might have happened in Iran already.
In my opinion? It looks a lot like the end of the world.
“The Clocktower, the River, the Bridges, the Wall, and smoke. All is drawn under a vast snow-flecked sky, an enormous cascade falling over the End of the World.”
— Haruki Murakami, Hard-Boiled Wonderland and the End of the World
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.