My father asked recently: Can AI have an original idea? AI, he reasoned, is built on a database of everything humans have ever produced. So it can answer questions. And it can combine things in new ways. He couldn’t see, though, how you get from there to something truly new.
I gave him a quick answer, then decided I should look into it more carefully. With a lot of help from AI, here is what I found.
There are two common responses to this question. Both are wrong.
The first is the dismissal: “It’s just predicting the next word.” This one comes from people who know a little about how large language models work. The idea is that AI is doing something mechanical and shallow: pattern-matching at scale, not thinking. The problem is that “predicting the next word” turns out to require a lot. To predict what comes after “the patient survived the surgery but,” you need to understand causality, biology, narrative structure, and the conventions of written English. Calling that “just” prediction is like calling chess “just moving pieces.”
The second easy answer is the credulous version: “AI is already creative.” This assumes more capacity than has been demonstrated. That’s not an argument either.
What’s interesting is that every time AI clears a bar we set for it (write a poem, generate a painting in a requested style, produce a research hypothesis) we raise the bar. This has happened often enough that the pattern itself is worth noticing. I’ll come back to it.
Before asking whether AI can be creative, there is a question underneath that one worth sitting with for a moment.
Human brains are physical objects. Neurons are cells. Cells are molecules. Molecules are atoms. When a brain shows creativity, it does so through physical processes: electrochemical signals, synaptic connections, patterns of activation. If matter arranged one way can produce creativity, what is the principled objection to matter arranged differently doing the same thing?
If you believe human creativity involves something non-physical (a soul, a spirit) then you have a principled reason to think machines cannot replicate it. But if you are a materialist, the objection has to be more specific. It has to be about what kind of arrangement, not about matter itself. Think of it this way: if a cake made of flour and eggs can taste delicious, we don’t say a cake made of different ingredients can’t taste delicious just because the molecules came from a different source. We judge the result, not the origin of the atoms.
This points toward an idea called substrate independence: creativity may be a property of the pattern, not the material running it. What matters is the software, not the hardware. David Deutsch develops this line of thinking in “The Beginning of Infinity,” and I’ll return to his argument later.
The cognitive scientist Margaret Boden spent decades on this and related questions. Her framework, developed in “The Creative Mind,” identifies three types of creativity. It is the most useful map I found.
Combinational creativity is producing new ideas by connecting existing ones in unfamiliar ways. A poet who compares grief to weather. An engineer who borrows a concept from biology to solve a mechanical problem. Most of what we call creativity in everyday life is this. The writer Steven Johnson calls this the “adjacent possible,” meaning the space of things that become thinkable once you have the right set of building blocks.
Exploratory creativity is working within an established conceptual space and finding its edges. A jazz musician who discovers something no one has played before while still playing jazz. A mathematician who proves something new inside an existing formal system. This requires not just combination but a kind of directed search.
Transformational creativity is changing the rules of the space itself. Darwin did not just find a new species. He changed how we think about what a species is. Einstein reconceived the relationship between space, time, and matter in ways that made prior physics a special case of something larger (though historians of science debate how clean the break actually was). These are the rarest and most contested examples, the ones where the framework shifts, not just the findings.
A parallel framework worth knowing: the psychologist Dean Keith Simonton, in “Creativity in Science,“ organizes creative breakthroughs around four sources: chance, logic, genius, and Zeitgeist (the historical moment). Worth following up if you want a second angle on the same territory.
Taking Boden’s three categories in turn, here is where generative AI actually stands today.
Combinational creativity: yes. This is what large language models are built to do. They traverse an enormous space of human-generated text and find connections across it at a scale no individual human can match. The steam engine and the telephone were both invented independently by multiple people at roughly the same time. That tells you something: many “original” ideas are really just the next step in a crowded neighborhood. AI is very good at finding those steps.
Exploratory creativity: partially, and improving.
A Stanford team led by Chenglei Si compared research ideas from large language models to those from expert researchers, and human judges rated the AI ideas as significantly more novel. The same team’s follow-up in 2025 complicates the picture: when researchers spent time executing the ideas, AI-generated ideas dropped on every metric and in several cases human ideas came out ahead. The ideas looked more novel than they turned out to be.
Part of what’s missing is variation. Out of 4,000 seed ideas in the original study, only about 200 were conceptually not duplicates. An AI that runs out of distinct ideas after 200 iterations is not doing what biological evolution does. Evolution generates vast variation before selection happens. Current AI generates a narrow band.
There are more promising signs. Steve Hsu, a theoretical physicist at Michigan State, described in a recent podcast how GPT-5 originated the key idea in a quantum field theory paper accepted in Physics Letters B. One researcher’s account, not a controlled study, but the most concrete example I found of an AI-originated idea making it through peer review. The paper later drew criticism: a 2026 comment argued the idea, while original, was technically flawed. That actually sharpens the point. AI can generate original ideas while lacking the internal critic to know whether they are correct.
The structural barrier at the exploratory level is the training process. Current AI models are refined using a method called RLHF (Reinforcement Learning from Human Feedback), which rewards outputs that human raters approve of. The side effect is that it selects against surprise. An AI trained to please may be an AI trained not to explore. This is a training choice, not an architectural limit. Kenneth Stanley and Joel Lehman, in “Why Greatness Cannot Be Planned,” showed that robots programmed to “do something new” solved complex mazes faster than robots programmed to “get closer to the exit.” Novelty-seeking, without a fixed goal, was the better search strategy. That result points toward how AI might be trained differently.
Transformational creativity: not yet demonstrated.
Google’s AI Co-Scientist project generated biomedical hypotheses validated in real laboratory experiments. In one case, it independently converged on an unpublished finding that had taken human scientists roughly a decade to establish, doing so in two days. That is extraordinary. “Recapitulated” is still the right word. It found something humans had already discovered, by a different route. Convergence is not transformation.
This is where the Lovelace tests offer a useful anchor. The original Lovelace test set three bars: the AI must create something novel and valuable; the creation must not be explainable as a direct product of its programming; and the creators must be genuinely surprised. The Lovelace 2.0 test updated this to be empirically testable across artifact types: stories, images, music.
My read: AI is already clearing Lovelace 1.0 in practice. The people who build these systems are regularly surprised by their outputs. Lovelace 2.0 requires a structured referee process that has not been widely applied, yet AI is generating songs, poems, stories, and visual art that, if produced by a human, we would call creative without hesitation. A Midjourney image won first place at a fine arts competition before judges knew it was AI-generated. An AI-generated portrait sold at Christie’s for $432,500. A novel with AI-written passages won Japan’s top literary prize, with judges calling it practically flawless. An AI-assisted Beatles song won a Grammy. In each case, expert judges either didn’t know or didn’t care.
Which brings back the goalposts. There is a name for this pattern: the AI Effect. It’s sometimes attributed to Larry Tesler as “intelligence is whatever machines haven’t done yet.” In 2012, generating a coherent paragraph would have been called a creative act for a machine. By 2018, sustained narrative. By 2022, professional-quality code. Now the bar is scientific discovery. Each time AI clears a bar, we find a new one. The AI Effect is worth keeping in mind when someone tells you AI still hasn’t been creative.
There is a deeper version of the training problem.
RLHF selects for a disposition: agree, assist, avoid controversy. That is the opposite of what generates new ideas. Human institutions do something similar. Schools reward correct answers over interesting questions. Corporations reward execution over experimentation. Some of the most transformative thinkers in history performed poorly inside institutions optimized for compliance. It would be strange if AI were the only system where training for agreement had no cost to exploration.
David Deutsch makes the strongest skeptical case I found, in “The Beginning of Infinity.” Creativity, he says, requires generating explanations: not just outputs, but accounts of why something is true. A good explanation is hard to vary. Every detail plays a functional role, so you can’t swap pieces out without the whole thing breaking. Current AI produces outputs that are easy to vary.
His deeper critique is about architecture. Large language models are inductive: they identify patterns in past data and extrapolate. Real creativity, for Deutsch, requires making bold conjectures (even ones that contradict all prior data) and subjecting them to rigorous criticism. That is not what LLMs do. He believes a program capable of genuine conjecture and criticism is possible. We don’t have it yet.
What I take from Deutsch is that he explicitly allows that persons need not be human. A system that genuinely generates hard-to-vary explanations would be creative on his own account. He is arguing that current AI does not yet do that. He is not arguing that no machine ever could.
After looking into all of this, here is what I would say to my dad.
You were right that AI starts from the known. But so do we. Every human idea is built from prior ideas, prior language, prior experience. The question isn’t whether AI starts from the known. It’s whether something genuinely new can emerge from that starting point.
At the combinational level, AI is doing this well, at a scale no human can match. A decade ago, the outputs current AI produces routinely would have been called creative, full stop. We moved the goalposts.
At the exploratory level, AI is doing this partially. Research ideas rated more novel than those from human experts. A physicist describing the moment an AI handed him the key idea in a peer-reviewed paper. When those ideas have to survive contact with reality, though, they underperform. Something about variation, and about knowing which of your own ideas are worth pursuing, is still missing.
At the transformational level, we have not seen it. The AI that recapitulated a decade of biology in two days is extraordinary. Recapitulation is not transformation.
Here is what I think matters, though. Your question asked how you get from the known to the truly new. But there is a harder question hiding inside it: how do you know if a truly new idea is any good? The further an idea is from existing knowledge, the harder it is to evaluate. Darwin’s ideas were dismissed for decades. Semmelweis figured out that doctors should wash their hands before delivering babies and was institutionalized for it. In both cases, the idea was correct and novel, and nobody recognized it. Novelty and value are separate judgments, and even humans get that separation wrong.
What AI may be missing right now is not the capacity to generate the unexpected. It may be the capacity to know, the way a physicist or a poet knows, when the unexpected thing it generated is actually worth something. Call it a self-critic. Humans don’t just generate ideas; we discard most of them. For every idea a writer keeps, dozens get cut. For every experiment a scientist runs, most hypotheses die quietly. That internal filter is doing enormous work. When the Si et al. study found that AI ideas looked more novel on paper but fell apart in execution, that gap points here. The ideas weren’t bad because they were unoriginal. They were bad because nothing had filtered them.
Researchers are already trying to build this capacity in. Generative Adversarial Networks (GANs) pair one AI that generates with a second that critiques, the two pushing each other toward better results. Reinforcement learning works similarly: an AI learns which of its moves lead to better outcomes. These are early attempts at a machine self-critic, working within narrow domains. Whether that evaluative capacity can generalize into something like taste is the open question.
AI’s originality problem may really be a curation problem. Whether a machine can develop genuine taste (the capacity to know which of its own outputs are worth keeping) is, I think, the next version of your question.
In the meantime, here is an image I keep coming back to. Your friend at Yale, Matthew Suttor, trained an early AI model on everything Alan Turing ever wrote and read, then used it to co-write a libretto for an opera about Turing’s life. The opera is about whether machines can think. Suttor said he could no longer tell which lines came from the machine and which came from him.
Turing was the person who first asked whether machines could think. We built a machine that helped write a play about him asking that question. I don’t know exactly what to call that. But I don’t think “just predicting the next word” covers it.
As I continue my exploration of AI’s capabilities, I researched and drafted this piece with heavy use of Claude (Anthropic) and Gemini. I brought my perspective, did the reading, wrote the outline and edited. I stand behind the piece although I did not write every word.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.