About three months ago I really got back into creative writing. Growing up I went through phases and obsessions like most kids do. One of the constants was my love for writing. I experimented with poems, screenplays, novels, short stories. Somewhere in my parents’ home there are probably old notebooks with page after page of teenage angst fuelled prose. OK — the purpose of this was not to make you cringe but to share some context on why this matters to me.
Until recently, creative tasks like writing fiction were something only humans could do well. A few years ago, LLMs appeared and challenged this long-held opinion. If you look at the best-selling books on Amazon and recent scandals, AI is already being used to write books (I’m not going to go deeper on this, but I think there are different shades here — using AI to edit or refine your work versus getting it to write everything are very different processes). But I digress. Social media posts from creative communities in particular argue that anything touched by a large language model is “slop”, instantly recognisable, instantly devalued. Used AI for research? Slop. Used AI for feedback? Slop. Extreme, but not as rare as you might think.
Others have made peace with AI-assisted writing and barely register the distinction any more. Why should they care, if it’s good? The rest of us are somewhere in between: aware that AI text has tells, but not entirely sure we’d catch them under real conditions, or what the implications are if we did.
A study published this week by Sears and Weisberg (2026) puts this exact question to the test using short fiction. Can people tell if something is written by AI? And do they prefer human or AI-written text?
The idea that people struggle to distinguish AI from human output isn’t new. It’s something researchers have been studying well before ChatGPT entered our lives. In 1972, researchers developed a program imitating a patient with schizophrenia, which successfully fooled a panel of judges nearly half the time (Colby et al., 1972). More recent work has repeated the pattern across domains that seem like they should be easy to judge: moral advice, student essays, advertising copy, Instagram captions. People frequently can’t tell, and when they do guess, they tend to rate the AI’s output more favourably on dimensions like clarity, trustworthiness, and even virtue.
Creative writing seemed like it might be the exception. Literary fiction, in particular, often rewards ambiguity. For example, a story that resists a tidy takeaway or resolution is frequently judged as more sophisticated, not less. Since AI-generated text tends to be unusually fluent and easy to process, and people generally prefer whatever requires less effort — a well-documented effect called processing fluency (Reber et al., 2004) — you might expect creative writing to be where the machine’s fluency stops being an advantage.
Well, it isn’t. Porter and Machery (2024) paired poems by Shakespeare and Plath against ChatGPT’s attempts to imitate them, and participants were, if anything, slightly worse than chance at spotting the human original. Köbis and Mossink (2021) found something similar for poetry, with one important caveat: people could tell the difference only when the AI poem was picked at random. Once researchers curated the output by letting ChatGPT generate a batch and then selecting its best output, the advantage disappeared.
Meanwhile, a separate body of work documents the opposite bias in people’s stated preferences: once you tell someone a piece was AI-generated, they mark it down, regardless of how it actually reads. This is a variant of algorithm aversion we’ve covered previously, the tendency to distrust automated systems even when they perform well (Dietvorst et al., 2015). This effect shows up reliably for art, advice, and writing alike. A recent large-scale study found this penalty so persistent that even when we directly tell participants how capable AI has become it doesn’t shift their ratings (Raj et al., 2026).
So previous research supports that people can’t detect AI writing, but they’ll punish it anyway once they’re told. Sears and Weisberg set out to test both claims at once, using fiction, a domain where human judgment is supposed to matter most.
The researchers ran three studies using six short stories (link to the stories), roughly 1,000 words each — three written by human authors and published in literary journals, and three generated by ChatGPT 4.0 to match the human stories’ themes. The prompts used specified things like a story about generational memory using koi fish as a symbol, or a government worker questioning his purpose during an air raid. Story quality was measured with two validated scales: one tracking absorption (how drawn in the reader felt) and one built specifically for this study, covering plot, character depth, theme, and use of literary devices.
Study 1: For the first study, the researchers recruited 1,682 participants and tested perception under deception. Participants read one story, but half were told the truth about who wrote it and half were lied to. Overall, participants rated AI-generated stories as more absorbing and higher quality than the human-written ones. They also rated whichever story they were told was written by a human more favourably, regardless of what they’d actually read. This confirmed previous research; people liked the AI’s output more, while still holding the belief that human authorship matters more.
Studies 2 and 3: 905 participants were recruited in two studies and this time and no deception occurred. Instead participants were given both a human and an AI story, and asked to pick which was which. This is the closest the study gets to a real Turing test! In Study 2, participants performed significantly worse than chance, about 39% correct, when 50% would be a coin flip. In Study 3, performance was at chance, around 52%. In other words, people couldn’t reliable tell the difference.
Participants were also asked to explain how they made their choice. People who reported basing their decision on “the story’s language” (i.e., how it read, how natural it sounded) were more likely to be wrong. The features people trusted as tells were, on average, actively misleading them.
One variable, however, did predict accuracy: self-reported AI literacy. People who used AI tools more, and scored higher on a validated AI literacy scale, were modestly better at spotting AI-generated fiction. Familiarity with fictional literature itself made no difference at all. Knowing what good writing looks like doesn’t help detect AI!
It’s true that AI text often has recognisable patterns at the sentence level — certain transitional phrases, a tendency toward balanced “on the one hand” constructions, an even, controlled emotional register. The Sears and Weisberg data, however, suggest that noticing a pattern and correctly attributing authorship are different skills.
There’s a deeper problem with treating “spot the AI” as a reliable skill: the detection tools built to do this systematically fail on certain kinds of human writing. Research on automated GPT detectors found they misclassified more than 60% of TOEFL essays written by non-native English speakers as AI-generated, while correctly identifying native-speaker essays almost every time (Liang et al., 2023). The mechanism is the same one at play in the human judgments above: text with lower lexical variety reads as more “AI-like” to a detector, and non-native writers, students, and people writing outside their first language often produce exactly that kind of text for reasons that have nothing to do with authorship. Having learned English as a second language myself, I'd add: a lot of what gets flagged as an "AI tell" is just what you're explicitly taught to do when writing formal essays in a non-native language!
Put together, these findings paint a slightly uncomfortable picture. We are not as good at detecting AI-generated writing as our confidence suggests. Our instincts about which features to trust are often unreliable, and the discomfort we feel about AI-generated creative work exists somewhat independently of whether we can actually identify it.
Sears and Weisberg suggest a mechanism for why AI writing might read as “better” in the first place: it’s often assembled from an enormous range of existing text, which may smooth out the idiosyncrasies and rough edges that mark individual human style. This is similar to how a face generated by averaging many real faces together tends to be rated as more attractive than any single real face (Langlois and Roggman, 1990). Familiar, competent, and averaged can look a great deal like “good,” at least on first read.
For anyone working with AI-assisted content a few things follow directly from this research:
“Sounding human” and “being written by a human” are not the same signal, but users conflate them constantly. If your product uses AI-generated content, don’t assume people can reliably tell, and don’t assume they’ll judge it worse on the merits if they don’t know its origin.
Disclosure changes evaluation independently of quality. If you label something as AI-generated, expect a rating penalty that has nothing to do with the actual content and its quality — this is consistent enough across studies.
Familiarity with AI, not general expertise, builds detection skill. If accurate identification matters for your use case (e.g., content moderation, academic integrity tooling, trust and safety), training people on what current AI systems actually tend to produce will do more than trusting general literacy or subject-matter expertise.
Like all research, the study discussed above had some limitations. In particular, the stories used in the study were short — about five minutes long — and all realistic fiction. Longer-form writing, or specific genres, might produce a different pattern entirely. The forced-choice design in Studies 2 and 3 also told participants upfront that one story was AI and one was human, which is more information than most of us have in the wild. This suggests that real-world detection could probably be even harder than what this study measured. Finally, the researchers used self reported measures of AI and literary expertise, which could be exaggerated/inaccurate.
Despite the limitations, researchers showed that across three studies and nearly 2,600 participants, people as a group were not reliably better than chance at telling human from AI-written fiction.
I’m curious to hear from you. What are your thoughts on this topic?
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.