Earlier this week, a highly selective philosophy journal knowingly published an article generated largely by AI. Claude, an LLM from Anthropic, wrote an essay in political philosophy for Philosophy & Public Affairs. Simon Goldstein, a professor at the University of Hong Kong, turned the crank.
The editorial board thought carefully about whether to consider submissions like this and decided—lol no, they just stumbled into it. Thanks to a poorly designed editorial management system (imagine!), the editor hadn’t read Goldstein’s cover letter describing his method, and Claude’s role came to light only late in the review process. By then, the paper had already been through multiple rounds of review. And according to Wiley, the publisher, AI use is not itself grounds for rejection. At this point, Jason Brennan, the journal’s editor-in-chief, thought what the hell, why not:
It’s plausible that the primary purpose of a journal is to publish work that advances knowledge, rather than to allocate professional credit… I don’t see an obvious reason to assume in advance that philosophical arguments or insights produced through substantial human-AI collaboration cannot merit publication.
Philosophers and other academics have to decide whether they want journals to publish articles produced substantially with AI and, if so, how much disclosure is required. Brennan describes the choice to publish Goldstein’s article as “an experiment” that will “force us to confront those questions.”
How was AI involved in this case? Goldstein explains his process in a guest post on the blog Daily Nous. He gave Claude “a one paragraph explanation of the thesis” plus “general instructions on writing philosophy papers.” Claude then produced “a series of outlines and drafts,” after which Goldstein offered detailed comments and corrections, nudging the essay in some directions and away from others. In sum, he “worked with Claude the way a heavy-handed supervisor writes with a PhD student in some fields.”
The paper was good enough to pass anonymous peer review at a journal that rejects about 95% of submissions. But early readers commenting at Daily Nous aren’t impressed. Some claim that the paper badly “mischaracterizes” the theories it discusses, others that the formalism is “sloppy” at best. Still others charge that the prose is weak (which Goldstein admits).
It’s hard to know what to make of these criticisms. Perhaps many published articles would buckle under the same scrutiny. And the criticism might be motivated—it comes from people who are aware of the paper’s provenance and are largely hostile toward AI. One friend says that the paper isn’t great—“too formal and goes on way too long”—but also that lots of philosophy papers are like that, and absent the provenance no one would bat an eye.
But suppose the critics are right. In that case, the paper’s flaws may stem from Goldstein’s lack of domain expertise. He specializes in AI and epistemology, not political philosophy. Plus, Goldstein “deliberately avoided intervening too much in Claude’s process…to see how far Claude could go to write something publishable.” Of course, you wouldn’t expect a co-authored paper led by most PhD students to be very good if their supervisor wasn’t an expert on the topic and often held back.
Yet this suggests a better kind of AI-assisted paper—one where the human supervisor does have domain expertise and intervenes heavily. (This is closer to Goldstein’s typical workflow.) Then the paper might actually achieve Goldstein’s stated goal, one that many other academics share: “to discover new, important truths.”
Goldstein’s AI-generated paper is the first to appear openly in a selective journal, but I bet others have flown in under the radar. Even so, we don’t know for sure whether those papers are high quality. (The review process is noisy.) And even if the authors confessed, you’d expect motivated criticism. We’d need blind taste tests—where critics don’t know which papers are AI-assisted and which aren’t.
But something like Goldstein’s process, guided by domain experts, has every chance of producing good work. After all, papers written this way with PhD students in Claude’s role are routinely good. The practice is standard across the sciences and increasingly in philosophy. (The closest analogue is bad mentorship: the student is first author but has no autonomy, and the supervisor supplies most of the ideas and decisions. Common, unfortunately.)
Like other LLMs, Claude has skills that PhD students lack, not least its ability to scan and synthesize a vast literature. Given the “jagged frontier” of AI capabilities, Claude is weaker in other ways, like being more prone to confident bullshit—but that’s where the heavy-handed supervisor comes in.
Put this in terms of a distinction between “centaurs” and “reverse centaurs.” Both are human-AI hybrids. The difference comes down to who is in control. A centaur is a human assisted by a machine; a reverse centaur is a machine assisted by a human. In this paper, Goldstein was arguably the bottom half of a reverse centaur. And my view is that we can expect better papers from a (regular) centaur, led by a human expert.
I’m an AI enthusiast, but I’m averse to reading centaur philosophy. I’m certainly not alone. And yet, if the goal is to discover new, important truths, is this aversion anything more than subjective taste? And as philosophical practice evolves, will our tastes change?
At Daily Nous, one of the most incisive criticisms of Goldstein’s project comes from USC philosophy professor Dmitri Gallow, who says that “discovering new, important truths” isn’t his only goal:
I have other goals too, like leaving behind a world capable of discovering, understanding, and competently evaluating philosophical ideas. And so I have the instrumental goal of training the next generation of a research community with these skills.
We have set up a system of carrots and sticks to produce the next generation of philosophers. It’s not a perfect system. But it is completely unclear what becomes of it when the sticks can’t easily discriminate between artificial and genuine ability, and the carrots can be purchased from Anthropic.
In such a world, will we still have people who dedicate their lives to understanding particular philosophical questions well enough to catch Claude’s mistakes and understand Claude’s contributions? How will these offices be allocated? There’s a fear that Simon’s vision for the future of philosophy is a modern recreation of Simony, with the office going to whoever pays Anthropic enough to get access to the best models. If so, then I would not expect human philosophical expertise to persist for long. And what value are the important truths being produced, if no one has the expertise to understand them?
Gallow’s worry assumes a world of reverse centaurs. But if the best journal articles are produced by centaurs, genuine (human) expertise will remain essential to philosophical practice. We won’t offload everything to AI. People will still dedicate their lives to philosophy.
Still, perhaps Daniel Muñoz is right that we’ll have to adjust our system of carrots and sticks. If AI supercharges productivity, hiring and promotion should prioritize research quality over quantity. Perhaps more importantly, committees will have to “put more weight on interviews, recommendation letters, job talks, and Q&As.”
Another possibility is that centaur philosophy will raise not only the floor but also the ceiling: the very best papers will come from the strongest domain experts collaborating with AI. Then the professional incentives to become an expert will remain.
What do you think of this argument? Are journal articles still worthwhile if they come from (regular) centaurs? Will we still be able to reliably apportion credit to the humans who create the best centaurs? Enough to preserve philosophical expertise?
Now—what if I told you that this very essay, the one I’ve been passing off as my own, was produced by Claude under my heavy-handed guidance?
It wasn’t, but why should that matter?
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.