This is the first in a series of three essays based on findings from the Organization Science AI Task Force. The full paper, “More versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review,” is available to download at the Organization Science website. Over the next two days, we will examine what is happening on the reviewer side as well as the institutional incentives driving these trends.
Something is up in academic research. Or in academic papers, at least.
If you are an editor or reviewer at a journal these days, you probably already know this. The manuscripts are arriving in greater volume, with a particular feel that is hard to pin down. On the surface, the papers look the same as ever, but the writing feels weightless in a way that rarely describes academic writing. The sentences connect smoothly, include citations to actual papers, and draw inferences from real (you hope) analyses. Yet you find yourself scratching your head at the meaning the words are trying to convey.
As the AI Task Force for Organization Science, we spent the past several months trying to put form to this feeling. What we find, in short, is that AI language models, in combination with strong publish-or-perish incentives, are pushing the field to produce more research rather than better research. The system is being overwhelmed by research that is, at a minimum, assisted by AI and, increasingly, substantially generated by it. This research is worse along two observable dimensions: writing quality (perhaps surprisingly) and overall quality, as measured by editorial outcomes (also perhaps surprisingly). Moreover, this research is not costless. It is imposing a burden on the volunteer labor of reviewers and editors that holds the peer review system together. Our report details what we found.
This Substack focuses on the first set of findings in our report: what is happening on the submission side of Organization Science. In Part II, we will consider what is happening on the review side. In Part III, we will take a deeper look at the institutional incentives that appear to be amplifying these trends, and our initial thoughts about what it might take to steer the system toward a different equilibrium. This last part is the most important and will need the ideas of our whole community to navigate. Peer review is a relatively recent institution, emerging in its current form in the 1970s. It is also an important one: underpinning hiring, promotion, and tenure decisions—in other words, academia as a profession. Yet it may be time for us to collectively rethink it for the age of AI. More on that in Part III.
Before we get into our findings, it’s important to state up front where we (the authors of the report) stand on AI. The short version is that we are in awe of the technology. These tools have fundamentally changed our own research and teaching over the past year. Scientists have gained a real superpower, similar to moving from the slide rule to Stata 19 overnight. In that sense, we are very excited about the promise of AI for science and we are open to experimentation with it. But we are also carefully studying how it’s being used in practice and what that means for evaluating and promoting scientific research.
The problem is that AI does not appear to be changing our field for the better so far. Instead, Organization Science, like many other journals, is overwhelmed by AI-generated research that is stressing its peer review system to the breaking point. First, let’s start with raw numbers. Submission volume at Organization Science has risen by 42% since the launch of ChatGPT in November 2022, compared to the prior two-year window. To put that in perspective, the COVID-19 pandemic, which sent many academics home with canceled conferences and more time to write, produced a 20% bump. The post-AI surge is on top of this base, and continues to increase.
There may be many possible explanations for this rise. Perhaps researchers suddenly became more productive on their own. Perhaps Organization Science’s reputation began to increase exactly in November 2022. Maybe the announcement of Taylor Swift’s Eras tour inspired a wave of new research. But when we decompose the volume increase by the degree of AI involvement in each manuscript, a picture begins to emerge.
The increase in submissions is almost entirely due to manuscripts with substantial AI-generated text. Submissions with little or no detectable AI (below 15% on Pangram, a widely used AI detection tool) have actually declined since late 2022. The gap between that decline and the 42% aggregate increase is filled by manuscripts in the higher AI categories: moderate AI collaboration (15-30%), substantial AI generation (30-70%), and what we consider mostly AI-written work (70%+). By early 2026, the majority of manuscripts submitted to Organization Science contain detectable AI-generated writing. The fastest-growing category is also the most concerning: manuscripts where 70% or more of the text appears to be AI-generated.
To be clear about what we are measuring: we used Pangram to score AI prevalence in writing on a continuous scale from 0 to 100% for each submission’s abstract. We validated this approach against full manuscripts and found that high AI scores in the abstract reliably signal heavy AI use throughout the paper, across the introduction, theory, methods, results, discussion, and conclusion sections. Based on our data, an abstract that reads like it was written by a language model is a decent, though not perfect, indicator that the rest of the manuscript was too.
A natural response to these numbers is: so what? If AI helps researchers write better papers faster, more submissions could be a good thing. The field has long worried about barriers to entry, about non-native English speakers facing disadvantages, and about the sheer drudgery of academic prose. If AI helps with these, the growth we are witnessing might carry good work in with it and more equitable opportunities. But that’s not what seems to be happening, at least from an exposition standpoint.
We calculated the average readability of submission abstracts over more than a decade. Writing quality was stable from the beginning of this period through COVID. Then, right at the launch of ChatGPT, readability scores began to drop. By January 2026, the average abstract’s Flesch Reading Ease score is 1.28 standard deviations lower than it was in January 2021. Submissions have become far harder to read.
This is counterintuitive. Most people assume that AI produces cleaner, more polished text. And in some narrow dimensions, it does: AI-assisted writing tends to be less hedging, less passive, and more specific (for instance, it includes more numbers). But on the measures that capture whether a reader can actually parse and absorb the prose, AI writing is worse. AI-generated academic writing uses longer words, more complex sentence structures, more jargon, and more nominalizations (i.e., those nouns that we never learned until we entered academia such as “conceptualization,” “problematization,” “contextualization,” and so forth).
The result is text that is superficially clear but substantively impenetrable.
To test this ourselves, Claudine took Lamar’s most recent publication (while Lamar was distracted by feeding Butterburger), stripped it of the abstract, title, and introduction, and asked ChatGPT 5.4 to write a title and abstract, aimed at a “top” journal with “expert” editors and reviewers as the audience. Here’s Lamar’s original followed by ChatGPT’s version.
Lamar’s version (Flesch Reading Ease: +22.55, 87th percentile of our sample, Pangram: 0.7%):
The Negative Consequences of Loss-Framed Performance Incentives
“Behavioral economists have proposed that incentive contracts result in higher productivity when bonuses are “loss framed”—prepaid then clawed back if targets are unmet. We test this claim by randomizing the pre- or post-payment of sales bonuses at 294 car dealerships. Although somewhat statistically imprecise, our analysis provides strong indications that the random assignment of loss framing had quantitatively important negative effects. We document that the negative effects of loss framing can arise due to an increase in incentives for “gaming” behaviors. Based on these claims, we reassess the common wisdom regarding the desirability of loss framing.”
ChatGPT’s version (Flesch Reading Ease: -2.19, 24th percentile of our sample, Pangram: 98.9%)
When Loss Framing Backfires: Experimental Evidence on Multitasking and Incentive Design
“Loss framing is widely assumed to enhance effort and performance, yet field evidence remains limited and theoretically incomplete. We report results from a large-scale randomized controlled trial involving automobile dealerships, in which a standard bonus contract was experimentally reframed from gains to losses via upfront payments with clawbacks. Contrary to prevailing expectations, loss framing weakly reduces overall performance, with substantial heterogeneity across task dimensions. Specifically, we find no effect on high-stakes sales categories but significant declines in lower-stakes categories, driven by shifts in the distribution of outcomes rather than uniform changes in effort. To explain these patterns, we develop a model of reference-dependent agents engaged in multidimensional tasks. The model shows that loss framing amplifies incentives to minimize loss exposure, which can induce inefficient reallocation of effort across tasks or toward lower-variance strategies. Consistent with this mechanism, distributional analyses reveal both increased target attainment in focal domains and offsetting declines elsewhere. Our findings challenge the generalizability of prior results on loss framing and highlight a critical boundary condition: when performance involves trade-offs across activities, loss-framed incentives may distort rather than enhance productivity.”
The ChatGPT version makes sense, is grammatically correct, and yet is a slog to get through and much harder to grasp than the original.
Now, extend these effects to the full manuscript. When AI scores reach above 30% and particularly above 70%, it means that the manuscript is similarly a slog to understand. It also implies that, during the research process itself, authors substantially delegate writing, and the thinking that goes along with it, to the models. The aha moments that come from the writing process are now gone.
If the new submissions consisted of strong papers, it would be a welcome problem. More good ideas competing for journal space is only a good thing. But the editorial process tells us that these AI-heavy submissions are, on the whole, weaker manuscripts.
Among manuscripts with 70%+ AI scores, nearly 70% are desk-rejected, meaning an editor determined they should not be sent out for review. For comparison, the desk rejection rate for low-AI manuscripts (below 15% AI) is 44 percent. By the end of the editorial funnel, only 3.2% of high-AI manuscripts receive a first-round Revise & Resubmit decision, compared to about 12% of low-AI submissions. It’s important to state here that these manuscripts are not being rejected because of their high AI scores. In fact, editors don’t know these scores when making their decisions. Instead, they are being rejected for the usual reasons — poor quality or fit — and these reasons are more prevalent in submissions with intensive AI use in the writing.
Editors are screening out many of these papers, so the system is working for now. But “working” means that editors, all of whom are volunteers with their own research and teaching obligations, are spending a growing share of their time assessing manuscripts that either don’t make it to review or are rejected once they get there. The same is true for the reviewers who evaluate the papers that make it past desk rejection. Every hour spent evaluating a paper that was fast and easy to create is an hour not spent on a paper that deserves careful attention. If reviewers and editors have less time and energy to evaluate strong submissions, then it is the diligent authors who suffer as well. The scarce input in peer review has always been expert human judgment.
That scarcity in judgment is increasingly bottlenecking the system and burning out reviewers, editorial boards, and editors. And the volume is still increasing. The editorial team at Organization Science has expanded in response, adding editors and reviewers to absorb the load. If submission rates continue on their current trajectory, however, and the AI share of those submissions keeps climbing, the system will eventually run out of willing volunteer experts to do the work.
We want to be careful about what our data shows and what it does not. These findings are descriptive. We are reporting patterns at a single journal, and while conversations with editors across the sciences suggest similar dynamics, we cannot speak for every field or every journal.
We also want to restate that we are deeply enthusiastic about AI’s potential to transform academic research, whether by enabling new methods, identifying new questions, or allowing researchers to tackle problems that were previously unreachable.
What we observe, however, is that so far, the dominant use of AI in academic writing appears to be making more not-so-great research. A paper that once took months to draft now takes days. But the papers being accelerated are, overwhelmingly, the ones that do not substantially advance the field. The result is a system where the bottleneck has shifted from production to evaluation: we can generate manuscripts faster than we can judge them.
In Part II of this series, we turn to the review side of this story. If AI is changing what gets submitted, is it also changing how submissions get evaluated? The answer, it turns out, is yes.
Members of the Organization Science AI Taskforce:
Claudine Gartenberg is an Associate Professor of Strategy at the Wharton School of the University of Pennsylvania and Senior Editor at Organization Science.
Sharique Hasan is an Associate Professor of Strategy at the Fuqua School of Duke University and Deputy Editor at Organization Science. He is the chair of the AI Task Force at Organization Science.
Alex Murray is an Associate Professor at the Lundquist College of Business at the University of Oregon and Senior Editor at Organization Science.
Lamar Pierce is the guy with the cat and bad music references. And the EIC of Organization Science.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.