Photo by Smithsonian on Unsplash
Lately, I’ve been guilty of getting distracted from writing my novel. I’ve been scrolling too much on social media, and I’m aware of the partnership between Pangram and Substack. I’m promising myself that this is the last post I’ll write until I reach the 30k-word milestone.
Yeah, because writing a novel isn’t easy. It isn’t even easy for the folks on the r/WritingWithAI subreddit, now that ChatGPT has removed the feature that mimics a favorite author’s style and voice–
let alone for a clueless person like me who spends at least two hours a day writing while listening to brown noise or the sound of a thunderstorm, because for some reason these sounds calm my mind and allow me to focus on the writing without being distracted by my own thoughts.
What I’m really struggling with is the realization that I could go faster. I’ve studied almost all the habits of the most famous writers, and if there’s one thing I’ve learned, it’s that each one has developed a lifestyle around their writing, but also, they made actual sacrifices.
Clearly, I have to come to terms with the fact that I can’t write more than 1,400 words a day. Not only that, but writing the novel is overshadowing everything else: since I don’t want to lose the habit of reading at least before sleeping, let alone exercising, I’ve had to set aside my direct study of Spanish and the writing of short stories and poetry. I never had much of an active social life, not in the spaniard or European sense at least.
I’m writing this post because I decided today to devote my time and energy to writing this post rather than to my novel. As you can see, I had to give something up to write this post.
All this to tell you that before using a tool like Pangram, even before em-dashes, and even before the rule of three, the first AI-detection has always been meta in nature for me. If I saw that the frequency of posts and the number of words were excessive, compared to the writer’s lifestyle or economic situation - yes, because if a writer can write full time, I expect them to be at least a bestseller. If they’re trust-fund kids, I don’t expect them to write about the best optimization habit for highly effective people and so on; that type of content is soooo underclass-, or if the writing itself was too verbose and utterly bland, I knew the author wasn’t human.
This is because I understand the nature and economics of things. For a large language model, it costs nothing to generate content, so the content will be verbose, rambling, and lacking in meaningful information.
Human nature is different. Since writing this post is taking time away from writing my book, I’ll make sure my writing is information-rich, thought-provoking, and I’ll avoid repetition because, quite simply, every repetition comes at a cost to me.
Returning to the original discussion about Pangram and Substack, two camps have immediately formed. One camp consists of skeptics who essentially argue that they don’t trust a company that uses AI to assess whether writing is human. In short, they doubt the tool’s effectiveness and overall ethics. They don’t trust Pangram’s true intentions either, as they claim the company might use their written texts to train its models without the writer’s actual consent.
Obviously, all of this had already been stated in the partnership’s FAQ: Pangram is the model with the lowest false-positive rate, and neither Pangram nor Substack will use the text to generate any models.
Therefore, the rest of this essay will assume the following:
Pangram is the best possible model for AI-detection in text.
Pangram will not steal data.
What disappoints me about this line of argument is that no one can see beyond their own perspective and understand why this move was made, or what the real reasons are (spoiler: ethics has nothing to do with it.)
The first thing to understand is that the reason Substack made this move is pure marketing and evaluation.
Substack has long announced that it’s open to sponsorships, and obviously, having an audience of humans and not bots, that’s cohesive and has a very clear identity, is something that’s absolutely appealing to sponsors.
Did you really think Substack could survive solely on the percentage of our paid subscriptions?
But what are the CONSEQUENCES of this partnership?
First of all, I’m not worried about a witch hunt. In fact, if anything, there’s a comical detail here: Substack offers creators the option to run Pangram on their drafts before publishing. Since users can use the Pangram model before publishing a post, we’re essentially giving them a free tool to self-censor and edit their posts so they publish once they pass the detector.
As a result, we’ll see less and less AI-generated content, but not because we’ve won the battle against slop, but because we’ve provided a tool for it to disguise itself.
In the long run, this could be positive in a way. More higher-quality content, because even the sloppers will realize that Human-certified writing is higher-status.
However, we no longer have any way of knowing whether the person used AI and then edited it until it passed the detector.
At least before, we knew it; the slop was obviously slop. Will this awareness lead to a double dose of paranoia? How can I be sure the person didn’t just manipulate the detector?
So what will we ask for to prove true authenticity: the first drafts? My voice memos with my Italian-Roman accent? A lie detector test? I have no idea. But I’ve already heard these absolutely ridiculous proposals.
Trust has already been undermined, and if you’ve ever dealt with someone who suffers from obsessive-compulsive disorder, you’ll understand: there’s no right number of checks that will put that person’s mind at peace about whether they locked the car or not. So if you’re already paranoid about whether a writer is using AI or not, I’m afraid the Pangram certification will not help you.
Consequently, in the long run, no test will ever put our doubts to rest. The fact is, we’ll never be certain, and that was never the point, in my opinion.
When I realize a post is AI-generated, I simply click “Not Interesting” and get on with my life. At the end of the day, I read with interest from probably no more than ten writers, because that’s the mental bandwidth I can devote to this social network.
There is, however, one absolutely disastrous consequence that no one is talking about, at least, not on my feed, and there’s no turning back.
I mentioned already that I assume both Pangram and Substack have good intentions.
Not long ago, a real struggle for AI companies was model collapse. Since the internet had begun to swarm in polluted slop, models were being retrained on the corpus they produced, like a snake biting its tail, thereby worsening the models’ quality. I obviously don’t have the details at hand, but I imagine this was the reason Anthropic began hiring poets, writers, and editors.
But now? Now they no longer have to pay these experts!
Substack will sit on gold: an indexed, pre-labeled corpus of high-confidence human text. Labeling, which is the expensive part of the pipeline, is done for free here by Pangram.
Moreover, nothing prevents third-party companies from hiring a few employees at ten dollars an hour to copy and paste directly from Substack posts (assuming the anti-scraping mechanism works). Such companies already exist, like Mercor or Alignerr. They use humans to verify the quality of AI model output. Toward the end of 2023, it emerged that Scale AI was recruiting poets and writers in many languages to produce and evaluate literary training text.
We’re basically giving away high-quality, labeled text for free! GENIUSES!!!
Obviously, Substack isn’t going to sell our data directly right now, but imagine how the company will be valued now that it’s sitting on a goldmine. Imagine Dario Amodei, instead of opening Pornhub in the evening, will open some random Substack post labeled “100% human certified.” MMMM….
In any case, don’t claim victory just yet, because nothing prevents Substack from changing its terms in the future. I mean, years of dealing with Meta, Amazon, and, in general, all those deep tech or Silicon Valley-backed companies should have already taught us something about how they treat their user base and how easily they change the platform’s terms of use from one day to the next.
You think you’ll move away? Did we move away from Meta? From Amazon? A few percentage of us did, like the number of false positives in Pangram’s model.
‿̩͙⊱༒︎༻♱༺༒︎⊰‿̩͙

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.