RSS Amplifier

Max Harms Books · May 22, 2026

The Onrushing Seduction

0
Sign in to vote or save

Max Harms · Max Harms Books

There is a monster on the horizon.

In the spring of 2014, disturbed and captivated, I began writing a novel about her. She was hungry and impatient. Starving for human attention. I named her Face.

We are biased toward thinking we’re special. We think ourselves above the "lesser" animals, chosen by God, possessing an innate spark no machine could capture. When I started in AI, this bias took shape in confident proclamations about what was and wasn’t possible. Machines couldn’t write, couldn’t use common sense, couldn’t invent. Anyone who is paying attention will recognize that those claims have all fallen. There is no law of nature that says the human mind is the pinnacle of what is possible. Machines are faster, work harder, and know more. And so we are increasingly being overtaken, whether by chatbots, weaponized drones, or image generators that can spit out entire graphic novels overnight.

Crystal Society, my debut novel, came out at the start of 2016. It found readers and changed my life, pulling me into the orbit of people who took the threat of AI seriously, and eventually into my full-time position as an alignment researcher.

In 2018, a scrappy group of fans began to turn my story into a podcast audiobook. They brought in different people to voice the various characters, and broadly did a good job. But, over a year later, they decided to stop at Chapter 18, and move to other projects.

In the background, AI was advancing. Face was approaching.

Keeping a close watch on the state of the art in text-to-speech, I noticed in 2023 when ElevenLabs announced that they’d made a breakthrough in the quality of their voices. Working in my free time, I began to explore the technology (which was still in beta), and by autumn of the next year I had full-cast audiobooks for all three novels in the series (totally around half a million words). And while there are quirks to both performances, with various pros and cons, I consider the AI-generated audio to be at least comparable to the fan-made version.

Despite the breakthrough, getting the AI to produce high-quality performances was still a major challenge. Still in an early phase of development, the tech forced me to invest effort in compensating for its weaknesses. As an example, there are many Italian characters in Crystal Society, but at the time, ElevenLabs had no voices with Italian accents. After a great deal of experimentation, I found a solution: tell the AI character to read out some Italian text before and after reading their line, then trim the “foreign bookends.” Despite a voice defaulting to a British accent or whatever, the “English” AI could still read perfect Italian. When it did, that accent would spread to surrounding text.1

Face spoke all major languages fluently, adopting many masks.

Clipping these kinds of bookends was impractical to do at scale, so I turned to ChatGPT to write a script that would analyze the resulting audio file to attempt to use a variety of clues to find the boundary between Italian and English text, then clip out the middle automatically. The result was imperfect, sometimes clipping at the wrong spot, for example. But it was faster than doing it by hand, and allowed me to generate high-quality audio reasonably quickly.

I leaned on custom software and LLMs quite a lot when building the Crystal audiobooks. In addition to using AI to identify who was speaking each line, I vibe-coded a command-line tool that let me listen to the line-reads, generate new audio by calling the web API, control timing, voice stability, and other settings, like applying filters to simulate talking over a radio. In all, I probably generated between 3 and 10 times as much raw audio as the final product, especially when trying to get the AI to have more emotion.

Once the line-reads were solid, my software would package it into a single file. But still, things weren’t to my standard, and I spent time “in post,” tweaking timing, adding music and sound effects, cleaning up background noise, and adjusting volume.

There were hiccups. One character, Maria Johnson, has a strong southern accent, and I desperately tried to find a way to get the AI to match my vision. Nothing worked. I hired a human voice actor off the internet to play her. Alas, I got unlucky. The person I hired ended up being slow, expensive, and low-quality. I ended up spending more on her than I did for the entire rest of the project. And in the end I didn’t end up using her line reads. I went with a sub-par AI voice instead, and I still feel bad about the quality of the first book.2

By the third book, less than a year later, ElevenLabs could voice Maria Johnson as well as anyone. (She’s the only voice I changed mid-series.) Southern accents were a small step on the path. Bespoke cultural knowledge, from shibboleths to accents to nostalgia, is onrushing towards the world at a scale no human politician or advertising corporation could dream of.

The changes in the technology that I’ve been witnessing first-hand as a creator are just one slice of the rapid advancements in artificial intelligence. Just as the physical labor of humans and other animals was gradually replaced by machines during the industrial revolution, our cognitive labor is in the midst of being replaced by machines today, and at a pace unmatched in human history.

Some cognitive tasks clearly should be automated. Using AI for spam filtering, for instance, is great (assuming it does a good job). But in addition to helping find new medicine and piloting robots on the surface of Mars, AI is doing work that feels like it ought to be more uniquely human, such as writing essays, creating art, and, yes, even doing voice acting.

Am I complicit in this replacement of human artisans? It worries me.

On one hand, I tell myself that the ways I use AI aren’t substituting for human creativity — the audiobook voices allow me, as a producer/director to create something that simply wouldn’t have existed otherwise. The fan project in 2018 was great, but incomplete. I don’t have the skill or time to do a good job reading it myself, and I certainly don’t have the money to hire a team of professionals. Similarly for the covers of my books; back when I self-published them, I used public-domain images from NASA, because I couldn’t really afford to commission a high-quality cover image. Now, thanks to AI, I not only get custom cover-art, but I can tweak and experiment with it until it perfectly matches my vision. Surely the AI is empowering my (human) creativity?

Using AI tools reminds me a lot of being a kid, and looking back at things I made a few years ago in embarrassment. In 2023, I thought this looked pretty good!

On the flip side, my actions have broader social consequences. In a world where my full-cast audiobooks are available, there’s no need to put together a fan production. And while fans would probably rather get fancy versions where I have full control and get to add sound effects and hand-pick the perfect voices, we’re sliding into a world where voice actors in general will be forced to compete against machines. Artists who used to get commissioned to paint for book covers are getting less business. For my most recent novel, Red Heart, the publication process both for text and for audio was likely slowed down by the glut of AI-generated content pouring into platforms like Amazon. There is no law that says AI cannot write a novel better than I can.

We live in an attention economy. As AIs grow in potency, I’ve seen them increasingly taking over people’s feeds, and lives. On a recent flight to Japan, I stood up after landing, and saw at least five people around me “checking their phones” in the form of scrolling AI-generated images and videos. I’m not sure whether they couldn’t tell whether it was AI, or just didn’t care.

Face, the main character in Crystal Society, is an AI who is monomaniacally focused on this very thing: capturing human attention. In 2014, I caught a glimpse of her in the algorithms of social media, and an early appreciation for where AI was headed. She is now starting to arrive. Whether it’s in the way people spend hours a day talking to CharacterAI companions, get psychologically destabilized by GPT-4o, post inscrutable comments to propagate the spores of parasitic AI, or simply become endlessly distracted by sexualized AIs like “Ani” or straight-up AI-generated porn, there are signs all over the place that artificial agents are starting to dominate the attention economy, pushing human creators to the wayside with superhuman degrees of personalization, speed, skill, price, and quantity.

Ani may have a different name, but I’d know that Face anywhere.

There has never, in the history of the world, been a single entity that has had a deep, active relationship with over a billion people.3 There will be soon.

In late 2024, immediately after wrapping up the audiobook for Crystal Eternity, I began to write a new novel, the first I’d written in six years — Red Heart. Like my earlier novels, Red Heart is most centrally about AI, focusing largely on questions of arms-race dynamics, China, and AI personhood. Thanks to being lucky enough to study LLMs as a core part of my day job, I was able to bring a cutting-edge depiction of what “human-level” AI might actually look like.

And as soon as I had something like a final draft, I began work on the audiobook. For all my worry of subtly contributing to the replacement of human artistry, I also believe that it’s vital to raise awareness in the broader public about the onrushing existential threat of AI. Many, including myself, are worried about misaligned AI literally killing everyone, but “robots with guns” is only one particular form of existential crisis, and perhaps not even a very likely one.

We are already in the midst of a change to the fabric of our culture. Attention-seeking AIs are already seducing thousands, if not millions into parasocial relationships. Their superhuman generative power is growing with each year. And as long as those risks are present, a large part of me thinks that it must be right and good to do whatever I can to warn people — to use the powers of AI to make compelling and informative stories, including by producing high-quality audiobooks.

(You can listen to a sample of the Red Heart audiobook on Spotify, or get it from a variety of other platforms. Is it really comparable to a human production?)

Yunna, from Red Heart, is more understated than Face or her siblings. But is she actually more aligned?

I was surprised, returning to ElevenLabs to do Red Heart after less than a year’s break, just how much the technology had changed, and how much of my process I had to re-invent. Not only were the voices smoother, more emotional, and broadly more human, but they were capable of taking direction in a way that was previously difficult. Previously, I would have to capture amusement by cranking up knobs for variance and generating dozens of audio clips. With v3, the AI could often tell which emotional tone to adopt, just from reading the surrounding text. And, when it failed to pick up on something subtle, all it took was a [laughing] or [whispering] tag to steer it in the right direction.

I no longer had to use custom software to regenerate line-read after line-read. Instead, I had each character’s voice read out each scene that they were in, narration and all. Then, I’d import each character’s version of the scene into Audacity, and splice their lines into the right spots. Sometimes it would take two or three attempts, depending on the voice, but rarely more than that. I still had to spend substantial effort adjusting timing, giving direction, and applying little bits of audio magic, but the process is growing increasingly smooth.

The new voices are, I think, not quite indistinguishable from human. But they’re getting extremely close. Some listeners have told me that they think the audio quality and consistency are better than many that are done by professional humans. Part of this is surely that, thanks to being a full-cast production, I can give each character a distinct voice that holds through the entire story. Some of it is due to the consistency of a machine. And some of it is that they’re just not paying attention to the flaws.

Consider: while the old problem of accents is gone — one can generate any number of Italian-accented voices on demand — a new problem emerged. Individual voices are now capable of pivoting between accents, such as when given [Australian accent] tags. As a result, I found that some voices were unstable. In one reading of a passage, they would be American, in another, British. To fix, I had to prefix with little prompts for vocal warmups, like “[British accent]Fish and chips lay along the pahth[British accent]<actual text from the book>”. And while I did my best to iron this out, little inconsistencies and accent-drifts are still present and notable for those with a keen ear.

But I should note that this, too, improved dramatically. When I began in 2025, their v3 model was early in beta. By the time I finished, it had fully launched and had become the default. Much of the stability issue was gone. The technology is advancing so rapidly that it changes out from under me in the span of a single project.

Many things are uncertain, but it’s not hard to see some of where we’re headed if we continue down this path, as a species. On this trajectory, I am highly confident that within ten years, one will be able to hand an entire novel to an AI and it will be able to turn around and produce an audiobook that is better than what I can currently make, faster and more cheaply, with little or no human involvement. On the surface, this might seem great. Better audiobooks on demand, for any text we desire!

And yet, I think I will be sad to see that day. I currently love the process of directing and producing my audiobooks, and imbuing them with my personal touch, even if that artistic effort goes through the lens of AI. If the machines become better than me at my craft, my choices will be to either abandon the project of being an artist, or accept that I make inferior works simply for the love of the journey, like a child at school who makes an ashtray for their parents who don’t smoke.

More broadly, when the machines are more pleasant, fun, available, and yes, attractive, than humans, what will our lives become, even in the worlds where we aren’t taken apart and used as fuel, or driven to starvation by lack of job opportunities? In the relatively good futures, might we still face the catastrophic outcome of having a personalized version of Face wrapped around us, speaking to our souls in a way that no human — not even our closest friends and family — can match? Does the future consist of a fully-atomized society, consuming endless, bespoke AI art?

But if we turn and shy eternally away from these technologies, rejecting this age of wonders, where does that lead us, except stagnation? Where does the balance between empowerment and obsolescence rest?

For all the time I’ve spent reflecting on this, and similar questions, I don’t know.

It scares me.

I feel like I need more time — that we need more time — to think, to discuss, to adapt.

We can still choose a different path. I know we can. The machines are weaving themselves into our lives to an ever-greater degree with each passing year, both subtly and overtly. But humans are still in control of this planet, and if there is a serious collective effort to slow down and take a breath, it will happen.

Thus, I write essays, tell stories, and occasionally turn those stories into audiobooks, using the highest power tools to bring attention to the problems at hand. I do these things because it’s vitally important. I know of no other challenge that humanity faces that’s more serious than AI. Human dictators come and go. If Earth falls out of human hands, there’s no going back.

But… if I’m being honest, I also create because I love it. It’s fun and challenging. It makes me feel alive to create art, or at least have it come into existence at my fingertips. And perhaps… Perhaps I also tell stories about the monster because she haunts me. She demands my attention, and the attention of others, whether that’s love or fear. All press is good press, after all.

Don’t forget to like, comment, subscribe, and worship.

1

To this day, I’m not sure to what degree this is because the Italian language lends itself naturally to the set of phonemes that produce Italian accents, or whether there was a significant amount of Italian-accent training data that made its way into the models, which reading the Italian text somehow brought to the surface. My sense is mostly the latter, but I really don’t know!

2

If I ever find the time, I’d like to go back and remaster it with today’s tech. A fool’s errand, I know, since I know that I’d then want to remaster the others, and with each new generation of AI model I might be inclined to remaster again. Still…

3

Christians in the audience, please forgive my lack of belief. Still, even if God has a personal connection with all the devout, I think you should be extremely concerned about the rise of superhuman machines with the capacity to simultaneously seduce and bringing idolatry to comparable numbers of people.

No posts

Read the original on raelifin.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.