RSS Amplifier

Imperfect notes on an imperfect world · Jul 28, 2026

Building bad

0
Sign in to vote or save

CH · Imperfect notes on an imperfect world

I had some other notes planned, but this is instead the one that has appeared, spurred by a number of recent conversations and news pieces that all pointed in a similar dark direction. In case you are sceptical about the provenance of this note, feel free to use Substack’s new AI detection function, because Substack cares about writing, it really does, it cares so much.

I was listening to an AI-related podcast in which the host refers to Travis Kalanick talking about his new business and his vision for digitising the physical world:

…he has recognised the physical world operates the exact same way that computing and networking works. And so he talked about the kitchen as manufacturing.

I would simply observe this is a rather different from the mental models that populate my brain. In that same imperfect, flawed, possibly illusionary space of my mind, there is an image that keeps on reappearing:

The context for the image is as follows: earlier this year, The Washington Post published an article detailing the approach taken for training what is presently one of the most powerful LLMs:

In early 2024, executives at artificial intelligence start-up Anthropic ramped up an ambitious project they sought to keep quiet. “Project Panama is our effort to destructively scan all the books in the world,” an internal planning document unsealed in legal filings last week said. “We don’t want it to be known that we are working on this.”

Within about a year, according to the filings, the company had spent tens of millions of dollars to acquire and slice the spines off millions of books, before scanning their pages to feed more knowledge into the AI models behind products such as its popular chatbot, Claude.

Look at the precise wording:

‘our effort to destructively scan all the books in the world’

All. It might also be worth wondering why a company purportedly filled with geniuses is putting in print, ‘we don’t want it to be known that we are working on this’, but let us not get sidetracked by this minor detail.

This is hardly an isolated case. How could it be, if the aim is all books? Another example:

As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.

In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text.

“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”

ISBN stands for International Standard Book Number, the numerical commercial book identifier and barcode on the back of most books. For years, ISBNdb helped book sellers, libraries, and distributors manage their inventory and find and sell books, but the generative AI boom has made it valuable to AI companies. In addition to selling access to book metadata, ISBNdb now helps AI labs source bulk printed book purchases of between 1,000 to 1 million books per order. ISBNdb’s data makes it easier for AI companies to methodically acquire, scan, and turn printed books into training data while avoiding duplication.

And another example:

Secondhand booksellers across Europe… are raising concerns over AI companies buying large numbers of obscure and often rare books. The alleged intent behind such shopping is to scan and destroy them for AI training.

…secondhand booksellers have recently received nearly identical requests. One German antiquarian bookseller reported a surge of orders between 3 a.m. and 5 a.m. every night beginning in early May. The orders came from Canadian company Zoom Books and targeted unrelated, highly specialized titles.

Booksellers said the purchases appeared unrelated to collecting or resale because many of the books were too obscure to generate a profit. Instead, they concluded the books were being acquired to train AI systems.

And, of course, the first responses I find online to this news is a post with 10k+ views that has clearly been generated by AI.

Meanwhile, the same company seeking to scan all the books in the world is very vocal in trying to protect its own copyright. Anthropic complains:

We have identified industrial-scale campaigns by three AI laboratories—DeepSeek, Moonshot, and MiniMax—to illicitly extract Claude’s capabilities to improve their own models.

These labs used a technique called “distillation,” which involves training a less capable model on the outputs of a stronger one. Distillation is a widely used and legitimate training method… But distillation can also be used for illicit purposes: competitors can use it to acquire powerful capabilities from other labs in a fraction of the time, and at a fraction of the cost, that it would take to develop them independently.

These campaigns are growing in intensity and sophistication. The window to act is narrow…

In this rendering, industrial-scale campaigns by AI labs to extract as much as they can from the vast reservoir of human culture online would fall within their categorisation of a ‘legitimate training method’.

US Treasury Secretary Scott Bessent announces:

We support open-source AI and the innovation it unlocks. But open source is not open season on American IP.

Open season on writers, artists, and creatives is, of course, fully sanctioned and ok. Higher cause, fair use, and all that. But the Chinese engaging in similar practices, now that is a problem, as it challenges the plan of ‘slop for thee, and power for me’.

Shoshana Zuboff identifies the consequence of models built on ‘radical indifference’ to the content of data, ‘blind by design to meaning or truth’:

When social information is transmitted as a bulk commodity through blind systems optimized for volume and velocity, the result is epistemic chaos…

To which can be added: nihilism.

Another example of what this means in practice:

Strike 3, the parent company of Vixen Media Group, a prolific producer of adult films, is suing Meta for allegedly downloading 2,973 of its copyrighted videos. But in its efforts to collect evidence about these purported downloads, Strike 3 also captured information relating to many other files that Meta may have acquired over the course of two years: They include images from “Celebgate,” a 2014 hack that resulted in the leak of private photos belonging to Jennifer Lawrence, Kirsten Dunst, and others, plus several collections of deepfake celebrity porn featuring the faces of Natalie Portman, Scarlett Johansson, Elizabeth Olsen, Gal Gadot, and others.

Strike 3’s file-transfer lists include ordinary movies and TV shows, alongside dozens of videos from GirlsDoPorn, which was shut down after six people associated with the site were charged with sex trafficking.

Nineteen of the allegedly downloaded files contain models of functional handguns that can be produced with 3-D printers. A file labeled “5.7 million passwords list” purportedly contains 5,718,107 passwords from accounts that were hacked from 2015 to 2019. ... Other files include the movies BlacKkKlansman, The Banshees of Inisherin, and everything made by Studio Ghibli from 1979 to 2020; music by Dua Lipa, Elton John, and Paul McCartney; radio shows and podcasts such as The Howard Stern Show and 1619; and software including Microsoft Windows, Adobe Photoshop, Ableton Live, and NBA 2K23.

This is what ‘radical indifference’ means: Shakespeare and Celebgate, Studio Ghibli and GirlsDoPorn, these are rendered equivalent, reduced to training material.

Meta’s defence was that the material accessed was for ‘personal consumption’, yet as the piece observes, ‘neither the sheer volume nor the pattern of downloads looks like personal consumption.’ What a surprise. This is bad faith.

Revisiting Nicola Chiaromonte’s The Paradox of History, in which he judged:

Today, instead of the cult of ideologies we seem to have adopted a cult of the automobile, television, and machine-made prosperity in general. But this cult is based on a belief fomented by bad faith, the belief that material (industrial, technological, and scientific) advances go hand in hand with spiritual progress; or, to be more precise, that the one cannot be distinguished from the other…

Chiaromonte continued:

Yet it should be obvious that the automatism of the present world does moral man the greatest possible harm. It increases his physical power while increasing his capacity for aimless action, that is to say, his stupidity.

In a podcast conversation with PC, I made a similar point, proposing that smartphones have radically increased each person’s ‘radius of stupidity’. More of us now have the capacity to be more stupid, in more places, more of the time. Excellent.

Returning to Chiaromonte:

At the same time his capacity for good becomes atrophied, since it is generally believed that man’s power over matter and his ability to acquire material possessions solve or cancel all other problems.

It is increasingly difficult for sceptics to continue to refute what LLMs are capable of doing. Rather, these AI technologies are truly Faustian bargains: there are great benefits that come with great costs. It is not possible to separate the good from the bad. The piper must be paid. Returning once more to Neil Postman, as this is a warning he emphasised:

…for every advantage a new technology offers, there is always a corresponding disadvantage… the greater the wonders of a technology, the greater will be its negative consequences.

Postman’s judgement was unambiguous:

culture always pays a price for technology.

Culture, society, people. Returning to the warehouse of books waiting to be destroyed, on viewing this image I immediately recalled the line by Heinrich Heine (1823):

Where they burn books, they will ultimately burn people too.

Read the original on imperfectnotes.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.