A recent article by Marit MacArthur caught my eye when it called out the terminology we use to describe some of the most important aspects of generative AI: prompt engineering and training data. I have written here that “prompt engineering is just doublespeak for question,” but MacArthur points out that what we’re really saying when we talk about the need to learn good prompt engineering is that AI users need to be good at writing, and that what students need to be studying is not “engineering” in computer science, but writing in the humanities.
It may be too late to move away from terminology like prompt engineering, but we need to keep reminding people that the skills needed to use AI are skills that have been taught in the humanities for generations. Namely, the ability to ask good questions and to write well. Which brings us to the other term that MacArthur critiques: training data.
[LLM] “training data”[is] a misleading euphemism for a vast trove of expert human writing.
(MacArthur, 2025)
This is such an important point. Indeed, calling it training data makes it sound like someone, mythical AI trainers, created data that was used to train LLMs, when in fact, we know that this data was not “created” for LLM training, but farmed from existing texts written by humans (not trainers) for other purposes (not training).
What’s deemphasized by tech companies who want to keep “training data” free is that LLMs and all future AI tools and platforms will necessarily continue to rely on new human expertise, captured in writing or deployed as critical reading and editing skills, in every field of human endeavor... Referring to human expertise captured in writing as “training data” obscures this fact.
(MacArthur, 2025)
Say it louder for the people in the back: prompt engineering is writing, and training data is writing that was written by humans. If we want to prevent AI slop from taking over, we need to make sure that humans are teaching other humans how to ask good questions, how to write well, and how to evaluate writing (their own and others) so they won’t be fooled by AI output that merely looks legit.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.