RSS Amplifier

Human•ities · Jan 26, 2026

I Required My Students to Use AI for Every Assignment in my History Class. Here’s What We Learned.

0
Sign in to vote or save

Pierce Salguero · Human•ities

“AI thinks with us, not for us.”

For 20 years, I’ve taught history students how to write research papers the traditional way: find sources in the library, cite them properly, read closely, develop an original argument, outline, draft, and revise. Last fall, on the first day of my upper-level seminar on Asian medical history, I abandoned that approach entirely.

“You must use AI for every stage of this assignment,” I told them. “And I have no idea how good your paper will be in the end.”

The premise was simple but radical. Students wouldn’t be graded on the quality of their papers. Instead, we would collectively grade the latest cutting-edge generative AI platforms on how well they could do the job. Could these tools actually perform scholarly research? Where would they excel? Where would they fail? My students and I would find out together.

In most classes, using AI to do every step of a research paper would be academic dishonesty. But I wanted to understand what we’re up against as both as educators and as students who will be graduating into an AI-saturated workplace. It seemed to me that first-hand experimentation was the only way to really find out.

Our class mantra was “AI thinks with us, not for us.” But learning exactly how best to collaborate with AI required a lot of trial and error.

The optimists in my class included Joe, a 37-year-old returning student. (All names have been changed.) He treated AI as a research assistant. “It helped me expand my ideas and put words on paper in a way I know I would not have been able to,” he said in our final session. Another student, Lori emphasized accessibility: “As someone with disabilities, not having to type everything out saved me a lot of physical pain.” For students who struggle with writing for whatever reason, AI seemed like an exciting equalizer of opportunities.

Other students were less sanguine. Yasmine put it succinctly when she said that she felt that working with an LLM was like doing a group project where the other student couldn’t be trusted to do their work properly. She quickly tired of constantly having to clean up the messes and sometimes wished she had just done the work herself.

Students quickly learned that different AI platforms were best for different tasks. The class experimented with multiple LLMs, each revealing distinct strengths and limitations. Three emerged as frontrunners:

  1. Notebook LM was praised as a general-purpose workhorse. The platform could answer questions about uploaded documents, create summaries and reports, draw mind maps, or even generate podcasts to help students understand the content. The critical limitation: it couldn’t reliably generate paper drafts, only outlines and generalizations.

  2. Claude became the preferred editor and writer, particularly appreciated for its ability to turn NotebookLM’s outlines into prose. Students found Claude more willing to admit errors and more capable of matching a specified writing voice—though it still tended to be overly wordy and required extensive oversight.

  3. ChatGPT was popular as a generator of “second opinions.” For most students, this model felt the warmest but it was less capable of doing serious academic work. It was primarily useful as an evaluator and cross-checker, providing feedback and alternative perspectives on work generated by other platforms.

Students developed highly individualized workflows, but common patterns emerged:

  • Source gathering was in the end done by most students manually. Although they first attempted to do so with AI, students quickly learned that reference lists generated by the platforms were filled with non-academic works of dubious quality or, worse, were altogether hallucinated. The most successful strategy was to search the university library manually, then upload PDFs to Notebook LM for analysis.

  • Role-playing and personality injection proved essential. Students instructed AI to adopt specific personas (”you are a mental health researcher,” “you are an excellent undergraduate writer who majors in history”) and adjusted for specific sophistication levels (“rewrite this for a nonspecialist”). They learned to improve responses by providing extensive context, pasting assignment prompts and grading rubrics from my syllabus, as well as examples of their own writing, into the LLM’s chat interface.

  • Platform hopping became standard practice. Some students used Notebook LM for outlining, then Claude for drafting. Others used NotebookLM for first drafts, then used Claude for editing. Still others generated drafts using two different platforms AI, and then a third to combine the best of both. Cross-platform evaluation provided a modicum of quality control. Students fed papers written with one platform into another for critique, used different AIs to assess the same work, and constantly asked for counterarguments and identification of weaknesses.

On the whole, most students found the AI platforms to be wildly insufficient for actual scholarship. Here are some things they discovered through painful trial and error:

  1. AI cannot handle citations. The consensus was that, at present, AI still has trouble handling sources. Even when the sources were manually acquired and uploaded to the platform, every student reported that AI still fabricated page numbers, scrambled formatting, quoted meaningless sentence fragments, or simply omitted citations entirely. “I would ask for quotes with page numbers and proper citations,” Yasmine reported, “but every single citation required manual verification against the original source.”

  2. AI is less effective when dealing with specialized knowledge. Many students chose to focus on niche topics such as medicine traveling along the Silk Road, female Asian American healers in the 19th century, or trepanation in religious versus medical contexts. While some scholarly work exists, these topics are not well-represented in the AI training data. Kaj, who wanted to write about Islamic medical texts translated in East Asia, found AI useless: “I already knew about the traditional medical books from another class, but AI didn’t. It was too niche for the platform to understand.”

  3. AI mangles the details. Numbers and lists failed repeatedly. Foreign characters caused particular confusion. When paraphrasing sources, AI routinely dropped crucial qualifiers. A statement about events occurring “during the height of the seventh century” became a general claim about all periods. Conditional statements became universal. Every passage generated by AI required verification to ensure meaning hadn’t been subverted, altered, or fabricated altogether.

  4. AI has trouble innovating. “Any new idea I tried to bring into my paper was soon filtered out,” Kaj explained. Another student, Martin, spent the entire semester fighting to focus his paper on his intended topic, while AI kept pulling him back toward a less interesting argument it seemed to insist he make instead.

The paradox we discovered throughout the semester was that students need more sophisticated scholarly skills to use AI effectively, not fewer.

Yasmine, a history major, put it plainly: “If you didn’t already have the academic fundamentals from other classes, it would be difficult.” Students who succeeded in the project drew on their own source evaluation skills to catch hallucinations, their own domain knowledge to recognize factual errors, their own writing competency to judge whether AI output met scholarly standards, and their own critical reading ability to detect when arguments lacked support.

AI didn’t eliminate the need for these competencies; it made them more critical than ever. Students who couldn’t write well struggled to edit AI effectively. Students unfamiliar with their subjects couldn’t catch AI’s missteps.

This raises a troubling question: if students need robust skills to use AI properly, but they develop those skills through the very practices AI promises to automate, how do we teach the next generation of students?

I’m still deciding what I will do when next year rolls around and I am scheduled to teach this course again. The experiment succeeded in its primary goal: students developed genuine expertise in AI’s capabilities and limitations, gained through their own direct experience. They can now critically evaluate the value of AI with confidence because they’ve wrestled with these tools extensively.

On the other hand, I’m troubled by what my students didn’t learn. Skills that previous cohorts of this class developed through traditional research—deep reading, careful synthesis, the patience required to sit with complex ideas—were largely bypassed this semester. “I didn’t read as much because we were using AI to summarize everything,” Joe admitted, “so I learned less about the topic.” Multiple students echoed this sentiment. The very efficiency AI promised came at a significant cost when it came to learning.

The papers were fine. Solid Bs all around. But when AI did the work, something essential was lost, even if the final product was good enough. As Kaj put it: “AI helps you gain some knowledge, but part of college is learning the skill of working with concepts bigger than yourself.” These students became sophisticated AI operators, and I believe they became better critical thinkers. Whether they became better writers or researchers, or whether they expanded their intellectual horizons, remains unclear.

As our final session concluded, Joe offered an ironic reflection: “For me, this class on AI highlighted how much the human matters.”

He’s right. After all, the humanities has never been primarily about producing documents. It’s about developing judgment, critical thinking, intellectual courage—the human capacities to engage deeply with ideas and form independent conclusions.

AI can assist with many tasks. But if we learned one thing from the semester’s experiment, it is that the technology cannot replace the fundamentally human work of learning to think.

No posts

Read the original on piercesalguero.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.