An experiment with AI-assisted writing

As in David’s most recent post, there’s been a lot in the news about finding proofs and counterexamples with AI. Last weekend, I decided to try an experiment with writing using AI. I learned a lot, and wanted to quickly discuss the experiment and my thoughts on it here. Lots of people are certainly already doing this, but I haven’t seen many people talking about it.

The starting point is that Victor Ostrik and I started a project back in 2017, generalizing a result of Kuperberg about quantum G2, from generic q to q a root of unity. Namely, we showed that for q a root of unity outside of a specific finite list, the Karoubi completion of the G2 spider category is equivalent to the category of tilting modules of the Lusztig form of the quantum group G2. At some point during those 9 years, we did a little bit of writing, and at some point I gave a talk on it, but otherwise we did very little writing. This was not for mathematical reasons, but rather for executive function reasons on my end, the global pandemic, and both of us becoming directors of graduate study. This suggested an interesting challenge: could I use LLMs (specifically ChatGPT 5.6 Sol work mode mostly at “very high” intensity, via IU’s “Edu” subscription) to write this paper that was essentially mathematically complete, but almost entirely unwritten, and how quickly could this be done. To some extent this was a free experiment, because realistically I don’t think we’d have ever finished the paper at this point, and so it’s not replacing a bespoke paper that could have existed.

After spending a decent chunk of the time from Saturday until now on it, I now have a draft that I’m pretty happy with. I want to emphasize although mathematically this is Victor and my joint work, and although Victor has allowed me to make this post, he has not signed off on the accuracy and all errors at this point should be blamed entirely on me. Also my work is supported under NSF DMS grant 2000093 and Simons Foundation grant MPS-TSM-00007608.

Ok, here’s what I did:

  1. First, I asked if Sol could one-shot the main theorem. The answer was yes, though for a somewhat simple reason: Bodish-Wu write “It is possible to adapt the approach from [1], which itself is based on [7], to prove that the Karoubi envelope of [the G2 web category] is equivalent to the category of tilting modules as long as $[2], [3] \neq 0$.” That is to say, Elijah already proved the same result for C2, and a similar argument will work for G2. So the robot supplied the similar argument. I asked it to write that argument up, and then to check it over for good references and to read it like a referee would and make edits. This took around 30 minutes. Here’s the resulting file.
  2. Second, I uploaded my talk slides (and the tiny file already written, which was mostly useless), and asked Sol to give a proof of the main results following the slides. Again I asked it to edit it. This took around 30 minutes. Here’s the resulting file.
  3. Then I looked at the files. As mathematical exposition, I consider both to be garbage.
  4. Then I spent several days giving feedback attempting to improve the second file based on my talk. At no point did I edit the source directly. Most of this was in what I would call the style of a (low executive function, see above) PhD advisor. That is, I would kinda skim the file, get annoyed about something, and tell it to fix it. While it was fixing the paper, I would skim some more to try to find something else that annoyed me. This was a long process! It took three days, nearly 100 prompts, 10-15 hours of reasoning, plus another 10-15 hours of non-reasoning computer time. This used nearly an entire week of my generous budget, and Sol estimates that this would cost around $100 (within a factor of 2) at metered rates. Eventually I got to a version of the paper that I’m pretty happy with. Here’s the resulting file.

I thought I’d distill some thoughts and some questions from the process, I’m of course very curious for your thoughts on the matter.

Comments:

  1. This was much faster than I could have written the paper myself, though slower than I thought it would be. I think the final product is comparable in quality to a typical math paper of mine. On the other hand, I think that compared to my fastest writing collaborators it was not orders of magnitude faster, and the quality is not close to the output of the best mathematical expositors. AI at this point is much worse at writing paper than finding counterexamples to conjectures.
  2. In this case, I was not very worried about errors, because I already had thought through the whole argument and was highly confident that it would work (modulo getting the exactly correct list of exceptions). Nonetheless, I felt like Sol did not make errors more frequently (or of a worse character) than I would expect of myself or a collaborator. Most errors were stuff like “Oh, forgot to check whether this theorem actually works at all roots of unity.” This is typical of my experience with 5.6, which is dramatically better at doing math accurately than previous ChatGPT models.
  3. In this case the vast majority of the ideas were already present from Victor and my work. In particular, the goal was not just to write a proof, but to write our specific proof. Nonetheless, I do think the model contributed mathematically in one key way: in my original sketch I always worked over each q individually, and the model preferred to work integrally, and this resulted in some very nice simplifications in Section 4.1. If and when we turn this into a real preprint, I will include a brief discussion of the intellectual contribution from the model.
  4. I was surprised when I printed out and read a near-final draft, that this feels to me like a paper I wrote. That is the voice is not different enough from what I would write with a human collaborator to feel like it’s not in large part mine.
  5. The experience is disconcertingly similar to advising a PhD student on a paper. That said, a PhD student would need less handholding on their second paper, but an LLM won’t really learn.
  6. I was surprised about how important “prompt engineering” remains, and I think that if I were to write another paper this way I would be able to write it faster and better. The key points are that the model is lazy and easily distracted (both properties I find highly relatable!). It’s lazy in the sense that if you ask it to do a lot of work all at once it will take shortcuts and not do a good job. At one point I had to be like “no, go look at exactly how I made TikZ diagrams, now make all your diagrams actually good like that.” It’s easily distractible in that if you’re not clear about the scope of your question and the document is long, it will start spending crazy amounts of time doing who knows what. Like it wrote the whole first draft in 20 minutes, but then when the paper was 50 pages long, I asked it to switch the order of two paragraphs and it took an hour. Make clear requests and not too many requests at once. Form a plan first and then implement the plan. Be specific about whether it should be editing the document, and if so in which sections. For simple tasks, medium intensity is better than very high.
  7. Starting again from sketch, I’d try to follow Terry Tao’s advice for writing and start with an outline and gradually flesh it out, rather than trying to start with a one-shot paper and then editing.

Questions:

  1. To what extent is this final paper adding any value to the original talk? Especially considering that readers themselves could use an AI model to flesh out points in the talk that they didn’t understand? Maybe we should just be focusing on talk-length digests and formal checking, rather than traditional papers?
  2. What should we do with this paper? I don’t want to make someone hand-referee it, because it doesn’t seem fair when it wasn’t hand-written. Probably we will put it on the arxiv once we’ve human-checked it fully and Victor has signed off on it, so that other people can use the results if they need to.
  3. Given the speed-up, when does it still make sense for me to write papers by hand? (Relevant here that I’m a very slow writer and don’t really enjoy it, the way I enjoy say preparing and giving a talk.)
  4. What does this mean for PhD advising? Many PhD students need a similar amount of guidance to what I gave the model in this project. But you can now remove the student from the loop (either intentionally, with the advisor just writing using LLM assistance rather than having students, or unintentionally, with the student just feeding all the suggestions to an LLM and reporting back to the advisor).
  5. Have any of you done better with AI-assisted paper writing? My points 6 and 7 above sounds like something where someone is going to say “blah, blah, scaffolding, blah, blah, multi-agent…”

What a strange world to live in…

The new counterexample to the Jacobian conjecture

As many of you have probably heard already, yesterday morning, Levent Alpöge tweeted that Fable had found a counterexample to the Jacobian Conjecture. Specifically, let

a=(1+xy)3z+y2(1+xy)(4+3xy),b=y+3x(1+xy)2z+3xy2(4+3xy),c=2x3x2yx3z,\begin{align*} a&=&(1+xy)^3z+y^2(1+xy)(4+3xy),\\ b&=&y+3x(1+xy)^2z+3xy^2(4+3xy),\\ c&=&2x-3x^2y-x^3z, \end{align*}

Then the Jacobian of (a,b,c) is easily checked to be -2. However, the map (a,b,c) is generically three to one, not bijective.

I’m sure many of you are playing with these polynomials to see what you can figure out about them. This is a place for us to share our observations. I’ll post a few minor observations of my own soon.

Continue reading

Congress proposes cutting of all funding to US academics who mentor Chinese students

I’m writing to point out a potential law which should be gathering more opposition and attention in math academia: The Securing American Funding and Expertise from Adversarial Research Exploitation Act. This is an amendment to the 2026 National Defense Authorization Act which has passed the House and could be added to the final version of the bill during reconcilliation in the Senate. I’m pulling most of my information from an article in Science.

This act would ban any US scientist from receiving federal funding if they have, within the last five years, worked with anyone from China, Russia, Iran or North Korea, where “worked with” includes joint research, co-authorship on papers, or advising a foreign graduate student or postdoctoral fellow. As I said in my message to my senators, this is everyone. Every mathematician has advised Chinese graduate students or collaborated with Chinese mathematicians, because China is integrated into the academic world and is one fifth of the earth.

This obviously isn’t secret, since you can read about it in Science, but I am surprised that I haven’t heard more alarm. Obvious people to contact are your senators and your representatives. I would also suggest contacting members of the Senate armed services committee, who are in charge of reconciling the House and Senate versions of the bill.

Moonshine over the integers

I’d been meaning to write a plug for my paper A self-dual integral form for the moonshine module on this blog for almost 7 years, but never got around to it until now. It turns out that sometimes, if you wait long enough, someone else will do your work for you. In this case, I recently noticed that Lieven Le Bruyn wrote up a nice summary of the result in 2021. I thought I’d add a little history of my own interaction with the problem.

Continue reading

Using discord for online teaching

On Wednesday, I asked several of my students what tools they use to collaborate online on their problem sets. Several of them mentioned Discord. I am currently trying to set up a Discord channel for my class. I imagine I am not the only one in this situation, so I am writing up my progress as I go here . If you have relevant knowledge, please leave an answer to this question or edit mine!

If you have questions to discuss, let’s do that in the comment thread here. And please promote this on twitter and wherever else math teacher’s gather!

I did a dry run with 6 of my students this afternoon. We spent 10 minutes introducing each other to the system, broke into two groups of 3 and spent 10 minutes solving a math problem (problem 19.2 here) and 30 minutes debriefing. Most students found it awkward but workable. Here are things I/we saw as good:

  • Connection was reliable; much less dropping out than the video conferencing tools they report their other courses are using.
  • The presence of the chat stream with integrated graphics and LaTeX encouraged people to write things down, just like we encourage students to write on blackboards when doing group work in class. This made it easier for me to jump in and out of conversations.
  • It was pretty easy for me to jump back and forth from one group to the other. (Not sure how I would do with 4 or more groups, though.)
  • We really did solve a nontrivial problem and have a nontrivial conversation about it!

Things we didn’t like

  • Both groups needed, at one point, to draw an equation or commutative diagram like this one:
    \mathbb{Q}(\cos \tfrac{2 \pi}{7} ) \cong \mathbb{Q}[x]/f(x) \mathbb{Q}[x] \cong \mathbb{Q}(\cos \tfrac{4 \pi}{7}).
    One group did the LaTeX, the other draw a commutative diagram in MS Paint and dropped it into the thread. Both expressed frustration about how much slower this was than drawing a picture in person.
  • Some students felt that it was awkward not seeing the people they were talking to.

Read Izabella Laba on diversity statements

There has been a dispute running through mathematical twitter about diversity statements in academic hiring. Prompted by that, Izabella Laba has just written an excellent post, which affirms the importance of diversity as a goal, but lays out the many tricky issues with diversity statements. It makes a lot of points I would like to make, and raises others that I hadn’t thought of but should have. In case there is someone reading this blog but not reading Professor Laba’s, go check it out.

Sadly, this blog is still dead. I have a list of things I would like to write on it, but I have no idea when or if I will find the time.

Why single variable analysis is easier

Jadagul writes:

Got a draft of the course schedule for next year. Looks like I might get to teach real analysis.

I probably need someone to talk me out of trying to do everything in R^n.

A subsequent update indicates that the more standard alternative is teaching one variable analysis.

This is my second go around teaching rigorous multivariable analysis — key points are the multivariate chain rule, the inverse and implicit function theorems, Fubini’s theorem, the multivariate change of variables formula, the definition of manifolds, differential forms, Stokes’ theorem, the degree of a differentiable map and some preview of de Rham cohomology. I wouldn’t say I’m doing a great job, but at least I know why it’s hard to do. I haven’t taught single variable, but I have read over the day-to-day syllabus and problem sets of our most experienced instructor.

Here is the conceptual difference: It is quite doable to start with the axioms of an ordered field and build rigorously to the Fundamental Theorem of Calculus. Doing this gives students a real appreciation for the nontrivial power of mathematical reasoning. I don’t want to say that it is actually impossible to do the same for Stokes’ theorem (according to rumor, Brian Conrad did it), but I never manage — there comes a point where I start waving my hands and saying “Oh yes, and throw in a partition of unity” or “Yes, there is an inverse function theorem for maps between n-folds just like the one for maps between open subsets of \mathbb{R}^n.” I think most students probably benefit from seeing things done carefully for a term first.

Below the fold, a list of specific topics much harder in more than one variable. If you have found ways not to make them harder, please chime in in the comments!

Continue reading

A grad student fellowship

My CAREER grant includes funding for a thesis-writing fellowship for graduate students who have done extraordinary teaching and outreach during their time as a grad student.  If you know any such grad students who are planning on graduating in the 2018-2019 school year please encourage them to apply.  Deadline is Jan 31 and details here.

While I’m shamelessly plugging stuff for early career people, Dave Penneys, Julia Plavnik, and I are running an MRC this summer in Quantum Symmetry.  It’s aimed at people at -2 to +5 years from Ph.D. working in tensor categories, subfactors, topological phases of matter, and related fields, and the deadline is Feb 15.

 

Fighting the grad student tax

I’m throwing this post up quickly, because time is of the essence. I had hoped someone else would do the work. If they did, please link them in the comments.

As many of you know, the US House and Senate have passed revisions to the tax code. According to the House, but not the Senate draft, graduate tuition remissions are taxed as income. Thus, here at U Michigan, our graduate stipend is 19K and our tuition is 12K. If the House version takes effect, our students would be billed as if receiving 31K without getting a penny more to pay it with.

It is thus crucial which version of the bill goes forward. The first meeting of the reconciliation committee is TONIGHT, at 6 PM Eastern Time. Please contact your congress people. You can look up their contact information here. Even if they are clearly on the right side of this issue, they need to be able to report how many calls they have gotten about it when arguing with their colleagues. Remember — be polite, make it clear what you are asking for, and make it clear that you live and vote in their district. If you work at a large public university in their district, you may want to point out the effect this will have on that university.

I’ll try to look up information about congress people who are specifically on the committee or otherwise particularly vulnerable. Jordan Ellenberg wrote a “friends only” facebook post relevant to this, which I encourage him to repost publicly, on his blog or in the comments here.

UPDATE According to the Wall Street Journal, the grad tax is out. Ordinarily, I thank my congress people when they’ve been on the right side of an issue and won. (Congress people are human, and appreciate thanks too!) In this case, I believe the negotiations happened largely in secret, so I’m not sure who deserves thanks. If anyone knows, feel free to post below.

… and Elsevier taketh away.

Readers may recall that during the 2013 “peak-Elsevier” period, Elsevier made an interesting concession to the mathematical community — they released all their old mathematical content (“old” here means a rolling 4-year embargo) under a fairly permissive licence.

Unfortunately, sometime in the intervening period they have quietly withdrawn some of the rights they gave to that content. In particular, they no longer give the right to redistribute on non-commercial terms. Of course, the 2013 licence is no longer available on their website, but thankfully David Roberts saved a copy at https://plus.google.com/u/0/+DavidRoberts/posts/asYgXTq9Y2r. The critical sentence there is

“Users may access, download, copy, display, redistribute, adapt, translate, text mine and data mine the articles provided that: …”

The new licence, at https://www.elsevier.com/about/our-business/policies/open-access-licenses/elsevier-user-license now reads

“Users may access, download, copy, translate, text and data mine (but may not redistribute, display or adapt) the articles for non-commercial purposes provided that users: …”

I think this is pretty upsetting. The big publishers hold the copyright on our collective cultural heritage, and they can deny us access to the mathematical literature at a whim. The promise that we could redistribute on a non-commercial basis was a guarantee that we could preserve the literature. If this is to be taken away, I hope that mathematicians will go to war again.

Hopefully Elsevier will soon come out with a “oops, this was a mistake, those lawyers, you know?” but this will only happen if we get on their case.

What to do:

  • Elsevier journal editors: please contact your Elsevier representations, and ask that the licence for the open archives be restored to what it was, to assure the mathematical community that we have ongoing access to the old literature.
  • Elsevier referees and authors: please contact your journal editors, to ask them to contact Elsevier. If you are currently refereeing or submitting, please bring up this issue directly.
  • Everyone: contact Elsevier, either by email or social media (twitter facebook google+).
  • Happily, as we have a copy of the 2013 licence, all the Elsevier open mathematics archive up to 2009 is still available for non-commercial redistribution under their terms. You can find these at https://tqft.net/misc/elsevier-oa/.