RSS Amplifier

Recent Questions - MathOverflow Meta · Mar 11, 2026

Possible updates to our AI generated content policy

0
Sign in to vote or save

MathOverflow Meta

45

$\begingroup$

It is becoming increasingly clear that artificial intelligence tools can be useful for math research; see for example recent projects led by mathematicians like the First Proof project. As the capabilities of AI evolve, we are considering revisions to MathOverflow's official policy on AI generated content.

While the current official policy does not outright say that AI generated content is banned, it comes close to this, stating in its first sentence: "If the mathematical component of your content is deemed to be generated by AI, it will likely be deleted, along with any reputation earned from it."

We are considering two, interrelated, possible policy updates here.

The first proposed update would be to say that it is sometimes okay to post an answer to a question where an AI tool was used to help with the mathematical ideas, as long as:

  • you have carefully checked that the mathematics is correct and affirmatively "vouch for it";
  • you write everything in your own words (i.e., don't just copy the output of the LLM);
  • you declare that and how you used the AI tool.

The second proposed update would be that, in addition to the above, we require that anyone who does use AI in this way:

  • is not anonymous,

i.e., posts under their real name as an established mathematician, and links to their professional website in their profile or does something similar to establish their "belonging" to the global math research community.

The idea behind the first proposal is that, if AI tools can be and are being used ethically in current math research, then they should be permissible on MathOverflow as well, with the appropriate guardrails. The idea behind the second proposal is that an anonymous user vouching for some AI output counts for very little, and MathOverflow is meant to be a question-and-answer site for experts, where the value is in the human expertise.

For example, if both of these proposals were adopted, then some recent answers by Mark Wildon would be totally acceptable. (I am not trying to single anyone out here, just giving some examples.)

Question: What does the MO community think about these possible updates to our AI policy?

V2Blast's user avatar

asked Mar 11 at 14:06

Sam Hopkins's user avatar

$\endgroup$

33

39

$\begingroup$

There is an almost unfathomable amount of money sloshing through the AI industry right now. The various AI companies have an enormous financial incentive to generate the perception among the general public that their tools can automate large chunks of the economy. One way they do this is by harnessing scientists and loudly publicizing/exaggerating every possible "achievement" of their models applied to scientific research.

Academic mathematicians are not used to being players in the financial-political nexus. I think we're often naive about how things work, and this makes us prime candidates for manipulation. I therefore think that in addition to the rules you propose, we should additionally require anyone posting answers using LLMs to disclose any financial relationship they have with AI firms (employment, consulting fees, grants/reimbursements, etc). Of course, I'm not asking that they disclose personal financial details, merely the existence of the relationships.

This should be a familiar process to anyone who works in academia and deals with grants. I have to fill out an annual financial disclosure form for my university, and there are additional financial disclosure forms whenever I apply for grants.

Of course, this is not easily enforceable, and depends on the good will of the people involved. Anyone who wants to lie could easily do so, though I think that academic mathematicians are generally honest about these kinds of things.

answered Mar 11 at 17:03

Andy Putman - no longer active's user avatar

$\endgroup$

5

30

$\begingroup$

Inspired by the now-deleted answer, I performed a small experiment to learn how capable current AI systems are for current Mathoverflow questions.

The linked answer was to an elementary question about a stochastic process. Since the capabilities of AI vary from one subfield of math to another, I decided to test AI on elementary questions about stochastic processes. The author of that question, Nate River, has asked a number of elementary questions about stochastic processes that I found interesting. I chose to test AI systems (two versions of ChatGPT) on two other questions of his which I had tried to answer and made some progress on but not completely resolved to my satisfaction.

These were the only two questions I tried. On both of them, the AI suggested ideas that led to substantial progress. I edited my answers to incorporate these ideas after verifying them myself. What I wrote is entirely mine in terms of words, equations, and organizational structure, with what I got from the AI equivalent to a 1-2 sentence suggestion. I now realize this violates the current AI policy, sorry. (If a new policy is adopted that allows this, I will edit my answers to conform to the new policy.)

For one question, Gambler's ruin following the martinagle betting strategy I had obtained a lower bound and found a plausible strategy to obtain an upper bound but wasn't quite able to make my calculations work. (I probably could have made them work if I put more effort in after failing the first time, but didn't.) ChatGPT (the latest Plus version) gave a better approach. I think its calculations were correct, but I didn't look at them: It was easy enough to do the calculations on my own, and I tried to look at the AI outputs as little as possible. Some steps were not fully justified, but I easily came up with my own justifications for these steps.

For another question The happy village problem, I had obtained an upper bound and then returned to the problem later and found a similar lower bound, and in fact my answer was expected. But I was bothered by, in addition to the gap between the upper and lower bounds, the fact that neither bound depended substantially on one of the problem's parameters. I wondered if the true asymptotic was independent of that parameter or not. I asked ChatGPT this question (the latest Pro version), uploading my existing arguments. ChatGPT was able to:

(1) give a better strategy for the upper bound, which I have checked and I am convinced is right

(2) point to a paper from which I suspect it is possible to derive a matching lower bound, showing that in fact the true asymptotic is independent of this parameter, which I haven't checked at all.

However:

(1) In computing the value of the new upper bound, ChatGPT made a basic error assuming independence of random variables that are not in fact independent, and computed the wrong bound.

(2) The explanation given for why the paper was relevant was not at all clear.

In summary, I believe that in some areas of mathematics, for the majority of questions which I could not resolve on my own after spending significant time thinking about it, I could do better by asking a current AI system. Thus the premise of the current policy that AI systems are not effective for Mathoverflow-level questions is not at all accurate. At the same time, current AI systems are not completely trustworthy. In particular, this means that an expert using an AI system can provide a non-expert a much better answer than they could obtain using the AI system for themself.

Therefore I would support some kind of new AI policy that allows MO users to use AI as a source of ideas for their answers, perhaps along the lines of Timothy Chow's suggested modification.

Currently, I don't get any financial benefit that I can think of from AI companies. My wife works in AI research at Princeton University's AI lab. She formerly worked in big tech. Her research is not related to AI for math.

answered Mar 14 at 14:45

Will Sawin's user avatar

$\endgroup$

1

24

$\begingroup$

It might be appropriate if we align our AI practices with the Leiden Declaration on Artificial Intelligence and Mathematics, a community initiative endorsed by the International Mathematical Union (summary).

The "AI generated content policy" proposed in the OP adheres to three principles of this declaration: Disclose tool use, Retain the responsibility for correctness, Affirm the humanity of authorship.

Two further principles might be considered: Put effort into proper attribution, Welcome new contributors.

answered Jun 2 at 13:31

Carlo Beenakker's user avatar

$\endgroup$

0

16

$\begingroup$

I am mostly a lurker on the MO meta, but I have been following the trend of AI discussions here. I am an occasional question asker, student of mathematics, and young adult. I have been using the stack exchange platform since I was old enough to do so (as evidenced by the quality of my questions improving over the years). And as all of these things, I would be less inclined to participate on this forum if AI is permitted. And I am not alone in this respect, I know others my age who are exhausted by AI. There are general economic, political, and ethical concerns about the AI industry's past and current behavior. I will not rehash them here; there are others who can exposit these things more eloquently than I. Andy Putnam's answer addresses some of these concerns, although I have many many more. But, as he mentioned, we are not used to being players in finance and politics, so that is not the point which I will focus on here.

My leading question is, if AI answers become common place, what motivation is there for me to ask a question here at all rather than just ask the machine myself? I see two reasons. The weaker case is that somebody with access to a proprietary model may ask on my behalf. In this case, our platform is due an extreme change in climate, and the democratization of the voting and point systems becomes more strongly tied to economic pressures and is just a game of "who can afford the fanciest model?" This is addressed in the current policy:

Users who ask questions on MathOverflow likely have access to the same AI services as anyone else who may be persuaded to use them. They likely want to ask questions on MathOverflow because they want feedback from an actual expert.

Now, the stronger argument I see is that I would ask a question here for the sake of someone else to vet the machines response. This would perhaps be reasonable for MSE, but here I am doubtful. Here questions are expected to be research-level, and most people who are asking research-level questions are well-versed enough in mathematics to be able to sit down and work out if an answer is valid. A second party providing a sanity check is good and is something I do seek out from time to time, but it isn't something that I would consider worth using the MO platform for.

All this is to say, no, I do not think this policy should be revised. This is a platform where I come for expert insights, not artificial ones.

answered Mar 18 at 18:23

tox123's user avatar

$\endgroup$

3

15

$\begingroup$

Since nobody has yet weighed in with a positive response, let me try to do so. I believe that the first proposed update is a good one. The second proposed update I would suggest modifying to say that either someone should use a verifiable identity or have 10k reputation (or both).

I think we should try to describe more explicitly the "baby" and the "bathwater" in question. In my view, a "baby" would be an answer by a subject-matter expert who has used generative AI as a tool to assist with crafting an answer, and has carefully checked that the content is correct. In such cases, almost all the concerns expressed in the current policy would be addressed. For example, Users who ask questions on MathOverflow expect to receive an answer authored and vetted by a human. This ensures that the answer is factual, relevant, and complete, up to the standards of another human. The only box that might not be ticked is the word "authored"; some might want to insist that any generative AI output be rewritten by a human. There are certainly cases where human rewriting would be desirable, and I'd support having the policy explicitly encourage rewriting, but I don't think it makes sense to insist on rewriting in every case. For example, if the answer is equation-heavy, how is one supposed to "rewrite" all the equations? What matters, I would argue, is that the human has checked the content for correctness, to the level that is expected for publication in a journal. If this has happened, then I don't see any good reason to disallow such content on MO.

What about "bathwater"? There is the obvious case of someone who posts generative AI content that is simply wrong, or even content that happens to be correct but that the person who posted the content has not personally confirmed to be correct. I certainly don't want this type of content on MO. It is not just anonymous users who do this. There has been extensive discussion of a comment by Iosef Pinelis in which he posted a link to a ChatGPT answer even though he admitted that he knew nothing about the subject and did not check it for correctness. Though he dug in his heels and insisted that his action was justified, I still don't agree. Unverified AI content can also show up in questions, not just answers; as far as I can tell, this question asked in part for a published description of an 18-vertex graph whose existence was apparently vouched for only by a questionable chatbot remark (and this provenance was not disclosed until I repeatedly pushed for it).

There are other cases that are more subtle. In a now-deleted answer, there was a case of a user posting an AI-generated solution to a problem, backed up with Lean code (also AI-generated). I didn't study the solution, but others have said that it was correct. However, the user did not respond to queries asking whether he had checked the correctness himself. (Outside of a few known "exploits," Lean can generally be trusted to correctly determine that an alleged formal proof does in fact correctly prove a given formal theorem, but a human still needs to check that the given formal theorem correctly expresses what it's supposed to express.) The lack of response was concerning, and would violate even the new proposed MO policy, but apart from that (and some editorial polemic inserted by the human author, about the role of AI in mathematics in general), the response would seem to be the kind of thing we want to see on MO. The Lean verification served as a guardrail against an incorrect answer, and the answer was not readily available otherwise; a few people had posted partial answers but not a full solution.

In another answer here, Andy Putman warns about the vested interests of companies who have a strong profit motive that threatens to distort objective truth. I might just be ignorant, but so far I don't think we've seen an influx of answers from "AI for math" companies that are potentially correct, but posted by employees who may have some mathematical training, but not enough to really personally confirm that everything is mathematically correct. But even if it hasn't happened yet, it may be coming, and we should be prepared for it. Perhaps the revised MO policy should require anyone who posts AI-generated content to disclose any vested interests.

answered Mar 14 at 13:21

Timothy Chow's user avatar

$\endgroup$

4

13

$\begingroup$

Unfortunately for me banning (or even restricting) anonymous contributions is a hard no.

Privacy and anonymity on the internet are under attack right now in many jurisdictions, often under the pretense of 'protecting the children'. I believe that MO should not take any actions that normalize the requirement for users to be non-anonymous.

answered Mar 16 at 11:26

Federico Poloni's user avatar

$\endgroup$

5

12

$\begingroup$

It is now a month since this question was posted. There appears to be consensus that if AI-assisted answers are permitted then:

  • Anonymous answers are permitted, but possibly with a reputation threshold (Frederico Poloni, Anurag Sahay).

  • Anyone posting an AI-assisted answer should disclose any financial relationship they have with AI firms (Timothy Chow as a suggestion in part of a longer post, Andy Putman).

There is a post from tox123 proposing no change to the existing near ban (11 votes) and posts from Will Sawin (22 votes) and Timothy Chow (10 votes and 8 votes, the latter aggregating all AI-assisted answers to date) that are more positive, while warning of the pitfalls. Here all vote totals are 'votes up' minus 'votes down'.

In a comment on my other answer, I suggested edits to the proposed policy as follows.

  • If you post a question or answer using AI-assistance then you must have carefully checked that the mathematics is correct and stand by it.

  • You do not just copy the output of the LLM, but instead edit it so that your answer is in your own words.

  • You declare in your question or answer how you used the AI tool. This will make it clear that you 'stand by' the mathematics. For example "I checked and edited the output of [some LLM] to get this proof and believe it is correct" will be a sufficient declaration.

URLs or chat transcripts are not required. Anyone following these points will not be pressured to make any further declaration.

Whether the same policy should apply to questions has not been discussed much. I propose that we apply exactly the same standards to questions as answers, so one policy applies throughout.

Please downvote this post if you want to stick with the existing policy which says that "If the mathematical component of your content is deemed to be generated by AI, it will likely be deleted, along with any reputation earned from it."

Please upvote this post if you think that the five bullet points above should be the new policy on AI-assisted questions and answers.

In the event that the policy is changed, it will be left to the team of moderators to decide whether a reputation threshold for anonymous posts is required.

answered Apr 15 at 17:22

Mark Wildon's user avatar

$\endgroup$

22

11

$\begingroup$

Since the OP has clarified why they want a provision against anonymous users using AI in the proposed manner, I want to write an answer explicitly opposing this.

If someone is using AI responsibly, viz., they are taking the time to digest the mathematics, rewriting it in their own words, and posting it with a clear acknowledgment that they used such tools, they should be allowed to do so anonymously. All the reasons there are to be anonymous do not suddenly go away just because an AI tool is involved.

Furthermore, the issue that this proposal is trying to solve -- namely that AI slop can seem correct while being garbage -- will not arise if someone really is using AI responsibly, since they are going to write the answer in their own words. The irresponsible/unethical anonymous user will use LLMs to write their answers without disclosing it anyway, and having this policy one way or the other is not going to change that.

answered Mar 13 at 3:30

Anurag Sahay's user avatar

$\endgroup$

21

9

$\begingroup$

It's always good to base decisions on data, so I think it would be helpful to have an answer here that links to examples of answers (or even questions) that would likely be acceptable under the new policy but seem to be unacceptable under the current policy. I'm going to focus on cases where an LLM contributed significantly to core mathematical content, rather than (say) literature searches or "soft" content.

Two examples by Mark Wildon were already mentioned above in the question; I repeat them here for convenience.

A series related to $\zeta(3)$
Few conjectural series for $\zeta(5)$ and $\zeta(6)$

Also from him there is

A connection between the hyperbolic cosine function and Riemann's zeta function

Similarly, I'm repeating Will Sawin's examples here for convenience.

Gambler’s ruin following the martingale betting strategy
The happy village problem

Two other answers were posted just recently, by Eugene Go.

Are there integers $x,y,z$ such that $(x+1)y^2−xz^2=x^3+2x+2$?
Are there integers $x,y,z$ such that $1 + x - x^3 + x^2 y^2 + z + z^2 = 0$?

Here are two examples by Terry Tao:

Is the least common multiple sequence $\text{lcm}(1, 2, \dots, n)$ a subset of the highly abundant numbers?
Sphere with bounded curvature

Here's one for which an LLM origin (of a deleted answer that was rewritten and reposted) is strongly suspected but not confirmed:

Invertible perturbations of matrices

I've made this answer community wiki in order to encourage others to add further examples in the same vein.


EDIT: Here's another example that is similar to the last example above (i.e., an apparently LLM-generated answer deleted but then examined and cleaned up by a Real Mathematician™ and re-posted):

Are margins of $L$-log -concave functions $L$-log-concave?


EDIT: One further example from Darij Grinberg:

Central idempotents from characters in Frobenius algebras (generalizing Lusztig arXiv:math/0208154v2 §19)


EDIT: Carlo Beenakker cited ChatGPT for a lower bound in the following answer, although ChatGPT's initial bound was not better than what Emil Jeřábek had already provided.

Prove that except for $n=1$, there are more topologies of size $n$ than groups of order $n$


EDIT: Darij Grinberg has posted a new answer to an old question, remarking that "This marks the first time that any LLM has given me a useful proof (I did get some interesting ones before)."

Rational congruence of binomial coefficient matrices

And one more answer from Darij Grinberg

Is there an elementary proof for this identity involving a simplex, quadratic forms and determinants?

$\endgroup$

8

9

$\begingroup$

I recently posted an answer (https://mathoverflow.net/a/513804/477593) with both a llm-written, human-verified/rewritten proof and a shareable, verifiable Lean proof. I think this is a good model, since it allows verification of the argument alongside clear exposition. It is also one concrete way to meet the first proposal's requirement to have carefully checked the mathematics and "vouch for it".

I built the site hosting the Lean verification (example of such verification), https://theoremdb.org, for collaborating on theorems and sharing Lean proofs. As far as I can tell nothing like it exists yet. I think something like this should exist and be used, whether through my site or another one if someone wants to build it.

answered Aug 1 at 23:27

Philip Weiss's user avatar

$\endgroup$

7

$\begingroup$

I was reluctant to express too much of an opinion on this question, as it is important that this policy be the organic consensus of the MO community rather than the result of undue influence. But I do have two possible variants of the proposed policy to put forward. The rationale is that I think one can make a distinction between two types of MO questions:

  1. Questions where the primary objective is to obtain a correct, verified, and explainable answer to the question at hand [and not to obtain answers that are incorrect, unverified, or inexplicable], regardless of the source of that answer;
  2. Questions where one is also additionally interested in secondary goals such as connecting the question to the literature, creating interaction between human MO participants, or obtaining "big picture" insights that go beyond the specific parameters of the question.

One can make the case that AI-assisted answers (when accompanied by sufficient verification) can be of value of questions of the first type, but could be detrimental to questions of the second type. So the variants I propose are to make the author of the question select between the proposed AI policies. There are two forms of this:

  1. "AI opt out": By default, the new policy permitting certain AI uses is in place, unless the author selects a checkbox "human-generated answers only" (similar to the existing "community wiki" checkbox), in which case the old policy applies instead.
  2. "AI opt in": By default, the existing policy stays in place, unless the author selects a checkbox "AI-assisted answers welcome (if compliant with site policies)", in which case the new policy is in effect.

Again, I do not wish to express an opinion on which of these options would be preferable to each other (or to the old or new policy implemented sitewide), but I thought it worth proposing them for discussion in any event.

EDIT: I am editing my answer to bump this topic back to the top of the Meta MO queue in view of recent developments, and specifically the MO thread "What is the unit distance exponent", which is currently improving the lower bound on the Erdos unit distance exponent in real time, almost entirely via heavy AI assistance. In an alternate universe where both the initial unit distance breakthrough, and the subsequent MO discussion to improve the exponent, were performed by humans (and traditional scientific computation packages) without any generative AI assistance, I believe this thread would be regarded as a success story for MO; but in our actual reality, this is a thread that is in violation of the existing policy on AI-generated content, though it would (perhaps with minor edits and additional disclosures) be permitted under the emerging proposal as summarized in Mark Wildon's answer. So the moderators here may soon be forced into a decision as to whether to enforce the old policy or the new one on this thread in particular.

Fundamentally, I see MO as having to make a difficult choice between speed and the human-centric nature of the site. I can make a transportation analogy. A road can be designed to be a pedestrian-only walkway, a paved road for automobiles with pedestrian sidewalks, or a freeway with no pedestrian access. All three options have their advantages and disadvantages, and there is room in a planned city for all three; but one has to make a choice between them (given that the alternative of a chaotic street filled with both pedestrians and vehicles is inferior to all three planned options).

In my mind both the old policy and new policy are defensible choices, but they come with real tradeoffs. One can make MO a slow "pedestrian street" in which almost no AI contributions are allowed, but for which real-time mathematical activity, such as the improvements to the distance exponent, would not happen on this site. One can make it a "paved road with sidewalks" in which certain specific "lanes" of AI activity are allowed, such as dedicated threads permitting AI assistance, while other "lanes" will be reserved for humans. Or one can make it a "freeway" in which regulated use of AI is by default permitted in all MO posts, which (when done correctly) will likely lead to the highest speed of correct mathematical answers generated, but at the cost of human interaction.

My "opt in" and "opt out" proposals were intended as something close to the "paved road with sidewalks" model. In an "opt-in" model, for instance, one could have most MO threads be governed by the old policy in which AI contributions are extremely limited, but a few threads such as the unit distance exponent thread might be flagged (by a suitable checkbox similar to the community wiki checkbox) to be explicitly permissive of content that is largely AI-generated. (This is actually my favored option, but so far it has not attracted much additional support.)

In any event, I hope that there is more discussion on this policy topic, as it will become increasingly urgent to reach a consensus (or at least a compromise) on this issue going forward.

answered Mar 13 at 15:56

Terry Tao's user avatar

$\endgroup$

13

6

$\begingroup$

I welcome the proposed shift from focusing on how mathematical content was produced to who takes responsibility for it.

To me, this seems to address the real issue.

Current LLMs are certainly capable of producing incorrect mathematics, fabricated references, or arguments whose validity even the person posting them has not carefully checked. Requiring that authors understand, verify, and stand behind what they post is therefore entirely reasonable.

However, I wonder about the underlying principle in the longer term.

Suppose AI systems eventually become capable of producing mathematical arguments that are independently verified to be correct (perhaps even formally verified), and suppose a user fully checks and endorses the result before posting it.

In that situation, would the AI origin itself still be considered relevant, or is the present policy primarily a response to the current state of the technology?

In other words, is the goal to exclude AI-generated mathematics as such, or simply to ensure that everything posted on MathOverflow is mathematically sound and that someone is accountable for it?

For me the mathematical correctness of a post is more important, not whether the source is a human brain or an LLM.

answered Jul 16 at 17:43

Marco Mantovanelli's user avatar

$\endgroup$

5

$\begingroup$

I followed this proposed updated policy in my answer to MO508945. My conclusion is that the policy will largely prevent responsible use of LLMs by professional mathematicians. My main problem is the requirement that you

'affirmatively "vouch for it"

On the fact of it, this seems reasonable. I do not want AI slop. But it goes too far. It should be implicit that as a professional mathematician I have carefully checked any mathematics that I post, and I believe it to be correct. This 'vouching' requirement goes significantly beyond our normal practices: when you submit a co-authored paper, are you required explicitly to vouch for it? Expecting me to stake my entire professional reputation on any answer

'where an AI tool was used to help with the mathematical ideas'

is too much. David Roberts the moderator, commented on this 'legalistic' requirement when he asked me to make explicit in the answer the comment I had left with details of the LLM help.

I do not yet have a counterproposal. Maybe the problem is insoluble. I think I will probably stop using LLMs for MathOverflow entirely, even though my answers to MO508663 and MO508718 and the further answer linked above (to a significantly harder question) show that at least in some areas, the state-of-the-art paid-for models can be very useful.

I'm sure we can all predict another possible reaction, which is that LLMs are used, but not acknowledged. Had I posted the answer to MO508945 without acknowledging LLM help, would anyone have detected it? I am not claiming I would do this. But by its legalistic and onerous nature, the proposed policy almost invites this behaviour. This, I think, is a great shame: it will create a false impression in the community about the capabilities of state-of-the-art LLMs.

answered Mar 17 at 9:26

Mark Wildon's user avatar

$\endgroup$

24

You must log in to answer this question.

Start asking to get answers

Find the answer to your question by asking.

Ask question

Explore related questions

See similar questions with these tags.

Read the original on meta.mathoverflow.net

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.