Twitter/X

@tonylfeng states that all the mathematics in the AI-produced work is…

Brief

@tonylfeng says the AI-produced mathematics (about ten separate results) is fundamentally correct and strong enough for top venues—Inventiones for math and STOC Best Paper for TCS—and would previously secure tenure‑track interviews or even tenure for PhD authors. He rejects 'Fields Medal level' claims, highlights #3 (existence of non‑sofic groups) as a longstanding landmark, and calls for overhauling hiring and publishing norms.

Why it matters

@tonylfeng states that all the mathematics in the AI-produced work is fundamentally correct.

Key details

  • The consensus is the corpus of ~10 results is strong enough for top venues (Inventiones for math papers; STOC Best Paper for TCS); prior to 2026 a PhD whose sole result was one of these would likely get a tenure‑line interview at a strong university and some results could even secure tenure.
  • Mathematicians consulted say the results are clever and beautiful but not 'deep' enough for Fields Medal status; #3 (existence of non‑sofic groups) meets traditional proxies of importance (longstanding, famous problem with an active community), and the author urges overhauling hiring and publishing practices and applying consistent standards to human work.
Source evidence

I am a mathematician but I avoided commenting on this because the problems are (far) outside my domain. Even if I sat down to read the technical details (which I have not), it would be hard to appreciate the context of prior work, etc. But I've gotten enough sense of things through colleagues to be confident about some aspects.

First of all, all the mathematics is fundamentally correct.

How impressive are the results? There is a spectrum of reactions (as with everything), but the consensus seems to be all the results could realistically land (by prior standards) in the very highest tier of publications, e.g., Inventiones for the math papers and "Best Paper Award" at STOC for the TCS papers.

What does this mean practically? Prior to 2026, a Ph.D. student whose only result was one of these 10 would likely land an interview for a tenure line job at a strong university. For some of the results, you could probably replace "tenure line" with "tenured". When Astra comes out, it will clearly be very disruptive to the existing system. I think it's good that the community has the time to react and prepare. Personally I think we should totally overhaul hiring and publishing practices.

Are the results "Fields Medal Level"? I have not heard a mathematician answer Yes to this, and the quoted reason is invariably that the solutions are very clever and beautiful but not "deep". This is consistent with what I can see from a cursory glance.

However, the same criticisms could be levied at certain human generated works which have won major prizes. Ultimately, comparing results is very subjective and also somewhat political, and I absolutely think that if a human had one of these results, and also the right influential senior advocate, then that advocate could make a strong case for the work to win major prizes. The fact is that most mathematicians, including prize committees, don't have the breadth or time to evaluate all the top results, so they rely on proxies for importance like: Is the problem longstanding? Is it famous? Was there a large community of people actively trying to solve it? For e.g. #3 (existence of non-sofic groups), the answer to each of these questions is undoubtedly Yes. I've always felt that people should pay less attention to these proxies and focus more on the actual mathematics; glad to see that is happening with the AI solutions, but the same standards should then also be applied everywhere.

wh (@nrehiew_)

The rate of progress is completely astonishing. Can any experts in these fields provide some insight as to how impressive or significant these problems are?

— https://nitter.net/nrehiew_/status/2083468738700152934#m