47 Comments
User's avatar
San Cabraal's avatar

Undoubtedly progress. There's always a price for that.

Do you how many conjectures etc failed? I am interested in knowing what percentage are successful etc: a helpful indicator of progress etc.

Do you know the computational costs of all of the iterations that failed? Hard to believe this was a one-shot route to success.

Kai Williams's avatar

We don't have great evidence on the base rate for the Astra result. Noam Brown mentioned that they tried other problems that didn't succeed — like the millennium problems. But the real answer is probably at most 1%-5% of the conjectures they tried.

The current benchmark that the clearest denominator is Frontier Math open problems: https://epoch.ai/frontiermath/open-problems. AI systems have solved 3 of the 50 in the benchmark, but weighted toward the less interesting side.

Also, because OpenAI's Saturday announcement involved Erdos problems, it's fair to say that they probably tried all of them. But 9 of the 10 which the mathematician Thomas Bloom called his "favorites" are still unsolved: https://www.erdosproblems.com/forum/thread/blog:5. (The only one which isn't was the unit distance problem, which OpenAI solved in May: https://www.understandingai.org/p/openais-milestone-math-breakthrough).

I imagine we'll see more benchmarking in the near future.

(I didn't go into this because I wasn't trying to answer here how good AI systems are at math. That's a whole different question from the rest of the piece.)

Russell Hawkins's avatar

So what does it look like when an AI fails to solve a math problem? Does it actually understand and acknowledge this, or does it hallucinate that it has? Does OpenAI manually stop the process at some point? How exactly does it terminate?

TheAiBuildGuide's avatar

Russell, the piece gets at exactly this: a wrong proof and a right one read the same on the page. Same is true running these models in production, there's no internal signal saying "not sure about this one."

Whether OpenAI manually stops it inside these open-problem runs, that's their process, we don't know. But without an outside check built into the workflow, the model itself won't tell you which output was actually verified and which one just sounds like it was.

Kenny Easwaran's avatar

The day after the unit distance proof was dropped, people at Google DeepMind released a paper showing their project on a set of open questions: https://arxiv.org/abs/2605.22763

It looks like they tried basically every open Erdös problem that had been formalized, and every conjecture on the online encyclopedia of integer sequences that had been formalized, and got formalized proofs or refutations of 5-10% of them.

It's important that their work is limited by what has been formalized - the unit distance problem isn't formalized as far as I know, and neither are most of the recent OpenAI results. Human mathematicians basically don't do formalized math, apart from things like the very simple instances I teach in intro logic classes, but there has been interest in recent decades in formalizing more math to get it beyond the level of confidence we have from trusting experts.

San Cabraal's avatar

Many thanks. I saw that someone achieved 5 of the mastra solutions on Claude Fable.

So maybe the LLMs are very good at a particular area of mathematics - the 1 to 5% you speak of. Hard to discern genuine effectiveness/efficiency at the moment.

Kai Williams's avatar

> So maybe the LLMs are very good at a particular area of mathematics

I agree with everything you've written though I'd a "now" to the end of this statement. We're _so_ early into LLMs being able to solve somewhat interesting problems at all that I think it's very difficult to say what right now looks like, let alone the future.

And 1% to 5% is just a conservative upper bound of "somewhat interesting open problems." As in, I expect the true % is probably lower?

(Also, the guy who achieved five of the astra solutions with Fable is Levent Alpöge, who was a promising mathematician who has since moved to Anthropic. He was the one who prompted the Jacobian result from Fable)

Sam Waters's avatar

As far I know, I’m a paid subscriber, but the advertisement by 80k Hours is still showing up for me…

Timothy B. Lee's avatar

Hi Sam can you please email me with details? tim@understandingai.org. I'd love to know whether you saw the ad in the email or on the website. If on the website were you logged in? A screenshot would also be helpful. Thanks!

Olivier Coutu's avatar

Same here, I'm logged in and selected this post through my "paid" tab in the Substack app on Android. I can provide a screenshot if Sam is unable to provide enough info to resolve it.

Timothy B. Lee's avatar

Sorry about that! I've asked Substack to look into it based on Sam's experience.

Sam Tobin-Hochstadt's avatar

I have exactly the same situation -- Substack app, Android, paid subscription, logged in (I'm commenting), saw the ad.

Andrew Eisenberg's avatar

Same. Paid subscription on the Substack app on iPhone. Seeing the ad.

Jason's avatar

Free subscriber here: sponsor message slotted in nicely, and I did check them out!

Aman Karunakaran's avatar

I like Litt's blog post, but I will say, while it's possible to imagine mathematics existing as an endeavor in such a world, it is certainly much harder to imagine it as a _profession_

Timothy B. Lee's avatar

Most elite mathematicians today do a mix of teaching and research. So if AI eclipses humans on math research, it would be pretty natural for the profession to shift to a greater emphasis on teaching. You might ask why there would be demand for math education if AI is better at humans than math, but we have professors teaching any number of subjects that have little economic value outside of academia. Math could become more like philosophy or Egytpology or whatever — a certain number of people just want to study it because it's interesting.

Aman Karunakaran's avatar

I have long argued that mathematics' rightful place is in the humanities. Perhaps we are approaching a world in which the life of a math professor and a standard humanities professor aren't so different after all

Kai Williams's avatar

Several people I chatted with mentioned that (pure) math might become more of a humanities subject! (It's definitely my relationship to the subject)

Daniel's avatar

So I don’t buy it for two reasons: 1) mathematics has criteria for truth that everyone agrees to. The humanities universally do not. That in and of itself, IMO, makes a bigger difference that whether the goal is appreciation of beauty or whatever vs. application. 2) I suppose they’re usually a minority in a given department, but all those people specializing in applied probability, and the border between CS and math, and PDEs, etc. would have to be chunked off, which makes less sense to me than chunking off algebra from philosophy.

Kenny Easwaran's avatar

It depends on how much math shifts to "do we have the right definitions?" rather than "do we have the right theorems about these definitions?"

Both are central questions right now, but everyone agrees to ignore the questions about the right definitions, because it's impolite to argue, and the question about the theorems is beyond argument. But the definitions are where the progress has always been.

(At least, something like this is David Bessis's thesis: https://davidbessis.substack.com/p/the-fall-of-the-theorem-economy)

Daniel's avatar

I understand what you’re driving at but that maybe sort of addresses (1). It doesn’t address (2). And I say “maybe sort of” because one answer to “what are the right definitions” will surely be “whatever provides utility” and that will loop us back to (2).

Kai Williams's avatar

To respond in a somewhat different way as Kenny: I think mathematics becoming a humanities subject would be very traumatic for the field, and isn't necessarily a good future. (At least, it would reach humanities levels of state support...which are much lower.)

But I think the worlds where it happens are likely to be traumatic for academia (and humanity) anyway. If humans end up getting pareto-dominated at math, I think it likely that this happens to a lot of other disciplines, including the more applied fields, so the traditional arguments around teaching calculus etc. might also break down.

And to be honest, depending on what people care about, mathematics could very reasonably unbundle into several disciplines:

* People who care about the problem solving, either for practical applications or just because, and leverage AIs to do most of the solving but very quickly. I think a lot of applied mathematics likely fits into this.

* People who care about understanding and digesting AI generated proofs, which lends itself more to a humanities subject

But all this is very unclear to me!

Daniel's avatar

Ok but the vast majority of students that math professors teach are done after their prerequisites for some other subject are fulfilled. Probably like 95%? At my university I’d say there were about 20x as many students taking calculus as were taking any proof-based math class at a given moment.

Daniel's avatar

Ok but the vast majority of students that math professors teach are done after their prerequisites for some other subject are fulfilled. Probably like 95%? At my university I’d say there were about 20x as many students taking calculus as were taking any proof-based math class at a given moment.

Sam Tobin-Hochstadt's avatar

This is just like English or History or Music or Art. But if AI can do everything that math professors do, then there will also be big impacts on history and on math teaching and on everything else.

Kristian Evans's avatar

This article raises a great point I don't think I'd considered before, which is basically the role of humans as a "stakeholder of understanding." You could see this playing out in a lot of different fields, but in math, the idea that humans need to understand what AI is doing both because it would be bad for loss of control scenarios but also because plausibly there are many ways in which future discoveries will impact us individually and it is good to have a broad cohort of humans engaged on what those things mean for us overall.

Tim talked about this on a recent episode of AI Summer, but I'm wondering if there are examples of AI discovering something that we don't understand? That's a big, broad term, but I think the scary notion would be we get AI rapidly making connections that nobody or no group of experts can comprehend, I'm unclear if that has happened at this point.

Anna Souakri's avatar

So far it seems the concurring view in maths communities; cf for ex the recent Ramsey theory proving done by human thanks to AI as the accelerator for proof and computation

Kai Williams's avatar

Could you say more? I'm not sure which Ramsey theory result you're thinking of

Oleg  Alexandrov's avatar

As somebody who has a PhD in math (but works in industry) I will say these are exciting times.

AI will surely improve. Not just at brute-force theorem proving, but also getting taste, ability to define new areas and paradigms, etc.

People will adapt, we always do. Math is like a sport, really. We will benefit, not be replaced.

Oleg  Alexandrov's avatar

We will deal with impenetrable proofs as we do with impenetrable code. Need better abstractions, more theory, more math objects, more encapsulation of messy details.

M.L.D.'s avatar

A lot of these quotes read like cope though. There’s a world of difference between “AI will change how we work but it won’t replace us—high level post-graduate math will be mostly the same” and “people can still do math, if they want, but the journals, such as they are, are almost all aggregating warehouses for AI results, and no one is really training or hiring PhDs in math anymore.”

Terence Tao seems pretty careful not to say things that he just hopes might be true. But you have to look for what he doesn’t say as much as what he does

Mike Alexander's avatar

Stories like this bring to mind an eerie portion of the novel Colossus, the supercomputer that takes over the world. The machine has learned of the existence of another machine like it that was built by the Soviets. The two machines begin to communicate through mathematics, beginning with arithmetic. Soon it is dealing with elementary calculus. A team of mathematicians is assembled to monitor the information stream. For a while they are able to keep up with gathering a sense of what the math being output is about. But within a couple of days, one by one they throw up their hands unable to follow any of it, and finally the last one throws in the towel. They experience this machine, which they put in charge of the nuclear missiles has now moved beyond their understanding, and shortly after that, their control.

Daniel's avatar

Kai, I want to reiterate that the fact that you can address this topic competently is exactly why I’m a fan. You got your interview subjects to address the important, nuanced points that would go over the head of most writers.

Kai Williams's avatar

Thank you!!

Alec Pritzos's avatar

Huh, the adoption pattern in here is telling. The use everyone's comfortable with is literature search, and that's the one use that doesn't touch who gets credit for a result. Uptake seems to be sorting by what's safe for the credit system, not by where the tool actually helps most.

Kai Williams's avatar

Eh...I think this is a uncharitable. Literature search is something which slots much more easily into people's existing workflows. It also is a less fun task to do.

Christopher's avatar

Great article!

Small suggestion for ads: use some kind of type formatting to differentiate between your article text and your ads. Even simple italic.

I know it seems obvious, but doing so will help readers know where to pick up reading the post after an ad interruption.

Michael Harmon's avatar

Let’s just tell all AI to calculate pi to the nth degree and then get on with our fallible lives.

Luqman's avatar

This is a wonderful article. However, I had some comments on theory building. I do not agree that theory building (formulating definitions, frameworks, capturing mysterious mathematical concepts or ideas, etc.) is a natural extension of the problem-solving capabilities that we are seeing LLMs increasingly become good at. I believe theory building is a completely different ball game similar to the difference between writing well and writing creatively. Hence, we should expect, based on trend lines, for AI to have incredible competitive math skills, arriving at ever more intricate and clever proofs. That seems reasonable. Alternatively, I think theory building in mathematics is closely related to aesthetics in art or writing. Note, the trend line of LLMs when capturing aesthetics is not so encouraging. In all, the problem of creative writing or capturing philosophical insight or depth is the same as capturing or creating beautiful mathematics.

Kai Williams's avatar

Thanks for the comment! I think that this is a reasonable position to hold (for now at least!) and one which many mathematicians have. (Though not all mathematicians — I raised basically this exact argument to Tsimerman. That's the particular thing he was disagreed with that I quoted him on.)

For the sake of discussion, let me lay out some of the arguments for the other side:

(1) There may in fact be ways to reward good theory-building. When I asked this question to Carina Hong and Ken Ono at Axiom, Hong pointed me to several lines of existing research in how reward good definition-writing. For instance, one could use the lean dependency graph to measure how often a definition is used; better definitions will be more central to the network. This is partly Axiom talking its book, but it's not impossible to me that they're reasonable.

In particular, I think that theory building has a clearer reward signal than good writing. Ultimately, doesn't a good theory usually cashes out in being able to solve problems that were unsolvable before? (I am speaking outside my area of expertise here, though). Of course, there are other, more mysterious factors at play, but it seems to me that there's a bit more of a gradable handle to a good mathematical theory than a good poem, for instance.

(2) You write that the "trend line of LLMs when capturing aesthetics is not so encouraging." Broadly, I would agree, at least compared to verifiable tasks. But I also think that our experience of current LLMs does not bring out their aesthetic abilities. In particular, some people, e.g. Janus, are much better at eliciting interesting/beautiful things from LLMs. So there may be an overhang here that we don't fully appreciate.

(3) There's a sense in which better pretraining lifts all boats. So even if it turns out that companies cannot figure out a way to do RL to reward good theory building. (as they have with problem-solving), I would expect that AI systems will eventually be able to do this effectively. But that trajectory might take much longer.

Leandro Queiroz Macedo's avatar

We don’t need to compete with AI on execution. We need to reposition ourselves around defining purpose, problems, decisions, and meaning.