16 Comments
User's avatar
San Cabraal's avatar

Undoubtedly progress. There's always a price for that.

Do you how many conjectures etc failed? I am interested in knowing what percentage are successful etc: a helpful indicator of progress etc.

Do you know the computational costs of all of the iterations that failed? Hard to believe this was a one-shot route to success.

Kai Williams's avatar

We don't have great evidence on the base rate for the Astra result. Noam Brown mentioned that they tried other problems that didn't succeed — like the millennium problems. But the real answer is probably at most 1%-5% of the conjectures they tried.

The current benchmark that the clearest denominator is Frontier Math open problems: https://epoch.ai/frontiermath/open-problems. AI systems have solved 3 of the 50 in the benchmark, but weighted toward the less interesting side.

Also, because OpenAI's Saturday announcement involved Erdos problems, it's fair to say that they probably tried all of them. But 9 of the 10 which the mathematician Thomas Bloom called his "favorites" are still unsolved: https://www.erdosproblems.com/forum/thread/blog:5. (The only one which isn't was the unit distance problem, which OpenAI solved in May: https://www.understandingai.org/p/openais-milestone-math-breakthrough).

I imagine we'll see more benchmarking in the near future.

(I didn't go into this because I wasn't trying to answer here how good AI systems are at math. That's a whole different question from the rest of the piece.)

Russell Hawkins's avatar

So what does it look like when an AI fails to solve a math problem? Does it actually understand and acknowledge this, or does it hallucinate that it has? Does OpenAI manually stop the process at some point? How exactly does it terminate?

San Cabraal's avatar

Many thanks. I saw that someone achieved 5 of the mastra solutions on Claude Fable.

So maybe the LLMs are very good at a particular area of mathematics - the 1 to 5% you speak of. Hard to discern genuine effectiveness/efficiency at the moment.

Kai Williams's avatar

> So maybe the LLMs are very good at a particular area of mathematics

I agree with everything you've written though I'd a "now" to the end of this statement. We're _so_ early into LLMs being able to solve somewhat interesting problems at all that I think it's very difficult to say what right now looks like, let alone the future.

And 1% to 5% is just a conservative upper bound of "somewhat interesting open problems." As in, I expect the true % is probably lower?

(Also, the guy who achieved five of the astra solutions with Fable is Levent Alpöge, who was a promising mathematician who has since moved to Anthropic. He was the one who prompted the Jacobian result from Fable)

Sam Waters's avatar

As far I know, I’m a paid subscriber, but the advertisement by 80k Hours is still showing up for me…

Timothy B. Lee's avatar

Hi Sam can you please email me with details? tim@understandingai.org. I'd love to know whether you saw the ad in the email or on the website. If on the website were you logged in? A screenshot would also be helpful. Thanks!

Olivier's avatar

Same here, I'm logged in and selected this post through my "paid" tab in the Substack app on Android. I can provide a screenshot if Sam is unable to provide enough info to resolve it.

Jason's avatar

Free subscriber here: sponsor message slotted in nicely, and I did check them out!

Aman Karunakaran's avatar

I like Litt's blog post, but I will say, while it's possible to imagine mathematics existing as an endeavor in such a world, it is certainly much harder to imagine it as a _profession_

Timothy B. Lee's avatar

Most elite mathematicians today do a mix of teaching and research. So if AI eclipses humans on math research, it would be pretty natural for the profession to shift to a greater emphasis on teaching. You might ask why there would be demand for math education if AI is better at humans than math, but we have professors teaching any number of subjects that have little economic value outside of academia. Math could become more like philosophy or Egytpology or whatever — a certain number of people just want to study it because it's interesting.

Aman Karunakaran's avatar

I have long argued that mathematics' rightful place is in the humanities. Perhaps we are approaching a world in which the life of a math professor and a standard humanities professor aren't so different after all

Kai Williams's avatar

Several people I chatted with mentioned that (pure) math might become more of a humanities subject! (It's definitely my relationship to the subject)

Kristian Evans's avatar

This article raises a great point I don't think I'd considered before, which is basically the role of humans as a "stakeholder of understanding." You could see this playing out in a lot of different fields, but in math, the idea that humans need to understand what AI is doing both because it would be bad for loss of control scenarios but also because plausibly there are many ways in which future discoveries will impact us individually and it is good to have a broad cohort of humans engaged on what those things mean for us overall.

Tim talked about this on a recent episode of AI Summer, but I'm wondering if there are examples of AI discovering something that we don't understand? That's a big, broad term, but I think the scary notion would be we get AI rapidly making connections that nobody or no group of experts can comprehend, I'm unclear if that has happened at this point.

Anna Souakri's avatar

So far it seems the concurring view in maths communities; cf for ex the recent Ramsey theory proving done by human thanks to AI as the accelerator for proof and computation

Stevan Fairburn's avatar

The part I keep coming back to is the gap between proof production and proof acceptance.

A system can generate a candidate proof, but the profession still needs provenance, replayability, adversarial review, and a way to decide which human or group owns the correction if a subtle step fails.

That maps to clinical AI more than it first appears: the hard product boundary is not just whether the model can produce a plausible answer, but what inspection artifact lets another expert safely inherit it.