What It Actually Means That OpenAI's Astra 'Solved' Ten Maths Problems

OpenAI says a prototype model called Astra produced solutions to ten longstanding open problems, but the claim leans on formalisation, human preparation and a definition of 'solved' worth scrutinising.

Portrait of Mara Ellison 8 min read
A researcher's desk covered in handwritten proof drafts beside an open laptop
The gap between a promising sketch and a verified proof is where most of the actual work happens.

OpenAI announced this week that Astra, a prototype model not available outside the company, had produced solutions to ten longstanding open mathematics problems. The announcement came with a familiar shape: a striking headline number, a handful of named problem areas in combinatorics and number theory, and an assurance that each solution had been formalised in the proof assistant Lean before being counted as solved. That last detail is doing a great deal of work in the claim, and it is worth taking apart carefully rather than accepting or dismissing the number ten at face value.

What 'solved by AI' is being asked to mean

In ordinary usage, saying a problem has been solved implies that a person or a machine arrived at the answer largely unaided. OpenAI's own account of the Astra results does not claim that. According to the company, Astra produced candidate arguments and constructions, which were then reviewed, tidied and in several cases substantially reworked by mathematicians before being translated into Lean's formal syntax for machine verification. That is a genuinely different claim from a model autonomously discovering and proving a theorem end to end, and the distinction has been blurred in most of the coverage that followed the announcement.

The role of formal verification

Lean checks that a sequence of logical deductions follows validly from stated axioms, with no ambiguity left for a human reader to misjudge. A Lean certificate that compiles is a strong guarantee that the logic of a specific written argument is sound. It is not, by itself, a guarantee about who or what generated the argument, how much of it required human intervention to state correctly, or whether the underlying mathematical idea was genuinely novel versus a known technique applied to a slightly different setting.

A compiler can tell you the steps are valid. It has no opinion on who worked out what the steps needed to be.

Verification by outside mathematicians has not happened yet

  • Astra itself remains unreleased, so independent researchers cannot reproduce the process that generated the candidate arguments.
  • OpenAI has not yet published the full statements of all ten problems in a form the wider mathematics community can evaluate line by line.
  • The degree of human editing between Astra's output and the final Lean-verified manuscripts has not been quantified publicly.
  • Several of the named problem areas are populated by many adjacent open questions of varying difficulty, and headline framing tends to favour the more tractable end of that range.

The credit and authorship question

Mathematics has faced authorship disputes before, but rarely ones involving a system that cannot be interviewed about its reasoning process and a company with a direct commercial interest in the answer looking as impressive as possible. If a mathematician spends substantial effort turning a rough machine-generated sketch into a rigorous, formally verified proof, the conventional attribution of that work becomes genuinely unclear. Journals and preprint servers have not settled on a standard for describing this kind of collaboration, and the ambiguity currently benefits whichever party controls the announcement.

Why the timing is not incidental

The announcement arrived during a period of intense competitive signalling among frontier labs, in which mathematical reasoning has become a preferred proxy for general capability because it is legible to a technical audience and difficult for the public to independently check. Anthropic made a comparable, smaller claim for a model called Fable earlier in the summer, and the pattern across both announcements is the same: a headline number, formal verification cited as the guarantee of rigour, and limited detail on the human labour involved in getting there.

What to watch

The claim will become genuinely assessable once OpenAI publishes the specific problem statements, the Lean certificates, and an honest account of how much human reworking separated Astra's raw output from the verified result. Until independent mathematicians have had the chance to examine that record, the responsible reading of 'Astra solved ten maths problems' is that a prototype model contributed usefully to ten proofs that mathematicians then finished and verified, which is a real result but a considerably narrower one than the headline implies.

Share:

Was this helpful?

Portrait of Mara Ellison

Technology Editor, Lonic

Mara has covered enterprise software for eleven years and spent two of them embedded with deployment teams shipping agent systems into production support desks.

  • Artificial intelligence
  • Enterprise software
  • Automation

Read our editorial standards or send a correction.