Inside the OpenAI Astra and Anthropic Fable Math Claims Nobody Can Verify Yet

OpenAI says an unreleased model solved ten hard maths problems and Anthropic says its own cracked five, but the proofs still depend on humans to check them.

Portrait of Mara Ellison 8 min read
A chalkboard covered in dense mathematical notation photographed at an angle in soft light
Formal verification can confirm a proof's logic is valid without confirming the claim behind it.

OpenAI said on 3 August that an unreleased model internally called Astra had solved ten hard, previously open mathematics problems, according to reports in India Today and the Indian Express. Anthropic followed with its own claim: a model it calls Fable had cracked five. Both companies say each solution was formalised as a Lean certificate, and both claims arrived alongside price cuts on existing OpenAI models named Luna and Terra. Treat the headline numbers with some caution, because verifying a mathematical claim of this kind is slower and more human-dependent than the announcements suggest.

What a Lean certificate actually proves

Lean is a formal proof assistant: a piece of software that checks whether a sequence of logical steps follows validly from a set of axioms, with no gaps a human reader might miss or gloss over. A Lean certificate is the machine-checkable encoding of a proof in that system. If a proof compiles in Lean, its logical steps are, by construction, valid. That is a genuinely strong guarantee, and it is why mathematicians have increasingly adopted Lean for verifying difficult proofs over the past several years.

Where humans are still required

According to the reporting, each argument from Astra and Fable was prepared into manuscripts by humans before formalisation. That detail matters more than it might appear. Translating an informal mathematical argument, however it was generated, into the precise syntax Lean requires is itself a substantial task involving human judgement about what the argument is actually claiming. A model producing a promising sketch is not the same as a model producing a Lean-verified proof, and the gap between those two things is where human mathematicians did the work in this instance.

Formal verification confirms the logic is airtight. It does not confirm that a human didn't have to figure out what the logic was supposed to say.

Why the race to claim credit matters

  • Mathematical reasoning is treated as a proxy for general reasoning capability, so claims here move investor and public perception disproportionately.
  • Neither Astra nor Fable has been released publicly, meaning outside mathematicians cannot yet reproduce or challenge the claimed solutions independently.
  • Price cuts on existing models announced alongside the claims suggest a commercial motive layered on top of the research one.
  • The specific problems solved have not been detailed publicly in a form the broader mathematics community can evaluate at the time of writing.

The pattern of maths claims in AI

This is not the first time a lab has announced a mathematical milestone before independent scrutiny could catch up. The recurring lesson from previous cycles is that initial announcements tend to describe the most favourable framing of a result, and that the framing sometimes narrows once mathematicians outside the lab examine the actual problems solved, the degree of human involvement in preparing them, and whether the problems were genuinely open or merely difficult and under-attempted.

What to watch

Watch for independent mathematicians publishing commentary on the specific ten and five problems once details become available, and watch whether OpenAI or Anthropic release the underlying models or at least the full Lean certificates for public inspection. Until that happens, the responsible reading of the Astra and Fable claims is that they describe promising results prepared with substantial human assistance, not autonomous proof discovery.

Share:

Was this helpful?

Portrait of Mara Ellison

Technology Editor, Lonic

Mara has covered enterprise software for eleven years and spent two of them embedded with deployment teams shipping agent systems into production support desks.

  • Artificial intelligence
  • Enterprise software
  • Automation

Read our editorial standards or send a correction.