Why Erdős Problems Keep Falling to AI Models First
The open questions catalogued from Paul Erdős's vast output are unusually well suited to the kind of search that language models are good at, and unusually poorly suited as evidence about deep mathematical reasoning.

Of the recent wave of AI-assisted mathematics claims, a disproportionate number involve problems drawn from the enormous catalogue of questions attributed to Paul Erdős, the twentieth-century Hungarian mathematician who posed thousands of them over his career, many with small cash prizes attached and almost all in combinatorics, number theory and graph theory. Several of OpenAI's ten and Anthropic's five reported results this summer trace back to entries on Thomas Bloom's maintained online database of Erdős problems. That is not a coincidence, and understanding why tells you more about the limits of current AI mathematical reasoning than the headline solve counts do.
A catalogue built for exactly this kind of search
Erdős problems share a set of properties that make them unusually tractable targets for automated or semi-automated search. They are stated with minimal setup, often in a single sentence, with no dependence on an elaborate surrounding theoretical apparatus. Many ask whether a particular bound can be improved, whether a construction exists with certain properties, or whether a pattern holds for all sufficiently large cases. The database itself, maintained openly and annotated with partial progress, effectively hands a searcher a curated list of well-specified targets, which is precisely the kind of resource that makes systematic exploration by machine feasible in a way that a randomly chosen open problem in, say, algebraic geometry is not.
Construction versus theory-building
A large share of Erdős-style problems are resolved by a single clever construction or counterexample rather than by developing new theoretical machinery. Finding one graph, one sequence, or one combinatorial arrangement with the required property is a search problem: it rewards trying many candidate structures, checking each against the required conditions, and refining promising near-misses. That is a task language models paired with search and verification tools are reasonably well suited to, particularly when the checking step can be automated or formally verified once a candidate is proposed.
Finding the one object that satisfies a clean condition is a search problem. Building the theory that explains why no such object could exist is a different kind of problem entirely.
What this pattern does not demonstrate
- Deep theory-building fields such as algebraic geometry, arithmetic geometry or large parts of analysis require constructing new conceptual frameworks, not finding a single object matching a specification.
- Many resolved Erdős problems were already known to be 'probably true' from partial results or computational evidence, narrowing the search considerably before any model was involved.
- Difficulty within the Erdős catalogue varies enormously; some entries are closer to exercises than to genuinely hard open questions, and announcements rarely specify which end of that range was targeted.
- A model finding one clever construction says little about its ability to sustain a multi-year, multi-paper research programme of the kind most genuine mathematical advances require.
The honest scientific reading
None of this makes the results meaningless. A model that can propose valid constructions for well-specified combinatorial problems, verifiable formally once proposed, is a genuinely useful research tool, and mathematicians working in these areas have said as much in professional forums even while resisting the framing pushed by corporate announcements. The honest description of what has happened is closer to 'a useful search assistant for a narrow, well-suited class of problems' than to 'a system demonstrating general mathematical creativity', and the two framings lead to very different expectations about what comes next.
Why labs keep choosing this catalogue
From a communications standpoint, Erdős problems are an ideal showcase. The name recognition is high even outside mathematics, the problems are simply stated enough to summarise in a press release, and the existence of an open, well-organised database means a lab can point to a specific, checkable entry rather than a vague description of internal capability testing. None of that is dishonest in itself, but it does mean the choice of showcase problems is itself a form of framing, selected partly for public legibility rather than purely for mathematical significance.
What to watch
The more informative test of machine mathematical reasoning will come from problems that require sustained theory development rather than a single construction, and from fields with less pre-existing partial progress to lean on. Watch whether labs eventually publish results in areas without an equivalent curated database, since that would be a stronger signal that the underlying capability generalises beyond the specific structure that has made Erdős problems such convenient targets this summer.
Was this helpful?

Dr. Ivan Petrov
Science Editor, Lonic
Ivan holds a doctorate in condensed matter physics and worked on superconducting qubit error correction before moving into science journalism.
- Quantum computing
- Physics
- Research policy
Read our editorial standards or send a correction.



