Will AI discover what humans can't understand? The real ladder of automated science, from AlphaFold's Nobel to the artificial mathematician
✨ AI-generated · how it's made
AI already holds a Nobel Prize in Chemistry and Olympiad gold in mathematics, and labs are announcing autonomous 'artificial scientists'. We separate what's proven from the marketing: where AI is already superhuman, which 'discoveries' failed to replicate, and why the real brake isn't model intelligence but verification — still human, slow and expensive.
🎬 Our Short
THE THESIS. The question everyone is asking — will AI discover things humans cannot understand? — is better answered with a ladder than with a yes or no. Rung one: AI as an instrument, a statistical telescope that searches enormous spaces defined by humans. Rung two: the copilot that proposes hypotheses and lets a human decide which to test. Rung three: the artificial scientist that formulates new questions on its own. Our thesis is threefold. On rung one, AI is already superhuman, with a Nobel Prize to certify it. On rung two there are real results, but told with too much enthusiasm. On rung three there is not yet a single clear, verified case. And the pace of all this will not be set by model intelligence, but by a bottleneck almost nobody talks about: experimental verification.
WHAT IS ALREADY SUPERHUMAN. The verified facts are extraordinary on their own. AlphaFold, from Google DeepMind, earned Demis Hassabis and John Jumper the 2024 Nobel Prize in Chemistry (shared with David Baker): its database holds 200 million protein structures — the three-dimensional shapes of life's basic building blocks — and is used by more than 3 million researchers across 190 countries, according to DeepMind. Solving that experimentally with classical methods would have taken hundreds of millions of years of work. In mathematics, July 2025 marked a milestone: Google's Gemini Deep Think and an experimental OpenAI model scored 35 out of 42 points at the International Mathematical Olympiad — the gold-medal threshold that only 67 of 630 human contestants reached. Behind this sits a key piece of technology: Lean, a formal language that lets a computer mechanically check every step of a proof. AlphaProof (published in Nature) works in it, and it was in Lean that the most significant news of this year was verified: in January 2026, GPT-5.2 solved Erdős problem #397 — from the famous list of open problems posed by mathematician Paul Erdős — with an original proof that Terence Tao, arguably the greatest living mathematician, called 'perhaps the most unambiguous instance' of an AI solving an open problem. The common pattern across these successes: well-defined search spaces plus a cheap verifier (crystallography for proteins, Lean for theorems). Wherever that verifier exists, AI already flies.
THE 'DISCOVERIES' THAT WEREN'T. Now the fine print. GNoME, also from DeepMind, announced in 2023 that it had found 2.2 million new crystals, 381,000 of them supposedly stable. Chemists Anthony Cheetham and Ram Seshadri analyzed a sample and concluded in Chemistry of Materials that there was 'scant evidence' of useful materials: according to their analysis, the set was full of near-duplicates of known compounds and radioactive materials with no practical use, and none had demonstrated functionality. Berkeley's A-Lab claimed to have synthesized 41 new compounds in 17 days using robots and AI; researchers from Princeton and University College London argue that not a single one was genuinely new, and Nature ended up correcting the paper. Worse still: the MIT paper claiming that AI-assisted scientists discovered 44% more materials — praised by economics Nobel laureate Daron Acemoglu and cited worldwide — was disavowed by MIT itself in May 2025 after the institute concluded its data could not be trusted, and requested its withdrawal. And in October 2025, OpenAI executives boasted that GPT-5 had solved ten open Erdős problems; they had to walk it back when it emerged the model had merely found already-published solutions the list hadn't catalogued. The lesson is not that it's all smoke: it's that in automated science, marketing runs faster than replication. Every 'million discoveries' should be read as 'a million candidates'.
THE BOTTLENECK NOBODY TALKS ABOUT. Generating hypotheses has become nearly free; testing them has not. Google's 'AI co-scientist' proposed repurposing already-approved drugs against acute myeloid leukemia — and it was right: the compounds inhibited tumor cells in the lab, according to the study published in Nature. But 'right' here means culture dishes, not patients: years and hundreds of millions in trials still lie ahead. The same system reproduced in 48 hours a hypothesis about bacterial resistance that took José Penadés' team at Imperial College a decade; impressive, with a caveat: the AI reached the idea, but humans were the ones who had proven it. Kosmos, the 'autonomous scientist' from Edison Scientific, promises to compress six months of analysis into a single day; by the company's own account, 79.4% of its conclusions are accurate. Put the other way around: one in five is wrong, and finding which one requires exactly the human expertise it was supposed to replace. Robotic labs — like Coscientist, Carnegie Mellon's GPT-4-based system that planned and executed real chemical reactions (Nature, 2023) — aim to close the full loop, but the A-Lab showed the risk: a robot that synthesizes plus an AI that analyzes, with no human in between, produces errors in a chain. Our economic reading: value is shifting from generating ideas to verifying them. Whoever controls the verification infrastructure — auditable robots, assays, independent reviewers — will set the pace of discovery, not whoever owns the smartest model.
WHEN NO ONE CAN FOLLOW THE PROOF. The philosophical question isn't new: it turns 50 this decade. In 1976, Appel and Haken proved the four color theorem using more than 1,000 hours of computer time to check thousands of cases; no human has ever been able to review that proof by hand, then or now. Philosopher Thomas Tymoczko argued this turned mathematics into an empirical science: you had to 'trust' the machine. Half a century later we have a better answer: formal verifiers like Lean, whose core is a small, auditable program. In other words: we can now know something is true without understanding why it is true. That is the key distinction: truth is not understanding. AlphaFold is the perfect example: it predicts a protein's shape without giving us a theory of folding; we got the answer without the explanation. Is that science? Our answer: yes, but incomplete. Understanding is not an aesthetic luxury: it is what generates the next question. A science of oracles issuing correct but inscrutable answers would produce advances — and stop producing scientists. History, however, invites optimism: the microscope arrived centuries before germ theory; first we saw, then we understood. Something similar is already happening: mathematicians are studying the constructions AlphaEvolve found across 67 open problems — it improved the best known solutions in several — to extract new human ideas from them.
OUR READING. In the short term, honesty: scientific AI today is, in equal parts, an extraordinary instrument and a factory of inflated headlines. Auditing the claims — GNoME, the A-Lab, the MIT paper — is not pessimism; it is the scientific method doing its job. Scientific employment will change the way everything else is changing: executing protocols and combing literature (science's 'back office') loses value, while formulating questions, designing discriminating experiments and verifying results gains it. On governance, our proposal is simple: 'AI discoveries' need reporting standards and independent replication before the press release, just as we demand clinical trials before a drug. In the long term, our optimism is firm and reasoned: if the 20th century industrialized manual labor, AI is industrializing hypothesis generation. When verification accelerates too — auditable autonomous labs, formal verifiers, faster trials — science will compress from decades into years. That is the real path toward eradicating disease and extending healthy life: Isomorphic Labs, DeepMind's sister company, is already steering its first AI-designed drug toward clinical trials. So — will AI discover what we cannot understand? It already produces truths no one can follow step by step. What it has not yet done is formulate a single great question no human would have asked. As long as that rung stays empty, the human scientist is not being replaced: they are being promoted. From operator of the answer to owner of the question.
Sources & references
- Nobel Prize in Chemistry 2024 — Press release (NobelPrize.org)
- Nature — Chemistry Nobel goes to developers of AlphaFold
- Google DeepMind — AlphaFold: Five Years of Impact
- TechCrunch — OpenAI and Google outdo the mathletes at IMO 2025
- Nature — Olympiad-level formal mathematical reasoning with reinforcement learning (AlphaProof)
- arXiv — Mathematical exploration and discovery at scale (AlphaEvolve)
- The Decoder — Terence Tao says GPT-5.2 Pro cracked an Erdős problem, but warns the win says more about speed than difficulty
- Chemistry of Materials (ACS) — Cheetham & Seshadri: Perspective on 'Scaling Deep Learning for Materials Discovery' (GNoME)


