Seven Hundred and Twenty-Two Manuscripts, and One Problem Nobody Has Solved for Two Hundred Years
On October 6, 2026, OpenAI dropped 722 mathematical manuscripts onto the world — the largest single event in the history of AI mathematical logic — and quietly suggested that one of them might contain a solution to a problem that has defeated every mathematician alive.
Not a clever problem. Not a hard homework problem. The Navier-Stokes existence and smoothness problem is one of seven Millennium Prize Problems, a list assembled in 2000 by mathematicians who essentially asked: what are the questions we most need answered and least know how to answer? Solve any one of them and you collect a million dollars. More importantly, you rewrite a chapter of human knowledge. Navier-Stokes describes how fluids move, how blood flows, how weather forms. Nobody has been able to prove, rigorously, that its equations always produce smooth, well-behaved solutions. The problem has been open for over two centuries in its essential form.
Here is the strange part. The manuscript is one of 722.
That number is not a typo. The release covers 372 distinct result sets, and 4,000 of the proofs inside have been checked using Lean, a programming language that functions as a ruthless grammar checker for mathematical logic: if a step is wrong, Lean refuses to compile it. This is not a scientist announcing a discovery at a press conference. This is a machine producing verified mathematics at a rate that has no historical precedent.
To calibrate the scale: the average mathematician produces perhaps a handful of significant proofs across a career. Hardy, Ramanujan, Erdős — the most prolific human mathematicians in history — worked for decades to amass outputs a fraction of this size. OpenAI's model appears to have done something structurally different in kind, not merely in speed.
The Institute for Advanced Study, home to Einstein and Gödel, is now advising OpenAI on how to release these results responsibly. That detail alone tells you something: even OpenAI understands it has arrived somewhere that requires outside scholarly judgment.
The Long Road to a Silver Medal: How AI Learned to Think in Mathematics
The International Mathematical Olympiad is not an exam you pass by remembering formulas. Six problems over two days, each one a small locked room with no obvious door. Professional mathematicians sometimes struggle with them. Which is exactly why the AI community adopted the IMO as its stress-test of choice for genuine reasoning: pattern-matching won't save you here.
For years, AI systems failed embarrassingly. Then, in the summer of 2024, Google DeepMind's AlphaProof solved four out of six IMO problems, earning a result that would translate to a silver medal. The research community had been quietly skeptical that such a score was years away. It was not years away. That single result shifted the conversation from "can AI reason mathematically" to "how far can it go, and how fast."
The honest answer is: unevenly. AlphaGeometry 2, released in February 2025, can solve 88% of IMO geometry problems from the past 25 years. Impressive — until you note that geometry is one structured corner of a much larger continent. Generalization remains the stubborn open question.
Meanwhile, OpenAI took a different route with its o1 series. Rather than specializing in geometry, the model was built around what the team calls reasoning tokens: dedicated computational steps where the model works through a problem internally before producing an answer. Think of it less as the model remembering a solution and more as the model doing scratchwork. The result was 83% on the IMO qualification exam. Not the olympiad itself — the qualification round. That distinction matters, and anyone reporting otherwise is selling something.
Three data points, three different architectures, three partial victories. The shape of progress is becoming visible.
The Grammar Checker for Truth: What Lean 4 Actually Does
Think of a word processor's spell-checker, then imagine it worked for logical truth rather than spelling. Lean 4 is something like that, except the stakes are higher and the mechanism is more fundamental. It is a programming language that will simply refuse to compile a proof that contains a logical error, in the same way a compiler refuses to run code with a syntax mistake.
That distinction matters enormously. A human-written proof can harbor a subtle gap for years before anyone notices. A Lean-checked proof cannot. The verification is not a second opinion from another mathematician who might share the same blind spot; it is a machine checking every deductive step against the axioms of mathematics itself, with no possibility of polite agreement.
This is where neuro-symbolic AI earns its hyphenated name. Neural networks are brilliant pattern-matchers and genuinely terrible logicians. Lean 4 is an unforgiving logician with no intuition at all. Neuro-symbolic systems graft the two together: the neural network generates candidate proof steps, and the symbolic engine either endorses them or kills them. Hallucinated mathematics, the kind that looks confident and is quietly wrong, gets caught before it survives a single step.
Lean 4 has become the industry standard for machine-checking proofs, supported by the Lean Focused Research Organization. It is no accident that OpenAI chose it as the verification layer for all 4,000 problems in Tuesday's release. Gartner, not known for early enthusiasm, placed neuro-symbolic AI on its 2025 Hype Cycle as a critical technology for logical reliability. Gartner's attention is a signal worth noting and a caveat worth keeping: it means the technology is real enough to track, and not yet settled enough to trust without question.
Solving a problem and understanding it are not the same operation.
The Trillion-Parameter Rival: Mistral Large 4 and the Architecture of Scale
The same week OpenAI's manuscripts landed, a French company quietly published documentation for a model with a peculiar double identity. Mistral Large 4 carries 1.05 trillion total parameters. But at any given moment, only 52 billion of them are actually working.
That gap is the trick. The architecture is called Mixture-of-Experts, and the intuition is surprisingly simple: not every question needs every expert in the room. A router — a lightweight traffic-directing system — reads each incoming problem and routes it to the relevant specialist sub-networks, leaving the rest dormant. The trillion-parameter universe sits waiting; the 52-billion active slice does the actual thinking. The whole is dramatically cheaper to run than the sum of its parts.
Mistral also built something else into the model: a 1.6 billion-parameter vision encoder. That number sounds modest beside a trillion, but its implication is not. Mathematics is not only text. Diagrams, geometric figures, hand-drawn proofs photographed from a blackboard — these are how human mathematicians actually work. A model that can read an image as fluently as a sentence has a fundamentally wider definition of what counts as a solvable problem.
What Mistral's release makes undeniable is that the race to build logically capable AI is not one company's project. OpenAI, Google DeepMind, and now Mistral are each approaching the same destination by different roads. The competitive landscape matters here, because it means the pressure to verify, to improve, and to honestly stress-test these systems is not going to slow down.
The Illusion of Understanding: Where the Machine Still Fails
Take a model that can solve four out of six International Mathematical Olympiad problems. Now move one disk. Apple researchers, in their 2025 paper aptly titled "The Illusion of Thinking," did something like that with the Tower of Hanoi: they changed one parameter, a single constraint in a puzzle the model had apparently "mastered," and performance collapsed. Not gracefully. Completely. What looked like reasoning turned out to be pattern recognition wearing a very convincing costume.
The comparison to earlier AI failures is instructive here. A decade ago, image classifiers would identify a school bus correctly, then misidentify it entirely when rotated slightly. We called that brittleness a growing pain. The new failure is subtler and therefore more dangerous, because the model's confidence does not waver when it is wrong.
Video generation offers a vivid parallel. The "World Models' Last Exam in Physics" benchmark gave AI systems sequences of real physical events and asked them to model what comes next. The best-performing model scored 57.76 out of 100 for physical consistency. It could produce footage that looked plausible to a human eye while violating the laws of thermodynamics underneath. Visually convincing. Physically wrong. The gap between those two things is precisely where lives could be lost in any system that acts on what it perceives.
The honest distinction to hold onto is this: solving a problem and understanding it are not the same operation. Lean 4 enforces logical correctness at every step, which is powerful. But correctness in a known structure is not the same as competence when the structure shifts. That gap is not a flaw to be patched in the next release. It is, right now, the actual frontier.
Teaching, Flipping Bits, and the Horizon of AI Mathematical Logic
Solving a problem and explaining it are two different cognitive acts. The Sherpa reinforcement learning framework takes that distinction seriously: by training AI tutors to adapt to different student archetypes, it improved tutoring performance by an average of 20.5 percentage points. The number is striking because it suggests that reasoning ability and pedagogical ability are separable muscles, and AI, like a brilliant mathematician with no patience for questions, can fail at the second even while mastering the first.
Meanwhile, as proofs grow logically airtight, the attack surface migrates downward — past the code, past the logic, all the way to the silicon. The BARE-AI framework exists to catch bit-flip attacks, where physical hardware perturbations corrupt AI model weights at the memory level. It detects such attacks with up to 98% accuracy in vision models, which is a reminder that a formally verified proof is only as trustworthy as the machine running the verifier.
Three questions remain genuinely open. Whether the 722 manuscripts survive peer review. Whether the Navier-Stokes solution holds under independent scrutiny. Whether an 83% IMO score reflects real reasoning or very expensive pattern-matching on training data. If the bottleneck in AI mathematical logic is shifting from finding proofs to deciding which problems deserve asking, we have not yet worked out who — or what — gets to make that choice.