One manuscript claims to settle how well fractions can approximate π. Another claims a major advance in the mathematics of multiplying large matrices. They are not standalone announcements. They sit inside a collection of 722 manuscripts that OpenAI has placed online for the mathematical community to examine.

The scale deserves a pause. This is a release of proposed research, not another percentage on an exam with a known answer key.

It also deserves a close reading. A manuscript is not the same thing as an independent discovery. A formal proof is not the same thing as an understandable explanation. And a result appearing in a company repository is not, by itself, a judgment that it is correct, original or important.

On October 6, OpenAI announced the collection, produced using an unreleased internal frontier model. Some of the underlying manuscripts carry earlier dates. The publication makes a remarkable possibility more tangible: AI systems may substantially expand the supply of research results, while the work of turning those results into shared knowledge changes shape.

To see why that matters, it helps to get past the number 722 and look at what is actually there.

A research collection, not 722 victory laps

The repository catalogue groups the manuscripts into 372 result families across 17 mathematical disciplines. A family can contain a main result, supporting arguments, consequences or alternative proofs. Counting every document as a separate solved problem would misdescribe the collection.

Vastkind independently reproduced the catalogue counts from its index and overview at the release commit. The two agree on all 372 families and 722 unique manuscript paths. That is a check of the published inventory, not an independent verification of its mathematics.

The published collection 372 families. Seventeen fields.

Number of result families in each discipline. A family may contain several manuscripts; these are catalogue counts, not confirmed breakthroughs.

Theoretical computer science40
Combinatorics37
Algebraic and complex geometry36
Number theory31
Differential geometry29
Probability and statistical mechanics29
Mathematical physics25
Operator algebras19
Algebra18
Topology18
Partial differential equations16
Real and complex analysis16
Convex and metric geometry15
Group theory14
Dynamical systems and ergodic theory12
Functional analysis11
Mathematical logic6
Source: OpenAI catalogue, release commit adc7f124, October 6, 2026. Counts reproduced by Vastkind; 372 families, 722 indexed manuscripts.
How this chart was counted

Each numbered family is counted once, in the discipline assigned by the source. The catalogue and overview agree on all family IDs and manuscript paths. All 17 categories are shown, ordered by family count; ties are alphabetical. Bar lengths start at zero and use a common maximum of 40. Family count is not a measure of importance, novelty, correctness or field-wide success.

Cross-check: original overview source. The repository has additional PDF files beyond the indexed manuscripts; this chart measures the catalogue, not every file.

The distribution matters because this is not a single-purpose demonstration confined to one narrow problem type. The catalogue extends from number theory and computer science to geometry, probability, topology and differential equations.

Theoretical computer science has the largest number of families, at 40. Probability and statistical mechanics contains 29 families but 105 manuscripts. Those are different measures of output, and neither tells us how important the results will prove to be. One influential theorem could matter more than a large pile of companion texts.

OpenAI says the model was posed approximately 4,000 problems during the evaluation. That does not make 722 divided by 4,000 a meaningful success rate: documents, result families and attempted problems are different units, and the published material is a selected collection. The README also describes different verification stages and explicitly warns that some unformalized results may contain issues.

The interesting question is not how large a victory counter we can construct. It is what kinds of advances the collection claims, and what evidence lets other people investigate them.

What a new result about π would actually mean

You do not need advanced mathematics to meet the first example. You probably encountered the fraction 22/7 as an approximation to π. A much better one is 355/113.

Neither is exactly π. Their usefulness illustrates the real question: how closely can fractions approach a number that is not itself a fraction?

There is a trade-off between the size of a fraction's denominator and the accuracy it buys. A larger denominator gives more possibilities, but different irrational numbers can behave very differently. The irrationality exponent describes how exceptionally accurate infinitely many of their fractional approximations can be.

OpenAI's September 24 manuscript claims that this exponent for π is exactly 2. More precisely, choose any exponent ν greater than 2. The claim is that, once a denominator q is sufficiently large, no fraction p/q gets closer to π than q to the power minus ν.

The phrase “sufficiently large” does real work. The threshold can depend on ν, and the argument does not supply an effective way to calculate it. Nor does the theorem claim a simple error floor of exactly 1/q² for every fraction. Exceptionally good individual approximations such as 355/113 do not refute it.

This would not be another calculation of π's digits, or a first proof that π is irrational. It would settle a structural question about how π relates to the rational numbers.

There is substantial history behind that distinction. A 2020 paper by Doron Zeilberger and Wadim Zudilin established an upper bound of approximately 7.1032. A September 2026 preprint by Yufei Bai claims a modest improvement to below approximately 7.1019. The proposed value 2 is therefore not another small decimal improvement to those upper bounds. It is a claim to identify the exact exponent.

It is also not a clean story of unaided humans being overtaken by the first machine. The earlier work used computational methods; Bai's preprint explicitly acknowledges AI assistance. Mathematics was already becoming a mixed human-and-machine activity.

The new result comes with a formalization description and a specified comparison target. That makes the claim more inspectable than a PDF alone. It does not make this article a proof review. We examined the stated theorem and the associated checking configuration, but did not independently compile or adjudicate the proof.

The distinction is worth preserving because the mathematical ambition is large enough without exaggeration.

Why 2.25 is a consequential number for computing

The second example concerns multiplying matrices: rectangular arrays of numbers that encode operations on other numbers. For two square matrices with n rows and n columns, the familiar method computes n² output entries, each by combining n pairs of inputs. Its multiplication count therefore grows like n³.

That is not the only possible method. By sharing intermediate calculations in clever ways, algorithms can reduce the rate at which the work grows. The matrix multiplication exponent, conventionally called ω, captures the best asymptotic growth rate. “Asymptotic” means the discussion concerns what happens as matrix size grows, rather than the speed of one particular implementation on one machine.

An August paper using modern optimization and AlphaEvolve reported an upper bound below 2.371177. It improved a sequence of bounds that had been inching downward. OpenAI's October 2 manuscript, released in the new collection, claims an upper bound of 9/4, or 2.25, over the complex numbers.

A mathematical limit, not a speed testA bigger step in the exponent bound

Selected reported upper bounds on the square-matrix multiplication exponent ω. Lower is stronger. The new 2.25 claim is not independently validated here.

2023
Earlier reported bound
< 2.371866
2024
Earlier reported bound
< 2.371552
2025
Earlier reported bound
< 2.371339
August 2026
Optimization + AlphaEvolve
< 2.371177
October 2026
New OpenAI claim
≤ 2.250000
Dot scale deliberately shown from 2.20 to 2.40. Values are bounds, not the exact unknown exponent or measured runtimes. The new claim concerns arithmetic operations over the complex numbers.
Sources and what the comparison means

2023–August 2026 values: Table 1 of Dupont et al., August 17, 2026. October value: Theorem 1.1 in OpenAI's October 2 manuscript, released October 6. These are selected published claims, not a complete history or uniform independent audit. Both recent approaches already involve machine assistance.

Each dot marks an upper-bound value on the same axis; the shortened axis is explicit because this is a dot plot, not proportional bars. The 9/4 result allows an extra positive ε in the exponent and ε-dependent constants. It supplies no measured GPU speedup, finite-size crossover, memory cost or numerical-stability result.

A change in an exponent is different from a small fixed percentage saving. In the simple reference functions n³ and n²·²⁵, doubling n multiplies their values by 8 and approximately 4.76 respectively. That elementary comparison explains why exponents attract attention. It is not a forecast for an actual algorithm: the theorem allows a positive extra exponent ε, and constants and implementation costs matter.

The exact claim is that, for every positive ε, matrix products over the complex numbers can be computed with a number of scalar arithmetic operations bounded by a constant times n to the power 9/4 + ε. The constant may depend on ε. It is not a promise that a practical program takes precisely n²·²⁵ steps.

The manuscript's proposed approach works with tensor representations of matrix multiplication, arranging and separating computational blocks before translating the construction into an exponent bound. That is a claim about the structure of a computation, not a new chip or a benchmark run.

Here is where an exciting mathematical result can become a misleading technology headline. A stronger exponent bound does not automatically give a faster GPU kernel, lower memory traffic, numerical stability or a useful speedup at the matrix sizes people actually use. The manuscript explicitly does not identify a competitive finite matrix size.

The formalization notes also make the mathematical setting specific. A statement over the complex numbers should not be casually promoted into an identical claim about every possible number system. The comparison target counts arithmetic operations in a defined model; bit-level costs and real hardware are separate questions.

If the proposed result stands, it would be a substantial theoretical advance. Converting that advance into an engineering improvement would be another achievement, not a detail we can assume has already happened.

What Lean can settle, and what it cannot

Much of the force of this release comes from the accompanying formal proof artifacts. They offer a different kind of evidence from a model confidently explaining that its own answer is correct.

Lean is a theorem prover and programming language. In formal verification, mathematical objects, assumptions and the desired conclusion are specified precisely. A proof supplies a chain of justification that can be checked against the system's logical rules. The standard is not whether the argument sounds convincing, but whether the formal claim follows in that framework.

That separation is powerful. The system proposing an argument need not also be the final judge of it. A machine-generated proof artifact can, in principle, be checked by other people running the appropriate tools.

But “the proof was checked” still needs an object. Which statement was checked? Under which definitions and assumptions? Does that statement faithfully express the theorem being advertised?

Imagine proving that a sorting program returns an ordered list while forgetting to require that it preserves the original elements. A program that always returns an empty list could satisfy the incomplete specification. The checker would be doing its job; the specification would be wrong. This is an illustrative software example, not a finding about OpenAI's proofs.

An explanatory map, not a completion scoreFive questions hiding inside “verified”

Formal proof checking addresses a precise logical task. A research result also raises questions that the checker alone does not answer.

StatementDoes the formal sentence capture the mathematical question being claimed?
ProofDoes that sentence follow from the allowed assumptions under the checker’s rules?
NoveltyWhat is new, and how does it relate to earlier results and methods?
UnderstandingCan researchers explain the ideas, not only certify the conclusion?
UseCan those ideas help answer another question or support a real application?
Vastkind synthesis based on Lean's verification model, the Leiden Declaration and AGMAI's release recommendations. These are different questions, not a claim that every paper passes five sequential gates.

OpenAI's repository includes comparison statements and configurations that associate them with solution modules and permitted axioms. We inspected those mappings for the π and matrix examples. They make the intended targets visible rather than leaving readers to infer them from titles.

Even within one paper, scope matters. The π manuscript also discusses a consequence for the Flint–Hills series. Its selected comparison statement targets the main π theorem, not that additional consequence. The solution code contains further material, but that does not expand the scope of the selected check. A headline saying simply “the paper is verified” can conceal distinctions like this.

There is no honest collection-wide verification percentage in this article. Families with formalization documentation, manuscripts listed as sources, and formal theorem targets are not interchangeable units. The repository's formalization catalogue itself labels its scope as partial progress and carries an unchecked review status. That metadata is not a verdict that every result is wrong, or that nobody has checked anything. It is another reason not to invent a universal seal of approval.

Correctness is also not novelty. OpenAI's collection includes a claim about Catalan's constant, but Zhi-Wei Sun had already posted a proof claim about the same constant on September 3. Neither claim is adjudicated here. The point is that determining priority and relationship to earlier work requires a literature check, not a successful compilation.

The Leiden Declaration on AI and mathematics makes the broader responsibility explicit: formal correctness does not by itself establish that the formalization captures the intended mathematics. Researchers still need to connect the precise statement to the question people care about.

The three-hour figure is striking, but not a price tag

OpenAI describes the average result as using roughly three hours of ChatGPT Pro thinking compute with the internal model. That comparison is arresting because it frames research-scale output in a familiar product unit.

It should not be read as an instruction to buy Pro and reproduce the collection this afternoon. The producing model is not publicly released. A compute equivalent is not a retail price, and it is not necessarily elapsed time from choosing a problem to obtaining an independently accepted paper.

The public material does not provide a complete research cost account covering model development, all unsuccessful attempts, curation, formalization, exposition and subsequent human scrutiny. Multiplying three hours by 722 would not repair that gap: “result” and “manuscript” are not established as the same counting unit.

The disclosure is useful, nevertheless. OpenAI has also released ten abridged reasoning summaries and information about attempted problems, giving researchers more process evidence than a polished announcement alone. Those records make better questions possible about what the system did, what humans supplied and where the expense actually sat. They are not a full reconstruction of every attempt.

This is the larger limitation of treating science like a model benchmark. A score rewards a specified output. A research program also needs problem selection, provenance, failures, usable methods and a community capable of extending the work.

Mathematicians are already changing the machinery around research

A direct response from outside OpenAI is neither a victory announcement nor a dismissal.

In its October 6 statement, the Advisory Group on Mathematics and Artificial Intelligence calls the publication an important event. It also says its advisory role should not be interpreted as an assessment of the results' impact or an endorsement of the process that produced them.

Its central distinction is sharp: “This release is the beginning, not the completion” of human understanding and the work's incorporation into mathematics.

The group's September recommendations go further than asking for cautious language. They call for clear exposition, proper attribution, inspectable formal artifacts, process disclosure and durable community-controlled repositories. They also object to laboratories testing advanced research problems on proprietary models unavailable to the wider scientific community.

OpenAI says it will fund workshops, conferences and special programs to help people understand major AI-produced results, and that it is working toward responsibly releasing the model. Those are announced commitments, not evidence that the funding has been delivered or the model is already available.

AGMAI's objection is not a request to hide existing results. It is an argument about who gets to participate in producing the next ones. A field in which outside mathematicians only explain discoveries delivered by a handful of labs would be different from one in which they can choose their own questions and use comparable tools to investigate them.

Two nearby developments show that these are operational questions already being acted on, not just anxieties about a distant future.

On October 6, mathematician Thomas Bloom announced changes to the Erdős Problems website. He described genuine progress alongside a flood of poorly explained AI proof claims. His response included a pause on new problem comments and proof claims, removal of solved-status counters, and a stronger emphasis on accounts of what is understood. The changes arose from a longer debate and feedback process, not from this one OpenAI release.

Meanwhile, mathematician Ben Antieau introduced Hexagon, a community repository that had begun accepting public submissions on September 28. It permits AI-generated work even when the submitters do not claim to understand it. The aim is to make such material durable, discoverable and citable instead of scattering it across social media. Hexagon explicitly does not provide peer review; inclusion is not certification of correctness, quality or novelty.

These are complementary responses. One protects a place for discussing and understanding problems. The other creates a place to store results that still need that work. Both resist the idea that a rising count of uploaded proofs is a sufficient measure of mathematical progress.

The opportunity is larger than replacing a proof writer

There is a more productive way to imagine this technology than a contest between an AI lab and displaced mathematicians.

In an October 3 account of her own research, Jennifer Taback describes AI helping repair a false lemma, improve the organization of a paper and open new directions. She also describes errors and the continuing responsibility to understand and check the work. This is one researcher's experience, not a productivity trial or a review of OpenAI's unreleased model.

It illustrates a possibility that the huge collection alone cannot measure: widening access to useful mathematical conversation. A researcher without a nearby specialist might explore a connection, test a line of attack or find an error earlier. The valuable outcome would not necessarily be another spectacular theorem. It could be a better question or a collaboration that becomes possible.

Mathematical explanation matters for the same reason. Grant Sanderson argues for rewarding deeper explanatory work, not just the existence of a proof. Knowing why a construction works, why someone would invent it and which parts can travel to another problem creates capabilities that a yes-or-no certificate alone does not.

Consider the difference between receiving a finished bridge and understanding bridge design. The structure may be useful. The transferable insight lets you adapt it to a different river, detect a bad assumption and build the next one. Mathematics is not civil engineering, but the distinction between an output and a reusable way of thinking is real.

This is where the optimistic reading becomes demanding rather than merely celebratory. More generated results could mean more opportunities to learn, combine ideas and attack questions that previously lacked enough attention. Realizing that opportunity requires investment in explanation, verification and access, not only in generating the next batch.

What this changes outside mathematics, and what it does not

It is tempting to draw a straight line from new theorems to new medicines, energy systems or an accelerating cycle of AI designing better AI.

The release does not establish those outcomes. Mathematical proof has a special advantage: once the statement, assumptions and rules are fixed, there is a precise object to verify. In empirical science, a beautifully reasoned result can still rest on a model that leaves out the behavior that matters.

Amit Sahai makes that distinction through a hypothetical fusion-plant example: a theorem about a design cannot, by itself, establish that its mathematical model faithfully captures the physical system. Experiment, assumptions and engineering judgment remain necessary. There is no new fusion plant in that example.

The same boundary matters for the matrix result. A mathematical complexity claim is one thing; running an efficient, stable implementation on a real workload is another. Evidence of the former can motivate work on the latter without being mistaken for it.

Still, the possibility is consequential. If systems can generate useful new results across many areas, researchers could have a much larger stock of constructions, techniques and hypotheses to work with. The payoff would depend not just on how fast the machines produce them, but on which results survive scrutiny and help people or other tools do something they could not do before.

That is not yet a measured rate of scientific acceleration. It is a plausible mechanism worth taking seriously and testing.

Sources and scope

This article uses the initial OpenAI repository commit, adc7f124, dated October 6, 2026. We independently counted catalogue entries, reconciled the index with the overview, inspected the selected π and matrix theorem statements and their formalization mappings, and compared the claims with the cited original papers and public statements. Repository counts describe published artifacts, not confirmed breakthroughs. The matrix chart compares reported upper bounds, not measured performance.

We did not rerun Lean or independently verify the mathematical arguments. The article distinguishes OpenAI's claims, our catalogue checks, researchers' published positions and our own analysis. Manuscript dates are separate from the collection's release date.

The next milestone is what survives and gets used

Vastkind's earlier coverage of OpenAI's geometry result asked what changes when a model produces candidate knowledge rather than merely explaining existing material. The new collection makes that question harder to leave at the level of a thought experiment.

The next useful signals are concrete: independently reproduced formal checks of named statements; corrections and revisions; clear accounts of which ideas are new; researchers explaining and extending important proofs; and access that lets people outside the producing lab ask different questions.

None of these requires pretending the release is ordinary. A collection spanning this many areas, with inspectable manuscripts and substantial formal material, is a serious event. Nor does taking it seriously require declaring that mathematics has been solved or that an intelligence explosion has been demonstrated.

The most ambitious outcome is not a warehouse of results that nobody understands. It is more people able to understand more mathematics, ask better questions and discover things that were previously out of reach.

The 722 manuscripts are a beginning of that test, not its final score.

For the next part of the story, read how AI changes the problem of trust in scientific literature.