Feature

Science in the age of AI

The science behind AI-generated proofs, protein predictions and automated experiments—and what these advances mean for the work of discovery.

Numerically computed Lorenz attractor, drawn in blue with a copper segment against a dark background.
The Lorenz attractor: an original numerical rendering of a system whose exact equations can produce sensitive, diverging trajectories. Figure 5 explores the mathematics.

A new kind of scientific work

A mathematical proof, a protein structure and a weather forecast all try to make something unknown accessible. They arrive by different routes. A proof establishes what follows from a set of assumptions. A structure describes the arrangement of a molecule. A forecast estimates how a physical system will change. Artificial intelligence is entering each of these activities, and the consequences extend well beyond the speed at which researchers can write or calculate.

The striking change is that a machine can now contribute the candidate itself: a possible argument, a molecular arrangement, an algorithm or a proposed next experiment. That changes where a researcher begins. Instead of facing an empty page or an impossibly large search, they may face a collection of plausible answers. Their work then includes deciding which answers deserve attention, finding out whether they survive scrutiny, and understanding what can be built from them.

There is real excitement here. A useful calculation may become cheap enough to repeat thousands of times. A question previously set aside may become tractable. A student may explore a mathematical construction that would once have required a specialist collaborator. Yet the importance of an answer still depends on the question it answers. A more accurate prediction of molecular geometry, for example, supports a different claim from an explanation of what that molecule does inside a cell.

This feature follows that distinction through mathematics, biology, chemistry and the physical sciences. It asks how AI changes the work of discovery, what the strongest examples actually demonstrate, and what still has to happen before a promising output becomes dependable knowledge.

The label “AI” covers several kinds of scientific tool. A language model can propose text or code. A structure predictor estimates spatial relationships. A search system generates alternatives and keeps those that score well. An experimental planner chooses which measurement would be most useful next. Some projects combine these components; others use a specialised model with no conversational interface.

Much of their value comes from an asymmetry: finding a good answer can be far harder than checking a proposed one. A search may have millions of possible routes, while a candidate solution has a relatively compact test. The test need not be perfect to be useful, but its meaning matters. A program can be checked against an exact mathematical identity. A predicted crystal can be checked by another calculation. A real sample can be measured in a laboratory. These checks answer different questions.

A mathematical landscape makes this visible. In Figure 1, each point represents a candidate pair of numbers, and the surface height represents its score. The lowest point is known exactly. The long, curved valley shows why a search can move towards apparently better answers while still having a difficult route to the best one. Real scientific objectives are usually less completely known, but the example exposes the role of the score in directing a search.

A three-dimensional Rosenbrock surface and matching contour map show a curved valley leading to the exact minimum at x equals 1 and y equals 1.
Figure 1. The shape of a search. Original surface and contour plots of the two-dimensional Rosenbrock function, R(x, y) = (1 − x)² + 100(y − x²)². Its unique minimum is R = 0 at (1, 1): both squared terms vanish there. Height and colour show log₁₀(1 + R), revealing the narrow valley without changing the location of the minimum. Function definition [16].
View full-size figure ↗

FunSearch, described by Bernardino Romera-Paredes and colleagues in Nature in 2023, gives a concrete example. A language model generated computer programs; an evaluator tested them; successful programs informed later proposals. Applied to the cap-set problem, the system found new constructions in combinatorics. It also developed useful heuristics for packing items into bins. Its strength came from the interaction between generation and evaluation: the system could discard many unhelpful proposals without requiring a mathematician to inspect each one. [1]

Such a system has two sources of intelligence: the method proposing possibilities and the way success has been defined. A researcher who chooses an inadequate score can obtain increasingly impressive solutions to the wrong task. Optimising the measured output of a reaction, for instance, may overlook whether the result is reproducible or whether the ingredients are practical. Better search makes the definition of the problem more consequential.

Mathematics: from argument to proof

Mathematics offers an unusually clear demonstration of the difference between an attractive answer and a reliable one. Consider the claim that the first n odd numbers add to n squared. Checking the first five cases suggests a pattern. The picture below explains why the pattern continues: enlarging a square of side n − 1 to one of side n requires a border containing 2n − 1 new cells. Repeating that construction accounts for every cell of the larger square.

A nine-by-nine square made from nine successive L-shaped borders containing the odd numbers 1 through 17. The total is 81.
Figure 2. Why the odd numbers make a square. Each colour marks one successive L-shaped border, from 1 cell to 17 cells. The copper outer border has 9 + 8 = 17 cells; the whole square has 81. The identity n² − (n − 1)² = 2n − 1 explains the general construction. The 9 × 9 example illustrates an argument that applies to every positive integer.
View full-size figure ↗

A proof can therefore do more than reassure us about an answer. It can expose a reusable idea. The square construction connects arithmetic to geometry, and the difference of successive squares appears in other settings. In advanced research, the useful idea may be a new representation, a surprising correspondence between fields, or a method that solves a whole class of problems. The length or difficulty of a proof does not by itself tell us how much understanding it will create.

Formal proof assistants add another way to check the logical chain. In Lean, a theorem is expressed in a precise language and its proof is checked by a small core program, the kernel. The guarantee concerns the formal statement, its definitions and the assumptions on which it depends. A reviewer still needs to establish that the formal statement captures the intended problem. Lean’s own validation guidance also explains why unfinished dependencies and additional axioms must be examined. A successful-looking build alone is an incomplete description of what has been verified. [2]

This becomes especially important when proposed proofs arrive in large quantities. On 6 October 2026, OpenAI released a collection of mathematical manuscripts produced by an internal model, with supporting materials and Lean formalizations for many results. It described the release as work for the mathematical community to examine and develop. [3] As checked on 10 October, the repository listed 719 manuscripts in 372 families and reported that approximately 42% of top-line results had been formalized. A family can contain companion arguments, consequences or alternative proofs. Those numbers should therefore not be read as 719 independently established solutions to open problems. The repository explicitly records differing verification states. [4]

It would be premature to derive a verdict on every manuscript from the release itself. The enduring development is the growing separation between the rate at which arguments can be proposed and the rate at which a community can understand them. In its September 2026 recommendations, the independent Advisory Group on Mathematics and Artificial Intelligence called for careful attribution, versioned releases, clear formalization status and support for community-led understanding. It also objected to testing advanced problems on proprietary models inaccessible to the wider field. [5]

The potential reaches scientific computing through quite direct routes. AlphaTensor, reported by Alhussein Fawzi and colleagues in 2022, searched for matrix-multiplication algorithms and produced exactly checkable constructions, including improvements for particular finite-field settings. Matrix multiplication is a basic operation in simulation and machine learning. An improvement there can become useful far beyond the original mathematical problem, although a lower operation count in a specific setting does not automatically produce faster performance on every computer. [6]

This is one credible route from mathematical progress to progress elsewhere: a new algorithm changes the cost of a calculation, which changes the scale of questions that can be explored. Another route is conceptual. A theorem may clarify which measurements can identify an unknown quantity, or when an optimisation method is guaranteed to converge. Such implications have to be worked out result by result. No collection of proofs, however large, automatically settles the scientific questions that depend on mathematics.

Biology: structure and the living system

A protein is a physical object as well as a sequence. The arrangement of its atoms constrains how it can move and interact, so a useful three-dimensional model can give researchers a much better starting point for asking biological questions. The rendering below comes from experimental coordinates of human ubiquitin deposited in the Protein Data Bank in 1987. Its compact fold gives a tangible sense of what a structure contains: positions and relationships that cannot be read directly from a list of amino acids. [7]

AlphaFold’s advance was demonstrated against experimentally determined structures. In the CASP14 assessment, the 2021 paper by John Jumper and colleagues reported a median backbone error of 0.96 ångström for AlphaFold, compared with 2.8 ångströms for the next-best method, using the paper’s specified measure across 87 protein domains. An ångström is one ten-billionth of a metre. This result made accurate structural hypotheses available for many questions that previously lacked them. The benchmark measured geometry; its score did not measure how well a model understood cellular function. [8]

A ribbon model of experimental ubiquitin, with blue helices, copper sheets and teal loops. Below it, a separate CASP14 benchmark shows median backbone errors of 0.96 angstrom for AlphaFold 2 and 2.8 for the next-best method.
Figure 3. Molecular geometry becomes easier to investigate. Above: an original ribbon rendering of ubiquitin, chain A of the experimental X-ray structure PDB 1UBQ. Blue marks helices, copper marks sheets and teal marks connecting loops. Below: reported CASP14 median Cα RMSD at 95% residue coverage, with 95% confidence intervals: AlphaFold 0.96 Å (0.85–1.16); next-best method 2.8 Å (2.7–4.0). The benchmark covers 87 domains and is separate from the ubiquitin example. Data: Jumper et al. (2021).
View full-size figure ↗

AlphaFold 3 extended the scope to complexes involving proteins, nucleic acids, small molecules and other components. Its 2024 paper also describes limitations, including clashes, errors in stereochemistry and incomplete treatment of alternative conformations. A particularly consequential limitation is that its output does not represent the full dynamic ensemble of a molecular system in solution. Sampling several predictions does not automatically recover that ensemble. These qualifications help researchers decide what the model can support. [9]

Imagine having a plausible model of two cellular components fitting together. That may suggest a useful question about their interaction. Establishing that the interaction occurs in the relevant cell, at the relevant time, and matters for a biological process requires evidence at those levels. Location, abundance, competing partners and environmental conditions can all change the interpretation. A model of molecular geometry is a valuable part of this reasoning, and its value grows when the experimental question is precise.

This changes the practical sequence of research. Structural possibilities can be compared earlier, before committing to a difficult measurement. An unexpected experimental observation can be examined against candidate explanations. Researchers can also identify places where a model is uncertain and where additional evidence would be especially informative. The scientific gain is the improvement in the questions that reach the laboratory, as well as the time saved in generating possible answers.

There is a second lesson in the ubiquitin figure. The coordinates predate today’s AI systems by decades. Reliable experimental records, public databases and careful curation are part of the infrastructure on which new models depend. A persuasive image of a protein is the visible end of a much longer chain of work. Understanding that chain helps explain why sustained investment in measurement remains essential even when prediction becomes dramatically faster.

Chemistry: the loop reaches the laboratory

Chemistry makes the transition from proposal to physical result particularly tangible. A candidate can look promising in a calculation yet be difficult to produce, unstable under useful conditions, or less effective than expected when measured. An automated research system becomes more powerful when it can learn from those encounters with the physical world.

In 2020, Benjamin Burger and colleagues reported a mobile robotic chemist that performed 688 experiments over eight days. A Bayesian search algorithm guided its exploration of a ten-variable space, and the search identified photocatalyst mixtures with six times the activity of the initial formulations. The robot carried out physical work in a laboratory; the measurements then guided subsequent choices. This is a specific demonstrated improvement in a defined experimental search, rather than evidence that every stage of chemistry had become autonomous. [10]

The attraction of such a loop is that it can spend effort selectively. One experiment may test a candidate expected to perform well. Another may resolve uncertainty about a region of the search. These purposes can compete: always repeating the most promising option may miss a better one elsewhere, while endless exploration may yield little usable progress. Choosing between them is part of the experimental strategy.

Before-and-after plots of a Gaussian-process model. A new observation near x equals 4.5 reduces the model uncertainty in the middle of the domain and brings its mean nearer the known curve.
Figure 4. An observation narrows the possibilities. A Gaussian-process model estimates the curve f(x) = sin(0.9x) + 0.25cos(2x) from four observations, then adds one where its uncertainty is greatest. Blue is the model mean; shading is ±1.96 model standard deviations; the dashed curve is the known function in this worked example. The copper point is the added observation. These are calculated values illustrating the selection principle, separate from the laboratory results above. Kernel, inputs and outputs are included in the figure data.
View full-size figure ↗

Materials discovery adds a useful intermediate stage. In the GNoME work reported by Amil Merchant and colleagues in 2023, learned models helped screen crystal candidates, followed by density-functional-theory calculations of their energetic stability. The study reported 381,000 new structures on its updated computed stability hull. That is a computational result about a large collection of candidates. It does not mean that 381,000 new materials were manufactured and shown to work in devices. [11]

A stability calculation is still useful. It can eliminate implausible candidates and focus laboratory effort. But the path from a favourable energy estimate to a practical material includes production, characterisation and testing under intended conditions. A battery material, for instance, would have to satisfy several demands at once. Optimising one property can leave the larger engineering problem unsolved.

Automation also increases the importance of recording failures. If a sample could not be prepared or a measurement was unreliable, treating that event as an ordinary low score may teach the system the wrong lesson. The experiment, the instrument and the record have to be interpretable together. A laboratory that can run more trials gains most when it also preserves enough context to learn correctly from them.

Physics: prediction and its limits

The looping form on the 99Science homepage comes from the Lorenz equations, a model introduced in Edward Lorenz’s 1963 work on deterministic nonperiodic flow. Three coupled equations can generate trajectories that remain within a recognisable shape while becoming highly sensitive to their starting points. It is an apt image for this subject because it separates knowing the rules from knowing exactly what will happen far into the future. [12]

For this article, we numerically integrated two trajectories with initial x values differing by one millionth. Both follow the same equations. Their separation initially remains small, then becomes much larger. The lower panel shows that distance on a logarithmic scale, where each major step represents a multiplicative change. The shape of the upper panel and the divergence in the lower panel describe the same system from different perspectives.

Two Lorenz trajectories form a three-dimensional double-lobed attractor. A logarithmic plot beneath shows their separation growing from a tiny starting difference to much larger distances.
Figure 5. Exact rules can coexist with limited predictability. Original numerical integration of the Lorenz system with σ = 10, ρ = 28 and β = 8/3. Starting states: (1, 1, 1) and (1.000001, 1, 1). The upper panel shows the first trajectory in blue and the final 11 model-time units of the second in copper; the lower panel shows their three-dimensional Euclidean separation. Time is in model units. Fourth-order Runge–Kutta integration, step 0.002. This small mathematical system helps explain sensitivity; it is not a quantitative model of Earth’s weather. Download the plotted data.
View full-size figure ↗

AI can improve forecasts within such constraints. GraphCast, described by Rémi Lam and colleagues in Science in 2023, learned from historical atmospheric reanalysis and generated global forecasts up to ten days ahead. The paper reported better performance than its operational deterministic comparator on about 90% of 1,380 evaluated targets. Those targets combined particular variables, levels and forecast times. The percentage describes that evaluation, rather than a probability that any individual forecast will be correct. [13]

Forecasting also benefits from describing possible futures and their probabilities. GenCast, published online in 2024 and in a 2025 issue of Nature, developed machine-learning ensemble forecasts and evaluated them against the European Centre for Medium-Range Weather Forecasts’ ensemble system. [14] An ensemble gives decision-makers something a single smooth prediction cannot: a view of how outcomes may vary. Its usefulness depends on whether the probabilities remain reliable when checked against events.

Physics and machine learning can also operate inside the same model. NeuralGCM, reported by Dmitrii Kochkov and colleagues in 2024, combines a numerical atmospheric dynamical core with learned components. Its work spans weather prediction and climate simulation. This approach preserves explicit physical structure while using learning for parts of the calculation that are difficult to represent adequately at the model’s resolution. [15]

For a farmer deciding when to plant, or a city preparing for heavy rain, the relevant question is local and practical. Does a forecast provide reliable information at the required place and lead time? A strong global benchmark is encouraging evidence, but decisions need evaluation at the scale at which consequences occur. The same principle applies to any AI system moved from an impressive demonstration into scientific or public use.

What a better model still needs

Across these fields, a recurring difficulty is that useful discoveries can be rare. Even a reasonably discriminating model can produce a shortlist dominated by false leads when the underlying task contains very few successes. The arithmetic is simple enough to examine directly.

Suppose a collection contains 1,000 candidates, of which 20 would actually meet the scientific objective. A screening model identifies 80% of those useful candidates, giving 16 true leads. Suppose it also incorrectly selects 5% of the other 980 candidates, giving 49 false leads. The shortlist contains 65 candidates, but only about a quarter are useful. These are chosen numbers for a worked example; they are not performance estimates for any system discussed above.

A grid of 1,000 candidates contains 20 useful cases. The model selects 16 of those plus 49 false leads, leaving about 25 percent useful cases in a 65-candidate shortlist.
Figure 6. The starting odds matter. Above: all 1,000 candidates, with useful cases in blue and selected cases ringed in copper. Below: the 65 selected candidates, with 16 useful cases in blue and 49 false leads in copper. Positive predictive value is 16 ÷ (16 + 49), approximately 25%. A model can substantially enrich a search—from 2% useful candidates initially to about 25% in the shortlist—while still leaving considerable verification work.
View full-size figure ↗

That can still be an excellent tool. In this example, the proportion of useful candidates has increased more than twelvefold. Whether the improvement is enough depends on the cost of following up a false lead and the value of finding a true one. Running a cheap calculation, occupying an expensive instrument and committing months of experimental work are very different consequences. A benchmark becomes scientifically meaningful when it is connected to those costs.

Researchers must also test what happens beyond the examples used to develop a model. A random split of closely related observations may exaggerate how well the system handles something unfamiliar. A stronger evaluation might hold out a later period, a different family of examples or a genuinely new experimental setting. The appropriate separation depends on the claim: future-weather performance, for example, poses a different test from interpolation within a previously studied dataset.

Finally, access to results and access to the system that produced them are separate matters. Public manuscripts may allow scrutiny while a proprietary model remains unavailable for independent use. Conversely, available code may be difficult to reproduce without its training data, computing resources or experimental apparatus. A claim of reproducibility needs to specify what another group can actually repeat. The AGMAI recommendations make the access problem explicit for mathematics; the underlying question of who can participate extends across science. [5]

The discovery that matters

It is tempting to count progress in outputs: more proofs, more structures, more candidates, more papers. Those counts tell us something about a system’s capacity. The deeper change appears when an output helps someone ask a question they could not previously ask, understand an observation that resisted explanation, or make a result useful in a new setting.

The examples in this article point to several plausible forms of that change. Mathematical search can supply constructions and algorithms. Structural prediction can put a concrete molecular hypothesis within reach. Experimental planning can make better use of a limited number of measurements. Forecasting can turn learned regularities into earlier or more informative predictions. Each contribution has its own route to reliability, and each can create work for other parts of science.

A research group adopting these tools should therefore be able to explain its intended gain in ordinary language. Perhaps it expects to reduce the number of experiments needed to reach a target. Perhaps it wants to expose a hidden mathematical relationship or identify where a measurement would resolve uncertainty. Stating that gain makes it possible to judge success against the scientific purpose, rather than against the novelty of the software.

The most consequential future will be built from many such changes. Some will be dramatic enough to alter an entire field. Others will be quieter: a reliable estimate delivered early enough to guide an experiment, an overlooked connection found in a large search, or an explanation made clear enough for someone else to extend. AI’s lasting contribution to science will be measured in the knowledge that people can test, understand and use.

Sources and figure methods

Research checked: 10 October 2026. This feature draws on the original research papers, official proof-assistant documentation and the dated mathematical release below. The OpenAI collection is an evolving research repository; its figures describe the version consulted, not a final assessment of all its claims. The figures were created for 99Science from the cited coordinates, reported statistics, explicit calculations and explanatory diagrams. Download figure code and data.

  1. Romera-Paredes, B. et al. (2023). Mathematical discoveries from program search with large language models. Nature 625, 468–475 (2024 issue). The FunSearch research paper.
  2. Lean language reference. Validating a Lean Proof. Official documentation on formal statements, dependencies, axioms and levels of checking; accessed 10 October 2026.
  3. OpenAI (6 October 2026). Sharing AI progress in mathematics. The release announcement and description of its supporting material.
  4. OpenAI. Mathematics research repository, release history and formal proof artifacts. Catalogue and verification status checked 10 October 2026.
  5. Advisory Group on Mathematics and Artificial Intelligence (29 September 2026). Responsible Release of AI-Generated Mathematics. Independent recommendations on attribution, verification, understanding and access.
  6. Fawzi, A. et al. (2022). Discovering faster matrix multiplication algorithms with reinforcement learning. Nature 610, 47–53.
  7. Vijay-Kumar, S., Bugg, C. E. & Cook, W. J. (1987). Structure of ubiquitin refined at 1.8 Å resolution. Journal of Molecular Biology 194, 531–544. Coordinates: PDB 1UBQ, chain A.
  8. Jumper, J. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. Figure 3 uses the reported CASP14 summary statistics; it does not reproduce a publisher figure.
  9. Abramson, J. et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500. See especially the model-limitations discussion.
  10. Burger, B. et al. (2020). A mobile robotic chemist. Nature 583, 237–241.
  11. Merchant, A. et al. (2023). Scaling deep learning for materials discovery. Nature 624, 80–85.
  12. Lorenz, E. N. (1963). Deterministic Nonperiodic Flow. Journal of the Atmospheric Sciences 20, 130–141. Figure 5 and the opening image are new numerical renderings of the model, not reproduced figures from the paper.
  13. Lam, R. et al. (2023). Learning skillful medium-range global weather forecasting. Science 382, 1416–1421.
  14. Price, I. et al. (2024, online; 2025, issue). Probabilistic weather forecasting with machine learning. Nature 637, 84–90.
  15. Kochkov, D. et al. (2024). Neural general circulation models for weather and climate. Nature 632, 1060–1066.
  16. SciPy reference. The Rosenbrock function. Definition of the optimisation test function rendered in Figure 1.

Author

Discussion

Leave a Reply