By Yunzhong, from Waifeisi
QbitAI | Official account QbitAI
OpenAI has recently been dropping bombshells on the math community one after another.
Last month, OpenAI announced that its AI model cracked one of humanity's seven great unsolved Millennium Prize Problems—the Navier–Stokes equations—in just 88 hours; yesterday, it released 722 math manuscripts, three of which concern Millennium Problems.
But in the PaperBenchX evaluation released by UniPat AI, facing 93 real paper reproduction tasks, the strongest performer, GPT-6 Astra, achieved a full reproduction rate of only 13.98%.
△ Overall and stage-by-stage performance of the 10 tested configurations on the 93 reproduction tasks
Think along with us!Current models have already achieved great breakthroughs like cracking Millennium math problems, yet they rarely complete a reliable end-to-end scientific reproduction.
(With a serious face) This raises some questions: in the broad field of AI for Science, what standard should we use to judge "doing science correctly"? With a rigorous and down-to-earth scientific attitude, what exactly is the first step toward an AI mathematician or even a General AI Scientist? And how should we form a set of verifiable, comparable standards to truly measure AI's progress in scientific research?
What ruler should we use to measure AI's progress in science?
Right now, let us savor the open letter jointly issued by 25 Fields Medal winners, including Terence Tao and Yu Deng.
They affirmed that large language models' mathematical ability has risen sharply and can already solve major problems in mathematics, but they also raised a more critical question:AI companies treat solving math problems as benchmark tests, which is harmful to mathematics as a science and to the mathematical community.
So this is not a letter against AI; it asks a simpler question:What ruler should we use to measure AI's progress in science?
The letter also contains an even weightier sentence: the misalignment of treating solving math problems as benchmark tests is only part of a broader misalignment, one that equally affects other scientific and creative professions, and even society as a whole.
Indeed, the questions raised above only become harder to answer when extended from mathematics to the natural sciences and even to society.
Mathematics at least has a last line of defense: Lean. Whether a proof is correct can be checked step by step by a machine, and the truth or falsity of an answer has a safety net.
But in other scientific fields, a reflection coefficient from an electromagnetic simulation, a band structure, or a converged band gap has no Lean. A number that matches the paper may hide a wrong physical model, an unconverged simulation, or invalid post-processing that happened to produce plausible values.
So, before measuring the General AI Scientist capabilities of AI autonomously posing problems, designing experiments, and making new discoveries, there is a more fundamental question:Can it reliably redo a piece of already published scientific work and prove that it did it right?
The logic of this question lies in the fact that, since reproduction has a uniquely correct answer, it can be evaluated accurately.
A recent work released by UniPat AI"PaperBenchX"targets exactly this—it built 93 reproduction tasks from 93 research papers, spanning 12 research directions including electromagnetics, photonics, chemistry, materials, biology, and robotics, plus 10 domain-native scientific environments, containing 3,168 expert-verified scoring items.It is the world's first Benchmark for multidisciplinary end-to-end paper result reproduction.
△ PaperBenchX's task distribution, source papers, and composition of scoring items
PaperBenchX's exam rules are simple:
The Agent received aReal papers, a pre-configured one with the corresponding scientific software installedContainerized environments, and an explanation of "which part to reproduce",Task Brief;
an agent hands back aAn executable reproduction workflow;
Scoring criteria, tolerances, reference answersCompletely hidden from the Agent。
An especially important point: the paper itself is public, so the Agent can of course see the result numbers in the paper. But what is clearly being tested here is not "guessing the answer" — it is testing whether the AI can reconstruct a scientifically valid workflow and run it on its own to produce evidence supporting that conclusion.
So what exactly does PaperBenchX test?
The answer is: it does not test "how many numbers you got"; it only accepts regenerated evidence. After the agent submits, the system deletes all outputs, disconnects the network, and reruns the entire workflow from scratch in an isolated environment; the scorer trusts only the replay artifacts.In a sense, you could say it builds a "Lean" for scientific simulation.
How did the results turn out?
The result: across 93 paper reproduction tasks, the strongest performer, GPT-6 Astra, achieved a full reproduction rate of only 13.98%.
This means that models can produce plausible-looking results on many tasks, but when it comes to actually running the complete workflow from scratch and generating evidence that matches the paper and is verifiable, the success rate is less than one in seven.
This is the gap between today's AI and a General AI Scientist that can cross the boundaries of different scientific fields.
The significance of PaperBenchX is that, for the first time, it has quantified this gap.。
Currently, the UniPat AI team has selected 1 representative task from each research direction, open-sourcing a total of12 test tasks. These tasks cover different research directions and scientific software stacks, and can be used to evaluate Agents, debug reproduction workflows, and study models' end-to-end reproduction capabilities in real scientific environments.
At the same time, the team maintains 81 closed test tasks to preserve PaperBenchX's discriminative power and long-term evaluation validity.
Behind 13.98%, where does reliable reproduction get stuck?
Next, let's take a concrete look at the 10 groups of frontier model and Agent framework configurations that PaperBenchX evaluated on the same set of 93 tasks, analyze the evaluation results, and figure out exactly where reliable reproduction gets stuck.
△ (a) shows the stage scores of the tested configurations; (b) shows the relationship between the number of interaction turns and scores; (c) shows the workflow time allocation of sampled trajectories
Partial completion is not full reproduction
First, look at figure a, the stage scores of the tested configurations. It can be seen that across all tested configurations, the average scores for Modeling and Execution were 62.70% and 61.84% respectively, while Validation was only 42.80%.
This means that Agents are able to complete part of the modeling, call solvers, and generate result files, but these results are not necessarily correct, nor do they necessarily prove that the results genuinely support the paper's conclusions. Overall, they still have great difficulty chaining the various stages together to complete a full, credible reproduction.
More interactions do not mean better reproduction results
Next, look at figure b, the relationship between the number of interaction turns and scores. Trajectory analysis shows that long runs often mean repeated attempts, fault recovery, or continued computation on incorrect model settings. This shows that more interaction turns do not necessarily mean higher scores.
Early decisions determine whether subsequent execution is worthwhile
Next, look at figure c, the workflow time allocation of sampled trajectories. Take Fable 5 as an example: it spent 37.1% of its time on paper reconstruction and code implementation, the highest proportion of early investment among the five models, yet it was able to achieve near-optimal partial scores with relatively few interactions.
In other words, whether early decisions—such as paper understanding, scientific modeling, parameter selection, and experiment organization—are reasonable largely determines whether subsequent computation turns into a waste of resources.
Because if the geometric structure, boundary conditions, mode definitions, or key parameters are already set incorrectly at the early stage, then even if the solver runs normally, it is only computing the wrong scientific system. One wrong early decision can render hours of subsequent computation worthless.
In summary, looking only at results that appear to run through or at the final outputs cannot distinguish between different causes of failure. Some Agents' experiments ran successfully but with incorrect scientific models or result analyses; other Agents failed to complete reruns due to a lack of credible solving processes or code errors.
The UniPat AI team further pointed out several key points for achieving "reliable reproduction":
The difficulty is not "whether it can run," but "whether what runs holds up scientifically" : a seemingly matching number may come from a wrong model, unsuitable parameters, a non-converged simulation, or invalid post-processing;
Evaluation of results must be based on newly generated evidence, not on the Agent's self-reports: delete outputs, cut the network, rerun, and accept only replayed artifacts;
Some progress is real, but complete reproduction remains rare: Agents can already write reasonable simulations and run them to completion, but making the entire process and final results withstand verification is still very difficult.
PaperBenchX's differentiated design
Establishing "whole-paper reproduction" as an Agent evaluation paradigm is not a UniPat AI first; OpenAI's PaperBench is similar work.
However, past tasks have concentrated on the machine learning paper domain. Once we truly enter scientific research scenarios such as physics, chemistry, and materials, reproducing a paper is no longer as simple as "running the code once," and faces three challenges:
Understanding the scientific problem: a paper is not a set of instructions that can be executed as-is. Facing real scientific problems, the Agent needs to understand the research goals on its own, judge how to set boundary conditions, how to handle dispersion, and which parameters will truly affect the conclusions, rather than simply running the code according to the README;
Correctly completing the scientific simulation: Successfully getting the code to run is only the first step. If parameters such as resolution, solver tolerance, and sampling duration are set incorrectly, the numerical results can be completely wrong. The Agent must also be able to read diagnostic messages, identify anomalies such as numerical divergence, and determine when parameters need to be adjusted or the mesh refined. None of these are generic code-execution abilities; they are judgments that depend heavily on domain-specific knowledge;
Judging whether the results are actually trustworthy: This is the most easily overlooked, and also the most fatal, step. Producing a number that matches the paper does not mean the science was done correctly. Wrong physical models, simulations that have not converged, or even unreasonable post-processing can all produce a result that 'looks right.' Therefore, scientific reproduction does not have a simple pass/fail criterion like ordinary coding tasks do; the Agent must also be able to judge whether the scientific process behind the results is reliable.
True scientific reproduction requires the Agent to complete a full pipeline: first understand the scientific problem, then execute the scientific computation, and finally verify the scientific conclusion. These three things are precisely the fundamental skills an AI Scientist must possess:reading science, executing science, and judging science.
So it can be understood that what PaperBenchX tests is not whether an Agent can run a piece of research code, but whether it can complete a relatively complete scientific workflow.
This also constitutes the core difference between PaperBenchX and existing similar benchmarks—behind this lie several key design decisions of this work:
First, the Agent must work inside real scientific software, not machine-learning tasks with a different set of disciplinary data.
PaperBenchX covers 12 research directions and 10 simulation platforms, each domain with its own parameter systems, computational workflows, and convergence criteria. Existing benchmarks (PaperBench, ScienceAgentBench, etc.) are essentially confined to the ML/Python stack; PaperBenchX is the first reproduction benchmark that lets Agents directly face domain-native solvers such as Ansys HFSS, Ansys Lumerical, Meep, PySCF/GPU4PySCF, and ABACUS across different disciplinary fields.
Second, the artifacts submitted by an Agent must be fully reproducible.
After an Agent submits, the system deletes all outputs, disconnects the network, and re-executes the entire workflow in an isolated environment. Only scientific artifacts that can be regenerated during this process proceed to subsequent scoring—each scoring item is anchored to a specific scientific requirement, determined by programmatic checks and LLM-based scoring.
Third, a limited time budget requires the Agent to make the trade-offs of the research process on its own.
Per-task budget:4 – 24 hours, with a median of 7 hours. Within this time window, the Agent must decide on its own how to allocate compute: when to check intermediate results, whether a failed computation needs to be re-run, which parameters are worth tuning, and where to focus its efforts within the limited time. This is closer to how real research is done, rather than a code-execution task with unlimited attempts.
Fourth, the verification method rules out any possibility of muddling through.
With typical code benchmarks, running unit tests tells you whether something is correct, but scientific reproduction cannot be judged by a single pass/fail. Therefore, PaperBenchX splits verification into three layers, each answering a question:
Can it be re-run:after the Agent submits a task, the system deletes all outputs it generated, disconnects from the external network, and executes the workflow from scratch in an isolated environment. Anything that cannot be re-run simply doesn't count. This step filters out all possibilities of manually patching together results, relying on caches, or copying answers from the internet;
Can the evidence be traced:it must be possible to trace back to the corresponding computation process. For example, if an Agent claims a simulation has converged, it must provide the corresponding convergence records; if it reports a physical quantity, it must be possible to locate the simulation output that generated this result along with the subsequent analysis process. This way, what is evaluated is not an isolated number, but the complete chain of evidence supporting that number.
Does it hold up scientifically:finally, the evaluation system judges whether the results satisfy the specific scientific requirements. Each scoring criterion corresponds to an explicit scientific requirement and the evidence needed. For questions that can be explicitly computed—such as numerical consistency, units, and error tolerances—programs check automatically; for questions requiring judgment informed by domain knowledge, LLM reviewers take over.
Moreover, text written by the Agent itself, such as claims of "successful reproduction," is never treated as evidence, nor as an instruction to the judges. Only this can prevent an Agent from "muddling through" by writing a report that looks complete without actually completing the underlying scientific computations.
So the question naturally deepens:How exactly can the scientific conclusions of a paper be decomposed into a set of requirements that can be executed, recomputed, and objectively judged?
PaperBenchX does not simply hand a paper to an Agent and compare final numbers. Every task must go through two stages—"building candidate tasks" and "validation and review"—with nine steps in total, before it formally enters the evaluation.
△ PaperBenchX's two-stage, nine-step task construction and validation process
In other words, PaperBenchX compiles the scientific conclusions of a paper into a new kind of object: a scientific workflow that can be executed by AI, replayed by the environment, recomputed from raw evidence, and finally verified item by item by a scoring system.
For more details on how to "turn a paper into an evaluable scientific reproduction task," you can check out the GitHub repo or the official Blog~
GitHub open-source link:
https://github.com/UniPat-AI/PaperBenchX
Official Blog link:
https://unipat.ai/blog/PaperBenchX
— End —
