Today for AI

量子位官网 · 10/8/2026, 17:28:44

UniPat Releases PaperBenchX: GPT-6 Astra Achieves Only 13.98% Reproduction Rate, Highlighting End-to-End Scientific Validation Gap

By 允中Original title: ChatGPT踢到铁板了!能破解千禧数学难题,但论文复现率低至13.98%?
78AI Score
Executive Summary

UniPat AI introduced PaperBenchX, a benchmark evaluating large models on 93 real-world paper reproduction tasks. Results show that even the top-performing GPT-6 Astra achieved only a 13.98% full reproduction rate, revealing a significant gap between solving specific math problems and reliable end-to-end scientific replication. The benchmark compiles scientific conclusions into executable, replayable, and verifiable workflows to establish new standards for measuring AI progress in science.

SOURCE COVERAGEOriginal coverage

OpenAI has recently been dropping heavy bombs on the mathematics community.

Last month, OpenAI announced that its AI model solved one of the seven Millennium Prize Problems—the Navier-Stokes equations—in just 88 hours, conquering a challenge that had stumped humanity for a century. Yesterday, they released 722 mathematical manuscripts, claiming solutions to three more Millennium Problems.

However, in the PaperBenchX evaluation released by UniPat AI, which tested 93 real-world paper reproduction tasks, the top-performing GPT-6 Astra achieved a complete reproduction rate of only 13.98%.

△ Overall and stage-wise performance of 10 tested configurations across 93 reproduction tasks

Think about it! Current models can achieve breakthroughs as grand as solving Millennium Prize problems, yet they rarely succeed at reliable end-to-end scientific reproduction.

(Serious face) This raises some questions: In the broad field of AI for Science, what standards should we use to judge whether science is being done correctly? What is the first step toward becoming an AI mathematician or even a General AI Scientist, grounded in rigorous and solid research attitudes? How do we establish verifiable and comparable standards to truly measure AI's progress in scientific research?

What yardstick should be used to measure AI's progress in science?

At this moment, let us carefully re-read the open letter co-signed by 25 Fields Medalists, including Terence Tao and Deng Yu.

They acknowledged the rapid improvement in Large Language Models' (LLMs) mathematical capabilities, noting their ability to solve major problems in mathematics. However, they raised a more critical issue: Using problem-solving in mathematics as a benchmark is harmful to the discipline of mathematics and the mathematical community.

Thus, this is not a letter opposing AI. It asks a simpler question: What yardstick should be used to measure AI's progress in science?

The letter also contains a heavier statement: The misalignment of using mathematical problem-solving as a benchmark is just part of a broader misalignment that affects other scientific and creative professions, and indeed society as a whole.

Indeed, extending the issues raised above from mathematics to natural sciences and social sciences makes them even harder to answer.

Mathematics at least has a final line of defense: Lean. Whether a proof is correct can be checked step-by-step by machines, providing a fallback for truth values.

But in other scientific fields, there is no Lean for electromagnetic simulation reflection coefficients, band structures, or converged band gaps. A number that matches a paper might hide behind an incorrect physical model, unconverged simulations, or invalid post-processing that coincidentally yields a plausible value.

Therefore, before measuring the capabilities of a General AI Scientist—which autonomously poses questions, designs experiments, and makes new discoveries—there is a more fundamental question: Can it reliably reproduce published scientific work and prove it did so correctly?

The logic here is that because reproduction has a single correct answer, it can be accurately evaluated.

UniPat AI recently released a work called "PaperBenchX", targeting exactly this task. It constructed 93 reproduction tasks from 93 research papers, spanning 12 research directions—including electromagnetics, photonics, chemistry, materials, biology, and robotics—and utilizing 10 domain-native scientific environments. With 3,168 expert-verified scoring items, it is the first international multi-disciplinary end-to-end benchmark for reproducing paper results.

△ Task distribution, source papers, and scoring item composition of PaperBenchX

The rules of the PaperBenchX exam are simple:

The Agent receives a real paper, a containerized environment pre-installed with the corresponding scientific software, and a task brief specifying "which parts to reproduce."

The Agent submits an executable reproduction workflow.

Scoring criteria, tolerances, and reference answers are hidden from the Agent.

It is crucial to note that while the papers themselves are public and Agents can see the result numbers within them, this exam clearly tests not "guessing the answer," but whether AI can reconstruct a scientifically valid process and generate evidence supporting those conclusions on its own.

So, what does PaperBenchX actually test?

The answer is: It does not test "how many numbers were obtained," but only recognizes regenerated evidence. After the Agent submits its work, the system deletes all outputs, disconnects from the network, and reruns the workflow from scratch in an isolated environment. The scorer trusts only the replayed artifacts. In a sense, this builds a "Lean" for scientific simulation.

What were the results?

The result is: Among the 93 paper reproduction tasks, the best-performing model, GPT-6 Astra, achieved a complete reproduction rate of only 13.98%.

This means that while models can produce seemingly reasonable results in many tasks, the success rate of running through the entire workflow from scratch and generating verifiable evidence consistent with the paper is less than one-seventh.

This represents the distance between today's AI and a General AI Scientist capable of crossing boundaries across different scientific domains.

The significance of PaperBenchX lies in quantifying this distance for the first time.

Currently, the UniPat AI team has selected one representative task from each research direction, open-sourcing 12 test tasks. These tasks cover various research directions and scientific software stacks, serving to evaluate Agents, debug reproduction workflows, and study models' end-to-end reproduction capabilities in real scientific environments.

Simultaneously, 81 closed test tasks are maintained to preserve PaperBenchX's discriminative power and long-term evaluation validity.

Where does reliable reproduction get stuck behind the 13.98%?

Next, we will specifically examine the 10 frontier model and Agent framework configurations evaluated by PaperBenchX on the same set of 93 tasks, analyzing the results to understand where reliable reproduction gets stuck.

△ (a) Stage scores for the tested configurations; (b) Relationship between interaction rounds and scores; (c) Workflow time allocation in sampled trajectories.

Partial Completion Does Not Equal Full Reproduction

First, let's look at Figure a, which shows the stage scores for the tested configurations. It is evident that across all tested configurations, the average scores for Modeling and Execution were 62.70% and 61.84%, respectively, while Validation scored only 42.80%.

This indicates that while Agents can complete partial modeling tasks, invoke solvers, and generate result files, these results are not necessarily correct, nor do they prove that the results truly support the paper's conclusions. Overall, it remains difficult for them to connect all stages into a complete and credible reproduction.

More Interactions Do Not Mean Better Reproduction Results

Next, consider Figure b, which illustrates the relationship between interaction rounds and scores. Trajectory analysis reveals that prolonged execution often implies repeated attempts, fault recovery, or continuing calculations based on incorrect model settings. This demonstrates that an increase in interaction rounds does not guarantee higher scores.

Early Decisions Determine the Value of Subsequent Execution

Now, look at Figure c, showing the workflow time allocation in sampled trajectories. Taking Fable 5 as an example, it allocated 37.1% of its time to paper reconstruction and code implementation—the highest proportion among the five models—yet it achieved near-optimal partial scores with fewer interactions.

In other words, the rationality of early-stage decisions—such as understanding the paper, scientific modeling, parameter selection, and experiment organization—largely determines whether subsequent computations become a waste of resources.

If geometric structures, boundary conditions, mode definitions, or key parameters are set incorrectly in the early stages, even if the solver runs normally, it is merely calculating an erroneous scientific system. A single wrong decision early on can render hours of subsequent computation worthless.

In summary, looking solely at seemingly successful runs or final outputs cannot distinguish between different causes of failure. Some Agents' experiments run successfully, but their scientific models or result analyses are flawed; others fail to re-run due to a lack of credible solving processes or erroneous code.

The UniPat AI team further highlighted several key aspects required to achieve "reliable reproduction":

The difficulty lies not in "whether it can run," but in "whether what runs is scientifically valid": A number that appears to match may stem from an incorrect model, inappropriate parameters, unconverged simulations, or invalid post-processing;

Evaluation of results must be based on regenerated evidence, not Agent self-reports: Delete outputs, disconnect from the network, re-run, and accept only replayed artifacts;

Some progress is real, but full reproduction remains rare: Agents can already write reasonable simulations and execute them to completion, but ensuring that the entire process and final results withstand verification remains very difficult.

Differentiated Design of PaperBenchX

Establishing "full paper reproduction" as an Agent evaluation paradigm was not pioneered by UniPat AI; OpenAI's PaperBench represents similar work.

However, previous tasks focused primarily on machine learning papers. When entering scientific research scenarios such as physics, chemistry, and materials science, reproducing a paper is no longer as simple as "running the code once." It faces three challenges:

Understanding Scientific Problems: A paper is not an instruction manual to be followed step-by-step. Facing real scientific problems, Agents need to independently understand research objectives, determine how to set boundary conditions, handle dispersion, and identify which parameters truly affect conclusions, rather than simply running code according to a README;

Correctly Executing Scientific Simulations: Successfully running the code is only the first step. Improper settings for parameters such as resolution, solver tolerance, and sampling duration can lead to completely erroneous numerical results. Agents also need to interpret diagnostic information, identify anomalies like numerical divergence, and judge when to adjust parameters or refine meshes. These capabilities rely heavily on specific disciplinary knowledge rather than generic code execution skills;

Judging Whether Results Are Truly Credible: This is the most overlooked yet critical step. Producing a number that matches the paper does not mean the science is correct. Incorrect physical models, unconverged simulations, or unreasonable post-processing can yield "seemingly correct" results. Therefore, scientific reproduction lacks a simple pass/fail standard like ordinary coding tasks; Agents must also assess the reliability of the scientific process behind the results.

True scientific reproduction requires Agents to complete an end-to-end pipeline: first understanding the scientific problem, then executing scientific computations, and finally validating scientific conclusions. These three elements constitute the fundamental competencies required of an AI Scientist: understanding science, executing science, and judging science.

Therefore, PaperBenchX tests not whether an Agent can run a piece of scientific code, but whether it can complete a relatively comprehensive scientific workflow.

This constitutes the core difference between PaperBenchX and existing similar benchmarks, driven by several key design principles:

First, Agents must operate within real scientific software, not just machine learning tasks with swapped domain data.

PaperBenchX covers 12 domains...

Per-task budget: 4–24 hours, with a median of 7 hours. Within this time window, the Agent must autonomously decide how to allocate computational resources: when to inspect intermediate results, whether to rerun calculations after failures, which parameters are worth tuning, and where to focus its efforts given limited time. This setup more closely mirrors real-world scientific workflows than code execution tasks that allow for unlimited retries.

Fourth, the verification methodology eliminates any possibility of cheating.

Standard coding benchmarks can determine correctness by simply running unit tests, but scientific reproduction cannot be judged solely by a pass/fail outcome. Therefore, PaperBenchX divides verification into three layers, each addressing a specific question:

Can it be reproduced? After the Agent submits a task, the system deletes all generated outputs, disconnects from the external network, and executes the workflow from scratch in an isolated environment. If the result cannot be reproduced, it is immediately disqualified. This step filters out any attempts to manually piece together results, rely on cached data, or copy answers from the internet.

Is the evidence traceable? Results must be traceable back to their corresponding computational processes. For example, if the Agent claims a simulation has converged, it must provide the relevant convergence logs; if it reports a physical quantity, there must be simulation outputs and subsequent analysis steps that generate this result. Thus, the evaluation assesses not just an isolated number, but the complete chain of evidence supporting that number.

Is it scientifically valid? Finally, the evaluation system determines whether the results meet specific scientific requirements. Each scoring criterion corresponds to clear scientific demands and required evidence. Issues involving numerical consistency, units, and error tolerances—which can be explicitly calculated—are checked automatically by programs. Questions requiring domain expertise for judgment are evaluated by LLM judges.

Moreover, text written by the Agent itself, such as "reproduction successful," is never treated as evidence nor as instructions to the judges. This prevents Agents from submitting seemingly comprehensive reports without actually performing the underlying scientific computations to "game" the system.

This leads to a further question: How exactly can scientific conclusions in a paper be decomposed into a set of executable, recomputable, and objectively verifiable requirements?

PaperBenchX does not simply hand a paper to an Agent and compare final numbers. Every task must go through two phases—"Candidate Task Construction" and "Verification & Review"—comprising nine steps in total, before it officially enters the evaluation pipeline.

△ The two-phase, nine-step task construction and verification process of PaperBenchX

In essence, PaperBenchX compiles the scientific conclusions of a paper into a new type of object: a scientific workflow that can be executed by AI, replayed by the environment, recomputed from original evidence, and finally verified item-by-item by scorers.

For more details on "how to turn a paper into an evaluable scientific reproduction task," please visit our GitHub or official Blog.

GitHub Open Source Link:

https://github.com/UniPat-AI/PaperBenchX

Official Blog Link:

https://unipat.ai/blog/PaperBenchX

— End —