Perspective: Material discovery is accelerating, but experiments’ reproducibility is not. We need to address that now
Below I explore why reproducibility is a pressing challenge and how I envision a solution framework, focusing on materials and chemical synthesis.
Material discovery is changing the game of innovation more rapidly than ever. Tools - such as machine learning and data driven methods, generative AI, LLMs, AI agents, and high-performance computing - can now predict specialised structures, properties, and process steps with unprecedented speed and accuracy.1-4 Together, these overcome the traditional trade-off in high-throughput computational prediction – where high accuracy can only be achieved over long timeframes, and vice versa.
And yet a more fundamental bottleneck remains unaddressed: the reproducibility of the predictions’ experimental trial, i.e. the reliability with which the same result is obtained when an experiment is repeated under identical or comparable conditions, across time, operators, instruments, and environments. Whether a validated finding can be independently replicated, and whether it holds at scale, are questions that no computational pipeline - however sophisticated - can answer straightaway. This is not due to a limitation of science or individual rigour, but because the scientific community has never built an infrastructure to systematically address - yet alone to measure - reproducibility.
Not having a clear understanding of how reproducible an experimental result is can have great consequences. For example, when a synthesis that worked on milligram scale fails over kilograms, the cost is not just wasted time, but a broken pipeline. In a field expected to deliver solutions and advanced manufacturing materials on tight timelines, this could rapidly place a ceiling on impact. If we do not address this, we risk decoupling discovery from utility - where AI generates thousands of theoretical leads that are fundamentally unmanufacturable because the underlying experimental protocols can’t be translated to the real world.
Reproducibility – challenges and approaches
In experimental physical and life sciences, reproducibility takes two forms for the same process protocol: repeatability under identical conditions, and robustness across varying ones. Failures are often driven by non-obvious variables - reagent source, ambient humidity, glassware geometry, instrument type, or operational settings. One persistent challenge is procedural reporting. Minor environmental factors - sometimes as unexpected and trivial as fluctuations in room temperature, or uncalibrated instruments due to the high level of vibration from noise - can silently alter outcomes. In computational work, analogous issues could arise from dataset heterogeneity, model choices, and hardware variation, although they can be minimised by targeted troubleshooting.
Recent efforts from the community have begun to identify specific failure modes across domains. For example, metal-organic frameworks, despite decades of large-scale use for gas capture, have only recently been examined through a reproducibility lens,6 with the most important factor contributing to variability traced to seemingly trivial factors such as reagent supplier.5 In contrast, in chemical vapour deposition synthesis of 2D materials like graphene, reproducibility issues have persisted for decades before it was shown that the key variable was the internal reactor environment.7 Although researchers had been reporting external parameters, it has been unclear which ones controlled the outcome.
These are important findings, but they remain relatively isolated. Part of the reason is that there is currently no shared framework or standardised format for capturing and comparing reproducibility data across experiments and domains. And beyond infrastructure, there is a deeper issue: no existing academic, funding, or publishing structure directly rewards the generation of reproducibility data as a research output. Furthermore, negative results and failed attempts are rarely reported. This means that the practical knowledge of what fails, and why, tends to disappear rather than accumulate, which in time leads to numerous repeats and reattempts of similar experiments that drive up the timeline and cost.
Mapping reproducibility has always been hard – until now
Addressing these challenges is not trivial. It requires a seamless integration at multiple levels - how researchers think about reporting discoveries, how institutions and funders value reproducibility data, and in how the community builds and contributes to shared infrastructure, (such as a potential unified ‘reproducibility database’). But four recent developments make a first, tractable step more feasible now than it was 5 years ago:
Automated8 and autonomous9 lab platforms can now generate richly annotated experimental data at scale and can be directed toward reproducibility probing rather than just discovery.
Since no standardised format exists for experimental reproducibility data, LLM-based retrieval tools10 could serve as a versatile solution - accommodating datasets that go beyond human-readable text, including spectral images, atomic coordinates, and microscopy colour intensities.
Bayesian optimisation11 in the design of experiments - increasingly augmented by AI - is mature enough to model complex multi-variable spaces with sparse data.
Open-source data bases, such as The Materials Project12 and CCDC13, have already shown that community-wide data deposition works when journals and funders incentivise it.
A general solution
Given these advancements, I see three main pillars that could remedy long term the issue of reproducibility in science:
The infrastructure – a ‘reproducibility database’. A reproducibility database should be a lightweight, shared platform for recording how reliably experiments can be repeated.14 Rather than enforcing a rigid ontology upfront, it would allow researchers to upload results and metadata that can be harmonised post hoc using modern ML and LLM-based tools. The main types of data include: (1) manual repeatability data, (2) distributed or blind studies across sites, and (3) results from automated or autonomous experiments that probe robustness across execution and scale.15 Importantly, non-reproducible and negative results should also be reported.16
Processing and interpreting the metadata. AI-related frameworks, already established in design of experiments, could identify which variables most strongly influence reproducibility. The workflow would rely not only on high-throughput automation, but also on existing reproducibility-tested results from the literature.
Where full standardisation is infeasible - such as in specialised characterisation techniques - reproducibility metrics from related, tractable methods could serve as proxies. In practice, researchers could query the database to estimate protocol robustness, parameter sensitivity, and expected failure modes, while new results feedback to continuously improve the system.
Clearer incentives for the community. Having a structured map of reproducibility would not only strengthen fundamental research but also enable more reliable translation to large-scale applications. For this approach to work, reproducibility data must be rewarded, as collecting and organising reproducibility data - especially negative results – can prove time and resource consuming. Journals and funding bodies could encourage deposition of repeatability and reproducibility metrics alongside publications, analogous to crystal structure or computational data submission. Beyond this, institutions could support dedicated efforts - such as shared automation facilities focused specifically on testing reproducibility across studies.
Hypothesis testing using a pilot experiment
I propose a minimal, testable initial proof-of-concept towards tackling the reproducibility issue. The underlying hypothesis is that experimental reproducibility in materials science is not inherently unpredictable, but structured: failure modes cluster around a limited set of hidden variables. This can be tested by systematically collecting standardised data across an initial small number of experiments, across some laboratories to determine whether these patterns are predictable. An overall schematic is shown in the figure below. I will discuss how I concretely envision this with examples from material science and chemistry, but a similar approach could be integrated for other physical and life sciences.
Schematic of the reproducibility pilot framework.
A tractable pilot will act as a stress test and focus on ~ 5 well-characterised, relatively simple material syntheses, selected to balance simplicity and known variability. Each synthesis will be executed across ~10 independent laboratories by ~3 scientists in each using a minimal shared protocol, as well as, if possible, using high throughput automation robots. Two different scales will be tested – small (milligram) and big (gram). Results will be standardised across unambiguously defined metrics such as yield and purity. Data collected will initially include reagent source, environmental conditions, instrumentation, operator experience, and scale, according to previously reported important ‘black box’ variables but can be expanded towards an additional ‘others’ category. Both successful and failed outcomes will be recorded.
The pilot will initially determine three aspects: whether reproducibility outcomes cluster across sites, if variability can be explained by a small number of known factors, and the timeframes required to collect and analyse this data. The experiment could fail in informative ways - variability may appear randomly, no consistent factors may emerge, or data collection may prove too burdensome. Any outcome would be meaningful: either revealing structure in reproducibility or defining the limits of its predictability. By trailing different experiments, this would assess how a ‘map of reproducibility’ would look like.
Having obtained this dataset, it would next be interesting to train a simple model to predict emergent patterns in the reproducibility of related materials and assess whether these patterns generalise across scales. If the pilot is successful, it could then be expanded towards including different types of experiments (e.g. different material properties testing), or even investigating other scientific domains, such as drug synthesis or biochemistry.
Concluding remarks
Reproducibility is a prerequisite for impact in material discovery and related experimental and computational fields. Without systems to measure it, advances remain difficult to translate into reliable real-world outcomes. A minimal, testable effort to map reproducibility could therefore lay the groundwork for more robust and scalable science. In the era of AI-driven discovery and automation, narrowing the gap between prediction and experimental execution is increasingly important to ensure that this progress can be effectively realised. By building the infrastructure to map these failures today, we move the field from a culture of empirical luck to one of engineering certainty.
References
[1] Artificial Intelligence and Generative Models for Materials Discovery - A Review, A. Merchant, S. Batzner, S. S. Schoenholz, N. Castro-Perez, E. D. Cubuk and B. Kozinsky, arXiv preprint, 2023, arXiv:2508.03278.
[2] Generative AI for crystal structures: a review, Y. Xie, C. Shao, M. Wen and J. Gao, npj Comput. Mater., 2024, 10, 15.
[3] From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery, J. Zhang, T. Liu, X. Tang, Y. Luo and H. Chen, arXiv preprint, 2024, arXiv:2407.01603.
[4] What Artificial Intelligence can do for High-Performance Computing systems?, P. Pochelu, H. Cartiaux and J. Schleich, Eng. Appl. Artif. Intell., 2026, 164, 113248.
[5] Reproducibility in research into metal-organic frameworks in nanomedicine, S. Wuttke and M. Lismont, Commun. Mater., 2023, 4, 52.
[6] Humans, machines, and reproducibility in materials chemistry, Advanced Science News, 2023 (accessed April 2026).
[7] Repeatability and Reproducibility in the Chemical Vapor Deposition of 2D Films: A Physics-Driven Exploration of the Reactor Black Box, S. A. Tawfik, J. J. S. Virdi, M. J. S. Spencer and I. S. Cole, Chem. Mater., 2022, 34, 6750–6762.
[8] Integrating autonomy into automated research platforms, B. P. MacLeod, D. G. Moore and C. P. Berlinguette, Digital Discovery, 2023, 2, 1198–1212.
[9] Inside the ‘self-driving’ lab revolution, R. Brazil, Nature, 2023, 618, 448–450.
[10] Opportunities for retrieval and tool augmented large language models in scientific facilities, A. S. Rosen, S. M. Blau and K. A. Persson, npj Comput. Mater., 2024, 10, 28.
[11] Bayesian optimization with adaptive surrogate models for automated experimental design, H. S. Stein and J. M. Gregoire, npj Comput. Mater., 2019, 5, 75.
[12] Materials Project, A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder and K. A. Persson, APL Mater., 2013, 1, 011002.
[13] Cambridge Crystallographic Data Centre (CCDC), (accessed April 2026).
[14] Reproducibility, sharing and progress in nanomaterial databases, N. Marchese and A. J. Crosby, Beilstein J. Nanotechnol., 2020, 11, 1350–1362.
[15] National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science, The National Academies Press, Washington, DC, 2019.
[16] Publishing negative results is good for science, D. Knight, Microbiol. Today, 2017, 44, 50.



