Submitted:
07 November 2025
Posted:
10 November 2025
Read the latest preprint version here
Abstract
Keywords:
Introduction & Motivation
Our Experience
- Loading of incorrect data files.
- Errors in the calculations of test statistics.
- Indexing errors (e.g., in vector operations, during data anonymisation) leading to row-shifted data, resulting in subjects being associated with incorrect data.
- Incorrect custom implementations of complex operations (i.e., nested cross-validation).
- Unintended use of default settings in (external) open-source toolboxes. In this particular case, issues in the code were detected after publication. Once the programming mistake was discovered, the journal was proactively contacted and informed of the situation, analyses were rerun with the correct settings, and the updated results were published as a Corrigendum[18].
A Proposal





Code Review and Analysis Plans
Benefits, Costs, and a Way Forward
Acknowledgements
References
- Soergel, D. A. W. Rampant software errors may undermine scientific results. Preprint at (2015). [CrossRef]
- McConnell, S. Code Complete. (Microsoft Press, Redmond, Washington, 2004).
- Arvanitou, E.-M., Ampatzoglou, A., Chatzigeorgiou, A. & Carver, J. C. Software engineering practices for scientific software development: A systematic mapping study. Journal of Systems and Software 172, 110848 (2021).
- Storer, T. Bridging the Chasm: A Survey of Software Engineering Practice in Scientific Programming. ACM Comput. Surv. 50, 1–32 (2018).
- Woolston, C. Why science needs more research software engineers. Nature (2022). [CrossRef]
- RSECon 2025. RSECon25 https://rsecon25.society-rse.org/ (2025).
- Nordic RSE. Nordic RSE https://nordic-rse.org/.
- nl-rse. nl-rse https://nl-rse.org/.
- Merali, Z. Computational science: ...Error. Nature 467, 775–777 (2010).
- Heaton, D. & Carver, J. C. Claims about the use of software engineering practices in science: A systematic literature review. Information and Software Technology 67, 207–219 (2015).
- Schmidberger, M. & Brügge, B. Need of Software Engineering Methods for High Performance Computing Applications. in 2012 11th International Symposium on Parallel and Distributed Computing 40–46 (2012). doi:10.1109/ISPDC.2012.14.
- Nangia, U. & Katz, D. S. Track 1 Paper: Surveying the U.S. National Postdoctoral Association Regarding Software Use and Training in Research. 4361681 Bytes (2017) doi:10.6084/M9.FIGSHARE.5328442.
- Wilson, G. Software Carpentry: Getting Scientists to Write Better Code by Making Them More Productive. Computing in Science & Engineering 8, (2006).
- Reynolds, M. Science Is Full of Errors. Bounty Hunters Are Here to Find Them. Wired.
- Knight, W. Sloppy Use of Machine Learning Is Causing a ‘Reproducibility Crisis’ in Science. Wired.
- He, X. et al. Author Correction: Temporal dynamics in viral shedding and transmissibility of COVID-19. Nat Med 26, 1491–1493 (2020).
- Ashcroft, P. et al. COVID-19 infectivity profile correction. Swiss Medical Weekly 150, w20336–w20336 (2020).
- Iglesias, S. et al. Hierarchical Prediction Errors in Midbrain and Basal Forebrain during Sensory Learning. Neuron 101, 1196–1201 (2019).
- Marcus, A. Doing the right thing: Psychology researchers retract paper three days after learning of coding error. Retraction Watch https://retractionwatch.com/2019/08/13/doing-the-right-thing-psychology-researchers-retract-paper-three-days-after-learning-of-coding-error/ (2019).
- Dolk, T., Freigang, C., Bogon, J. & Dreisbach, G. RETRACTED: Auditory (dis-)fluency triggers sequential processing adjustments. Acta Psychologica 191, 69–75 (2018).
- Henson, K. E. et al. RETRACTED: Inferring the Effects of Cancer Treatment: Divergent Results From Early Breast Cancer Trialists’ Collaborative Group Meta-Analyses of Randomized Trials and Observational Data From SEER Registries. JCO 34, 803–809 (2016).
- In Sickness and in Health? Physical Illness as a Risk Factor for Marital Dissolution in Later Life - Amelia Karraker, Kenzie Latham, 2015. https://journals.sagepub.com/doi/abs/10.1177/0022146515596354.
- Mandhane, P. J. Notice of Retraction: Hahn LM, et al. Post–COVID-19 Condition in Children. JAMA Pediatrics. 2023;177(11):1226-1228. JAMA Pediatr 178, 1085–1086 (2024).
- Reinhart, C. & Rogoff, K. Errata: “Growth in A Time of Debt”. Harvard University https://carmenreinhart.com/wp-content/uploads/2020/02/36_data.pdf (2013).
- Wiegers, K. Humanizing Peer Reviews. https://web.archive.org/web/20060315135514/http://www.processimpact.com/articles/humanizing_reviews.html (2006).
- Atwood, J. Code Reviews: Just Do It. Coding Horror https://blog.codinghorror.com/code-reviews-just-do-it/ (2006).
- Miller, G. A Scientist’s Nightmare: Software Problem Leads to Five Retractions. Science 314, 1856–1857 (2006).
- MIT. The Missing Semester of Your CS Education. Missing Semester https://missing.csail.mit.edu/.
- Balaban, G., Grytten, I., Rand, K. D., Scheffer, L. & Sandve, G. K. Ten simple rules for quick and dirty scientific programming. PLOS Computational Biology 17, e1008549 (2021).
- Green, R. How To Write Unmaintainable Code. https://www.doc.ic.ac.uk/~susan/475/unmain.html.
- Motivation — Reproducible research documentation. https://coderefinery.github.io/reproducible-research/motivation/.
- Martin, R. C. Clean Code: A Handbook of Agile Software Craftsmanship. (Pearson, Upper Saddle River, NJ, 2008).
- Lynch, M. & published. Rules for Writing Software Tutorials. https://refactoringenglish.com/ (2025).
- Software Carpentry Lessons. Software Carpentry https://software-carpentry.org/lessons/.
- CodeRefinery. https://coderefinery.org/lessons/.
- UNIVERSE-HPC. Byte-sized RSE. UNIVERSE-HPC http://www.universe-hpc.ac.uk/events/byte-sized-rse/.
- Hastings, J., Haug, K. & Steinbeck, C. Ten recommendations for software engineering in research. GigaSci 3, 31 (2014).
- Lynch, M. How to Do Code Reviews Like a Human (Part One). https://mtlynch.io/human-code-reviews-1/ (2017).
- Lynch, M. How to Do Code Reviews Like a Human (Part Two). https://mtlynch.io/human-code-reviews-2/ (2017).
- Lynch, M. How to Make Your Code Reviewer Fall in Love with You. https://mtlynch.io/code-review-love/ (2020).
- Tatham, S. Code review antipatterns. https://www.chiark.greenend.org.uk/~sgtatham/quasiblog/code-review-antipatterns/ (2024).
- Ivimey-Cook, E. R. et al. Implementing code review in the scientific workflow: Insights from ecology and evolutionary biology. Journal of Evolutionary Biology 36, 1347–1356 (2023).
- Rokem, A. Ten simple rules for scientific code review. PLOS Computational Biology 20, e1012375 (2024).
- Vable, A. M., Diehl, S. F. & Glymour, M. M. Code Review as a Simple Trick to Enhance Reproducibility, Accelerate Learning, and Improve the Quality of Your Team’s Research. American Journal of Epidemiology 190, 2172–2177 (2021).
- Gee, T. What to Look for in a Code Review. (Leanpub, 2015).
- What to look for in a code review. eng-practices https://google.github.io/eng-practices/review/reviewer/looking-for.html.
- Common pitfalls and recommended practices. scikit-learn https://scikit-learn.org/stable/common_pitfalls.html.
- Barton, J. awesome-code-review. GitHub https://github.com/joho/awesome-code-review/blob/main/readme.md.
- Bennett, D., Silverstein, S. M. & Niv, Y. The Two Cultures of Computational Psychiatry. JAMA Psychiatry 76, 563 (2019).
- Breiman, L. Statistical Modeling: The Two Cultures (with comments and a rejoinder by the author). Statistical Science 16, 199–231 (2001).
- Nosek, B. A., Ebersole, C. R., DeHaven, A. C. & Mellor, D. T. The preregistration revolution. Proceedings of the National Academy of Sciences 115, 2600–2606 (2018).
- Brodeur, A., Cook, N. M., Hartley, J. S. & Heyes, A. Do Pre-Registration and Pre-Analysis Plans Reduce p-Hacking and Publication Bias? https://www.econstor.eu/handle/10419/262738 (2022).
- Pornprasit, C. & Tantithamthavorn, C. Fine-tuning and prompt engineering for large language models-based code review automation. Information and Software Technology 175, 107523 (2024).
- Yu, Y. et al. Fine-Tuning Large Language Models to Improve Accuracy and Comprehensibility of Automated Code Review. ACM Trans. Softw. Eng. Methodol. 34, 14:1-14:26 (2024).
- Haider, M. A., Mostofa, A. B., Mosaddek, S. S. B., Iqbal, A. & Ahmed, T. Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation. Preprint at (2024). [CrossRef]
- Lu, J., Yu, L., Li, X., Yang, L. & Zuo, C. LLaMA-Reviewer: Advancing Code Review Automation with Large Language Models through Parameter-Efficient Fine-Tuning. in 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE) 647–658 (2023). doi:10.1109/ISSRE59848.2023.00026.
- Kwon, S., Lee, S., Kim, T., Ryu, D. & Baik, J. Exploring LLM-based Automated Repairing of Ansible Script in Edge-Cloud Infrastructures. Journal of Web Engineering 889–912 (2023) doi:10.13052/jwe1540-9589.2263.
- Li, Z. et al. Automating code review activities by large-scale pre-training. in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering 1035–1047 (Association for Computing Machinery, New York, NY, USA, 2022). doi:10.1145/3540250.3549081.
- Fan, L. et al. Exploring the Capabilities of LLMs for Code Change Related Tasks. Preprint at (2024). [CrossRef]
- Martins, G. F., Firmino, E. C. M. & De Mello, V. P. The Use of Large Language Model in Code Review Automation: An Examination of Enforcing SOLID Principles. in Artificial Intelligence in HCI (eds Degen, H. & Ntoa, S.) 86–97 (Springer Nature Switzerland, Cham, 2024). doi:10.1007/978-3-031-60615-1_6.
- Tang, X. et al. CodeAgent: Autonomous Communicative Agents for Code Review. in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 11279–11313 (Association for Computational Linguistics, Miami, Florida, USA, 2024). doi:10.18653/v1/2024.emnlp-main.632.
- Zhao, Z. et al. The Right Prompts for the Job: Repair Code-Review Defects with Large Language Model. Preprint at (2023). [CrossRef]
- Gibney, E. AI tools are spotting errors in research papers: inside a growing movement. Nature (2025). [CrossRef]
- Krag, C. H. et al. Large language models for abstract screening in systematic- and scoping reviews: A diagnostic test accuracy study. 2024.10.01.24314702 Preprint at (2024). [CrossRef]
- Nashaat, M. & Miller, J. Towards Efficient Fine-Tuning of Language Models With Organizational Data for Automated Software Review. IEEE Transactions on Software Engineering 50, 2240–2253 (2024).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).