Submitted:
05 July 2026
Posted:
07 July 2026
You are already at the latest version
Abstract
Keywords:
Introduction
- Was the model trained and validated on high-quality, consistent datasets?
- Are the statistical analyses appropriate for the data's dynamic range?
- Was the chemical similarity between the training and test sets rigorously reported?
The Pitfalls of Legacy Benchmarks
A New Generation of Realistic Benchmarks
Comparing Data Distributions
| Dataset | LogS IQR |
|---|---|
| Training Set | 2.75 |
| Huuskonen | 2.59 |
| Antiviral | 0.83 |
| Biogen | 0.55 |
Evaluating the Similarity Between Training and Test Sets
Comparative Analysis: Characterizing Realistic Consistently Generated Datasets

Results and Implications
Recommendations for the Field
- Mandatory Similarity Reporting: Authors must move beyond reporting aggregate metrics across the entire test set. We recommend reporting the distribution of maximum Tanimoto similarity between test and training compounds. High-performance claims must be caveated if the test set is highly proximal to the training data. Ideally, performance should be stratified by similarity buckets (e.g., performance on "near-neighbors" vs. "remote scaffolds") to clearly define the model's domain of applicability.
- Transition to Certified Benchmarks: Journals and reviewers should discourage the continued use of flawed legacy datasets (e.g., Delaney, Huuskonen) as the primary evidence for model utility. Instead, models should be evaluated against certified datasets, such as those provided by the Polaris initiative or the ASAP Discovery consortium. These datasets ensure consistent experimental conditions, represent realistic lead-like chemical space, and provide a more rigorous assessment of a model's ability to handle the "twilight zone" of property prediction.
- Standardized Reviewer Checklists: Editorial boards should implement checklists requiring proper statistical rigor. This includes reporting confidence intervals for all metrics, performing null-model comparisons (e.g., against a simple LogP correlation), and verifying that the test set's dynamic range is not artificially inflated. A model that cannot outperform a simple baseline on a narrow-range, realistic dataset should be framed as a specialized tool rather than a general breakthrough.
- Integration of Uncertainty Quantification: In a real-world discovery setting, knowing when a model is likely to be wrong is as important as the prediction itself. Future benchmarks should require models to provide calibrated uncertainty estimates. We recommend that authors demonstrate that their model’s error correlates with its predicted uncertainty, allowing medicinal chemists to prioritize compounds where the model has high confidence and exercise caution where it operates outside its learned chemical space.
Conclusion
Data Availability Statement
References
- Delaney, J.S. ESOL: estimating aqueous solubility directly from molecular structure. J. Chem. Inf. Comput Sci. 2004, 44, 1000–1005. [Google Scholar] [CrossRef] [PubMed]
- Panapitiya, G.; Girard, M.; Hollas, A.; et al. Evaluation of deep learning architectures for aqueous solubility prediction. ACS Omega 2022, 7, 15695–15710. [Google Scholar] [CrossRef] [PubMed]
- Ali, M.; Vanderheiden, S.; Grathwol, C.W.; et al. Advancing aqueous solubility prediction: A machine learning approach for organic compounds using a curated data set. J. Chem. Inf. Model 2025, 65, 8426–8434. [Google Scholar] [CrossRef] [PubMed]
- Huuskonen, J. Estimation of aqueous solubility for a diverse set of organic compounds based on molecular topology. J. Chem. Inf. Comput Sci. 2000, 40, 773–777. [Google Scholar] [CrossRef] [PubMed]
- Fang, C.; Wang, Y.; Grater, R.; et al. Prospective validation of machine learning algorithms for absorption, distribution, metabolism, and excretion prediction: An industrial perspective. J. Chem. Inf. Model 2023, 63, 3263–3274. [Google Scholar] [CrossRef] [PubMed]
- Griffen, E.J.; Boulet, P. ASAP Discovery Center, COVID Moonshot (2024) Enabling equitable and affordable access to novel therapeutics for pandemic preparedness and response via creative intellectual property agreements. Wellcome Open Res. 9, 374. [CrossRef] [PubMed]



Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).