Preprint
Article

This version is not peer-reviewed.

Data Quality and Benchmarking Rigor in Machine Learning for Drug Discovery: A Perspective on Aqueous Solubility Modeling

Submitted:

05 July 2026

Posted:

07 July 2026

You are already at the latest version

Abstract
As machine learning (ML) becomes increasingly integrated into drug discovery, reliance on legacy datasets and superficial performance metrics threatens to stall genuine progress. This perspective examines common pitfalls in solubility modeling, specifically overreliance on flawed public datasets and insufficient similarity analysis between training and test sets. By comparing performance on "real-world" datasets with consistent experimental conditions, specifically the Biogen and ASAP Discovery sets, we demonstrate that inflated correlations can mask poor generalizability. We propose new guidelines for authors, reviewers, and journals to elevate the standard of ML validation.
Keywords: 
;  ;  

Introduction

Aqueous solubility is a cornerstone of successful drug development, a primary determinant of a compound’s pharmacokinetic profile and clinical efficacy. In early discovery, poor solubility often leads to low oral bioavailability because a drug must be in solution to be absorbed through the gastrointestinal tract. Beyond absorption, it also significantly affects downstream processes, including the reliability of in vitro bioassays, where insoluble compounds can yield misleading "false negative" results or nonspecific aggregation. From a formulation perspective, high solubility simplifies the transition from lead optimization to clinical trials by reducing the need for complex delivery vehicles. Consequently, maintaining adequate solubility is not merely a physicochemical requirement but a strategic necessity to avoid late-stage attrition in the drug development pipeline.
The literature on aqueous solubility prediction is extensive, reflecting a steady evolution from parsimonious empirical models to complex architectures. Early pioneering work, such as the_ESOL method developed by Delaney [1], demonstrated that remarkably effective estimates could be derived from simple linear group-contribution models based on basic molecular descriptors, including LogP and molecular weight. However, as computational power and data availability have increased, the field has pivoted toward_more sophisticated deep learning approaches [2]. These modern methods use Graph Neural Networks (GNNs), transformers, and ensemble techniques to capture intricate structural nuances and nonlinear relationships that simple models often overlook. While these advanced algorithms offer the potential for higher accuracy, they also require high-quality, diverse training data to ensure they generalize across the complex chemical space typical of modern drug discovery.
The rapid proliferation of machine learning papers in drug discovery has not always been matched by a corresponding increase in scientific rigor. When evaluating new predictive models, three fundamental questions must be addressed:
  • Was the model trained and validated on high-quality, consistent datasets?
  • Are the statistical analyses appropriate for the data's dynamic range?
  • Was the chemical similarity between the training and test sets rigorously reported?
A recent study in the Journal of Chemical Information and Modeling (JCIM) [3] highlights a concerning trend. The study harmonized solubility data from four disparate sources to create a 17,935-compound training set, which was then validated against the legacy Huuskonen dataset [4]. While the reported performance metrics were high, with R2 = 0.92 and MAE = 0.40 (Figure 1), these results often fail to translate into practical drug discovery applications due to fundamental flaws in the benchmarking strategy.

The Pitfalls of Legacy Benchmarks

The Huuskonen dataset, while a staple in the field, presents two significant issues that inflate performance metrics:
Unrealistic Dynamic Range: In practical drug discovery, solubility datasets typically span 2-3 log units. The Huuskonen set spans 12 logs. This extreme range mathematically inflates R2 values, making regression tasks appear easier than they are in real-world scenarios.
High Training-Test Proximity: Despite removing duplicates, most compounds in the Huuskonen test set remain highly similar to training set neighbors. At least 50% of the test compounds have a Tanimoto similarity >0.5 with the training set, indicating the model is largely performing interpolation rather than genuine prediction.
Furthermore, assembling "franken-datasets" from dozens of independent papers—each using different experimental protocols introduces significant noise. Inconsistency in assay values across publications is pervasive and undermines ML training, obscuring true model performance.

A New Generation of Realistic Benchmarks

To address the limitations of legacy datasets, there is a critical need for a new generation of realistic benchmarks that better reflect the practical challenges of medicinal chemistry. Unlike historical sets that often feature exaggerated chemical diversity or extreme property distributions, modern benchmarks must span a reasonable dynamic range of solubility values, typically 2 to 3 logs rather than the artificial 10+ log ranges seen in older collections. These benchmarks should strictly reflect the values typically observed in active drug discovery programs, focusing on the "twilight zone" of solubility where decision-making is most difficult. Furthermore, these datasets should contain molecules typical of those encountered in drug discovery programs, featuring complex scaffolds, diverse functional groups, and lead-like molecular weights characteristic of contemporary therapeutic candidates, rather than the simpler, commercially biased compounds prevalent in legacy sets.
The Biogen dataset is derived from the Biogen ADME collection [5] and represents a high-quality industrial dataset curated to evaluate ADME prediction models. It comprises a diverse set of compounds purchased from commercial suppliers and measured within a single, consistent experimental framework at Biogen. Solubility measurements are reported as log-transformed molar values (LogS), ensuring that the variance in the data reflects chemical structure rather than inter-laboratory experimental noise. Notably, the Biogen set exhibits a significantly narrower dynamic range (IQR = 0.55), reflecting the practical challenge of distinguishing compounds with subtle structural differences in a lead-optimization context.
The Antiviral dataset was derived from the ASAP [6] (AI-driven Structure-enabled Antiviral Platform) Discovery consortium, specifically the 2025 Antiviral ADMET Blind Challenge. This dataset stems from a global open-science effort to accelerate the discovery of oral antivirals for pandemic-ready targets, such as the SARS-CoV-2 main protease (Mpro). The compounds were synthesized and tested as part of an active drug discovery campaign, providing a rigorous snapshot of realistic chemical space. With a LogS IQR of 0.83, the Antiviral set serves as a critical benchmark for evaluating whether a model can generalize to novel, therapeutically relevant scaffolds absent from legacy training sets.

Comparing Data Distributions

Figure 2 offers a striking visual comparison of the LogS distributions across the four datasets, highlighting the stark contrast between legacy collections and modern benchmarking standards. The training set assembled by Ali and the Huuskonen dataset exhibit unrealistically broad dynamic ranges, spanning approximately 15 and 10 log units, respectively; such broad distributions often include extreme values rarely encountered in practical settings. In contrast, the Antiviral and Biogen datasets are more realistic, spanning a narrower, more relevant range of 3 to 5 log units. These distributions better reflect subtle property variations and the "twilight zone" solubility challenges typically managed within active drug discovery programs.
Table 1. Interquartile Range (IQR) Comparison.
Table 1. Interquartile Range (IQR) Comparison.
Dataset LogS IQR
Training Set 2.75
Huuskonen 2.59
Antiviral 0.83
Biogen 0.55

Evaluating the Similarity Between Training and Test Sets

Figure 3 illustrates the structural relationships among the evaluation datasets and the training set used in the Ali paper, quantifying chemical overlap using Tanimoto similarity. In cheminformatics, a Tanimoto similarity above 0.35 is typically considered a threshold for significant structural similarity. As shown in the figure, the Huuskonen dataset exhibits substantial overlap with the training set, maintaining a mean Tanimoto similarity of 0.4, suggesting that its performance may be inflated by data redundancy. In contrast, the Antiviral and Biogen datasets are considerably more distinct from the training data, with mean Tanimoto similarities of only 0.25 and 0.28, respectively. These lower scores indicate that the modern benchmarks represent novel chemical space, providing a much more rigorous test of a model's ability to generalize to new therapeutic scaffolds.

Comparative Analysis: Characterizing Realistic Consistently Generated Datasets

To demonstrate the discrepancy in model performance, an XGBoost model using RDKit 2D descriptors was trained on the JCIM dataset and tested against three distinct sets: the Huuskonen set and two "certified" datasets from the Polaris collection (Biogen and Antiviral). The Polaris datasets use consistent experimental methodologies and capture realistic dynamic ranges in lead optimization.
Figure 4. Comparing the performance of a machine learning model trained on the Ali dataset and tested on the Huuskonen, Antiviral, and Biogen datasets. While the model performs well on a dataset with a large dynamic range and similar compounds, its performance on more realistic datasets is poor.
Figure 4. Comparing the performance of a machine learning model trained on the Ali dataset and tested on the Huuskonen, Antiviral, and Biogen datasets. While the model performs well on a dataset with a large dynamic range and similar compounds, its performance on more realistic datasets is poor.
Preprints 221771 g004

Results and Implications

When the XGBoost model was applied to these realistic sets, the correlation between predicted and experimental values vanished, yielding results no better than a simple null model (dataset mean). This collapse in performance illustrates that high R2 values on legacy sets are often a byproduct of a broad dynamic range and high chemical similarity rather than genuine predictive power.
The results in Figure 2 reveal a stark divergence in model performance, underscoring the deceptive nature of legacy benchmarks. When evaluated on the Huuskonen dataset, the model achieves a deceptively high correlation (R2 = 0.92). However, as shown in the structural similarity analysis in Figure 3, this "success" is likely a byproduct of the dataset's unrealistic 10-log dynamic range and its high structural redundancy with traditional training sets. In such cases, the model is not necessarily learning the underlying physics of solubility but instead performing a sophisticated form of pattern matching against near-neighbors.
In stark contrast, the correlation virtually vanishes when the same model is applied to the Antiviral (R2 =−3.16) and Biogen (R2 =−1.59) datasets. These results, which perform no better than a null model (predicting the dataset mean), indicate a "performance collapse" that is far more representative of the challenges in real-world drug discovery. Because these datasets span a realistic 3–5 log range and feature novel, lead-like chemical space, they expose the limitations of models overfit to the broad, simple distributions of the past. This discrepancy implies that many currently published solubility methods may lack the resolution needed to guide medicinal chemistry decisions in the "twilight zone" of lead optimization, where distinguishing compounds with subtle structural differences is paramount.

Recommendations for the Field

To bridge the gap between retrospective benchmark performance and prospective discovery success, we propose the following actions for the research community:
  • Mandatory Similarity Reporting: Authors must move beyond reporting aggregate metrics across the entire test set. We recommend reporting the distribution of maximum Tanimoto similarity between test and training compounds. High-performance claims must be caveated if the test set is highly proximal to the training data. Ideally, performance should be stratified by similarity buckets (e.g., performance on "near-neighbors" vs. "remote scaffolds") to clearly define the model's domain of applicability.
  • Transition to Certified Benchmarks: Journals and reviewers should discourage the continued use of flawed legacy datasets (e.g., Delaney, Huuskonen) as the primary evidence for model utility. Instead, models should be evaluated against certified datasets, such as those provided by the Polaris initiative or the ASAP Discovery consortium. These datasets ensure consistent experimental conditions, represent realistic lead-like chemical space, and provide a more rigorous assessment of a model's ability to handle the "twilight zone" of property prediction.
  • Standardized Reviewer Checklists: Editorial boards should implement checklists requiring proper statistical rigor. This includes reporting confidence intervals for all metrics, performing null-model comparisons (e.g., against a simple LogP correlation), and verifying that the test set's dynamic range is not artificially inflated. A model that cannot outperform a simple baseline on a narrow-range, realistic dataset should be framed as a specialized tool rather than a general breakthrough.
  • Integration of Uncertainty Quantification: In a real-world discovery setting, knowing when a model is likely to be wrong is as important as the prediction itself. Future benchmarks should require models to provide calibrated uncertainty estimates. We recommend that authors demonstrate that their model’s error correlates with its predicted uncertainty, allowing medicinal chemists to prioritize compounds where the model has high confidence and exercise caution where it operates outside its learned chemical space.

Conclusion

To address the limitations of legacy datasets, there is a critical need for a new generation of realistic benchmarks that better reflect the practical challenges of medicinal chemistry. Unlike historical sets that often feature exaggerated chemical diversity or extreme property distributions, modern benchmarks must span a reasonable dynamic range of solubility values, typically covering 2 to 3 logs rather than the artificial 10+ log ranges seen in older collections. These benchmarks should strictly represent the values typically observed in active drug discovery programs, focusing on the "twilight zone" of solubility where decision-making is most difficult. Furthermore, it is essential that these datasets contain molecules typical of those encountered in drug discovery programs, featuring the complex scaffolds, diverse functional groups, and lead-like molecular weights characteristic of contemporary therapeutic candidates, rather than the simpler, commercially biased compounds prevalent in legacy sets.
The path toward more reliable machine learning models in drug discovery requires a fundamental shift in focus from algorithmic complexity to rigorous data quality and contextual relevance. By adopting benchmarks that mirror the chemical space and property distributions found in the laboratory, we can move beyond "solved" legacy problems and develop predictive tools that offer genuine utility to medicinal chemists. Ultimately, the goal is to ensure that high performance on a benchmark translates into successful decision-making in the clinic, rather than remaining a computational artifact of unrepresentative data.
Publishing high R2 values on legacy datasets is no longer a sufficient indicator of model utility. By adopting more rigorous validation standards and using high-quality, narrow-range datasets, the computational chemistry community can ensure that machine learning genuinely advances our ability to predict complex molecular properties such as aqueous solubility.

Data Availability Statement

All datasets and code used in this paper are available on GitHub at https://github.com/PatWalters/solubility_evaluation. This perspective is dedicated to the memory of Terry Stouch, a passionate protector of good data.

References

  1. Delaney, J.S. ESOL: estimating aqueous solubility directly from molecular structure. J. Chem. Inf. Comput Sci. 2004, 44, 1000–1005. [Google Scholar] [CrossRef] [PubMed]
  2. Panapitiya, G.; Girard, M.; Hollas, A.; et al. Evaluation of deep learning architectures for aqueous solubility prediction. ACS Omega 2022, 7, 15695–15710. [Google Scholar] [CrossRef] [PubMed]
  3. Ali, M.; Vanderheiden, S.; Grathwol, C.W.; et al. Advancing aqueous solubility prediction: A machine learning approach for organic compounds using a curated data set. J. Chem. Inf. Model 2025, 65, 8426–8434. [Google Scholar] [CrossRef] [PubMed]
  4. Huuskonen, J. Estimation of aqueous solubility for a diverse set of organic compounds based on molecular topology. J. Chem. Inf. Comput Sci. 2000, 40, 773–777. [Google Scholar] [CrossRef] [PubMed]
  5. Fang, C.; Wang, Y.; Grater, R.; et al. Prospective validation of machine learning algorithms for absorption, distribution, metabolism, and excretion prediction: An industrial perspective. J. Chem. Inf. Model 2023, 63, 3263–3274. [Google Scholar] [CrossRef] [PubMed]
  6. Griffen, E.J.; Boulet, P. ASAP Discovery Center, COVID Moonshot (2024) Enabling equitable and affordable access to novel therapeutics for pandemic preparedness and response via creative intellectual property agreements. Wellcome Open Res. 9, 374. [CrossRef] [PubMed]
Figure 1. A misleading correlation between experimental and predicted aqueous solubility reported in a 2025 paper by Ali (reference 1). The high correlation is primarily attributable to the 10-log dynamic range and the high similarity between the training and test sets.
Figure 1. A misleading correlation between experimental and predicted aqueous solubility reported in a 2025 paper by Ali (reference 1). The high correlation is primarily attributable to the 10-log dynamic range and the high similarity between the training and test sets.
Preprints 221771 g001
Figure 2. A comparison of the data LogS distributions for two widely used literature datasets and two more realistic drug discovery datasets.
Figure 2. A comparison of the data LogS distributions for two widely used literature datasets and two more realistic drug discovery datasets.
Preprints 221771 g002
Figure 3. A comparison of Tanimoto similarity between the Huuskonnen, Antiviral, and Bigoen datasets and the Ali training set. The boxplot shows that the Huuskonnen dataset is very similar to the training set, whereas the others are not.
Figure 3. A comparison of Tanimoto similarity between the Huuskonnen, Antiviral, and Bigoen datasets and the Ali training set. The boxplot shows that the Huuskonnen dataset is very similar to the training set, whereas the others are not.
Preprints 221771 g003
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings