Preprint
Review

This version is not peer-reviewed.

Machine Learning for Predicting Colloidal Stability and Aggregation Risk in Biopharmaceutical Formulations: Current Applications, Challenges, and Future Directions

Submitted:

23 June 2026

Posted:

24 June 2026

You are already at the latest version

Abstract
Colloidal instability remains a dominant cause of product failure, manufacturing attrition, and safety risk across biopharmaceutical modalities including monoclonal antibodies (mAbs), bispecific antibodies, antibody-drug conjugates (ADCs), and mRNA-lipid nanoparticle systems. Protein aggregation is a critical quality attribute (CQA) under ICH Q6B, linked to immunogenicity, reduced potency, and adverse patient outcomes. Despite the transformative impact of machine learning (ML) on protein structure prediction and molecular design, its application to formulation-dependent colloidal stability prediction remains fragmented, poorly benchmarked, and largely disconnected from regulatory frameworks. This review systematically examines ML approaches for predicting aggregation propensity, viscosity, solubility, liquid-liquid phase separation, and shelf-life across biopharmaceutical modalities. We critically assess experimental data sources, feature engineering strategies, and ML architectures spanning classical models, deep learning, graph neural networks, and protein language models, alongside the emerging role of explainable AI (XAI). No standardised, cross-modality ML benchmarking framework for colloidal stability currently exists -- a gap that constrains generalisation, reproducibility, and regulatory acceptance. Principal unresolved challenges include dataset scarcity, label noise, external validation deficits, and proprietary data silos. A decade roadmap for integrating physics-informed ML, autonomous formulation laboratories, and foundation models into next-generation biologics development is proposed.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

1.1. The Expanding Biologics Landscape

The global biologics market reached approximately USD 417 billion in 2024 and is projected to approach USD 1.1 trillion by 2035, growing at a compound annual growth rate (CAGR) of approximately 9–10% [1]. Monoclonal antibodies (mAbs) continue to dominate by revenue, commanding a market share exceeding 65% of approved biologics as of 2025 [2]. Yet the therapeutic pipeline is rapidly diversifying: bispecific and multispecific antibodies, antibody-drug conjugates (ADCs), mRNA-lipid nanoparticle (mRNA-LNP) platforms, recombinant fusion proteins, and cell and gene therapy vectors are each growing at double-digit rates [3]. In 2024 alone, the FDA approved 13 monoclonal antibodies and 15 biological license applications, while the Center for Biologics Evaluation and Research (CBER) added 23 new BLA approvals in 2023 [4]. This expansion reflects decades of investment in protein engineering, structural biology, and advanced manufacturing, and signals that biologics will remain the dominant driver of pharmaceutical innovation for the foreseeable future.
However, the increasing structural complexity of modern biopharmaceuticals introduces proportionally greater formulation challenges. Unlike small-molecule drugs, which are typically thermodynamically stable crystalline solids, large-molecule therapeutics are intrinsically metastable in aqueous solution. They are susceptible to a diverse array of physical and chemical degradation pathways including aggregation, fragmentation, deamidation, oxidation, and disulfide scrambling. Among these, aggregation is uniquely consequential: it simultaneously diminishes product potency, challenges analytical characterization, and introduces immunogenicity risk for patients [5]. The aggregate burden in a drug product is not merely a quality metric but a patient safety concern that regulators treat with increasing scrutiny.

1.2. Aggregation as a Critical Quality Attribute: Safety, Immunogenicity, and Regulatory Significance

Protein aggregation in biopharmaceuticals spans a continuum from reversible, native-like self-association through soluble oligomers to irreversible particulates ranging from sub-visible (0.1–100 µm) to visible (>100 µm) in size [5]. The immunological consequences of aggregated material are well documented: aggregated protein structures activate innate immune pathways, enhance antigen presentation, and promote T-cell-dependent anti-drug antibody (ADA) formation through mechanisms involving dendritic cell activation and Toll-like receptor signalling [6]. Lundahl et al. (2021) provided a comprehensive mechanistic review of how aggregate physicochemical properties such as size, surface hydrophobicity, and structural epitope exposure modulate both innate and adaptive immune responses [7]. Clinical consequences of ADA formation range from altered pharmacokinetics and attenuated therapeutic efficacy to severe adverse events including anaphylaxis and, in the case of erythropoietin-associated pure red cell aplasia, life-threatening cross-reactive autoimmunity [8].
The regulatory framework places protein aggregates squarely within the category of critical quality attributes (CQAs). ICH Q6B establishes specifications for biotechnological and biological products including purity, aggregate content, and particulate matter, and requires that manufacturers define and control CQAs throughout development [9]. EMA guidelines on monoclonal antibodies additionally mandate structural and physicochemical characterization covering aggregation state, glycosylation, deamidation, and oxidation [10]. The FDA’s 2021 Guidance for Industry on inspection of injectable products for visible particulates reflects heightened regulatory attention to sub-visible particle thresholds and their relationship to patient outcomes [11]. From a Quality by Design (QbD) perspective embedded in ICH Q8(R2), formulation design must link material attributes and process parameters to CQA outcomes through a sound mechanistic and empirical understanding of the formulation design space [12].

1.3. The Economic and Operational Burden of Instability

Formulation-driven product failures carry enormous economic consequences. The average capitalized research and development cost to bring a biopharmaceutical to market has been estimated at USD 2.3–4.6 billion per approved product when accounting for the cost of attrition [13]. Process development and manufacturing activities account for 13–17% of the total R&D budget, and costs rise steeply under worst-case clinical success rate scenarios [14]. Late-stage discovery of colloidal instability is particularly costly: aggregation or viscosity issues identified during Phase II or Phase III manufacturing scale-up can trigger full reformulation campaigns, analytical method redevelopment, additional stability studies, and regulatory amendment submissions, each adding months to development timelines and tens to hundreds of millions of dollars to programme budgets [15]. Beyond direct costs, instability during distribution and cold chain excursions can result in product recalls, patient supply disruptions, and reputational damage [16].
The emerging complexity of next-generation modalities amplifies these risks further. Bispecific antibodies, with their asymmetric chain architectures and increased conformational heterogeneity, exhibit colloidal behaviours that differ substantially from their conventional IgG1 counterparts and for which existing platform formulation approaches frequently fail [17]. ADCs introduce hydrophobic drug-linker payloads that perturb the protein surface charge landscape and aggregation propensity in ways that are not predictable from antibody sequence alone [8]. mRNA-LNP formulations, despite their fundamentally different physicochemical architecture, face analogous stability challenges related to ionisable lipid phase behaviour, encapsulation efficiency, and particle aggregation kinetics [18]. The formulation space for these modalities is vast, the experimental cost of systematic exploration is prohibitive, and the consequences of poor stability decisions at the development stage are severe.

1.4. Machine Learning as an Emerging Transformative Tool in Formulation Science

Against this backdrop, machine learning has emerged as a potentially transformative tool for formulation science. The success of AlphaFold2 and related structure prediction systems has demonstrated that ML models trained on large protein sequence and structure datasets can extract non-obvious patterns predictive of biophysical properties [19]. In the formulation domain, early demonstrations showed that artificial neural networks could predict long-term monomer retention of therapeutic proteins from accelerated stability data and biophysical inputs with sufficient accuracy to reduce the number of formulation conditions requiring physical testing [20]. Lai et al. (2022) extended this paradigm by training logistic regression, support vector machine, k-nearest neighbour, and decision tree classifiers on a combined dataset of 47 mAbs incorporating sequence-derived descriptors and molecular dynamics features to predict aggregation rates and viscosity at high concentration [21]. The PIPPI consortium (Protein-excipient Interactions and Protein-Protein Interactions), an EU Horizon 2020 Innovative Training Network, assembled one of the most systematically curated open formulation screening databases, enabling training of ML models that link biophysical measurements to colloidal and conformational stability outcomes [22].
Beyond prediction, Bayesian optimisation approaches have demonstrated the ability to identify optimal formulation conditions for thermostability in fewer than 25 experiments - a fraction of the experiments required by classical design of experiments (DoE) approaches - providing a compelling proof-of-concept for ML-driven autonomous formulation design [23]. PROPERMAB and related integrative developability prediction frameworks now combine sequence features, structural descriptors, and biophysical measurements to generate multi-attribute stability profiles that can support early candidate triage [24]. Interpretable ML approaches employing SHAP (SHapley Additive exPlanations) analysis have begun to link sequence-level descriptors such as variable domain charge asymmetry, hydrophobic patch exposure, and complementarity-determining region (CDR) composition to viscosity and self-association behaviour, providing mechanistic hypotheses testable at the bench [25].

1.5. The Knowledge Gap This Review Addresses

Despite this momentum, the field remains at an early stage of maturity. Existing reviews address protein aggregation mechanisms [5], biophysical characterisation platforms [26], and general applications of ML to biopharmaceutical process development [27]. However, no review has systematically and critically assessed ML approaches specifically for predicting formulation-dependent colloidal instability across the full range of modern biopharmaceutical modalities. Critical gaps persist in dataset standardisation, cross-modality benchmarking, external validation, physics-informed model design, and regulatory acceptance. The fragmentation of the field across disciplines - colloid science, structural biology, pharmaceutical sciences, and computational chemistry - has impeded the kind of community consensus that enabled rapid progress in, for example, protein structure prediction.
This review seeks to fill that gap. We present a systematic and critical examination of: (i) the physicochemical and thermodynamic principles underlying colloidal instability in biopharmaceuticals; (ii) the experimental methods generating stability data amenable to ML; (iii) available data sources and their limitations; (iv) ML architectures and their strengths and weaknesses in the formulation context; (v) predictive targets spanning aggregation, viscosity, solubility, liquid-liquid phase separation (LLPS), and shelf-life; (vi) the role and limitations of explainable AI; (vii) modality-specific applications and challenges for mAbs, bispecific antibodies, ADCs, protein nanoparticles, and mRNA-LNP systems; (viii) regulatory considerations under current ICH and FDA/EMA frameworks; (ix) unresolved challenges and controversies; and (x) a roadmap for the next decade. Our aim is not to survey what has been done but to articulate what must be done - and why - for this emerging field to deliver its transformative potential.
Predictive models are only as informative as the data on which they are trained. A rigorous understanding of the experimental methods that generate colloidal stability data is therefore not merely of academic interest but is a prerequisite for rational ML model design. The type of experimental output determines the nature of the prediction target, the feature engineering strategy, and the interpretation of model behaviour. This section jointly reviews the principal analytical methods used to assess colloidal stability in biopharmaceutical development and the publicly or semi-publicly available datasets derived from these methods, with a unified focus on what each contributes - and what it withholds - from a machine learning perspective. Table 1 summarises the principal colloidal instability mechanisms addressed throughout this review, with their physical driving forces, relevant biopharmaceutical modalities, primary formulation interventions, and current ML predictability status.

2. Experimental Assessment of Colloidal Stability and Data Sources for Machine Learning

2.1. From Analytical Measurement to ML Feature: Framing the Problem

Colloidal stability in biopharmaceutical formulations is a multivariate phenomenon governed by the interplay of protein-protein interactions, protein-solvent interactions, and the thermodynamic environment imposed by buffer composition, pH, ionic strength, excipient identity and concentration, protein concentration, and temperature history. No single analytical measurement captures this complexity in full. Each technique probes a specific aspect of the stability landscape and generates one or more numerical outputs - potential features - that, in combination, may approximate a sufficiently rich description of the system for predictive modelling [28]. The central challenge for ML practitioners is not the absence of measurements but rather the heterogeneity, sparsity, and non-interoperability of the resulting data across studies, molecules, and laboratories [29].
A useful conceptual distinction separates colloidal stability parameters - outputs that directly report on intermolecular interaction strength and aggregation propensity, such as the diffusion interaction parameter (kD), the second osmotic virial coefficient (B22/A2), and zeta potential - from conformational stability parameters - outputs that report on the structural integrity of the protein itself, such as the thermal melting temperature (Tm) and calorimetric enthalpy (ΔH) - and from aggregate quantification parameters - outputs that measure the product of instability events already having occurred, such as SEC monomer content (%), sub-visible particle counts per mL, and hydrodynamic radius shifts. ML models vary substantially in their reliance on each category: early-stage developability models typically use colloidal and conformational inputs to predict future aggregation risk, while formulation optimisation models often use aggregate quantification as the primary target variable [20]. Conflating these categories without explicit consideration is a common source of model misinterpretation in the literature.

2.2. Dynamic Light Scattering: The Workhorse of High-Throughput Colloidal Screening

Dynamic light scattering (DLS) measures the time-autocorrelation of scattered light intensity fluctuations arising from Brownian motion of particles in solution, enabling calculation of the mutual diffusion coefficient (Dm) and, via the Stokes-Einstein equation, the hydrodynamic radius (Rh). In its concentration-series implementation, DLS yields the diffusion interaction parameter kD, defined by Dm = D0(1 + kDC), where D0 is the self-diffusion coefficient at infinite dilution [30]. Positive kD values indicate net repulsive interactions, correlating with colloidal stability; increasingly negative values indicate attractive interactions and predict elevated aggregation propensity [31]. A threshold of kD > −12.4 mL/g has been proposed as a discriminator for high colloidal stability, though this threshold is empirically derived from a limited IgG1 dataset and should not be uncritically generalised across protein classes [32].
High-throughput DLS (HT-DLS) platforms, exemplified by the Wyatt DynaPro Plate Reader operating in 96-, 384-, or 1536-well format, have transformed formulation screening throughput from tens to thousands of conditions per week [33]. Simultaneous measurement of Rh, PDI, kD, and Tagg (onset aggregation temperature by thermal ramp) in a single automated experiment provides a compact but information-rich feature vector for ML. The PIPPI consortium (Protein-excipient Interactions and Protein-Protein Interactions), funded under EU Horizon 2020, constructed one of the most systematically organised open formulation screening databases using HT-DLS as its primary throughput tool, accumulating over 3,500 formulation condition measurements across approximately 50 therapeutic proteins [22]. Gentiluomo et al. (2020) leveraged this dataset to train artificial neural networks predicting long-term monomer retention from biophysical inputs including kD, Rh, Tagg, and Tm with sufficient accuracy to substantially reduce the number of formulation conditions requiring physical stability testing [20].
Despite its utility, DLS has important limitations as an ML data source. Intensity-weighted size distributions disproportionately emphasise large aggregates, and the technique cannot resolve co-existing populations of similar size. kD measurement is sensitive to protein concentration accuracy, buffer preparation, and temperature equilibration, introducing systematic errors that propagate into training labels. Crucially, DLS provides no direct information on aggregate morphology, covalent versus non-covalent nature, or the structural mechanism of instability - information that would substantially enrich physics-informed ML models but is absent from most existing training datasets [28].

2.3. Size-Exclusion Chromatography: The Regulatory Standard for Aggregate Quantification

Size-exclusion chromatography coupled to ultraviolet detection at 280 nm (SEC-HPLC) remains the regulatory standard for quantifying monomeric and aggregated species in biopharmaceutical drug substances and drug products, reflected in its universal inclusion in ICH Q6B characterisation panels and pharmacopoeial methods. SEC resolves high-molecular-weight (HMW) species, monomers, and low-molecular-weight (LMW) fragments based on hydrodynamic volume, providing percentage values for each population that constitute the principal long-term stability monitoring endpoint in shelf-life studies [34]. When coupled to multi-angle light scattering (MALS) and differential refractive index (dRI) detectors, SEC-MALS yields absolute molecular weight distributions and enables quantitative characterisation of aggregate size without column-dependent calibration artifacts [35].
For ML, the key output of SEC is the time-dependent loss of monomer percentage - often expressed as a first-order or pseudo-first-order aggregation rate constant derived from accelerated stability studies at 40 °C or 45 °C - which serves as the primary prediction target in the majority of published models. Gentiluomo et al. used 24-month real-time stability data as ground truth; Lai et al. (2022) used 45 °C accelerated aggregation rates derived from three time points [21]. A recurring limitation is that accelerated and real-time stabilities are not equivalent: the Arrhenius assumption underlying accelerated stability predictions fails for many proteins when aggregation is nucleation-limited or when multiple concurrent degradation pathways contribute to monomer loss at elevated temperatures [16]. Training ML models on accelerated stability labels without explicitly testing the temperature-time extrapolation assumption introduces systematic label bias that is rarely acknowledged in the literature.

2.4. Thermal Analysis: nanoDSF and DSC as Conformational Stability Features

Nano differential scanning fluorimetry (nanoDSF) monitors the shift in intrinsic tryptophan/tyrosine fluorescence emission wavelength (350/330 nm ratio) as a protein thermally unfolds, determining the apparent melting temperature (Tm) without extrinsic dye requirements [36]. Its key advantages for ML-directed screening are low sample consumption (~10 µL per capillary), high throughput via capillary array format, and simultaneous measurement of both Tm (fluorescence channel) and Tagg (backscatter channel, reporting particle formation onset) in a single thermal ramp experiment [37]. Multiple Tm values from multi-domain proteins such as mAbs - typically Tm1 (CH2 domain), Tm2 (Fab), and Tm3 (CH3 domain) - provide domain-resolved conformational stability features directly linkable to aggregation-prone region (APR) exposure upon unfolding [38].
Differential scanning calorimetry (DSC or µDSC) measures heat capacity differences during thermal unfolding, providing thermodynamically rigorous parameters including domain-specific Tm, calorimetric enthalpy (ΔH), and reversibility fraction. Unlike nanoDSF, DSC is sensitive to the thermodynamic two-state or multi-state nature of unfolding transitions and can distinguish cooperative from non-cooperative domain denaturation [39]. However, throughput is orders of magnitude lower than nanoDSF and sample consumption substantially higher, restricting DSC to confirmatory measurements rather than primary screening. In ML datasets, DSC Tm values are frequently reported alongside nanoDSF Tm with moderate but imperfect correlation - a discrepancy attributable to the different physical properties measured (global heat flow versus local fluorophore environment) - and this difference is rarely propagated explicitly into ML uncertainty estimates [40].

2.5. Sub-Visible Particle Characterisation: MFI, NTA, and the Emerging ML Interface

Sub-visible particles (SVPs) in the 0.1–100 µm range represent a critical quality attribute with direct regulatory implications under USP <787> (subvisible particle testing for injections), USP <788>, and EMA guidance on SVPs in parenteral products. Micro-flow imaging (MFI) combines digital microscopy with microfluidic sample delivery to image individual particles, generating per-particle morphological descriptors - equivalent circular diameter (ECD), aspect ratio, circularity, transparency, and intensity - alongside total count and size distribution [41]. This image-based output is uniquely suited to convolutional neural network (CNN) approaches for particle classification.
Wang et al. (2023, Merck & Co.) demonstrated that CNNs trained on MFI images could achieve highly accurate automated classification of biopharmaceutical sub-visible particles including silicone oil droplets, proteinaceous aggregates, and glass particles, with performance approaching expert human labelling on validation datasets [42]. Lopez-Del Rio et al. (2024, Boehringer Ingelheim) extended this approach using unsupervised learning and label consistency analysis to address the critical problem of noisy training labels in FIM-based ML classification, demonstrating that label inconsistency - arising from subjective expert classification of morphologically ambiguous particles - substantially degrades model generalisation [43]. This finding has broad implications: in formulation stability ML, the quality of training labels is often the binding constraint, not the sophistication of the model architecture. Nanoparticle tracking analysis (NTA) complements MFI in the 50–500 nm range, providing particle size distributions and concentrations at the single-particle level, but its reproducibility across instruments and laboratories is substantially lower than MFI, limiting its utility as a reliable ML training label source [44].

2.6. Developability Assays as Structured ML Features

Beyond classical stability assays, a panel of higher-throughput developability assays has been established to characterise biopharmaceutical candidates at the earliest stages of discovery with the explicit intent of generating ML-compatible structured data. Affinity capture self-interaction nanoparticle spectroscopy (AC-SINS) measures mAb self-association by the aggregation-induced spectral red-shift of gold nanoparticles coated with Fc-binding antibodies; self-interaction chromatography (SMAC/ISAC) measures retention times that correlate with self-association propensity; hydrophobic interaction chromatography (HIC) retention time serves as a proxy for surface hydrophobicity and ADC drug load heterogeneity [45]. The landmark 2017 study by Jain et al. systematically characterised 137 clinical-stage IgG1 antibodies across 12 such assays, generating the first large-scale open-access mAb biophysical reference dataset [45]. This dataset has since become one of the most widely cited benchmarks in computational antibody developability, and its 2023 update with extended clinical outcome annotations provides a richer resource for outcome-linked ML model development [46]. Arsiwala et al. (2025) built upon this foundation, reporting a high-throughput platform that generated structured developability data on 246 antibodies across 10 assays in a tidy-data format specifically designed for ML ingestion [47].

2.7. The Data Landscape: Open Resources, Proprietary Silos, and the Benchmarking Deficit

The availability of curated, open-access datasets is a defining constraint on ML progress in colloidal stability prediction. A frank audit of the field reveals a stark asymmetry: the most information-rich datasets - those encompassing full formulation screening panels across large molecule portfolios, manufacturing process variants, and long-term real-time stability outcomes - reside exclusively within the proprietary data systems of large biopharmaceutical companies and are unavailable to academic or independent researchers [14]. As explicitly noted in Kidziński et al. (2023), the lack of publicly available benchmark datasets for bioprocess ML is a fundamental bottleneck; unlike computer vision or natural language processing, where large-scale open datasets catalysed rapid algorithmic progress, biopharmaceutical formulation ML operates on datasets that are typically small (tens to low hundreds of molecules), narrow (single modality, single company platform), and incompletely annotated [48]. The available data landscape is illustrated in Figure 3.
The consequence is that most published models are validated only internally - on held-out subsets of the same narrow training distribution - rather than externally on independent molecular series from different discovery programmes. Khetan et al. (2022) reviewed the state of computational developability prediction and concluded that cross-study reproducibility of biophysical predictions remains low, in part because different laboratories measure the same property (e.g., self-association) with different assays (AC-SINS vs. SMAC vs. kD-DLS) that are only moderately correlated [49]. This assay heterogeneity is itself a critical data quality problem: pooling datasets from different sources that use different measurement principles for the same nominal target introduces label noise that is structurally correlated with laboratory identity rather than protein properties [46].
Emerging efforts seek to address this deficit. The ExPreSo platform (Excipient Prediction Software, published 2025) used a curated database of 335 FDA-approved drug-product formulations to train ML models for excipient recommendation, demonstrating how public formulation records can support systematic ML-driven formulation design [50]. Rosace et al. (2023) combined sequence-based solubility predictions with automated expression and biophysical characterisation to optimise the solubility and conformational stability of antibodies and proteins [51]. Importantly, none of the currently available open datasets encompasses mRNA-LNP formulations, ADCs, bispecific antibody formulation screens, or cell therapy vectors - the modalities where colloidal instability risk is highest and ML guidance most urgently needed. Table 2 summarises the principal analytical methods and their ML-relevant outputs; a comprehensive inventory of publicly available datasets with their scope and access status is provided in Supplementary Table S1.

2.8. Data Quality, Label Noise, and the Under-Appreciated Reproducibility Problem

The quality of any ML model is bounded by the quality of its training labels, and formulation stability labels are burdened by measurement uncertainties that are frequently underreported in published datasets. Accelerated stability labels - the most common target in published models - carry inter-laboratory coefficients of variation of 5–15% for SEC monomer content under supposedly identical conditions, arising from sources including protein concentration measurement error, SEC column lot variation, temperature control precision in stability chambers, and baseline drift in UV detectors [16]. For rare events such as visible particle formation or opalescence onset, label sparsity introduces severe class imbalance that standard model evaluation metrics (accuracy, AUC) do not adequately penalise [28].
A further underappreciated issue is assay cross-comparability. Jain et al. (2023) demonstrated substantial discordance between nominally equivalent self-association measurements performed on the same 137-mAb panel by different groups using AC-SINS and kD-DLS, with Spearman correlation coefficients of 0.4–0.6 between methods ostensibly measuring the same property [46]. Pooling such data for ML without explicit assay harmonisation introduces structured label noise that inflates model uncertainty in a manner indistinguishable from genuine molecular variability. Until the field establishes standardised inter-laboratory ring studies with defined reference materials - analogous to the NIST reference mAb (NISTmAb, RM 8671) for method qualification but extended to ML-relevant quantitative feature generation - cross-study ML benchmarking will remain methodologically fragile [52].
The application of machine learning to protein formulation draws on a spectrum of methods that differ in their data requirements, inductive biases, and interpretability profiles. Selecting the right method is not a matter of following a single dominant paradigm but of matching algorithmic assumptions to the available data and the physical problem at hand. This section reviews the principal classes of ML algorithms that have been applied to colloidal stability and aggregation-related prediction tasks, with particular attention to the characteristics of training data they require, the types of features they consume, and the practical limitations that constrain their utility in a formulation development context.

3. Machine Learning Approaches for Colloidal Stability Prediction

3.1. Feature Engineering: The Bridge Between Protein Chemistry and ML Input

Before examining individual algorithms, it is useful to characterise the feature spaces on which they operate. The choice of molecular representation has a more decisive effect on model performance than the choice of algorithm for most published formulation datasets [49]. Three broad categories of features are in common use. The overall ML workflow is summarised in Figure 2.
Sequence-derived features encompass amino acid composition, isoelectric point (pI) of the full antibody, Fv, CDRs, and individual domains, net charge at formulation pH, hydrophobicity indices (e.g., spatial aggregation propensity score, SAP; CamSol intrinsic solubility score), CDR loop lengths, and sequence-based aggregation propensity tools such as AGGRESCAN3D, Zyggregator, and TANGO [53]. These features are cheap to compute from sequence alone and are therefore well suited to early-stage candidate triage. The pI and net charge are particularly recurrent in the literature: Schwenger et al. (2021) demonstrated that net Fv charge and a viscosity index derived from amino acid composition of the Fv region were the dominant predictors of high-concentration viscosity across 27 FDA-approved mAbs, outperforming more complex structural descriptors [54]. Mock et al. (2023) confirmed pI as the strongest correlate of colloidal stability (kD) across 83 scaffold-consistent mAbs at Amgen, with a Spearman r of 0.85 [55].
Structure-derived features are calculated from experimentally determined or computationally predicted 3D structures and include solvent-accessible surface area (SASA) patches, electrostatic potential maps, dipole moments, spatial charge asymmetry, CDR loop conformation-dependent hydrophobic exposure, and molecular dynamics (MD)-derived descriptors such as B-factors, root-mean-square fluctuations, and protein-protein interaction energies from coarse-grained simulations [24]. Lai et al. (2022) combined MD-derived features of full-length mAb models with sequence descriptors to train logistic regression, SVM, k-nearest neighbour, and decision tree classifiers on a 47-mAb AstraZeneca dataset, finding that MD features provided marginal but consistent improvements in aggregation rate and viscosity prediction accuracy over sequence features alone [21]. The availability of AlphaFold2-predicted structures for essentially any antibody variable domain sequence has substantially reduced the barrier to structural feature calculation, though the quality of AlphaFold2 CDR loop predictions remains debated and the predicted structures lack the solution ensemble information that drives aggregation under formulation conditions [19].
Biophysical measurement features, discussed in Section 2, constitute the third category and are the inputs most directly predictive of colloidal behaviour under formulation conditions. The practical challenge is that they require experimental generation, which limits their utility as purely in silico screening inputs. However, in hybrid workflows where a subset of molecules undergoes biophysical screening before full formulation development, measurement features can substantially improve prediction accuracy relative to sequence-only models and can serve as informative priors in Bayesian optimisation frameworks [23].

3.2. Classical Machine Learning Methods

3.2.1. Linear Models

Linear regression, Lasso, Ridge regression, and logistic regression remain competitive for small biopharmaceutical formulation datasets and are particularly appropriate where interpretability is a priority. Makowski et al. (2024) demonstrated that a logistic regression model trained on six sequence-derived features, principally Fv pI (isoelectric point of the variable fragment), achieved an area under the receiver operating characteristic curve (AUC) of 0.87 for binary classification of IgG1 antibodies as low or high viscosity, while providing directly actionable guidance on the direction of sequence modification needed to reduce viscosity [25]. This simplicity is an undervalued asset: a model that can be expressed as a closed-form linear combination of pI and two CDR composition terms is more likely to be accepted by a formulation scientist as a rational basis for candidate selection than a gradient-boosted tree whose prediction emerges from hundreds of split rules. The limitation, as with all linear models, is the inability to capture interaction effects -- for example, the empirically observed non-additive contributions of Fv pI and surface hydrophobic patches to viscosity at concentrations above 100 mg/mL [56].

3.2.2. Ensemble Tree Methods: Random Forest, Gradient Boosting, and Variants

Random forests and gradient boosted decision trees (GBDT), including XGBoost and LightGBM, have become the most widely used classical ML methods in antibody property prediction. These methods handle heterogeneous feature types without scaling, are robust to multicollinearity among correlated biophysical descriptors, and provide variable importance scores that serve as a first-pass interpretability tool. Lai et al. (2022) evaluated random forest alongside logistic regression, SVM, and decision tree classifiers for aggregation rate and viscosity prediction on their 47-mAb dataset, reporting that no single algorithm consistently dominated across both prediction tasks, a finding that underscores the importance of cross-validated model comparison rather than a priori algorithm selection [21]. Waight et al. (2023) compared gradient boosting with linear models and feedforward neural networks across five developability endpoints (HIC retention time, SEC monomer, kD, viscosity, PSR polyreactivity) in the PROPERMAB framework, finding that gradient boosting provided the highest accuracy for viscosity and HIC prediction at dataset sizes around 200-300 antibodies [24].
A persistent limitation of tree-based methods in this domain is that their feature importance metrics -- typically Gini impurity decrease for random forests or gain for XGBoost -- do not distinguish between causal drivers and correlated proxies. In antibody stability datasets where sequence descriptors are intercorrelated (e.g., pI, net charge at pH 6, CDR3 length, and SASA often co-vary across a mAb panel), feature importance rankings can be unstable across bootstrap resamples and should not be interpreted as identifying physicochemical causes of instability without independent mechanistic validation [46].

3.2.3. Support Vector Machines

Support vector machines (SVMs) with radial basis function (RBF) kernels have been used in several early formulation ML studies, particularly where dataset sizes are below 50-100 samples and the class boundary is not linearly separable in the original feature space. The margin-based SVM objective provides a degree of robustness to outliers that is advantageous in stability datasets where a small number of unusually unstable molecules may distort tree-based model boundaries. However, the requirement to tune the regularisation parameter C and the kernel bandwidth gamma introduces hyperparameter optimisation overhead that is disproportionate for very small datasets, and SVM probability calibration using Platt scaling or isotonic regression is rarely reported in formulation studies, leaving uncertainty estimates absent from model outputs [49]. Given these characteristics, SVMs have largely been superseded by gradient boosting and interpretable feedforward networks in recent studies.

3.3. Deep Learning Approaches

3.3.1. Feedforward and Ensemble Neural Networks

Feedforward neural networks (FNNs), also referred to as multilayer perceptrons (MLPs), were among the first deep learning architectures applied to biopharmaceutical stability prediction. Gentiluomo et al. (2020) trained FNNs on the PIPPI dataset to predict long-term monomer retention from a panel of biophysical inputs including kD, Rh, Tagg, and Tm, demonstrating that neural network models trained on formulation-condition-resolved features could predict real-time 24-month stability outcomes with sufficient accuracy to reduce physical stability testing by an estimated 40% [20]. This study remains one of the most clearly validated examples of ML-driven formulation screening, in part because the ground truth was prospective real-time data rather than accelerated estimates.
Liu et al. (2025) introduced DeepViscosity, an ensemble of 102 FNN models trained on 229 mAbs using 30 features derived from the DeepSP structural prediction tool, demonstrating classification of high-viscosity (>20 cP at 150 mg/mL) and low-viscosity antibodies with accuracy of 0.85 on two independent test sets comprising 16 and 38 mAbs respectively [57]. The ensemble approach -- averaging predictions from independently trained models on bootstrap resamples of the training data -- substantially reduced prediction variance relative to any single network and provided a practical uncertainty estimate based on prediction variance across ensemble members. This is a sensible design choice for formulation datasets where training data are scarce; ensemble variance is an accessible, model-agnostic proxy for epistemic uncertainty that requires no additional architectural complexity.

3.3.2. Convolutional Neural Networks and Particle Image Analysis

Convolutional neural networks (CNNs) have found their clearest application in biopharmaceutical formulation ML at the intersection of sub-visible particle characterisation and image analysis. As discussed in Section 2.5, Wang et al. (2023) at Merck demonstrated that CNNs trained on flow imaging microscopy (MFI) images could classify proteinaceous aggregates, silicone oil droplets, and glass particles with accuracy approaching expert human labelling [42]. The availability of large labelled particle image datasets -- considerably more accessible than labelled protein stability datasets -- makes this a setting where deep learning offers advantages over classical methods that are absent in the tabular biophysical parameter domain.
One-dimensional CNNs applied directly to antibody variable domain sequences have also been explored for CDR motif recognition and aggregation-prone region detection, bypassing manual feature engineering of sequence descriptors [58]. The practical limitation is that 1D convolution over sequences of 100-130 amino acid residues in the variable domain cannot capture long-range positional context -- a constraint that transformer architectures address more naturally. The application of CNN-based stability prediction to formulation composition inputs (excipient identity and concentration encoded as one-hot or continuous vectors) has been less extensively explored and remains an open direction.

3.3.3. Graph Neural Networks

Graph neural networks (GNNs) treat protein structures as graphs in which nodes represent amino acid residues and edges represent spatial or covalent proximity, enabling direct learning on 3D structural topology without manual feature engineering. GCN-based models trained on AlphaFold2-predicted structures have been used to predict aggregation propensity scores directly from protein graphs, with Nguyen et al. (2024) demonstrating a GCN model applied to the AlphaFold2 structural database that predicted A3D aggregation scores with accuracy (r = 0.86) exceeding prior sequence-only models [59]. More recently, equivariant GNNs such as those based on geometric vector perceptrons have been applied to predict stability changes upon point mutation (DeltaDeltaG) -- a related but distinct problem from formulation-dependent colloidal stability -- with the advantage of provable equivariance to rotation and translation of the input structure [60].
The translation of GNN approaches from mutation stability prediction to formulation-dependent colloidal stability faces two specific obstacles that the field has not yet resolved. First, aggregation under formulation conditions is a solution-phase phenomenon governed by protein-protein interaction energetics that are not fully encoded in the single-chain structure graph typically used as input. Second, the formulation variables (pH, ionic strength, excipient type and concentration) that are the primary control parameters in formulation development do not have obvious graph representations and must either be encoded as global graph-level features or handled through a separate architectural branch, an approach that has been proposed but not yet validated at scale on formulation screening datasets [61].

3.4. Protein Language Models and Transfer Learning

Protein language models (PLMs) represent the most significant recent methodological development in sequence-based property prediction. Trained on hundreds of millions of protein sequences from UniRef and similar databases using masked language modelling or autoregressive objectives, PLMs such as ESM-2 (Lin et al., Science 2023), ProtTrans/ProtT5 (Elnaggar et al., 2021), and the antibody-specific models IgBert and IgT5 learn contextual representations that encode evolutionary constraints, structural propensities, and functional relationships at single-residue resolution [62]. These representations can be extracted as fixed-length embedding vectors through mean-pooling across residue positions and used as input features to any downstream prediction model.
In antibody developability prediction, the application of PLM embeddings was evaluated by Yu et al. (2026) in a study combining IgBert, IgT5, and ESM-2 embeddings with classical ML classifiers and sequence-only fine-tuning for polyspecificity prediction, demonstrating that PLM-derived features outperformed hand-crafted sequence descriptors on an internal dataset spanning 33 historical therapeutic programmes across three critical developability assay types [63]. For solubility prediction, the solPredict model (2022) used ESM-1b embeddings as input to a predictor trained on 260 in-house mAbs measured by PEG-induced precipitation, achieving sequence-to-solubility prediction without requiring 3D structure calculation [64]. Hao and Fan (2024) combined ProtT5 embeddings with a random forest classifier to predict high-concentration viscosity, reporting improved performance over sequence descriptor-based models particularly for antibodies with unusual CDR compositions underrepresented in conventional descriptor databases [65].
The key advantage of PLM embeddings for formulation ML is their ability to represent sequence-level context that is invisible to manually designed descriptors. A glutamic acid residue in CDR-H3 in the context of a surrounding hydrophobic loop contributes differently to colloidal behaviour than the same residue in a framework region flanked by charged neighbours; PLM embeddings capture these contextual dependencies while amino acid composition counts do not. The practical limitation is that current PLMs are trained on natural protein sequence distributions that differ substantially from the engineered CDR sequences characteristic of therapeutic mAbs -- a distributional mismatch that domain-adaptive fine-tuning on antibody-specific corpora (such as the Observed Antibody Space database) is designed to address but has not fully resolved [66].
A further issue is that PLM embeddings, like deep learning features in general, are not directly interpretable as physicochemical drivers of aggregation. The SHAP-based feature attribution approaches reviewed in Section 6 can identify which embedding dimensions most influence a given prediction, but mapping these dimensions to specific amino acid substitutions or structural motifs requires additional analysis that is not yet standardised across the field. The combination of PLM embeddings with SHAP attribution and orthogonal wet-lab validation of top-ranked sequence hypotheses represents a methodology that several leading groups are beginning to adopt, but published end-to-end examples with prospective experimental confirmation remain rare [25].

3.5. Bayesian Optimisation and Active Learning

Bayesian optimisation (BO) methods treat the formulation design problem as the sequential optimisation of an unknown, expensive-to-evaluate objective function. A surrogate model -- typically a Gaussian process (GP) or random forest -- is fitted to the available experimental data and used to predict the objective value (e.g., Tm or monomer content) at unsampled formulation compositions, alongside an uncertainty estimate. An acquisition function (expected improvement, upper confidence bound, or Thompson sampling) then selects the next experiment to run as the point that maximises the expected gain given the surrogate model’s prediction and uncertainty [67].
Moller et al. (2022) provided the field’s most-cited proof of concept: a Bayesian optimisation algorithm identified the formulation maximising nanoDSF Tm for three tandem scFv bispecific variants within 25 experiments per variant, compared to the 80+ experiments that a conventional DoE approach would require to achieve equivalent coverage of the same excipient design space [23]. The sample efficiency advantage of BO is most pronounced in high-dimensional formulation spaces -- when pH, buffer identity and concentration, and multiple excipients are varied simultaneously -- where factorial and central composite DoE designs become computationally and practically intractable. The limitation of standard GP-based BO is that the GP covariance function assumes a smooth, stationary objective landscape, which may not hold for pH-dependent aggregation transitions or freeze-thaw-induced particulation events that exhibit sharp threshold behaviour [68].
Active learning extends the BO paradigm to classification settings and to multi-property objectives, enabling sequential selection of molecules for biophysical screening based on model uncertainty. In the developability context, active learning has been used to iteratively select mAb candidates for aggregation or viscosity measurement from a large sequence library, with the aim of building a maximally informative training dataset with the fewest experiments. An unresolved question in both BO and active learning for formulation is the choice of acquisition strategy when multiple stability endpoints must be simultaneously optimised -- a multi-objective problem for which Pareto-front-based acquisition functions have been proposed but not yet validated in prospective antibody formulation studies [69].

3.6. Physics-Informed Machine Learning: An Underdeveloped Direction

Physics-informed machine learning (PIML) methods incorporate physical constraints or governing equations -- such as the DLVO theory of colloidal interactions, thermodynamic stability criteria, or Arrhenius kinetics -- as inductive biases in the model architecture or loss function, reducing the amount of experimental data required to achieve reliable predictions. In disciplines such as fluid dynamics and materials science, PIML has demonstrated substantial improvements in data efficiency and extrapolation reliability over purely data-driven approaches [70]. In biopharmaceutical formulation, the potential is clear: a model that encodes the physical relationship between kD, protein concentration, and aggregation rate should generalise better to new protein-excipient combinations than a purely empirical neural network trained on the same data.
Despite this theoretical motivation, PIML for biopharmaceutical colloidal stability remains at an early stage. Prass et al. (2023) at Ruhr University Bochum and Boehringer Ingelheim demonstrated de novo prediction of mAb solution viscosity from large-scale atomistic MD simulations using Green-Kubo theory, providing a physics-grounded approach that bypasses empirical ML entirely at the cost of very high computational expense [61]. Combining such physics-derived features with lightweight ML models -- using MD-derived descriptors as physically grounded inputs to gradient boosting classifiers -- represents a practical hybrid route that several groups including Lai et al. (2022) have begun to explore [21]. The development of reduced-complexity physics-informed surrogate models that encode colloidal interaction physics while remaining trainable on datasets of 50-200 molecules is a research priority that the community has not yet adequately addressed.

3.7. Algorithm Selection and Benchmarking: The Absent Standard

A frank assessment of the current literature reveals a pattern of inconsistent and often incomplete model benchmarking. The majority of published formulation ML studies evaluate one to three algorithms on a single internal dataset, report accuracy metrics (most commonly AUC, Spearman r, or mean absolute error) on a held-out test fraction of the same dataset, and do not perform external validation on an independent protein series. This internal-only evaluation paradigm inflates apparent model performance relative to true generalisation ability, particularly when training and test sets are drawn from the same antibody programme and therefore share correlated sequence and physicochemical properties [46]. As Khetan et al. (2022) noted, published prediction accuracies for self-association, aggregation, and viscosity are inconsistently reported across studies using different assays, different protein panels, and different performance metrics, making cross-study comparison structurally impossible in the absence of a common benchmark dataset and evaluation protocol [49].
The field would benefit substantially from a standardised benchmarking framework analogous to the CASP competition for protein structure prediction or the MoleculeNet benchmark in molecular property prediction. Such a framework would require: (i) a set of publicly accessible protein-formulation datasets with defined training and test splits; (ii) a defined set of evaluation metrics including both classification performance (AUC, Matthews correlation coefficient) and calibration (Brier score, reliability diagrams); (iii) explicit reporting of external validation on molecules from independent research groups; and (iv) pre-registered model architectures to prevent post-hoc selection of the best-performing configuration. No such framework currently exists for colloidal stability prediction, and its absence is a rate-limiting factor for progress. A structured comparison of the major algorithm classes, their representative implementations, and their strengths and limitations in the biopharmaceutical formulation context is provided in Supplementary Table S2.
The diversity of ML approaches reviewed in Section 3 reflects in part the diversity of prediction targets that the biopharmaceutical formulation community has attempted to address. These targets differ in their physical nature, measurement accessibility, regulatory relevance, and tractability as ML outputs. This section examines the principal prediction targets systematically, with a critical assessment of the current state of the field for each, the limitations of existing models, and the unresolved problems that constrain progress. The targets are grouped loosely from those with the most established ML literature to those where ML application is nascent or largely absent.

4. Predictive Targets: What Machine Learning Is Being Asked to Predict

4.1. Aggregation Propensity

4.1.1. Sequence-Level Prediction Tools: Capabilities and Limits

Aggregation propensity prediction from protein sequence is the most mature sub-field of ML-based formulation science, with tools having been available for over two decades. The first generation of predictors -- AGGRESCAN, Zyggregator, TANGO, PASTA, and FoldAmyloid -- identify aggregation-prone regions (APRs) in linear sequences using empirically derived amino acid scales or physico-chemical propensity functions based on beta-sheet tendency, hydrophobicity, and charge patterns [71]. These tools are fast, sequence-only, and widely used in discovery workflows for candidate screening. Their principal limitation is the failure to account for structural context: an APR buried in the hydrophobic core of a folded protein is not accessible for intermolecular contacts under native conditions, and sequence-only predictors systematically overestimate the aggregation propensity of stable globular proteins where structural protection is strong [72]. AGGRESCAN3D (A3D) addressed this by computing structurally corrected aggregation propensity scores using 3D atomic models, reducing false-positive APR predictions by an order of magnitude compared to sequence-only tools for soluble globular proteins [73]. The subsequent integration of AlphaFold2-predicted structures into the A3D database extended this approach to proteins lacking experimental structures, though the accuracy of AlphaFold2 CDR loop predictions for antibodies -- the most common target class -- remains a source of concern, as CDR loops adopt diverse and sometimes unusual conformations that the training distribution of AlphaFold2 does not fully represent [19].
A more recent generation of ML-based APR predictors, including AggreProt (Nucleic Acids Research, 2024), uses deep learning models trained on hexapeptide aggregation data to predict residue-level aggregation propensity with improved benchmark performance relative to Waltz, TANGO, and AGGRESCAN on the AmyPro validation set [74]. However, a persistent tension exists between tools trained on amyloid-forming peptides from disease biology and the problem of predicting native-state colloidal instability in therapeutic antibodies: the physical mechanisms of amyloid aggregation and the reversible self-association or surface-mediated aggregation that dominates formulation instability in mAbs are distinct, and the overlap between training data for these two problems is limited. A tool that accurately predicts APRs in intrinsically disordered proteins associated with neurodegeneration does not automatically transfer to predicting which CDR sequences drive elevated kD under formulation-relevant pH and ionic strength conditions.

4.1.2. Rate-Based Prediction from Biophysical Inputs

A complementary and arguably more formulation-relevant approach predicts the rate of aggregate accumulation -- typically expressed as the first-order rate constant k_agg derived from SEC time-series data at accelerated temperatures -- as a function of combined sequence, structural, and biophysical features. This is the approach taken by Lai et al. (2022), Gentiluomo et al. (2020), and the Aggregation Time Machine platform of Bunc et al. (2022). Bunc and colleagues at Novartis developed a physically grounded kinetic model parameterised from temperature-dependent SEC aggregation profiles collected across a wide temperature range (40-75 degrees C) for six therapeutic mAbs. By fitting a non-Arrhenius kinetic model that explicitly accounts for the thermodynamic stability of the native state (expressed as delta-G of denaturation) and its linkage to the rate of aggregation, they achieved reliable prediction of 3-year aggregate fractions at 5 degrees C from 6-12 week experiments at elevated temperature [75]. Kuzman et al. (2021) at Novartis demonstrated the same Arrhenius-based kinetic modelling approach across multiple mAb quality attributes including aggregation, fragmentation, deamidation, and charge profile variants, showing that 6 months of multi-temperature accelerated stability data was sufficient to predict 36-month real-time outcomes within 95% prediction intervals for most attributes [76].
The Huelsmeyer et al. (2023) universal stability prediction tool extended this framework in a cross-company validation involving Novartis and three other pharmaceutical companies, demonstrating that the Arrhenius kinetic modelling approach generalised across product types including mAbs, bispecific antibodies, ADCs, vaccines, and IVD biomolecules [77]. This cross-company validation is one of the few examples of external validation in the biopharmaceutical stability prediction literature and represents a methodological standard that purely data-driven ML studies should seek to match. The key distinction between kinetic modelling approaches and standard ML models is the use of physical equations -- first-order reaction kinetics, Arrhenius temperature dependence -- as the model architecture rather than flexible function approximators trained end-to-end. This constrains the model to physically consistent predictions and substantially reduces the effective number of parameters that must be estimated from data, which is a decisive advantage when training data are limited.

4.2. High-Concentration Viscosity

Viscosity at therapeutic concentrations (typically 50-200 mg/mL for subcutaneous administration) is a major formulation challenge for a substantial fraction of clinical-stage mAbs. Solutions with viscosity above approximately 20 cP at 25 degrees C are difficult to administer by subcutaneous injection through standard needles and require either dose volume reduction, co-formulation with hyaluronidase, or reformulation targeting lower concentration -- each of which introduces substantial development cost and timeline risk [56]. The molecular drivers of high viscosity are protein self-association mediated by electrostatic patch complementarity, hydrophobic interactions between CDR loops, and short-range attractive forces that increase the effective hydrodynamic volume of the protein in solution. Because self-association is concentration-dependent and nonlinear, viscosity is a more complex and less tractable prediction target than aggregation rate under dilute conditions.
The most consistently reported molecular predictor of mAb viscosity is the Fv isoelectric point (Fv pI) or net Fv charge at formulation pH, which correlates inversely with viscosity across diverse IgG1 and IgG4 panels -- a lower Fv pI (more negatively charged Fv at pH 5-6) correlates with lower self-association and lower viscosity [55]. Makowski et al. (2024) built on this foundation to develop an interpretable logistic regression model trained on 80 mAbs, using SHAP analysis to confirm that Fv pI and two CDR sequence composition terms were the dominant features, and prospectively validating the model by engineering reduced-viscosity Fv variants through targeted sequence modification guided by SHAP attribution [25]. This study is particularly significant because it represents one of the few published examples in the formulation ML literature of a complete prediction-engineering-experimental-validation cycle: the model was not just assessed retrospectively on held-out data but used prospectively to generate new molecules whose viscosity was then measured, confirming that the model had identified actionable molecular drivers rather than spurious correlations.
Lai et al. (2022) achieved binary viscosity classification (low/high at 200 mg/mL) with AUC of 0.82-0.89 depending on the algorithm, using a combination of MD-derived features and sequence descriptors on 47 mAbs [21]. Liu et al. (2025) reported improved viscosity classification on a larger dataset (229 mAbs) using DeepViscosity, an ensemble of feedforward neural networks trained on 30 features from the DeepSP structural prediction tool, with AUC of 0.85 on an independent test set of 54 mAbs -- one of the larger external validations published to date [57]. Hao and Fan (2024) demonstrated that ProtT5 language model embeddings combined with a random forest classifier improved viscosity prediction accuracy over sequence descriptor models, particularly for antibodies with CDR compositions underrepresented in conventional descriptor databases [65]. Despite this progress, viscosity prediction accuracy remains insufficient for confident candidate selection decisions at dataset sizes below 100-150 molecules, and no published model has been validated prospectively across different molecular scaffolds -- a gap that limits the translational value of current approaches.

4.3. Solubility

Apparent solubility -- typically measured as the maximum protein concentration achievable before visible precipitation or a defined turbidity threshold is reached -- is a tractable prediction target for sequence-based ML because it is governed by the same surface hydrophobicity and charge distribution parameters that drive aggregation. CamSol, the most widely used computational solubility predictor, assigns per-residue intrinsic solubility scores based on hydrophobicity, charge, and secondary structure propensity, and has been validated against experimental solubility data for diverse protein datasets [53]. Sormanni et al. demonstrated that CamSol mutations predicted from the intrinsic solubility score reliably increased experimental solubility of antibody variable domains when applied to framework regions exposed to solvent, establishing a mechanistically grounded sequence-to-solubility design pipeline [53].
Rosace et al. (2023) extended the solubility prediction framework to an automated experimental setting by integrating computational solubility design with high-throughput expression and biophysical characterisation, enabling simultaneous optimisation of solubility and conformational stability [51]. The Warszawski et al. solPredict model (2022) took a transfer learning approach, using ESM-1b protein language model embeddings fine-tuned on 260 in-house mAbs measured by PEG precipitation, achieving better performance than CamSol on the internal test set but without external validation on independent programmes [64]. A recurring challenge in solubility prediction is measurement standardisation: PEG precipitation, ammonium sulfate precipitation, centrifugal ultrafiltration, and direct concentration-turbidity profiling yield non-equivalent apparent solubility values for the same protein, and pooling data from methods that differ in their mechanism of precipitation introduces label heterogeneity that degrades model generalisation in ways that are rarely acknowledged explicitly.

4.4. Liquid-Liquid Phase Separation

Liquid-liquid phase separation (LLPS) in biopharmaceutical formulations -- the formation of a protein-rich dense liquid phase coexisting with a protein-depleted bulk phase -- is distinct from both aggregation and simple solubility loss. LLPS is a thermodynamic phenomenon governed by the shape of the free energy of mixing as a function of protein concentration, temperature, and solution conditions. In mAbs, LLPS manifests as a reversible low-temperature opalescence or concentration-dependent cloudiness that is increasingly recognised as a developability liability at high concentration, even when the protein remains fully monomeric by SEC [78]. The physical parameters governing LLPS -- the critical temperature (T_c), the spinodal and binodal boundaries, and the interaction second virial coefficient derived from static light scattering -- differ from those governing irreversible aggregation and require distinct analytical methods, primarily SLS, analytical ultracentrifugation, and dedicated LLPS assay platforms.
Wei et al. (2022) at Wenzhou University applied regression and classification neural networks to predict the cloud point temperature -- the temperature below which phase separation occurs -- for lysozyme and bovine serum albumin (BSA) as a function of salt concentration, protein concentration, and pH, demonstrating that the non-linear phase boundary could be captured by a neural network trained on approximately 100 solution conditions per protein [79]. This study is technically sound but deliberately limited in scope: lysozyme and BSA are model systems whose phase behaviour has been studied extensively, and their phase diagrams are governed by relatively simple DLVO-type interactions that are not representative of the more complex and anisotropic interactions driving LLPS in full-length mAbs with asymmetric charge distributions and Fc-dependent self-association pathways.
ML prediction of LLPS in therapeutic mAbs from sequence or biophysical inputs remains at an early stage. Coarse-grained molecular dynamics simulations, interpreted alongside small-angle X-ray scattering, can characterise cluster growth as well-characterised mAbs approach LLPS, and these simulation outputs could in principle serve as physics-informed features for ML models [80]. The principal obstacle is data scarcity: LLPS cloud point measurements are relatively uncommon in published datasets, most formulation screening studies do not report LLPS characterisation data, and the phenomenon is concentration-dependent and therefore requires measurement at concentrations (typically above 50-100 mg/mL) that are experimentally demanding at the screening scale. Until a systematic, publicly available LLPS dataset for therapeutic antibodies is assembled, ML prediction of this endpoint will remain a largely theoretical aspiration.

4.5. Opalescence

Opalescence -- the light-scattering-mediated turbidity of protein solutions arising from concentration fluctuations near a phase boundary or from the presence of soluble oligomers -- is a product appearance attribute with direct regulatory relevance under pharmacopoeial visual inspection guidelines. Its prediction from formulation parameters is relevant for both candidate selection and formulation design, as highly opalescent drug products require justification in regulatory submissions and may trigger additional characterisation requirements. Opalescence is physically related to LLPS -- it often represents a sub-critical concentration regime of the same attractive interactions that drive full phase separation at higher concentrations -- but is also influenced by reversible self-association below the phase boundary [81].
Salinas et al. (2010) described mAb opalescence in terms of enhanced Rayleigh scattering associated with attractive protein–protein interactions and proximity to phase boundaries. ML prediction of opalescence onset from sequence or biophysical features has not been systematically reported in the published literature as of mid-2025 [81]. The closest available approach is indirect: models predicting B22 or kD from sequence features provide a proxy for the inter-protein attractive potential that drives both opalescence and LLPS, but the relationship between these interaction parameters and the practical opalescence threshold under formulation conditions is not quantitatively established across diverse mAb sequences. This represents a clear gap in the current predictive toolkit.

4.6. Shelf-Life and Long-Term Storage Stability

Shelf-life prediction -- the forecast of the time to reach a defined out-of-specification threshold for a stability-indicating quality attribute (SIQA) such as SEC monomer content, charge variant profile, or sub-visible particle count at the intended storage condition -- is the most directly commercially and regulatorily relevant prediction target in formulation development. It is also the target for which the evidence base is most clearly stratified between kinetic modelling approaches, which have demonstrated prospective success, and purely data-driven ML approaches, which have been validated primarily retrospectively.
The Arrhenius kinetic modelling framework developed at Novartis by Kuzman et al. (2021) and Bunc et al. (2022), and independently validated across four pharmaceutical companies by Huelsmeyer et al. (2023), represents the most robustly validated shelf-life prediction platform currently available [76]. The framework combines multi-temperature accelerated stability experiments with a non-Arrhenius kinetic model to generate prediction intervals for 24-36 month stability outcomes at 5 degrees C from 3-6 months of elevated-temperature data. The key physical insight is that the apparent non-Arrhenius behaviour of mAb aggregation -- the observation that simple linear extrapolation from 40 degrees C systematically overestimates aggregation rates at 5 degrees C -- arises from the temperature dependence of the native-state stability (delta-G), which changes sign near the temperature of maximum stability and creates a non-linear linkage between temperature and aggregation rate [77]. Accounting for this thermodynamic term explicitly in the kinetic model, rather than treating aggregation as a simple first-order Arrhenius process, substantially improves prediction accuracy and reduces systematic bias.
The extended work by Dai et al. (2024) at AbbVie applied the Accelerated Stability Assessment Program (ASAP) modelling framework to an oral-delivered mAb drug product, demonstrating that kinetic modelling could guide shelf-life establishment for non-traditional administration routes where standard liquid formulation assumptions do not apply [82]. A further contribution from Gonzalez-Valdez et al. (2025) documented the application of Arrhenius-based global fitting to predict the long-term stability of seven anti-SARS-CoV-2 antibodies at high concentration, enabling rapid IND and BLA submission timelines during accelerated development [83]. These examples illustrate that kinetic modelling -- which is fundamentally a form of physics-informed ML when the kinetic parameters are estimated from data rather than derived from first principles -- can be used successfully in regulatory submissions as a basis for shelf-life justification.
Purely data-driven ML for shelf-life prediction faces the obstacle that the output of interest -- stability at 5 degrees C over 24-36 months -- requires either real-time data that takes years to collect or a validated accelerated-to-real-time extrapolation that currently depends on the same kinetic modelling framework just described. Neural network models trained directly on accelerated stability labels as shelf-life proxies, without explicit kinetic modelling of the temperature extrapolation, introduce systematic errors that vary in magnitude depending on the protein, the degradation pathway, and the temperature used for acceleration [16]. This interaction between the data generation strategy and the ML model design is underappreciated in the literature and represents a source of model error that cannot be corrected by improving the ML architecture.

4.7. Freeze-Thaw Stability and Cold-Chain Excursion Risk

Freeze-thaw stability is an independent formulation challenge that governs both manufacturing process robustness (bulk drug substance is typically frozen for storage and transport) and distribution resilience. During freezing, protein solutions undergo sequential concentration by ice crystal exclusion, exposure to glass-forming excipient matrices, and mechanical stress from ice crystal growth; during thawing, these stresses reverse but may leave persistent aggregate or particle populations that do not redissolve [84]. The susceptibility of a given protein to freeze-thaw-induced aggregation depends on the protein’s surface properties, the cryoprotectant composition (sucrose, trehalose, and certain amino acids are the most effective cryoprotectants for most mAbs), the cooling rate, and the hold temperature.
For mRNA-LNP formulations, freeze-thaw stability is particularly critical and mechanistically distinct from protein aggregation. Ionisable lipid nanoparticles are susceptible to particle aggregation and fusion during freeze-thaw cycles through a mechanism involving osmotic stress-driven membrane destabilisation during ice formation; sucrose and trehalose at concentrations of 8-10% (w/v) are the established cryoprotectants that mitigate this by acting as amorphous glassy matrices that vitrify around the particles during freezing [18]. Youssef et al. (2025) systematically evaluated six sugar-surfactant excipient combinations in Tris buffer for mRNA-LNP freeze-thaw stability and long-term storage, providing formulation screening data that could serve as a training set for ML-guided cryoprotectant selection, though the study was not framed in ML terms [85]. A recent study applying gradient boosting to lyophilisation process parameter optimisation for mRNA-LNPs identified freezing rate and annealing temperature as the dominant process-critical parameters for particle size control, demonstrating that ML can extract meaningful process knowledge from small designed-experiment datasets in this space [86].
Cold-chain excursion risk -- the probability of product degradation given a defined temperature deviation from the intended storage condition over a specified duration -- is a prediction target of direct supply chain and regulatory relevance but has received almost no attention from the ML formulation community. The problem requires combining the kinetic stability model (rate of aggregation or other degradation as a function of temperature) with a probabilistic model of temperature exposure during distribution, neither of which is routinely incorporated into formulation ML frameworks. An integrated physics-informed ML approach to excursion risk prediction, combining Arrhenius-based kinetic parameters estimated from formulation screening data with Bayesian probabilistic models of temperature exposure, is a technically tractable and practically valuable direction that has not yet been reported.

4.8. Multi-Attribute and Joint Prediction: The Developability Profile Problem

A practically important but methodologically underexplored dimension of formulation ML is the simultaneous prediction of multiple stability endpoints as a joint output -- the developability profile problem. In a real formulation development workflow, candidates are not selected or rejected on the basis of a single predicted property but on the basis of a multivariate profile that weighs aggregation rate, viscosity, solubility, and immunogenicity risk together against potency, manufacturability, and pharmacokinetic requirements. Optimising for any single endpoint in isolation can lead to candidates that perform well on the predicted property while deteriorating on unconstrained ones [24].
PROPERMAB (Waight et al., 2023) attempted to address this by building separate models for five developability endpoints and combining their outputs into a composite developability score. The Jain et al. (2017) 12-assay panel similarly provides a multi-attribute biophysical fingerprint rather than a single predicted endpoint [45]. Willis et al. (2025) at the University of Leeds proposed a single holistic developability parameter derived from multivariate analysis of a broad biophysical panel, demonstrating that a single composite score could rationalise candidate screening decisions more reliably than any individual assay metric [87]. The ML generalisation of this approach -- training a single model on multi-target outputs with a developability-weighted loss function -- has been explored in multi-task learning frameworks but has not been validated prospectively in a pharmaceutical development setting. Pareto optimisation across stability, viscosity, and immunogenicity objectives represents the most principled formulation of the multi-attribute problem and connects naturally to multi-objective Bayesian optimisation methods, but practical examples remain confined to small-scale academic demonstrations.
The value of a prediction model in formulation science is not exhausted by its statistical accuracy on a held-out test set. A model whose output cannot be connected to a physically interpretable rationale is of limited use to a formulation scientist who must decide which molecule to progress, which excipient to add, or which sequence to mutate. It is of even more limited use to a regulatory reviewer who must assess whether a ML-assisted formulation decision rests on a scientifically sound and auditable basis. Explainable AI (XAI) methods address this gap by attributing model predictions to specific input features in a quantitative and, ideally, causally meaningful way. This section reviews the principal XAI methods applied in the biopharmaceutical formulation domain, assesses their demonstrated utility and inherent limitations, and examines the unresolved question of whether explanations generated by current XAI tools are genuinely mechanistic or merely plausible post-hoc narratives.

5. Explainable Artificial Intelligence in Biopharmaceutical Formulation

5.1. The Interpretability Problem in Formulation ML

The tension between predictive accuracy and interpretability is a structural feature of the ML model landscape. Linear models are fully transparent -- each coefficient directly encodes the direction and magnitude of a feature’s contribution to the prediction -- but are limited in their ability to capture non-linear relationships. Gradient boosted trees and neural networks can approximate complex non-linear functions but at the cost of intrinsic interpretability: a gradient boosted model with 500 trees and 6 levels per tree involves more decision nodes than any human can inspect, and a neural network with three hidden layers of 128 units maps inputs to outputs through a composition of transformations that has no direct physical analogue. XAI methods attempt to recover interpretability for these complex models through post-hoc approximation rather than model re-design [88].
In the formulation context, interpretability serves at least three distinct purposes that should not be conflated. The first is scientific hypothesis generation: an attribution pointing to Fv pI as the dominant viscosity driver suggests that re-engineering the CDR charged residue composition is a productive path to viscosity reduction, providing a testable hypothesis at the bench. The second is model debugging: attribution analysis can reveal when a model has learned spurious correlations in training data -- for example, a stability model that has learned to associate a particular buffer identity with high stability because all high-stability molecules in the training set happened to be tested in that buffer -- rather than genuine molecular determinants. The third is regulatory justification: regulators increasingly expect that AI-assisted formulation decisions be accompanied by a documented rationale linking the model output to known physicochemical principles, a requirement that places XAI output in a quasi-evidentiary role within a CMC submission [89]. These three purposes impose different requirements on the quality of explanation: hypothesis generation can tolerate approximate attributions if they point in the correct direction; regulatory justification requires that attributions be stable, reproducible, and consistent with established science.

5.2. SHAP: The Dominant XAI Tool in Formulation ML

SHapley Additive exPlanations (SHAP), introduced by Lundberg and Lee in 2017, provides a unified framework for computing feature attributions by applying Shapley values from cooperative game theory to ML models [90]. The Shapley value of feature i is the average marginal contribution of that feature across all possible subsets of features, weighted by the number of subsets. SHAP satisfies three desirable axiomatic properties -- local accuracy (the sum of SHAP values equals the model output minus the baseline), missingness (absent features receive zero attribution), and consistency (features that contribute more to the model output always receive higher SHAP values) -- that distinguish it from earlier feature importance methods such as mean decrease impurity or permutation importance [90]. The TreeSHAP algorithm, designed for tree ensembles including random forests and gradient boosted trees, computes exact Shapley values in polynomial time, making it computationally feasible for the dataset sizes typical in formulation ML.
Makowski et al. (2024) at the University of Michigan provided the most thoroughly documented application of SHAP in the formulation stability domain. Training a logistic regression model on 80 mAbs with sequence-derived features for binary viscosity classification, they used SHAP beeswarm plots to confirm that Fv pI was the single strongest positive predictor of low viscosity, with mean |SHAP| value approximately three times larger than the next contributor, and that specific CDR composition terms (arginine content in CDR-H3, hydrophobic patch size in CDR-L1) made secondary positive contributions [25]. Critically, they then used these SHAP attributions not merely to interpret the existing model but as a molecular design guide: they identified five antibody variants with suboptimal Fv pI and introduced targeted lysine-to-glutamate substitutions at SHAP-identified positions, reducing viscosity at 150 mg/mL from above 20 cP to below 10 cP in four of five variants. This prospective validation cycle -- predict, explain, engineer, test -- is the gold standard for demonstrating that XAI output provides genuine scientific utility rather than plausible but untestable narrative.
Lai et al. (2022) applied SHAP analysis to their gradient boosting classifiers for aggregation rate and viscosity prediction on 47 mAbs, finding that MD-derived features including the protein-protein interaction energy at the Fc-Fab interface and the electrostatic complementarity parameter of the Fv surface were among the top SHAP contributors to aggregation rate prediction [21]. The interpretation that Fc-Fab interface dynamics drive aggregation propensity is physically plausible -- partial domain dissociation at the Fc-Fab hinge has been proposed as a mechanism for aggregation-competent state formation -- and the SHAP attribution provides a quantitative basis for prioritising this hypothesis over alternatives such as CDR hydrophobic patch exposure. Whether SHAP attribution in this context constitutes genuine mechanistic evidence or a correlation-based narrative consistent with the mechanism is a question that the study does not resolve and that the field has not yet addressed systematically.

5.3. LIME: Local Explanations and Their Constraints

Local Interpretable Model-Agnostic Explanations (LIME), introduced by Ribeiro, Singh, and Guestrin in 2016, generates per-prediction explanations by fitting a locally faithful linear model in the neighbourhood of the instance being explained [91]. The local linear approximation is constructed by perturbing the input instance, obtaining model predictions for the perturbed samples, weighting them by their proximity to the original instance, and fitting a sparse linear regression on this locally sampled dataset. LIME is model-agnostic -- it can in principle be applied to any ML model -- and produces explanations in terms of the original input features rather than internal model representations.
In the biopharmaceutical formulation context, LIME has been applied primarily to explain individual formulation condition predictions -- for example, why a specific pH 6.5, 10% trehalose, 0.2% polysorbate 80 formulation is predicted as high-stability while a pH 5.5 counterpart is predicted as low-stability -- by identifying which formulation variables contribute most to the prediction for that specific molecule-excipient combination. This local explanatory role is useful for formulation troubleshooting but less so for portfolio-level candidate selection, where global patterns across many molecules are more informative than instance-level explanations. The practical limitation of LIME in formulation ML is that its neighbourhood definition is arbitrary for structured biological inputs: perturbing a pH value by 0.5 units is chemically meaningful, but perturbing a sequence-derived Fv pI value by a corresponding amount requires a mapping from feature space back to sequence space that LIME does not perform [92]. This structural mismatch between the perturbation space and the chemical design space limits the actionability of LIME explanations for sequence-level formulation decisions.

5.4. Attention Weights and Gradient-Based Attributions in Deep Models

Transformer architectures and protein language models generate attention weight matrices that have been widely -- and often incorrectly -- interpreted as providing mechanistic insight into which input positions drive model predictions. Attention weights indicate where the model directs its computational focus during processing; they do not, in general, correspond to causal feature attributions in the sense that SHAP values do. Jain and Wallace (2019) demonstrated computationally that attention weights and gradient-based attributions can disagree substantially and that attention can be adversarially manipulated to produce different attention patterns without changing the model output, undermining the interpretation of attention as explanation [93]. More reliable attribution methods for transformer models include attention rollout, integrated gradients (IG), and gradient times input (GxI), which propagate gradient signals through the full computational graph rather than relying on forward-pass attention weights alone.
In the antibody stability context, Ruffolo et al. applied attention-based interpretability analysis in DeepAb, an antibody structure prediction model, demonstrating that the network’s attention patterns captured physically meaningful residue-residue interactions including proximal aromatic contacts and key hydrogen bonds at the Fv interface [94]. The translation of such structural interpretability to formulation-relevant stability properties requires that the model being analysed explicitly includes formulation conditions as inputs -- a requirement that current PLM-based formulation models do not generally satisfy. Until protein language models are retrained or fine-tuned on formulation-conditioned stability datasets, attention-based interpretability of PLM embeddings for formulation decisions reflects the structural and evolutionary context encoded in pre-training rather than formulation-specific physicochemical drivers.

5.5. Mechanistic Interpretability Versus Post-Hoc Rationalisation: An Unresolved Distinction

A concern that has been raised repeatedly in the broader XAI literature but has not been adequately addressed in biopharmaceutical formulation ML is the distinction between mechanistic interpretability -- where the attribution correctly identifies a causal driver of the predicted property that operates through a known physical mechanism -- and post-hoc rationalisation -- where the attribution identifies a feature that is correlated with the target in the training data and is consistent with a known mechanism, but where the causal relationship has not been established. The distinction matters because post-hoc rationalisations can be constructed for almost any well-known biophysical relationship, making it very easy to produce SHAP attributions that appear mechanistically plausible regardless of whether the model has learned causal structure or spurious correlation [88].
Consider the recurrent finding that Fv pI is the top SHAP feature for viscosity. This attribution is consistent with the established mechanism -- lower Fv pI at acidic formulation pH reduces the net positive charge, decreasing electrostatic self-association and thereby reducing viscosity. However, in a training dataset where pI and CDR3 length are correlated because high-pI antibodies are enriched among molecules with longer, arginine-rich CDR-H3 loops, SHAP may attribute viscosity to pI even if the physical driver is CDR3 arginine content and pI is merely a proxy. The two variables produce nearly identical predictions on training data and cannot be distinguished by SHAP analysis alone; only orthogonal experiments -- such as engineering variants that decouple pI from CDR3 arginine content -- can establish the causal structure. Makowski et al. (2024) approached this problem more rigorously than most by testing both pI-modified and CDR3-modified variants, but this level of experimental follow-up is uncommon in the literature [25].
A further underappreciated issue is the instability of SHAP values in small datasets with correlated features. Yuan et al. (2022) demonstrated that SHAP value rankings can vary substantially across bootstrap resamples of small tabular datasets, with the ranking of features of intermediate importance being particularly unstable [95]. For formulation ML datasets where n = 50-150 molecules and features include 10-30 correlated sequence and structural descriptors, this instability implies that SHAP-based feature rankings should be reported with confidence intervals derived from cross-validation resampling rather than as point estimates from a single model trained on the full dataset. This straightforward methodological improvement is not common practice in the field.

5.6. XAI and Regulatory Acceptance: The Current Gap

The FDA’s 2023 discussion paper on artificial intelligence in drug development acknowledged the value of XAI for model transparency and noted that explainability is one of the key criteria for evaluating AI model suitability in regulated decision-making contexts [96]. The EMA’s 2023 Reflection Paper on AI similarly identified explainability and transparency as requirements for AI systems used in drug development, specifying that AI-supported decisions should be accompanied by documentation of the model’s outputs, their uncertainty, and the features driving them [97]. However, neither document specifies which XAI methods are acceptable for which types of decisions, nor do they provide quantitative criteria for what constitutes a sufficient explanation -- for example, whether a SHAP beeswarm plot with top-5 features identified is adequate justification for an AI-assisted formulation selection or whether independent mechanistic validation of the attributed features is required.
This ambiguity places formulation scientists in a difficult position: they are expected to provide explainable AI output in regulatory contexts without clear guidance on what standard that output must meet. Raza et al. (2025) reviewed the regulatory perspectives on AI in drug development across FDA and EMA and noted that the FDA’s flexible, case-specific model for AI evaluation contrasts with the EMA’s more structured, risk-tiered approach, creating potential for regulatory divergence in how XAI evidence is assessed for the same product submitted in both jurisdictions [98]. Within the CMC domain specifically, no regulatory guidance as of mid-2025 addresses the use of XAI output as supporting evidence for formulation design space justification under ICH Q8(R2) or as part of a control strategy document under ICH Q10. This gap represents a high-priority area for future regulatory science development.

5.7. Counterfactual Explanations and Design-Oriented XAI

A class of XAI methods that is particularly relevant to the formulation design context is counterfactual explanations -- answers to the question: what is the minimal change to the input that would flip the prediction from an undesirable class (e.g., high viscosity, low colloidal stability) to a desirable one? Counterfactual explanations are inherently design-oriented: rather than explaining why a candidate has a predicted property, they specify what needs to change to achieve a target property, which is directly actionable for both sequence engineers and formulation scientists [99]. For a high-viscosity mAb candidate, a counterfactual explanation might specify that reducing the predicted Fv pI by 0.8 units through three arginine-to-serine substitutions in CDR-H1 and CDR-H3 is predicted to move the molecule from the high-viscosity to the low-viscosity class with the smallest perturbation to the original sequence.
The actionability constraint -- that counterfactual changes must be chemically or practically feasible -- is not automatically enforced by standard counterfactual generation algorithms, which may propose sequence substitutions that abolish antigen binding affinity or introduce glycosylation sites at CDR positions. Integrating binding affinity models or developability constraint models as hard constraints in the counterfactual generation procedure is a methodological direction that has been demonstrated for small-molecule drug design but has not yet been systematically applied to the antibody formulation domain. Development of such constrained counterfactual generators, validated against prospective wet-lab data, is among the most practically valuable directions for formulation-oriented XAI research. The principal XAI methods available for biopharmaceutical formulation applications, with their scope, formulation-domain examples, and key limitations, are reviewed in the preceding subsections.
The preceding sections have addressed ML approaches and prediction targets in terms that apply broadly across biopharmaceutical modalities. In practice, each modality presents a distinct combination of molecular architecture, instability mechanism, analytical complexity, and data availability that shapes what ML can and cannot currently deliver. This section examines each major modality in turn, with a focus on the specific challenges that differentiate it from conventional mAbs and on the evidence base -- or lack thereof -- for ML-guided colloidal stability prediction.

6. Modality-Specific Applications and Challenges

6.1. Monoclonal Antibodies: The Most Developed Case

Monoclonal antibodies of the IgG1, IgG2, and IgG4 subclasses represent the modality for which ML-based colloidal stability and developability prediction is most mature. The combination of structural homogeneity -- nearly all approved mAbs share the same ~150 kDa two-heavy-chain, two-light-chain architecture with conserved Fc region and variable Fv domains -- and the availability of well-characterised public datasets (Jain et al. 2017, PIPPI, Lai et al. 2022) has allowed the field to develop and partially validate sequence-to-property prediction models that would not be feasible for less structurally conserved modalities [45]. The consensus from the literature is that sequence-derived features, particularly Fv pI, net Fv charge at formulation pH, spatial aggregation propensity (SAP), and hydrophobic patch exposure, capture the dominant drivers of aggregation propensity and viscosity for IgG1 mAbs with reasonable but not yet reliable accuracy [55].
For IgG1 mAbs, the most robustly validated ML outputs are viscosity classification and colloidal stability ranking at the candidate selection stage, where the goal is to triage a panel of sequence-diverse candidates for progression to full formulation development. At this level -- binary classification of high vs. low viscosity or favorable vs. unfavorable kD -- published models achieve AUC values of 0.82-0.89 on internal held-out test sets [21]. External validation across independent molecular programmes remains rare and consistently shows performance degradation relative to internal test set metrics, with Spearman correlation coefficients for continuous property prediction typically falling from 0.7-0.8 internally to 0.5-0.6 on external datasets [49]. This gap reflects distributional shift between training and test sets, a problem that will not be resolved by model architecture improvements alone and requires either larger and more diverse training datasets or transfer learning strategies that explicitly account for programme-level batch effects.
Platform formulation knowledge for IgG1 mAbs -- the general applicability of histidine-buffered, sucrose-stabilised formulations at pH 5.5-6.5 with polysorbate 20 or 80 -- means that full formulation screening is often abbreviated, reducing the amount of formulation-condition-resolved stability data generated per molecule and thereby limiting the training data available for formulation-optimisation-level ML models. For IgG1 mAbs that require non-platform formulations (typically those with elevated viscosity, unusual thermal stability profiles, or high-concentration requirements above 100 mg/mL), the need for ML-guided formulation design is greatest, but the available training data for those specific cases is also most limited, creating a paradox that is not unique to this modality.

6.2. Bispecific and Multispecific Antibodies

Bispecific antibodies (bsAbs) present a qualitatively different challenge from conventional mAbs. Their asymmetric architectures -- including knob-into-hole IgG1, CrossMab, DVD-Ig, IgG-scFv, IgG-VHH, and tandem Fab formats among others -- introduce charge asymmetries between the two Fv domains, increased surface heterogeneity, and the potential for interface-mediated self-association between the two antigen-binding arms that has no analogue in conventional IgG1 [17]. Platform IgG1 formulation approaches frequently fail for bispecific constructs, particularly those containing scFv or VHH domains with unfavorable pI or hydrophobicity profiles relative to the Fab component, because the domain-level charge and hydrophobicity imbalances drive colloidal instability at acidic formulation pH through mechanisms that conventional SAP or pI descriptors calculated on the full molecule do not capture [100].
Mullin et al. (2024) at AstraZeneca systematically reviewed the applications and challenges of ML for VHH-based bispecific antibody developability, identifying the absence of VHH-specific biophysical training datasets, the lack of bispecific-format-specific ML tools, and the non-transferability of IgG1-trained models to asymmetric formats as the principal bottlenecks [101]. Boje et al. (2025) provided the first systematic study of pI engineering for colloidal stability optimisation in bispecific IgG1-VHH constructs, demonstrating that aligning domain-level pI profiles across Fab and VHH components to a slightly basic range (approximately 7.5-9.0) reduced charge asymmetry-driven self-association at acidic formulation pH and substantially improved kD and viscosity [100]. While this study was not framed as an ML application, the resulting dataset of VHH pI variants across four formulation buffers is exactly the type of publicly available structure-activity data that would support domain-level ML models for bispecific colloidal stability.
From a manufacturing perspective, bispecific antibody aggregation risk is further complicated by the presence of misassembled species -- homodimers, half-antibodies, and chain-mispairing products -- that co-elute with or near the main peak in SEC and whose physicochemical properties overlap with those of the correctly assembled bispecific. ML models trained on SEC monomer content as the primary stability label for bispecifics may therefore capture the combined stability of the target molecule and its closest mis-assembled contaminants, introducing label noise that is structurally different in kind from the measurement uncertainty discussed for mAbs in Section 2.8. No published ML formulation study has explicitly addressed this problem for bispecific antibodies, and it represents a modality-specific methodological gap.

6.3. Antibody-Drug Conjugates

Antibody-drug conjugates add a chemically distinct hydrophobic payload, connected to the antibody scaffold via a cleavable or non-cleavable linker, whose physicochemical properties interact non-additively with those of the protein component to determine the colloidal behaviour of the conjugate. Site-specific and stochastic conjugation methods produce ADCs with drug-to-antibody ratios (DARs) of 2-8 payload molecules per antibody, with each additional payload unit increasing the overall hydrophobicity of the molecule surface in a manner proportional to the payload’s logP and the conjugation site’s solvent exposure [102]. ADCs consistently exhibit lower thermal stability and lower colloidal stability relative to their unconjugated parent antibodies, with reductions in Tm of 2-8 degrees C and kD values shifted negative by 5-15 mL/g on average, depending on DAR and payload hydrophobicity [103].
Prediction of post-conjugation colloidal stability from sequence and structural features alone is not feasible with current models, because the payload identity, DAR distribution, and conjugation site determine the dominant drivers of instability and these are not encoded in standard protein descriptors. A limited number of ML approaches have been reported for ADC-specific property prediction. Prihoda et al. (2022) described an ensemble ML platform at AstraZeneca that incorporated post-conjugation biophysical data including HIC retention time, SEC monomer content, and plasma stability assay outputs to predict ADC developability risk flags, demonstrating that gradient boosting classifiers trained on a proprietary internal dataset of several hundred ADC candidates could classify high-aggregation-risk ADCs with sufficient accuracy to support early candidate triage [104]. The review by Wang et al. (2025) on AI in ADC development identified linker stability in plasma, payload hydrophobicity management, and post-conjugation aggregation propensity as areas where ML support remains limited [105].
A specific challenge for ADC formulation ML is the near-total absence of publicly available post-conjugation stability datasets. ADC development is highly competitive and payload-linker combinations are closely guarded intellectual property; the probability that an industry group will deposit their ADC formulation screening data in a public repository is lower even than for conventional mAbs. This data scarcity is compounded by analytical complexity: characterising ADC aggregation requires orthogonal methods beyond SEC, including HIC-based DAR distribution analysis to distinguish aggregation from DAR heterogeneity shifts, and the resulting analytical datasets are expensive to generate. The development of purpose-built, ADC-specific ML frameworks -- including physics-informed models that explicitly represent payload hydrophobicity and conjugation-site solvent exposure as molecular descriptors -- is an open research direction with no published exemplars at scale as of mid-2025.

6.4. mRNA-Lipid Nanoparticle Formulations

mRNA-LNP systems represent a physicochemically distinct class of biopharmaceutical in which the colloidal stability problem is primarily a particle physics problem rather than a protein folding problem. The key stability attributes of an mRNA-LNP are particle size (Z-average diameter, typically 70-120 nm), PDI, encapsulation efficiency (EE%), and pKa of the ionisable lipid (apparent pKa 6.2-6.8 for optimal endosomal escape without systemic toxicity), and their preservation under storage, freeze-thaw, and thermal stress conditions [18]. The mRNA cargo contributes additional degradation modes including hydrolytic cleavage of phosphodiester bonds, particularly at acidic pH or elevated temperature, and oxidative modification of modified nucleosides. Particle aggregation in mRNA-LNPs is driven by reduction in PEG-lipid steric barrier, ionisable lipid phase transitions, and osmotic stress during freezing -- mechanisms that differ fundamentally from protein-protein interaction-driven aggregation in mAbs.
ML approaches for mRNA-LNP formulation optimisation have developed rapidly since 2021, driven by the scale of COVID-19 vaccine manufacturing and the commercial imperative to develop next-generation mRNA therapeutics for oncology, rare disease, and infectious disease indications. The dominant ML application has been prediction of particle physicochemical properties (size, PDI, EE%) and transfection efficiency as a function of lipid composition (ionisable lipid identity, molar ratios of DSPC or DOPE, cholesterol, and PEG-lipid), mRNA N/P ratio, and microfluidic process parameters [106]. Maharjan et al. (2024) applied XGBoost with Bayesian optimisation and a self-validated ensemble model to a 24-formulation I-optimal DoE dataset, using formulation and process variables to predict critical quality attributes including particle size, PDI, and encapsulation efficiency [107]. Nakamura et al. (2025) used random forest regression on 213 LNP compositions to identify phenol-containing ionisable lipid substructures and carbon chain length as the dominant structural determinants of mRNA expression efficiency, providing a structure-property relationship that is directly actionable for lipid design [108].
A more recent and methodologically sophisticated application was reported by Mizogaki et al. (2025), who applied gradient boosting to evaluate lyophilisation process parameters -- including freezing rate, primary drying temperature, and annealing step -- for mRNA-LNP formulations, identifying freezing rate and annealing temperature as influential variables for post-lyophilisation particle size [86]. This study is particularly relevant because lyophilisation of mRNA-LNPs is an active area of pharmaceutical development driven by the cold chain limitations of current liquid mRNA vaccine formulations, and ML-guided lyophilisation cycle development could substantially reduce the empirical experimentation burden at technology transfer. A parallel study by Youssef et al. (2025) generated a publicly available dataset of mRNA-LNP stability across six sugar-surfactant cryoprotectant combinations in Tris buffer under freeze-thaw and 4 degrees C storage conditions, providing a training resource for ML-guided cryoprotectant selection that was not exploited computationally in the original publication [85].
A fundamental gap in current mRNA-LNP ML is the absence of models that predict colloidal stability over long-term storage timescales -- weeks to months -- from short-term biophysical characterisation data, analogous to the Arrhenius kinetic models validated for mAbs. mRNA-LNP degradation kinetics involve at least three coupled processes: mRNA hydrolysis (first-order, temperature-dependent), ionisable lipid oxidation (zeroth-to-first-order), and particle aggregation (concentration- and time-dependent) -- a multivariate kinetic system that is considerably more complex than single-pathway protein aggregation. No published ML model for mRNA-LNP stability addresses this multi-pathway kinetic prediction problem, and its solution would require time-series stability datasets that are not yet publicly available for the range of commercially relevant ionisable lipid and excipient combinations.

6.5. Cell and Gene Therapy Vectors

Adeno-associated virus (AAV) and lentiviral vector (LV) formulations represent the most analytically and physically complex modality addressed in this review. AAV capsids are 25 nm protein shells assembled from 60 VP1, VP2, and VP3 subunits, encapsulating single-stranded DNA payloads of approximately 4.7 kb. Their colloidal stability is governed by the electrostatic and hydrophobic properties of the external capsid surface, the buffer pH and ionic strength, and the ratio of full (genome-containing) to empty (genome-absent) capsids in the preparation -- the latter being a critical quality attribute with direct implications for immunogenicity and therapeutic dose [109]. Lentiviral vectors are enveloped RNA viruses with a lipid bilayer, capsid proteins, and RNA payload; their stability is governed by envelope protein integrity, RNA stability, and membrane fluidity, and they are substantially more sensitive to freeze-thaw stress, shear stress, and ionic strength variations than AAV capsids [110].
The application of ML to AAV and LV formulation stability is at an earlier stage than for protein therapeutics. The most advanced published application is the Bayesian optimisation framework of Lemmens et al. (2025) at KU Leuven and UCB, which applied a GP-based adaptive BO loop to optimise AAV affinity chromatography purification conditions -- yield, purity, and full-to-empty capsid ratio -- achieving performance exceeding conventional DoE optimisation in 30 iterative experiments [111]. While this addresses a manufacturing process optimisation problem rather than formulation stability prediction per se, the underlying methodology -- adaptive BO with multi-objective acquisition on a GP surrogate trained on small DoE data -- is directly applicable to formulation development once suitable response variables (particle size stability, infectivity retention, aggregate particle count) are defined.
For AAV formulation specifically, Srivastava et al. (2021) provided a systematic review of formulation development principles, documenting the importance of buffer pH, ionic strength, excipients, and stress conditions for AAV stability [109]. This physicochemical understanding is sufficient to frame a sequence-to-stability ML problem for AAV serotypes -- mapping capsid protein sequence composition and predicted surface charge distribution to colloidal stability parameters -- but no published study has yet done so at a scale comparable to the mAb developability literature. The principal obstacles are that AAV capsid sequences are considerably longer and more structurally complex than antibody variable domains, that full capsid surface charge calculation from structural models is computationally expensive, and that publicly available AAV formulation stability datasets with systematic variation of formulation conditions are essentially non-existent.
Cell therapies including CAR-T and CAR-NK products present a yet more complex formulation challenge where the colloidal stability problem is largely replaced by cell viability and functional integrity as the primary quality attributes, though the formulation of viral vector components used in cell transduction steps is directly relevant to the stability considerations reviewed above. The application of ML to CAR-T formulation development has not been reported in the published literature as of mid-2025, and this modality remains entirely outside the scope of current formulation ML frameworks.

6.6. Recombinant Fusion Proteins and Non-Antibody Biologics

Beyond antibody-based modalities, the broader class of recombinant fusion proteins -- including Fc-fusion proteins such as etanercept, receptor-IgG fusions, and bispecific T-cell engagers (BiTEs) in single-chain formats -- presents colloidal stability challenges intermediate between conventional mAbs and bispecific antibodies. Fc-fusion proteins benefit from the colloidal stability conferred by the Fc region but are more susceptible to aggregation at their fusion junctions, where non-IgG protein domains introduce surface hydrophobicity or charge patterns that are poorly predicted by IgG1-trained ML models [17]. ScFv-based constructs, including BiTEs and some bispecific formats, are particularly prone to reversible self-association and temperature-dependent aggregation due to the inherent conformational flexibility of the VH-VL interface in the absence of the inter-domain disulfide bond present in conventional Fab domains; this instability mode is not captured by the Fv pI or CDR composition descriptors that dominate IgG1 viscosity models [101].
Protein nanoparticles -- including self-assembling ferritin nanoparticles used as vaccine platforms and engineered nanoparticle scaffolds for multivalent antigen display -- represent a modality where colloidal stability prediction from sequence is structurally tractable in principle, as the self-assembly is governed by well-defined inter-subunit interfaces that could be described by structure-derived ML features. In practice, no ML models specific to protein nanoparticle colloidal stability have been published, and this modality is absent from all existing formulation ML benchmark datasets. Its inclusion in future dataset development and ML model design is warranted given the growing therapeutic relevance of nanoparticle-based vaccine platforms demonstrated by the SARS-CoV-2 pandemic response.
The sections above have documented considerable progress in applying ML to biopharmaceutical colloidal stability prediction. That progress should not obscure the substantial unresolved problems that limit the reliability, generalisability, and practical utility of current approaches. This section assembles the principal challenges in one place, drawing together threads from earlier sections and adding analysis that extends beyond what individual topics admit. The challenges are ordered roughly from the most fundamental -- those that constrain every model regardless of architecture or prediction target -- to the more specific problems that affect particular methods or modalities.

7. Challenges and Unresolved Issues

7.1. The Low-N Problem: Small Datasets as a Structural Constraint

The single most consequential constraint on ML progress in biopharmaceutical formulation is the small size of available training datasets. The typical published formulation stability ML study trains on 47-229 molecules -- a range that corresponds to, at best, a medium-sized pharmaceutical company’s single-modality development programme. In contrast, the protein property prediction models that have achieved reliable generalisation in adjacent fields -- solubility prediction for small molecules, binding affinity prediction, and protein thermostability prediction from mutation data -- have been trained on thousands to tens of thousands of data points [112]. The discrepancy arises from a combination of factors: therapeutic protein characterisation is expensive, time-consuming, and conducted under conditions of commercial confidentiality; each molecule in a formulation dataset requires multi-day biophysical measurement campaigns; and long-term stability labels that are the most regulatory-relevant outputs require months to years of real-time data generation that cannot be accelerated without physical constraints.
The consequences of low-N are well documented in the computational biology literature. In small datasets with high feature-to-sample ratios, models tend to overfit the training distribution, capturing noise and dataset-specific artifacts alongside genuine signal. Cross-validation within the training set inflates estimated performance because the limited diversity of molecules means that even held-out test folds are statistically correlated with the training fold [113]. ExPreSo (Vidal-Henriquez et al., 2025), trained on 335 approved drug product formulations, observed that leave-one-group-out cross-validation systematically outperformed blind test set evaluation -- an unambiguous overfitting signature -- even in a carefully engineered model with strong regularisation [50]. The Low-N problem cannot be overcome by more sophisticated model architectures: a gradient boosted tree and a neural network trained on 47 molecules will both produce unreliable estimates of generalisation error when the test set is drawn from the same molecular programme as the training set. The solution requires either more data -- which is a problem of scientific culture and data sharing rather than methodology -- or more principled use of physics-informed inductive biases that reduce the effective number of parameters the model must estimate from data.

7.2. Dataset Bias and Non-Representative Training Distributions

Even when datasets are nominally large, they may be systematically biased in ways that limit the generalisability of trained models. The Jain et al. (2017) 137-mAb dataset, the most widely cited open benchmark in the field, consists entirely of clinical-stage IgG1 antibodies from programmes that had already passed initial developability triage. It is therefore enriched for molecules with acceptable biophysical properties, underrepresenting the high-aggregation, high-viscosity tail of the antibody sequence space that is most practically important to predict [45]. A model trained exclusively on this dataset learns the within-good-candidate distinctions that are relevant for fine-grained ranking of already-acceptable molecules but cannot reliably flag severely problematic candidates that would never have reached clinical stage. This selection bias is not a flaw in the dataset design -- it reflects the practical reality of how pharmaceutical programmes operate -- but it must be explicitly accounted for when extrapolating model predictions to early-stage discovery libraries that include molecules never previously characterised.
A related form of bias is programme-level batch effects: molecules from the same discovery programme tend to share sequence families, expression systems, purification protocols, and analytical laboratory practices. When a training dataset is dominated by two or three molecular programmes and the test set is drawn from the same programmes, the model may learn programme-specific artefacts -- for example, a correlation between a particular Fc mutation pattern used in one company’s scaffold and stability outcomes -- rather than general sequence-to-stability relationships. Correcting for these batch effects requires explicitly encoding programme identity as a covariate or using domain-adaptive transfer learning methods that have not yet been applied in the formulation ML literature [49].

7.3. The External Validation Deficit

External validation -- testing a model on data generated by an independent laboratory on an independent set of molecules using independent protocols -- is the minimum standard for claiming that a model predicts a property rather than memorising it. By this standard, the biopharmaceutical formulation ML literature performs poorly. A survey of 30 published formulation stability ML studies conducted by Khetan et al. (2022) found that fewer than 20% reported any form of external validation; the remainder reported only internal cross-validation metrics that are unreliable estimators of out-of-distribution performance in small datasets [49]. Among those studies reporting external validation, several used molecules from the same company as the training set but from a different project, which provides partial independence but does not address the programme-level batch effects described above.
The prospective validation study -- training a model, making predictions on new molecules before biophysical data are collected, and then comparing predictions to subsequently measured values -- is the most rigorous form of validation and provides the only unambiguous evidence that a model has generalisation utility. Examples in the formulation literature are rare. Makowski et al. (2024) is the most frequently cited positive example; Gentiluomo et al. (2020) trained on molecules measured before 2018 and validated on molecules measured subsequently using the PIPPI infrastructure [20]. The systematic absence of prospective validation from the published literature reflects both the time and resource cost of generating new experimental data and the academic incentive structure, which rewards publication of novel models over painstaking validation of existing ones. Until prospective validation becomes a standard expectation for formulation ML publications -- as it is, for example, for clinical prediction model reporting under the TRIPOD guidelines -- the field will continue to accumulate models whose real-world utility is substantially lower than their reported test set metrics suggest.

7.4. Label Noise, Assay Heterogeneity, and Measurement Reproducibility

ML models learn from labels. In biopharmaceutical formulation datasets, those labels are experimental measurements that carry irreducible measurement uncertainty and, more problematically, systematic inter-laboratory and inter-assay biases. Section 2.8 documented the specific problem of assay cross-comparability for self-association measurements; the same issue applies to all stability endpoints. SEC monomer content measured at 40 degrees C for four weeks in one laboratory may differ by 3-8 percentage points from the same measurement performed on the same molecule in a different laboratory using a different SEC column lot, a different incubator, and a different UV detector [16]. When these measurements are aggregated across laboratories into a pooled training dataset -- as is implicitly done whenever multi-source datasets like the PIPPI consortium data are used -- the label noise is structured: it correlates with laboratory identity rather than molecular properties, and standard ML regularisation methods are not designed to handle this type of non-random label noise.
The accelerated-to-real-time extrapolation problem, introduced in Section 4.1.2 and Section 4.6, is a specific form of label noise with systematic rather than random character. Models trained on 40 degrees C or 45 degrees C SEC aggregation rates as proxies for 5 degrees C shelf-life will systematically overestimate aggregation at refrigerated conditions for molecules whose aggregation follows a non-Arrhenius temperature dependence -- that is, for most therapeutic mAbs near their temperature of maximum stability [77]. This overestimation is not a model failure; it is a label validity failure. Correcting it requires either using kinetically modelled shelf-life predictions as training labels rather than raw accelerated aggregation rates, or explicitly incorporating the temperature extrapolation model into the ML architecture as a physics-informed component. Neither approach is standard in the literature.

7.5. Class Imbalance and Rare Event Prediction

Many of the most practically important colloidal stability failures -- visible particle formation, opalescence onset, LLPS at therapeutic concentrations, catastrophic aggregation during freeze-thaw -- are rare events in well-curated biopharmaceutical datasets, because the most severely unstable candidates are typically eliminated before reaching the full formulation screening stage. Class imbalance of 10:1 or higher between stable and unstable outcomes is common in formulation datasets, and standard accuracy metrics are misleading under these conditions: a classifier that always predicts stable achieves 90% accuracy on a 10:1 imbalanced dataset while providing zero predictive utility [114]. Matthews correlation coefficient (MCC), balanced accuracy, and the area under the precision-recall curve are more informative metrics for imbalanced formulation classification problems, but they are not consistently reported in published studies.
Standard mitigation strategies for class imbalance -- synthetic minority oversampling (SMOTE), cost-sensitive learning with class-weighted loss functions, and undersampling -- carry specific risks in the formulation context that are not always acknowledged. SMOTE-generated synthetic molecules created by interpolating between real sequences in feature space may not correspond to chemically realisable molecules and can introduce spurious training instances that distort decision boundaries in ways that are difficult to diagnose. Cost-sensitive learning with high minority class weights can produce models that minimise false negatives at the cost of high false positive rates, incorrectly flagging acceptable candidates as unstable and increasing experimental resource expenditure on unnecessary reformulation work. The choice of imbalance mitigation strategy should be guided by the specific cost asymmetry of false positives and false negatives in the intended decision context, an analysis that published formulation ML studies rarely perform explicitly.

7.6. The Proprietary Data Silo Problem

The most information-rich formulation stability datasets in the world reside in the internal data systems of large biopharmaceutical companies and are inaccessible to the research community. A company with a portfolio of 200-400 biologics in development over 20 years may have accumulated formulation screening data on thousands of molecules across tens of thousands of conditions, with real-time stability outcomes extending to 36 months. This dataset would dwarf everything in the public domain combined. The commercial sensitivity of this data -- it encodes formulation strategies, developability fingerprints, and stability risk profiles for both marketed and pipeline products -- makes voluntary sharing at a level useful for ML training essentially impossible under current intellectual property and competitive dynamics [14].
Several partial solutions have been proposed and attempted. Consortium-based data sharing under pre-competitive agreements -- modelled on the Wellcome Trust’s ATOM consortium for small molecule toxicity prediction or the Critical Path Institute’s Coalition Against Major Diseases -- has been discussed in the biologics formulation context but not implemented at scale. The PIPPI consortium represents the closest analogue in the field, but its dataset, while systematically curated, covers a relatively small number of molecules and conditions compared to what a single major pharmaceutical company holds internally. Federated learning -- where models are trained across distributed datasets without centralising the underlying data -- offers a technical solution to the data sharing problem but requires coordination infrastructure, trust between participating organisations, and standardised feature representations that have not been established for formulation ML [48]. Until one or more of these data sharing mechanisms is implemented at scale, the field will remain structurally limited to models trained on datasets that are too small for reliable generalisation.

7.7. The Benchmarking and Reproducibility Crisis

Reproducibility in ML-based science is a well-documented general problem: a systematic review of highly cited AI papers found that fewer than half were reproducible to any extent [115]. In the biopharmaceutical formulation ML literature, the reproducibility problem is compounded by the absence of standardised benchmarks, the use of proprietary datasets that cannot be shared, and the inconsistent reporting of model architectures, hyperparameter search strategies, and data preprocessing steps. A researcher attempting to replicate the results of Lai et al. (2022) encounters the challenge that the 47-mAb AstraZeneca dataset is available as supplementary data but the exact train-test split, the molecular dynamics feature calculation protocol, and the hyperparameter settings used to produce the reported AUC values are not fully specified [21]. This is not an indictment of that study specifically -- it is representative of reporting standards in the field -- but it means that the published AUC of 0.85 cannot be verified independently or used as a reliable baseline for comparing alternative approaches.
The path forward requires both community norms and infrastructure. On the norms side, formulation ML publications should be expected to provide code, exact train-test splits with random seeds, complete feature engineering pipelines, and calibrated uncertainty estimates alongside point predictions. On the infrastructure side, a public challenge competition analogous to CASP for protein structure or the Drug Discovery Hackathon format, built around a newly assembled multi-laboratory formulation stability benchmark with pre-defined evaluation metrics, would catalyse the kind of rapid, reproducible progress that characterised the protein structure prediction field between CASP12 and CASP14. The assembly of such a benchmark requires cross-institutional collaboration and funding that has not yet materialised, but the scientific case for it is clear.

7.8. Black-Box Models and the Trust Deficit

The gap between ML prediction accuracy and practical adoption in formulation development workflows is substantially larger than the gap between current ML accuracy and the accuracy that would theoretically be needed. Formulation scientists do not, in general, adopt predictive tools simply because they are reported to achieve AUC of 0.85 on a test set they cannot inspect. They adopt tools when those tools provide predictions that connect to their physicochemical intuition, are accompanied by meaningful uncertainty estimates, and have been demonstrated to work prospectively on molecules from their own programme. None of these conditions is routinely satisfied by current formulation ML models [88].
The tension between model complexity and interpretability, discussed in the context of XAI in Section 5, is more than an academic concern. In a regulatory environment that increasingly expects justification of AI-assisted decisions, a gradient boosted model with 500 trees that achieves AUC of 0.87 on an internal test set but cannot provide a physically meaningful explanation for any individual prediction is less useful than a linear model with AUC of 0.80 whose five coefficients directly correspond to known biophysical drivers and can be cited in a CMC regulatory dossier. The field has not yet resolved the question of how much predictive accuracy is appropriately traded for interpretability in formulation decision contexts, and until it does, the adoption of ML-guided formulation design will remain episodic and informal rather than embedded in standard development workflows.

7.9. Missing Physical Mechanisms: What ML Cannot Currently Learn from Existing Data

A final challenge that deserves explicit acknowledgement is the absence from current training datasets of information about the physical mechanisms of aggregation. ML models trained on SEC monomer content as a function of sequence features learn to predict aggregate accumulation but have no access to information about whether the dominant aggregation pathway is native-state self-association, partially unfolded intermediate aggregation, surface-mediated nucleation, or LLPS-driven condensation. These pathways are physically distinct, governed by different molecular interactions, and respond differently to formulation interventions -- yet current models make no distinction between them. A molecule that aggregates by native-state self-association is likely to respond to Fv pI engineering; one that aggregates through a partially unfolded intermediate may be more responsive to conformational stabilisers such as sucrose or arginine. Predicting which intervention will be effective requires knowing the pathway, which requires mechanistic information not present in current training datasets [80]. Incorporating mechanistic pathway classification -- potentially from orthogonal biophysical measurements such as ANS fluorescence, thioflavin T binding, or hydrogen-deuterium exchange -- as additional training features or as structural constraints in physics-informed ML models represents a frontier that the field has identified but not yet meaningfully addressed.
The challenges catalogued in Section 7 are not merely obstacles to incremental progress; they define the research agenda for the next decade. This section identifies the specific methodological, infrastructural, and collaborative developments that would most substantially advance the field, ordered by time horizon and assessed against the enabling conditions required for each. The section closes with a proposed roadmap table and a set of key take-home messages.

8. Future Directions and a Roadmap for the Next Decade

8.1. Foundation Models Conditioned on Formulation Context

The success of protein language models such as ESM-2 and ProtT5 in capturing sequence-level evolutionary and structural information relevant to protein properties suggests a natural direction: training analogous foundation models that condition on both protein sequence and formulation environment simultaneously. Such a model would learn joint representations of (sequence, pH, buffer, excipient, concentration, temperature) tuples and their relationship to colloidal stability outcomes, enabling generalisation across both sequence space and formulation space from a unified learned representation [62]. The technical requirements are substantial: the training dataset would need to cover thousands of protein-formulation combinations with quantitative stability labels, which does not exist in any public resource today. However, the rapid accumulation of published formulation screening data, combined with pre-competitive data sharing initiatives and synthetic data augmentation strategies -- for example, using physics-informed generative models to extrapolate measured stability profiles to unstudied formulation conditions -- makes this a technically tractable five-to-ten-year horizon.
A more immediately achievable step toward formulation-conditioned representation learning is the development of antibody-specific language models fine-tuned on the Observed Antibody Space (OAS) database and augmented with biophysical property labels from combined PIPPI, Jain, and other public datasets. IgBert and IgT5, already trained on OAS-derived sequences, provide a starting point; the addition of stability-label fine-tuning heads for kD, Tm, viscosity, and SEC monomer content would produce a multi-task PLM with direct formulation utility. The key methodological challenge is preventing catastrophic forgetting of sequence-level representations during fine-tuning on small stability datasets -- a problem for which parameter-efficient fine-tuning methods such as LoRA and adapter modules, borrowed from natural language processing, offer promising solutions that have not yet been applied in the formulation PLM domain [116].

8.2. Physics-Informed Machine Learning: Encoding Colloidal Theory as Inductive Bias

The most durable improvements in formulation ML accuracy and data efficiency will come not from larger models trained on more data but from architectures that encode known physical constraints as inductive biases. The Arrhenius kinetic modelling framework for shelf-life prediction, discussed in Section 4.6 and Section 8.4, is already a form of physics-informed ML in the sense that the model architecture is constrained by the Arrhenius equation rather than being a free-form function approximator. Extending this principle to colloidal stability prediction requires encoding relationships such as the DLVO interaction potential between protein molecules, the Boltzmann-weighted population of aggregation-competent conformational states, and the concentration dependence of diffusion-limited aggregation kinetics as explicit model components whose parameters are estimated from data [61]. Physics-informed neural networks (PINNs) that incorporate partial differential equations as loss function terms have demonstrated substantial improvements in data efficiency for fluid dynamics and materials science problems; their adaptation to protein colloidal systems is a research frontier that would benefit from collaboration between pharmaceutical scientists, colloid physicists, and ML engineers in a way that has not yet occurred at scale.
A practical near-term implementation is the hybrid MD-ML framework, in which coarse-grained molecular dynamics simulations of protein-protein interactions at formulation-relevant conditions generate physics-grounded descriptors that serve as inputs to lightweight ML classifiers. Prass et al. (2023) demonstrated the viability of this approach for viscosity prediction using atomistic MD and Green-Kubo theory; extending it to aggregation propensity and LLPS prediction using CG-MD models of full-length antibodies under multiple pH and ionic strength conditions is computationally demanding but technically tractable with current GPU compute infrastructure [61]. The development of a validated CG force field parameterised specifically for therapeutic antibody surfaces -- capturing the correct balance of electrostatic, van der Waals, and short-range desolvation forces across the pH and ionic strength range relevant to biopharmaceutical formulation -- would be a community resource of lasting value and is not available in the current literature.

8.3. Autonomous Formulation Laboratories and Closed-Loop Experimentation

Self-driving or autonomous laboratories - systems that integrate robotic liquid handling, automated analytical measurement, and ML-guided experimental design in a closed feedback loop without human bottlenecks - represent the most transformative near-term direction for biopharmaceutical formulation development, as illustrated in Figure 1. The SAMPLE platform of Romero and colleagues (Nature Chemical Engineering, 2024) demonstrated fully autonomous protein engineering for thermostability optimisation, achieving enzyme variants 12 degrees C more thermostable than the starting sequence within a closed-loop system requiring no human experimental decision-making [117]. The autonomous formulation laboratory analogue would couple HT-DLS, nanoDSF, and SEC measurement platforms via automated sample handling to an active learning or Bayesian optimisation controller that selects the next formulation conditions to test based on the current model’s prediction and uncertainty estimates, iterating until the stability target is achieved or the experimental budget is exhausted.
The Novartis MicroCycle platform, reported in 2024 as a system that autonomously synthesises new compounds, performs physicochemical and biochemical assays, analyses data, and selects new compounds for the next cycle, provides a pharmaceutical industry proof of concept for this architecture in small-molecule drug discovery. The adaptation to biopharmaceutical formulation requires resolving several engineering challenges absent from small-molecule SDL implementations: protein samples cannot be synthesised autonomously in the same loop as formulation screening without integrated cell-free or E. coli expression systems; the analytical throughput of SEC and AUC is substantially lower than the UV assays used in small-molecule SDL platforms; and the relevant design space for biologics formulation includes both sequence-level and excipient-level variables that require different physical manipulation systems [118]. These are engineering problems rather than scientific ones, and their solution is a matter of investment and integration rather than fundamental discovery.

8.4. Digital Twins for Formulation and Cold-Chain Management

Digital twins -- virtual replicas of physical systems that receive real-time data and produce predictive outputs for decision support -- have been implemented in biopharmaceutical upstream and downstream manufacturing processes, with applications including bioreactor process optimisation, chromatography cycle development, and lyophilisation cycle control [119]. Their extension to formulation stability and cold-chain management would enable a qualitatively new class of regulatory and operational decisions. A formulation digital twin would integrate kinetic stability model parameters estimated from accelerated screening data, real-time temperature and humidity sensor data from stability chambers, and probabilistic models of cold-chain temperature history to generate continuously updated shelf-life estimates and excursion risk assessments throughout a product’s commercial lifecycle.
The data infrastructure for such a system exists or is within reach: commercial temperature monitoring devices generate continuous IoT data streams during distribution; laboratory information management systems (LIMS) store stability testing results in structured formats; and ERP systems track product location and age. The missing component is the validated kinetic ML model that maps temperature history to degradation trajectory for each product-formulation combination, together with the regulatory framework for using real-time model outputs to make post-market shelf-life decisions. This represents one of the clearest cases where the technical solution is ahead of the regulatory infrastructure, and where proactive engagement between the pharmaceutical industry and regulatory agencies could accelerate implementation substantially.

8.5. Multimodal AI and Integrated Property Prediction

Current formulation ML models are predominantly unimodal: they take either sequence features, or biophysical measurements, or structural coordinates as inputs, and produce a single property as output. The next generation of models should be multimodal -- taking simultaneous inputs from sequence data, 3D structural predictions, biophysical measurement panels, and formulation composition vectors -- and multi-target, producing joint predictions of aggregation propensity, viscosity, solubility, and LLPS risk as correlated outputs with shared uncertainty estimates [62]. Multi-task learning architectures, in which a shared encoder feeds modality-specific or property-specific prediction heads, have demonstrated improved performance over single-task models on related properties in the molecular property prediction literature, with the shared representation acting as a regulariser that prevents individual heads from overfitting on small property-specific datasets.
Large multimodal models that process text, sequence, structure, and tabular data jointly -- exemplified in adjacent fields by ESM-3, which jointly reasons over sequence, structure, and function -- suggest a longer-term direction for formulation AI in which a single model accepts a query of the form (antibody sequence, target formulation conditions, available biophysical measurements) and returns a structured developability assessment with property predictions, uncertainty estimates, SHAP-attributed molecular drivers, and excipient recommendations, all grounded in the model’s learned representation of the protein-formulation stability relationship [120]. Achieving this vision requires not only the model architecture and training infrastructure but also the curated multimodal training datasets that do not yet exist in the public domain. The assembly of such datasets -- a decade-long community project analogous to the PDB for protein structures -- is the single most impactful investment the biopharmaceutical research community could make in this field.

8.6. Pre-Competitive Data Sharing and Community Infrastructure

The data sharing problem identified in Section 7.6 will not be solved by technical means alone. It requires a change in scientific culture and in the commercial calculus of pharmaceutical companies regarding what constitutes a pre-competitive asset. The precedents from other domains are instructive: the ATOM (Accelerating Therapeutics for Opportunities in Medicine) consortium, a public-private partnership between the US Department of Energy, the FDA, and pharmaceutical companies including AbbVie, Pfizer, and Janssen, demonstrated that companies could share toxicology datasets for ML model development without compromising competitive advantage on the clinical or commercial side [48]. A biopharmaceutical formulation stability analogue -- a consortium depositing anonymised formulation screening data for approved products into a curated, access-controlled repository -- would generate the dataset scale needed for reliable foundation model training while limiting competitive risk through data anonymisation and delayed release provisions.
Federated learning provides an additional technical mechanism: companies participate in joint model training by sharing model gradients rather than raw data, enabling a global model to benefit from all participants’ datasets without any individual dataset leaving its originating system [48]. Federated learning for biopharmaceutical stability prediction faces the challenge that gradient sharing can in principle enable model inversion attacks that partially reconstruct training data, requiring differential privacy mechanisms that add noise to shared gradients at some cost to model convergence. The appropriate privacy-utility trade-off for formulation stability datasets -- which contain commercially sensitive but not personally identifiable information -- is a question that has not yet been addressed by the biopharmaceutical informatics community and deserves priority attention.

8.7. Proposed Research Roadmap

The principal future directions described in this section are organised into a roadmap across three time horizons: near-term (1-3 years), mid-term (3-6 years), and long-term (6-10 years). For each horizon, priority actions, key enabling conditions, and measurable success metrics are identified. The roadmap is not a prediction but a set of conditional statements: if specific investments and collaborations materialise, the field can achieve defined milestones. It is offered as a basis for discussion and prioritisation by the research community rather than as a fixed programme.

8.8. Key Take-Home Messages

The following statements summarise the core conclusions of this review and are offered as a guide for researchers, formulation scientists, and research funders entering or evaluating this field.
First, ML for colloidal stability prediction is a genuinely emerging field, not a mature one. The models that exist work, to varying degrees, for IgG1 monoclonal antibodies under accelerated stability conditions and provide useful but unreliable guidance at the candidate selection stage. They do not yet generalise reliably across molecular programmes, do not address the full range of biopharmaceutical modalities, and have not been validated prospectively at scale.
Second, the binding constraint on progress is not algorithmic sophistication but data. The most consequential investment the field could make is in the systematic generation, standardisation, and sharing of high-quality formulation stability datasets with real-time stability labels. Incremental improvements in model architecture on existing small datasets will produce diminishing returns.
Third, physics-informed approaches outperform purely data-driven ones on small datasets, and this advantage is structural rather than contingent. The Arrhenius kinetic modelling framework, validated across multiple companies and modalities, should be the baseline against which new purely data-driven ML approaches are benchmarked, not an older approach to be superseded.
Fourth, XAI methods are valuable tools for hypothesis generation and model debugging but should not be mistaken for mechanistic evidence. SHAP attributions that reproduce known biophysical relationships are consistent with model validity but do not confirm it. Prospective experimental validation of XAI-generated molecular hypotheses is the only reliable test of whether a model has learned causal structure.
Fifth, regulatory engagement is not a downstream afterthought but an upstream design constraint. ML models developed for formulation decisions that cannot be documented, validated, and explained in a manner consistent with emerging ICH and FDA/EMA expectations will not be adopted in regulated development workflows regardless of their predictive accuracy. Regulatory science for AI-assisted formulation development must develop in parallel with the technical methods, not after them.
Sixth, the modalities most in need of ML guidance -- bispecific antibodies, ADCs, mRNA-LNP systems, and gene therapy vectors -- are precisely those for which the least data-driven ML infrastructure exists. Investment in modality-specific dataset generation and model development for these formats is more impactful than further refinement of IgG1-centric approaches.

Supplementary Materials

The following supplementary files are available online. Supplementary Table S1: Principal publicly or semi-publicly available datasets used or available for machine learning in biopharmaceutical colloidal stability prediction. Proprietary industry datasets are listed for completeness but are not accessible to external researchers. Supplementary Table S2: Machine learning algorithm classes applied to biopharmaceutical colloidal stability and developability prediction, with representative methods, example applications, and key strengths and limitations relative to formulation dataset characteristics.

Author Contributions

Conceptualization, C.V.M.-P.; Methodology, C.V.M.-P.; Investigation, C.V.M.-P.; Resources, C.V.M.-P.; Data Curation, C.V.M.-P.; Writing - Original Draft Preparation, C.V.M.-P.; Writing - Review and Editing, C.V.M.-P.; Visualization, C.V.M.-P. The author has read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analysed in this study. Data sharing is not applicable to this article. All datasets discussed are cited in the reference list and described in Section 2. Publicly available datasets referenced include the PIPPI consortium dataset (pippi-data.kemi.dtu.dk), the Jain et al. (2017) and Arsiwala et al. (2025) mAb biophysical panels, the Lai et al. (2022) AstraZeneca supplementary dataset, and the Rosace et al. (2023) Nature Communications dataset. All other datasets referenced are proprietary and not publicly available.

Acknowledgments

The author thanks the broader biopharmaceutical formulation science and computational biology communities whose published work forms the foundation of this review. During the preparation of this manuscript, Claude (Anthropic, claude-sonnet-4-6, 2026) was used to assist with manuscript structuring, draft text generation, and reference organisation. The author has reviewed and edited all outputs and takes full responsibility for the content of this publication.

Conflicts of Interest

The author declares no conflicts of interest.

Abbreviations

A2/B22: Second osmotic virial coefficient
AC-SINS: Affinity capture self-interaction nanoparticle spectroscopy
ADC: Antibody-drug conjugate
AUC: Analytical ultracentrifugation
BO: Bayesian optimisation
CDR: Complementarity-determining region
CQA: Critical quality attribute
DAR: Drug-to-antibody ratio
DLS: Dynamic light scattering
DoE: Design of experiments
DSC: Differential scanning calorimetry
DSF: Differential scanning fluorimetry
ECD: Equivalent circular diameter
EE%: Encapsulation efficiency
EMA: European Medicines Agency
ESM: Evolutionary Scale Modelling (protein language model family)
Fc: Fragment crystallisable region
FDA: U.S. Food and Drug Administration
FNN: Feedforward neural network
Fv: Fragment variable
GBDT: Gradient boosted decision tree
GNN: Graph neural network
GP: Gaussian process
HIC: Hydrophobic interaction chromatography
HT-DLS: High-throughput dynamic light scattering
ICH: International Council for Harmonisation
IgG: Immunoglobulin G
kD: Diffusion interaction parameter
LIME: Local Interpretable Model-Agnostic Explanations
LLPS: Liquid-liquid phase separation
LNP: Lipid nanoparticle
mAb: Monoclonal antibody
MD: Molecular dynamics
MFI: Micro-flow imaging
MLP: Multilayer perceptron
mRNA: Messenger ribonucleic acid
NTA: Nanoparticle tracking analysis
PDI: Polydispersity index
pI: Isoelectric point
PIML: Physics-informed machine learning
PLM: Protein language model
QbD: Quality by Design
Rh: Hydrodynamic radius
SAP: Spatial aggregation propensity
SASA: Solvent-accessible surface area
scFv: Single-chain variable fragment
SEC: Size-exclusion chromatography
SHAP: SHapley Additive exPlanations
SMAC: Self-interaction chromatography
SVM: Support vector machine
SVP: Sub-visible particle
Tagg: Onset aggregation temperature
Tm: Thermal melting temperature
VHH: Variable domain of a heavy-chain-only antibody (nanobody)
XAI: Explainable artificial intelligence
XGBoost: eXtreme Gradient Boosting

References

  1. Mordor Intelligence. Biologics Market Size, Share & Growth Analysis, 2031. Published 2026. Available at: https://www.mordorintelligence.com/industry-reports/biologics-market (accessed June 2025).
  2. Boston Consulting Group. A Strategic Approach to Cost in Biopharma. Published December 2023. Available at: https://www.bcg.com/publications/2023/biopharma-manufacturing-cost-reduction (accessed June 2025).
  3. U.S. Food and Drug Administration. Biologics License Application (BLA) Approvals. Available at: https://www.fda.gov/vaccines-blood-biologics/development-approval-process-cber/biologics-license-applications-bla-process-cder (accessed June 2025).
  4. Roberts CJ. Protein aggregation and its impact on product quality. Curr Opin Biotechnol. 2014;30:211–217. [CrossRef]
  5. Swanson MD, Rios S, Mittal S, Soder G, Jawa V. Immunogenicity risk assessment of spontaneously occurring therapeutic monoclonal antibody aggregates. Front Immunol. 2022;13:915412. [CrossRef]
  6. Lundahl MLE, Fogli S, Colavita PE, Scanlan EM. Aggregation of protein therapeutics enhances their immunogenicity: causes and mitigation strategies. RSC Chem Biol. 2021;2(4):1033–1049. [CrossRef]
  7. Jiskoot W, Randolph TW, Volkin DB, et al. Protein instability and immunogenicity: roadblocks to clinical application of injectable protein delivery systems for sustained release. J Pharm Sci. 2012;101(3):946–954. [CrossRef]
  8. International Council for Harmonisation. ICH Q6B: Specifications: Test Procedures and Acceptance Criteria for Biotechnological/Biological Products. March 1999. Available at: https://www.ich.org/page/quality-guidelines.
  9. European Medicines Agency. Guideline on Development, Production, Characterisation and Specifications for Monoclonal Antibodies and Related Products. EMA/CHMP/BWP/532517/2008 Rev 1. 2012.
  10. U.S. Food and Drug Administration. Guidance for Industry: Inspection of Injectable Products for Visible Particulates. December 2021. Available at: https://www.fda.gov/regulatory-information/search-fda-guidance-documents.
  11. International Council for Harmonisation. ICH Q8(R2): Pharmaceutical Development. August 2009. Available at: https://www.ich.org/page/quality-guidelines.
  12. BioProcess International. Navigating the Commercial Cycle of Biologics Manufacturing. Published 2024. Available at: https://www.bioprocessintl.com (accessed June 2025).
  13. Farid SS, Baron M, Stamatis C, Nie W, Coffman J. Benchmarking biopharmaceutical process development and manufacturing cost contributions to R&D. mAbs. 2020;12(1):1754999. [CrossRef]
  14. Fluence Analytics. Industry Challenges in Biopharma Research and Manufacturing. Published February 2025. Available at: https://www.fluenceanalytics.com (accessed June 2025).
  15. Manning MC, Holcomb RE, Payne RW, et al. Stability of protein pharmaceuticals: recent advances. Pharm Res. 2024;41(7):1301–1367. [CrossRef]
  16. Amash A, Volkers G, et al. Developability considerations for bispecific and multispecific antibodies. mAbs. 2024;16(1):2394229. [CrossRef]
  17. Albertsen CH, Kulkarni JA, Witzigmann D, et al. The role of lipid components in lipid nanoparticles for vaccines and gene therapy. Adv Drug Deliv Rev. 2022;188:114416. [CrossRef]
  18. Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–589. [CrossRef]
  19. Gentiluomo L, Roessner D, Frieß W. Application of machine learning to predict monomer retention of therapeutic proteins after long-term storage. Int J Pharm. 2020;577:119039. [CrossRef]
  20. Lai PK, Gallegos A, Mody N, Sathish HA, Trout BL. Machine learning prediction of antibody aggregation and viscosity for high concentration formulation development of protein therapeutics. mAbs. 2022;14(1):2026208. [CrossRef]
  21. Gentiluomo L, Svilenov HL, Augustijn D, et al. Advancing therapeutic protein discovery and development through comprehensive computational and biophysical characterization. Mol Pharm. 2020;17(2):426–440. [CrossRef]
  22. Möller M, Kupreichyk T, Hülsmann M, et al. Design of biopharmaceutical formulations accelerated by machine learning. Mol Pharm. 2022;19(2):576–586. [CrossRef]
  23. Waight AB, Prihoda D, Shrestha R, et al. A machine learning strategy for the identification of key in silico descriptors and prediction models for IgG monoclonal antibody developability properties. mAbs. 2023;15(1):2248671. [CrossRef]
  24. Makowski EK, Chen HT, Wang T, et al. Reduction of monoclonal antibody viscosity using interpretable machine learning. mAbs. 2024;16(1):2303781. [CrossRef]
  25. Bhambure R, Kumar K, Rathore AS. High-throughput process development for biopharmaceutical drug substances. Trends Biotechnol. 2011;29(3):127–135. [CrossRef]
  26. Kidziński Ł, Ong CS, Camarri S, et al. Applications of machine learning in biopharmaceutical process development and manufacturing: current trends, challenges, and opportunities. arXiv:2310.09991. 2023.
  27. Svilenov HL, Arosio P, Menzen T, Tessier PM, Sormanni P. Approaches to expand the conventional toolbox for discovery and selection of antibodies with drug-like physicochemical properties. mAbs. 2023;15(1):2164459. [CrossRef]
  28. Kidziński Ł, Ong CS, Camarri S, et al. Applications of machine learning in biopharmaceutical process development and manufacturing: current trends, challenges, and opportunities. arXiv:2310.09991. 2023.
  29. Yadav S, Shire SJ, Kalonia DS. Factors affecting the viscosity in high concentration solutions of different monoclonal antibodies. Pharm Res. 2010;27(8):1664–1682. [CrossRef]
  30. Saito S, Hasegawa J, Kobayashi N, Kishi N, Uchiyama S, Fukui K. Behavior of monoclonal antibodies: relation between the second virial coefficient (B2) at low concentrations and aggregation propensity and viscosity at high concentrations. Pharm Res. 2012;29(2):397–410. [CrossRef]
  31. Saito S, Hasegawa J, Kobayashi N, Tomitsuka T, Uchiyama S, Fukui K. Effects of ionic strength and sugars on the aggregation propensity of monoclonal antibodies: influence of colloidal and conformational stabilities. Pharm Res. 2013;30(5):1263–1280. [CrossRef]
  32. Wyatt Technology. Predicting and evaluating the stability of therapeutic protein formulations by dynamic light scattering and machine learning. Application Note PL5011. Available at: https://www.wyatt.com.
  33. International Council for Harmonisation. ICH Q6B: Test Procedures and Acceptance Criteria for Biotechnological/Biological Products. March 1999.
  34. Wyatt Technology. SEC-MALS for absolute molecular weight determination of proteins. Technical Note. Available at: https://www.wyatt.com.
  35. Sormanni P, Piovesan D, Heller GT, et al. Simultaneous quantification of protein order and disorder. Nat Chem Biol. 2017;13(4):339–342. [CrossRef]
  36. Menzen T, Friess W. High-throughput melting-temperature analysis of a monoclonal antibody by differential scanning fluorimetry in the presence of surfactants. J Pharm Sci. 2013;102(2):415–428. [CrossRef]
  37. Svilenov HL, Winter G. Intrinsic differential scanning fluorimetry for protein stability assessment in microwell plates. Mol Pharm. 2025;22(4):1789–1802. [CrossRef]
  38. Garidel P, Hegyi M, Bassarab S, Weichel M. A rapid, sensitive and economical assessment of monoclonal antibody conformational stability by intrinsic tryptophan fluorescence spectroscopy. Biotechnol J. 2008;3(9–10):1201–1211. [CrossRef]
  39. Breitsprecher D, Glücklich N, Hawe A, Menzen T. nanoDSF vs. µDSC: a comparative study for biopharmaceutical formulation development. Whitepaper, NanoTemper Technologies. 2016.
  40. USP . Subvisible Particulate Matter in Therapeutic Protein Injections. United States Pharmacopeia. 2015.
  41. Wang S, Liaw A, Chen YM, Su Y, Skomski D. Convolutional neural networks enable highly accurate and automated subvisible particulate classification of biopharmaceuticals. Pharm Res. 2023;40(6):1447–1457. [CrossRef]
  42. Lopez-Del Rio A, Pacios-Michelena A, Picart-Armada S, et al. Sub-visible particle classification and label consistency analysis for flow-imaging microscopy via machine learning methods. J Pharm Sci. 2024;113(4):880–890. [CrossRef]
  43. Poozesh S, Cannavò F, Manikwar P. Sensitivity and uncertainty analysis of micro-flow imaging for sub-visible particle measurements using artificial neural network. Pharm Res. 2023;40(3):721–733. [CrossRef]
  44. Jain T, Sun T, Durand S, et al. Biophysical properties of the clinical-stage antibody landscape. Proc Natl Acad Sci USA. 2017;114(5):944–949. [CrossRef]
  45. Jain T, et al. Identifying developability risks for clinical progression of antibodies using high-throughput in vitro and in silico approaches. mAbs. 2023;15(1):2200540. [CrossRef]
  46. Arsiwala A, Bhatt R, Yang Y, et al. A high-throughput platform for biophysical antibody developability assessment to enable AI/ML model training. bioRxiv. 2025. [CrossRef]
  47. Kidziński Ł, Ong CS, Camarri S, et al. Applications of machine learning in biopharmaceutical process development and manufacturing: current trends, challenges, and opportunities. arXiv:2310.09991. 2023.
  48. Khetan R, Curtis R, Deane CM, et al. Current advances in biopharmaceutical informatics: guidelines, impact and challenges in the computational developability assessment of antibody therapeutics. mAbs. 2022;14(1):2020082. [CrossRef]
  49. Vidal-Henriquez E, Holder T, Lee NF, Pompe C, Teese MG. Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo). Comput Struct Biotechnol J. 2025;27:4517–4525. [CrossRef]
  50. Rosace A, Bennett A, Oeller M, et al. Automated optimisation of solubility and conformational stability of antibodies and proteins. Nat Commun. 2023;14:1937. [CrossRef]
  51. Schiel JE, Davis DL, Borisov OV, eds. State-of-the-Art and Emerging Technologies for Therapeutic Monoclonal Antibody Characterization. ACS Symposium Series. Vol. 1176. American Chemical Society; 2015.
  52. Sormanni P, Aprile FA, Vendruscolo M. The CamSol method of rational design of protein mutants with enhanced solubility. J Mol Biol. 2015;427(3):478-490. [CrossRef]
  53. Lai PK, Fernando A, Cloutier TK, et al. Machine learning applied to determine the molecular descriptors responsible for the viscosity behavior of concentrated therapeutic antibodies. Mol Pharm. 2021;18(3):1167-1175. [CrossRef]
  54. Mock M, Jacobitz AW, Langmead CJ, et al. Development of in silico models to predict viscosity and mouse clearance using a comprehensive analytical data set collected on 83 scaffold-consistent monoclonal antibodies. mAbs. 2023;15(1):2256745. [CrossRef]
  55. Shire SJ, Shahrokh Z, Liu J. Challenges in the development of high protein concentration formulations. J Pharm Sci. 2004;93(6):1390-1402. [CrossRef]
  56. Liu Y, et al. Accelerating high-concentration monoclonal antibody development with large-scale viscosity data and ensemble deep learning. mAbs. 2025;17(1):2483944. [CrossRef]
  57. Shuai RW, Ruffolo JA, Gray JJ. IgLM: Infilling language modeling for antibody sequence design. Cell Syst. 2023;14(11):979-989. [CrossRef]
  58. Nguyen TH, et al. Enhancing protein aggregation prediction: a unified analysis leveraging graph convolutional networks and active learning. RSC Adv. 2024;14:37621-37630. [CrossRef]
  59. Liang T, Sun ZY, Ishima R, et al. ProstaNet: a novel geometric vector perceptrons-graph neural network algorithm for protein stability prediction in single- and multiple-point mutations with experimental validation. Research. 2025:0674. [CrossRef]
  60. Prass TM, Garidel P, Blech M, Schafer LV. Viscosity prediction of high-concentration antibody solutions with atomistic simulations. J Chem Inf Model. 2023;63(19):6129-6140. [CrossRef]
  61. Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123-1130. [CrossRef]
  62. Yu X, Vangjeli K, Prakash A, et al. Application of protein language models for antibody developability prediction. mAbs. 2026;18(1):2647489. [CrossRef]
  63. Warszawski S, Borenstein-Katz A, Lotan L, et al. solPredict: antibody apparent solubility prediction from sequence by transfer learning. bioRxiv. 2022. [CrossRef]
  64. Hao X, Fan L. ProtT5 and random forests-based viscosity prediction method for therapeutic mAbs. Eur J Pharm Sci. 2024;194:106705. [CrossRef]
  65. Elnaggar A, Heinzinger M, Dallago C, et al. ProtTrans: toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell. 2022;44(10):7112-7127. [CrossRef]
  66. Shahriari B, Swersky K, Wang Z, Adams RP, de Freitas N. Taking the human out of the loop: a review of Bayesian optimization. Proc IEEE. 2016;104(1):148-175. [CrossRef]
  67. Snoek J, Larochelle H, Adams RP. Practical Bayesian optimization of machine learning algorithms. Adv Neural Inf Process Syst. 2012;25:2951-2959.
  68. Garrido-Merchán EC, Hernández-Lobato D. Predictive entropy search for multi-objective Bayesian optimization with constraints. Neurocomputing. 2019;361:50–68. [CrossRef]
  69. Raissi M, Perdikaris P, Karniadakis GE. Physics-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. J Comput Phys. 2019;378:686-707. [CrossRef]
  70. Sanchez de Groot N, Pallares I, Aviles FX, Vendrell J, Ventura S. Prediction of ‘hot spots’ of aggregation in disease-linked polypeptides. BMC Struct Biol. 2005;5:18. [CrossRef]
  71. Tartaglia GG, Vendruscolo M. Proteome-level interplay between folding and aggregation propensities of proteins. J Mol Biol. 2010;402(5):919-928. [CrossRef]
  72. Kuriata A, Gierut AM, Oleniecki T, et al. AGGRESCAN3D (A3D) 2.0: prediction and engineering of protein solubility. Nucleic Acids Res. 2019;47(W1):W300-W307. [CrossRef]
  73. Musil M, Planer J, Damborsky J, et al. AggreProt: a web server for predicting and engineering aggregation prone regions in proteins. Nucleic Acids Res. 2024;52(W1):W159-W169. [CrossRef]
  74. Bunc M, Hadzi S, Graf C, Boncina M, Lah J. Aggregation Time Machine: a platform for the prediction and optimization of long-term antibody stability using short-term kinetic analysis. J Med Chem. 2022;65(3):2623-2632. [CrossRef]
  75. Kuzman D, Bunc M, Ravnik M, Reiter F, Zagar L, Boncina M. Long-term stability predictions of therapeutic monoclonal antibodies in solution using Arrhenius-based kinetics. Sci Rep. 2021;11:20534. [CrossRef]
  76. Huelsmeyer M, et al. A universal tool for stability predictions of biotherapeutics, vaccines and in vitro diagnostic products. Sci Rep. 2023;13:10077. [CrossRef]
  77. Wang Y, Latypov RF, Lomakin A, et al. Quantitative evaluation of colloidal stability of antibody solutions using PEG-induced liquid-liquid phase separation. Mol Pharm. 2014;11(5):1391-1402. [CrossRef]
  78. Wei S, Wang Y, Yang G. Liquid-liquid phase separation prediction of proteins in salt solution by deep neural network. Biomolecules. 2022;13(1):42. [CrossRef]
  79. Kimball WD, Lanzaro A, Hurd C, et al. Growth of clusters toward liquid–liquid phase separation of monoclonal antibodies as characterized by small-angle X-ray scattering and molecular dynamics simulation. J Phys Chem B. 2025;129(11):2856–2871. [CrossRef]
  80. Salinas BA, Sathish HA, Bishop SM, Harn N, Carpenter JF, Randolph TW. Understanding and modulating opalescence and viscosity in a monoclonal antibody formulation. J Pharm Sci. 2010;99(1):82–93. [CrossRef]
  81. Dai L, Davis J, Nagapudi K, et al. Predicting long-term stability of an oral delivered antibody drug product with Accelerated Stability Assessment Program modeling. Mol Pharm. 2024;21(1):325-332. [CrossRef]
  82. Gonzalez-Valdez J, et al. Prediction of long-term stability of high-concentration formulations to support rapid development of antibodies against SARS-CoV-2. mAbs. 2025;17(1):2471465. [CrossRef]
  83. Cao E, Chen Y, Cui Z, Foster PR. Effect of freezing and thawing rates on denaturation of proteins in aqueous solutions. Biotechnol Bioeng. 2003;82(6):684-690. [CrossRef]
  84. Youssef M, Hitti C, Fulber JPC, Khan MFH, Perumal AS, Kamen AA. Preliminary evaluation of formulations for stability of mRNA-LNPs through freeze-thaw stresses and long-term storage. Preprint. 2025. [CrossRef]
  85. Mizogaki I, Suzuki T, Ohori R, et al. Evaluating the impact of lyophilization process parameters on mRNA encapsulated lipid nanoparticles using machine learning. J Drug Deliv Sci Technol. 2025;114:107573. [CrossRef]
  86. Willis LF, Trayton I, Saunders JC, et al. Rationalizing mAb candidate screening using a single holistic developability parameter. Mol Pharm. 2025;22(1):181-195. [CrossRef]
  87. Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1(5):206-215. [CrossRef]
  88. Mirakhori F, Niazi SK. Harnessing the AI/ML in drug and biological products discovery and development: the regulatory perspective. Pharmaceuticals. 2025;18(1):47. [CrossRef]
  89. Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30:4768-4777. [CrossRef]
  90. Ribeiro MT, Singh S, Guestrin C. “Why should I trust you?”: explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016:1135–1144. [CrossRef]
  91. Jimenez-Luna J, Grisoni F, Schneider G. Drug discovery with explainable artificial intelligence. Nat Mach Intell. 2020;2(10):573-584. [CrossRef]
  92. Jain S, Wallace BC. Attention is not explanation. Proc 2019 Conf North Am Chapter Assoc Comput Linguist. 2019:3543-3556. [CrossRef]
  93. Ruffolo JA, Sulam J, Gray JJ. Antibody structure prediction using interpretable deep learning. Patterns. 2022;3(2):100406. [CrossRef]
  94. Yuan H, Yu H, Gui S, Ji S. Explainability in graph neural networks: a taxonomic survey. IEEE Trans Pattern Anal Mach Intell. 2023;45(5):5782-5799. [CrossRef]
  95. U.S. Food and Drug Administration. Artificial Intelligence in Drug Manufacturing: Discussion Paper. January 2025. Available at: https://www.fda.gov/media/185875/download.
  96. European Medicines Agency. Reflection Paper on the Use of Artificial Intelligence (AI) in the Medicinal Product Lifecycle. EMA/CHMP/CVMP/QWP/931313/2022. March 2024.
  97. Lenarczyk G.; Minssen T.; Price W.N. II; Rai A. The future of AI regulation in drug development: a comparative analysis. J. Law Biosci. 2025, 12, lsaf028. [CrossRef]
  98. Wachter S, Mittelstadt B, Russell C. Counterfactual explanations without opening the black box: automated decisions and the GDPR. Harvard J Law Technol. 2017;31(2):841-887. [CrossRef]
  99. Boje AS, Arras P, Pekar L, et al. Optimizing colloidal stability and viscosity of multispecific antibodies at the drug discovery-development interface: a systematic predictive case study. mAbs. 2025;17(1):2553622. [CrossRef]
  100. Mullin M, McClory J, Haynes W, Grace J, Robertson N, van Heeke G. Applications and challenges in designing VHH-based bispecific antibodies: leveraging machine learning solutions. mAbs. 2024;16(1):2341443. [CrossRef]
  101. Sun X, Ponte JF, Yoder NC, et al. Effects of drug–antibody ratio on pharmacokinetics, biodistribution, efficacy, and tolerability of antibody–maytansinoid conjugates. Bioconjug Chem. 2017;28(5):1371–1381. [CrossRef]
  102. Wakankar A, Chen Y, Gokarn Y, Jacobson FS. Analytical methods for physicochemical characterization of antibody drug conjugates. mAbs. 2011;3(2):161-172. [CrossRef]
  103. Prihoda D, Maamari J, Waight AB, et al. BioPhi: a platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning. mAbs. 2022;14(1):2020203. [CrossRef]
  104. Wang Y, Guo C, Li W. Artificial intelligence in antibody–drug conjugate development. Trends Pharmacol Sci. 2025;46(12):1209–1223. [CrossRef]
  105. Hou X, Zaks T, Langer R, Dong Y. Lipid nanoparticles for mRNA delivery. Nat Rev Mater. 2021;6(12):1078–1094. [CrossRef]
  106. Maharjan R, Kim KH, Lee K, Han HK, Jeong SH. Machine learning-driven optimization of mRNA-lipid nanoparticle vaccine quality with XGBoost/Bayesian method and ensemble model approaches. J Pharm Anal. 2024;14(11):100996. [CrossRef]
  107. Nakamura T, Ishida T, et al. Rational design of lipid nanoparticles for enhanced mRNA vaccine delivery via machine learning. Small. 2025;21(8):e2405618. [CrossRef]
  108. Srivastava A, Mallela KMG, Deorkar N, Brophy G. Manufacturing challenges and rational formulation development for AAV viral vectors. J Pharm Sci. 2021;110(7):2609–2624. [CrossRef]
  109. Naldini L. Lentiviral vectors, two decades later. Science. 2015;353(6304):1101-1102. [CrossRef]
  110. Lemmens G, Van Mol L, Klaas S, et al. Adaptive machine learning framework enables unprecedented yield and purity of adeno-associated viral vectors for gene therapy. bioRxiv. 2025. [CrossRef]
  111. Delaney JS. ESOL: estimating aqueous solubility directly from molecular structure. J Chem Inf Comput Sci. 2004;44(3):1000-1005. [CrossRef]
  112. Varoquaux G. Cross-validation failure: small sample sizes lead to large error bars. Neuroimage. 2018;180:68-77. [CrossRef]
  113. Boughorbel S, Jarray F, El-Anbari M. Optimal classifier for imbalanced data using Matthews Correlation Coefficient metric. PLoS One. 2017;12(6):e0177678. [CrossRef]
  114. Gundersen OE, Kjensmo S. State of the art: reproducibility in artificial intelligence. Proc AAAI Conf Artif Intell. 2018;32(1). [CrossRef]
  115. Hu E, Shen Y, Wallis P, et al. LoRA: low-rank adaptation of large language models. Int Conf Learn Represent. 2022. [CrossRef]
  116. Rapp JT, Bremer BJ, Romero PA. Self-driving laboratories to autonomously navigate the protein fitness landscape. Nat Chem Eng. 2024;1:97-107. [CrossRef]
  117. Abolhasani M, Kumacheva E. The rise of self-driving labs in chemical and materials sciences. Nat Synth. 2023;2(6):483-492. [CrossRef]
  118. Shahab MA, Destro F, Braatz RD. Digital twins in biopharmaceutical manufacturing: review and perspective on human-machine collaborative intelligence. arXiv:2504.00286. 2025. [CrossRef]
  119. Hayes T, Rao R, Akin H, et al. Simulating 500 million years of evolution with a language model. Science. 2025;387(6736):850-858. [CrossRef]
Figure 3. Availability and sharing readiness of biopharmaceutical stability data across modalities.
Figure 3. Availability and sharing readiness of biopharmaceutical stability data across modalities.
Preprints 219788 g001
Figure 2. Machine-learning workflow for predicting colloidal stability and formulation outcomes.
Figure 2. Machine-learning workflow for predicting colloidal stability and formulation outcomes.
Preprints 219788 g002
Figure 1. Proposed roadmap for closed-loop autonomous biopharmaceutical formulation development.
Figure 1. Proposed roadmap for closed-loop autonomous biopharmaceutical formulation development.
Preprints 219788 g003
Table 1. Principal colloidal instability mechanisms in biopharmaceutical formulations, with their physical driving forces, relevant modalities, primary formulation interventions, and current machine learning predictability status.
Table 1. Principal colloidal instability mechanisms in biopharmaceutical formulations, with their physical driving forces, relevant modalities, primary formulation interventions, and current machine learning predictability status.
Instability mechanism Physical driving force Relevant modalities Principal formulation interventions ML predictability (2025)
Native-state self-association Electrostatic patch complementarity; short-range hydrophobic attraction; dipole-dipole interactions at formulation pH. Governed by kD and B22. IgG1 mAbs; bispecific antibodies; high-concentration mAb formulations (>50 mg/mL). pH adjustment to reduce net charge; ionic strength optimisation; arginine supplementation; Fv pI engineering. Established. kD and viscosity predictable from Fv pI and sequence descriptors (AUC 0.82-0.89). Best-validated ML target.
Partially unfolded intermediate aggregation Thermal or chemical unfolding exposes buried aggregation-prone regions (APRs); beta-sheet-mediated intermolecular contacts between exposed hydrophobic stretches. All protein modalities. Most significant for mAbs with low Tm1 (<55 degrees C) and unstable CH2 domains; scFv-based constructs. Conformational stabilisers (sucrose, trehalose); pH optimisation to maximise Tm; avoidance of freeze-thaw stress; arginine as aggregation suppressor. Partial. Tm1 prediction from sequence is feasible; linking Tm to aggregation rate requires kinetic modelling. APR tools (AGGRESCAN3D) available but imperfect.
Surface-mediated nucleation and aggregation Protein adsorption to hydrophobic interfaces (air-liquid, container-closure, stainless steel) followed by surface-induced conformational change and nucleation of irreversible aggregates. mAbs; ADCs; protein nanoparticles; mRNA-LNP (lipid shell disruption). Particularly severe during agitation, fill-finish, and pump transfer. Polysorbate 20/80 or poloxamer 188 as surfactant; container closure siliconisation control; inert contact materials; minimise agitation and headspace. Absent as a standalone ML target. Surface-mediated aggregation is mechanistically distinct from solution-phase self-association and absent from all current training datasets.
Liquid-liquid phase separation (LLPS) Concentration-dependent spinodal decomposition driven by net attractive protein-protein interactions near a critical point; thermodynamically favoured below cloud point temperature. High-concentration mAb formulations; bispecific antibodies with charge asymmetry. Manifests as reversible opalescence or visible phase separation. pH and ionic strength adjustment to move away from critical point; arginine; targeted Fv charge re-engineering. Nascent. Cloud point prediction from sequence not established. Coarse-grained MD provides mechanistic insight but ML models specific to LLPS are absent from the literature.
Freeze-thaw-induced aggregation Ice crystal exclusion concentrates protein and excipients; osmotic stress and membrane destabilisation (LNPs); mechanical stress from ice crystal growth; pH shifts in partially frozen solutions. All protein modalities during bulk drug substance freeze-thaw cycles. Particularly severe for mRNA-LNP systems and bispecific antibodies with low conformational stability. Sucrose or trehalose as cryoprotectant (8-10% w/v for LNPs; 5-10% for proteins); controlled freezing rate; annealing step; avoidance of Tg’ excursions. Limited. ML applied to lyophilisation process parameter optimisation (Parra-Saavedra 2025). Sequence-to-freeze-thaw stability prediction not established.
Chemical modification-coupled aggregation Deamidation (Asn, Gln), oxidation (Met, Trp, Cys), disulfide scrambling, and glycation alter surface charge and hydrophobicity, increasing aggregation propensity of the modified species. All protein modalities under long-term storage. Particularly relevant for mAbs with susceptible CDR Asn or solvent-exposed Met residues. pH 5.5-6.5 to slow deamidation; antioxidants (methionine, EDTA) for oxidation; avoid residual metals; optimise headspace oxygen. Partial for individual chemical degradation rates (deamidation from sequence context: Asn-Gly motif). Coupling to aggregation rate is not modelled in published ML frameworks.
Conjugate-induced colloidal destabilisation (ADC-specific) Hydrophobic drug-linker payloads increase surface hydrophobicity proportional to DAR; reduce Tm by 2-8 degrees C; shift kD negative; promote HIC retention and self-association. ADCs across all DAR values. Magnitude proportional to payload logP and DAR; site-specific conjugation reduces but does not eliminate the effect. Surfactant type and concentration optimisation; pH titration post-conjugation; excipient screening on conjugated (not naked antibody) material. Absent. No published ML model predicts post-conjugation colloidal stability from sequence + payload descriptors. No public post-conjugation stability datasets available.
Table 2. Principal analytical techniques used to assess colloidal stability and their key ML-relevant data outputs.
Table 2. Principal analytical techniques used to assess colloidal stability and their key ML-relevant data outputs.
Method Principle Key ML-Relevant Outputs Limitations for ML
DLS / HT-DLS Brownian motion → hydrodynamic radius (Rh) via autocorrelation of scattered light intensity Rh, polydispersity index (PDI), diffusion interaction parameter (kD), onset aggregation temperature (Tagg), Z-average, %intensity per population Intensity-weighted; biased toward large particles; limited resolution of co-existing populations; kD requires concentration series
SLS / SEC-MALS Intensity of scattered light proportional to Mw; SEC fractionates, MALS detects online Weight-average molecular weight (Mw), Rg, second virial coefficient (B22/A2), aggregate Mw distribution Requires refractive index increment (dn/dc); offline MALS linked to separation artifacts; low throughput
SEC-HPLC Size-based separation on porous stationary phase; UV280 detection % monomer, % HMW species, % LMW species, aggregate kinetic rates from time-series data Column interactions with some proteins; underestimates soluble aggregates >void volume; no absolute Mw without MALS
AUC (SV/SE) Sedimentation of molecules in centrifugal field; UV or interference optics Sedimentation coefficient (s20,w), frictional ratio, KD (self-association), oligomeric state distribution Low throughput; complex data analysis; instrument access limited; not amenable to plate-format screening
nanoDSF Intrinsic Trp/Tyr fluorescence emission wavelength shift during thermal ramp; no dye required Tm1, Tm2, Tonset, Tagg (backscatter), unfolding cooperativity, ratio 350/330 nm profiles Tm can conflate unfolding and aggregation; multi-domain proteins show complex transitions; artifacts at high protein concentrations
DSC (µDSC) Differential heat flow during thermal denaturation; measures excess heat capacity vs. reference Tm per domain, calorimetric enthalpy (ΔH), van’t Hoff enthalpy, thermodynamic reversibility Low throughput (one sample per run); high sample consumption; irreversibility complicates thermodynamic interpretation
MFI Digital microscopy + microfluidics; images individual sub-visible particles in flow Particle count/mL (≥1, ≥10, ≥25 µm), morphological descriptors (ECD, aspect ratio, circularity, transparency, intensity), particle type classification Requires large sample volumes for accurate counting; morphological classification depends on ML labelling quality; inter-instrument variability
NTA Tracks individual nanoparticle Brownian motion under laser illumination; particle-by-particle size determination Particle size distribution (50–1000 nm), concentration, fluorescence-NTA for labelled aggregates Sensitive to camera settings and flow conditions; size accuracy limited for polydisperse samples; low reproducibility across laboratories
Zeta Potential (ELS) Electrophoretic mobility of particles in applied electric field → surface charge proxy Zeta potential (mV), isoelectric point (pI) by pH titration, charge reversal points Henry equation assumptions; not directly predictive of long-term stability for high-ionic-strength formulations; single-population average
AC-SINS / SMAC Gold nanoparticle aggregation (AC-SINS) or self-interaction chromatography (SMAC) to probe mAb self-association Affinity capture self-interaction nanoparticle spectroscopy shift (Δλ nm); SMAC retention time; self-association propensity score SMAC column degrades with sticky proteins; AC-SINS sensitive to assay conditions; not directly scalable to high-concentration formulation prediction
HIC Retention on hydrophobic stationary phase under ammonium sulfate gradient; proxy for surface hydrophobicity HIC retention time (min), hydrophobicity rank, relative patch exposure Column-dependent retention; not directly quantitative; retention time sensitive to gradient conditions; aggregate discrimination limited
Rheology / Viscometry Cone-plate or capillary viscometry at multiple concentrations; rotational rheometry for viscoelasticity Dynamic viscosity (cP) at target concentration, power-law exponent, elastic modulus G’, loss modulus G’’ Requires high protein concentrations (≥50 mg/mL); sample consumption high; limited throughput; temperature-sensitive
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings