Preprint
Review

This version is not peer-reviewed.

From Screening to Generative Design: Advances in ML-Assisted MOFs for Carbon Capture

Submitted:

27 August 2026

Posted:

31 August 2026

You are already at the latest version

Abstract

The ongoing climate crisis, caused by the annual release of 37 billion metric tons of CO2 emissions, is putting pressure on the advancement of Carbon Capture and Storage (CCS) and Direct Air Capture (DAC) technologies. Metal–Organic Frameworks (MOFs), with their high surface areas and modular pore topologies, present a very attractive class of sorbents for CO2 capture; however, it currently remains computationally prohibitive to explore their extensive chemical design space. Herein, we provide a thorough evaluation of how machine learning (ML) (as an emerging technology) has played an increasing role in furthering our understanding of CO2 capture from MOFs. Through an organized investigation, we provide evaluations of the latest generation of models across four key areas: application at a process level, mechanistic interpretable modelling; physically relevant descriptors, and predictive performance metrics. Recent work with Machine Learning Interatomic Potentials (MLPs) shows that traditional assumptions about rigid frameworks are being challenged by the fact that diffusion properties and adsorption thermodynamics are heavily influenced by the flexibility of the framework. The use of physics-informed descriptor engineering yields R2 values of 0.81-0.97 across gas species and pressure regimes, while the generative nature of Deep Reinforcement Learning and transformer-based architectures has been shown to allow for the inverse design of frameworks with high affinities for gas species. The trend in this sector is moving towards optimization of multiple scales simultaneously and integrating processes to achieve an optimized property prediction. Current work with machine learning is focusing on using a combination of material properties and operational indicators (such as how much gas is recovered through pressure swing adsorption) to make predictions. As these techniques improve, there will be a similar need for a design that is both physically informed and understandable, thus allowing for a link between molecular discoveries and water-stable materials that have been experimentally verified and are suitable for use in commercial applications.

Keywords: 
;  ;  ;  ;  ;  
Introduction
The Earth’s overall temperature has continuously become warmer over time due to manmade factors. Last year, scientists determined that the global average temperature was approximately 1.2 °C higher than it was at the start of the Industrial Revolution, when man began using more fossil fuels to generate energy. The primary contributing factor to the Earth’s increased warmth is a rise in the number of Greenhouse Gases (GHGs) produced by human activities over time. CO2 accounts for approximately 60% of this increase. Human-induced CO2 emissions in 2022 topped approximately 37.15 billion metric tons and continue to increase unless robust international action is implemented that results in the attainment of net-zero GHG emissions by the year 2050. Emerging technological solutions to this problem include CCS (carbon capture and storage) and DAC (direct air capture). The objective of these methods is therefore to either capture carbon emissions from industrial sources or do so independently of any point source (as in DAC), so as to ultimately decrease atmospheric CO2 concentrations overall. [1]
Finding materials that have a strong affinity for carbon dioxide (CO2) while being selective from other gases like ambient water vapor (H2O) and nitrogen (N2) is a major challenge for these types of technologies. [2] Due to their extremely high surface areas, tuneable pore diameters, and programmability, Metal-Organic Frameworks (MOFs)—crystalline porous materials made of metal cluster components connected by organic ligands—offer a potential avenue for achieving this goal. MOFs provide a unique engineering capability not available with traditional adsorbents. In addition to having engineering options available for including open metal sites (OMS) in their structure, they also provide opportunities for using chemical functional group types (e.g., diamines) to improve their performance in CO2 binding. [3]
Chemical design (the process by which new compounds are discovered) will result in MOF design increasing exponentially, as there are already over one hundred thousand known structures synthesized, with nearly an infinite number of additional synthesized ones (trillions) predicted using in silico (computer) tools alone. The vast majority of the known MOFs and the proposed in silico MOFs would take too long to test and synthesize using traditional methods of laboratory synthesis or computational methods (such as DFT and GCMC). [4]
This bottleneck makes it necessary to switch from traditional screening techniques to data-driven computational screening and generative design for a speedy identification of next-generation carbon capture materials. As a result, machine learning (ML) has become an innovative means for CO2 remediation due to its ability to provide a balance between the quickness exhibited by classical force fields and the precise results obtained with ab initio methods. ML also plays a significant role in CO2 remediation through four different means in MOFs.
Machine Learning (ML) models such as Random Forests (RF) and Artificial Neural Networks (ANN) enable researchers to rapidly predict the CO2 working capacity and selectivity of tens of thousands of materials in only a fraction of the time of traditional simulations. [5]
Researchers can model the flexibility of frameworks through advanced Machine Learning Interatomic Potentials (MLPs). MLPs reveal that structural vibrations can increase CO2 diffusivity up to an order of magnitude versus what is predicted by stiff models. The combination of high-throughput computational screening and artificial intelligence has also improved predictive modelling and scientists’ ability to target experimental work at the best materials to enhance the performance of CO2 adsorption by designing core-shell metal-organic frameworks (MOFs). [6]
In addition, methods for designing metal-organic frameworks (MOFs) include large language models (LLMs) and deep reinforcement learning (DRL), which are utilized to traverse large, multidimensional space and discern materials with very low affinity to CO2 but are rarely included in existing databases. Machine learning has been employed to predict process-level variables (e.g., CO2 purity, CO2 recovery, and energy efficiency) related to the characteristics of materials used in pressure swing adsorption to bridge the gap between materials’ properties and practical operational cycle efficiencies.
By using understandable machine learning models to create SHAP and PDP analyses, researchers have the capability to quantify the effects of power material properties (Lewis acidity) on CO2 removal. This presents a theoretical framework for future CO2 capture experiments. [7]
The following fundamental standards served as the basis for writing the reviews:
Each evaluation outlines both the predicted quantities (for example, uptake, selectivity, or TOF) and the performance metrics (R2 and RMSE) that resulted when compared to pre-specified “ground truth” labels. Models were evaluated based on the sources of their descriptor inputs (like pore-limiting diameters, partial charges on metals, or energy-based radial distribution functions) and whether they are based on either fundamental physics or chemical intuition.
The reviews address the methods used to evaluate the models (such as k-fold cross-validation) and whether or not they showed transferability to new material classes or external experimental data.
Each evaluation highlights how the model provides a new scientific perspective on things like an ideal-sized pore having “two branches” or how dopants interact with micropore volume via a “coupling effect” in addition to how accurate reporting is. Every evaluation also highlights places where the model lacks precision, such as not being able to account for chemical reactions occurring in moist streams, having structural restrictions based on a rigid framework, or failing to include the presence of open metal sites.
This methodical foundation produces an exceedingly appropriate summary of the model’s value through its ability to be scaled up to produce large quantities and to be used in advancing computational material science.
The overall workflow of machine learning-assisted MOF design for CO2 capture, spanning from descriptor engineering to generative inverse design and process-level optimization, is illustrated in Figure 1
Preprints 230370 i001
Schematic representation of the machine learning workflow for MOF-based CO2 capture, illustrating the progression from descriptor engineering and high-throughput prediction to process-level optimization and generative inverse design. The closed-loop framework highlights iterative refinement of materials for enhanced carbon capture performance.
Review Methodology: The literature search for this review was conducted using Google Scholar, selected for its broad, multidisciplinary coverage of the chemistry, materials science, and machine learning literature relevant to this topic. The search covered publications from 2022 to 2026, a window chosen deliberately as it captures the period in which ML-assisted MOF screening matured from early proof-of-concept studies into the mechanistic, generative, and process-level approaches that form the focus of this review. Search terms combined “machine learning,” “metal-organic frameworks,” “computational screening,” and “experimental validation” in various combinations. Only English-language, peer-reviewed publications were considered. This initial search yielded approximately 50 candidate studies, which were subsequently narrowed to 25 based on direct relevance to ML-assisted MOF screening, mechanistic interpretability, generative design, or process-level optimization for CO2 capture; studies without a clear ML component, or falling outside the CO2/gas-capture scope, were excluded.
A note on attribution: unless otherwise indicated, each subsection below discusses a single cited study, and all quantitative values, datasets, and simulation methods reported within that subsection are drawn from the reference cited at its conclusion (or, where identified explicitly by author name, from that specific citation). Sentences framed as “we note,” “future work,” or similar forward-looking language reflect the authors’ own synthesis and critique rather than findings of the cited study.
Machine learning has also informed MOF and related nanomaterial research outside the CO2-capture scope of this review; for example, ML and nanoparticle engineering have been applied to gold-nanoparticle-loaded ZIF photocatalysts for photo-redox reactions [31], hydrogen-bonding nanotrap design in cage-like MOFs for natural gas upgrading and MTO product separation [32], and catalyst design for sustainable hydrogen production from plastic waste [33]. These studies illustrate the breadth of ML-assisted MOF and catalyst research, but address catalytic and separation objectives distinct from CO2 capture. This review instead focuses specifically on the predictive, generative, and process-integrated ML approaches applied to CO2 capture performance in MOFs, synthesizing this particular body of work with a structure and scope not, to the authors’ knowledge, previously assembled elsewhere.

1. Physics-Informed Descriptor Engineering

1.1. Descriptor Engineering Strategies

Deng et al. [8] showed that to comprehend the inner workings of porous materials’ complex potential energy surfaces (PES), an important step is to include energy descriptors that are spatially aware of their surroundings. Researchers used a combination of surface energy histogram, radial distribution function (RDF), and traditional geometric features plus Henry’s constant data to achieve better than an R2=0.97 prediction accuracy for nitrogen isotherms and R2=0.87 prediction accuracy for carbon dioxide; this provides the ML model with a physically meaningful surface representation, effectively solving the intermediate pressure bottleneck where binding energy and spatial inhomogeneity control the adsorption process.
To support this theoretical finding, the XGBoost model trained with over 10,000 different molecular structures has been used to expose that the way that the interaction site decays in space (measured through the radial distribution function, RDF) is the largest contributor to the capacity of different materials with essentially the same energy, while the distribution of affinities (or affinities vs. energy, f(E)) serves as the energetic baseline. However, the poor predictions from our model to explain the CO2 chemisorption results demonstrate the limitations of the current generation of measurements of one-point charges. Future machine-learning techniques informed by physics will address the current lack of usable data for predicting CO2’s orientation-containing packing (i.e., packing orientation and electrostatic multipoles) using dipole and quadrupole probes so that we can build bridges between simplified physisorption models and practical implementation through selective separation as is required to obtain industrial-use reliability. [8]
Zuo et al. [9] showed that, by moving beyond standalone pore or doping engineering, researchers have identified a critical micropore-dopant coupling mode that determines CO2 transport in carbon-based adsorbents. Using a Random Forest architecture trained on multi-scale simulation data (R2=0.934), this study introduces free volume (Vf) as a primary descriptor (accounting for 25% relative importance) to quantify the steric effects induced by surface functionalization. This approach reveals a significant thermodynamic shift: while basic dopants (e.g., NH2) utilize Lewis’s acid-base interactions to optimize adsorption at 7 Å; physisorption-dominant dopants (e.g., oxygen groups) occupy available nano space, necessitating an enlarged optimal pore size of 8–10 Å to maximize capacity. These results were validated experimentally, as adsorbents designed with this particular coupling mode attained a leading-level capacity of 4 mmol/g, a 130% improvement over unoptimized frameworks. These findings highlight the need for physics-based descriptor engineering to settle long-standing disputes over the promoting vs. inhibitory effects of heteroatom doping in porous carbons. [9]
Using descriptor engineering techniques, Teng et al. [10] showed that adding 700 or more “calculated” molecular descriptors in addition to the traditional structural model approaches improved predictions of CO2 adsorption by approximately 15%-20% when using advanced models compared to conventional structural model approaches. Their XGBoost framework shows how difficult it is for models to predict the CO2 adsorption properties in low-pressure environments (0.01 bar) based on their sensitivity at low pressures; however, it is successfully used to predict CO2 adsorption properties in industrially relevant pressure environments (2.5 bar) with very high R2 (0.94) values. Their SHAP explainability analysis indicates that the primary thermodynamics governing the transition at low pressures is caused by the changeover of atomic mass and number (representing van der Waals forces) as predictive factors to charge distribution and electronegativity (which account for Coulombic forces) as pressure increases. As with much of this literature, the potential for real-world performance oversights exists with the use of hypothetical datasets (hMOF, the widely used computationally generated hypothetical MOF database [29]) and simulated Grand Canonical Monte Carlo (GCMC) labels rather than experimental measurements. We note that this model cannot consider structural defects or competing adsorption in complex flue gas without experimentation, and this will greatly hinder the model from being scaled effectively throughout the industry. [10]

1.2. Hybrid Textural-Optimization and Outlier-Aware Predictive Modelling

Longe et al. [11] demonstrated that transitioning to hybridized optimization frameworks from traditional machine learning models was a critical advancement for addressing many of the current limitations associated with “black-box” types of adsorptions modelling. Through an integration of the Growth Optimization (GO) technique together with Least Square Support Vector Machines (LSSVM), researchers eliminated common hyperparameter tuning errors that lead to underestimating uptake at high levels when using solo modelling, achieving a R2 of 0.9798 through a collaborative modelling approach. Evaluation using a database of 475 experimental data points showed that the primary structural component influencing capture is SBET, evidenced by greater sensitivity of predictions compared to pore volume and Langmuir surface area.
Crucially, Williams’s plot analysis provides the level of statistical validity necessary for a quick screen using the fact that 94.95% of all experimental data is contained in the applicable domain of the model. Because the current model uses textural measurements of experiments, there remains an inherent gap between these high-fidelity experimental models and low-fidelity, theoretical proxies (high throughput) used in computational discovery processes. Future research into physics-informed machine learning will need to close this gap to improve the industrial scale-up of the “metal affinity” and “pore filling” dynamics shown here in the de novo design of next-generation carbon capture materials [11].

1.3. Ensemble Interaction Mapping and Node-Affinity Hierarchies

Iyiola et al. [12] showed that the successful use of unique stacking ensemble structures to close the accuracy gap between different machine learning models and complex experimental adsorption data is a significant step forward, due to the performance of the stacked learners being far superior to that of the underlying models alone. A record-setting R-squared of 0.9833 on CO2 uptake was achieved with a combined total of 1,212 experimental data points by combining the predictive power of a tree-based model, a kernel-based model, and a neural network model, thus demonstrating the effectiveness of this stacking ensemble structure. By employing an ablation study to separate the effects of thermodynamics from that of the architecture of the ensembles, and by using a technique to compute permutation importance (i.e., the degree of influence of each variable on the output variable), this study moves beyond simply comparing the performance of ensembles relative to their constituent models to provide an interpretability framework for comparing multiple different criteria. A comparative summary of the model families discussed throughout this section, including their relative strengths, weaknesses, and typical performance, is provided in Figure 2.
The analysis provides evidence for a metal-affinity hierarchy via partial dependence plots, indicating that copper (Cu) and magnesium (Mg) metal centers outperform traditional zinc-based nodes (Zn). However, the overrepresentation of Zn-based materials (70%) in previous experimental databases raises concerns about data imbalances and the potential for model overfitting in well-studied materials. Future physics-based (or physics-guided) machine learning and modelling approaches must include stratified sampling and synthetic oversampling (SMOTE) as they continue to identify high-performing MOFs in the underrepresented chemical subspace of rare or complex metal centers if they are to achieve commercial scalability. [12]
Physics-Informed Feature Engineering defines the fundamental language by which MOF structures are quantitatively represented, with large-scale prediction capabilities as the primary value of these techniques. When analytically defined chemical feature descriptors are constructed, supervised machine learning can then quickly make predictions on adsorption parameters for a sufficiently wide range of chemical spaces. The transition from constructing feature descriptors to making high-throughput predictions of CO2 remediation is the first main acceleration point within machine learning-driven research on CO2 remediation.

2. High-Throughput Screening and Universal Property Prediction

2.1. High-Throughput Screening and Discovery

Wan et al. [13] showed that using hybrid physics and machine learning (ML) models allows for better screening of composite materials than using traditional means of molecular simulations. Researchers were able to confirm that transfer learning could accurately predict how to develop composite materials, such as 6FDA-DAM, which had never been seen before, and that the method could be used for high-performing membranes when developed from building an updated regression-based model (stacked ensemble) on over 54,000 hybrid membranes with a very high degree of accuracy (R2 = 0.96). Additionally, mining the data, the research team found that when using a MOF (metal organic framework) filler, it becomes the limiting factor to achieve elite performance, and when evaluating membranes at the top end of performance, both PLD (Pore Limiting Diameter) and LCD (Largest Cavity Diameter) [30] had an importance rating greater than 30%, meaning that polymer selection becomes less of a consideration and that once the initial base polymer has a good level of permeability established, pore size becomes the most important part of composite engineering.
The constructive evaluation relies heavily on the ideal interface assumption of the Maxwell model. In many cases, the interfacial gaps or rigid polymer chains in the polymer materials result in the degradation of the theoretical performance of the “Robeson Limit” and their ability to demonstrate scalability in actual industrial applications. Future physics-driven machine learning studies must incorporate descriptors of interfacial morphology to describe the complexity of bonding between the organic linkers and polyimide matrix to ensure that the materials produced via high-throughput screening will maintain their separation efficiencies under the mechanical stresses associated with flue gas streams. [13]
Achour et al. [14] moved from using a lot of simulation-based data when developing machine learning models to using an experiment-based data set as a more accurate basis for benchmarking CO2 capture. They demonstrated an improvement in RMSE of 15% using ensemble learning on 236 different experimental metal-organic frameworks (MOFs), achieved using the CATBoost framework rather than a more traditional gradient-boosting framework. When combined with the SHAP interpretability of the model, this results in a better understanding of how chemical structure governs atomic environments and electronic states as they relate to optimizing adsorption, even though the model also shows that pressure and surface area continue to be the primary factors driving thermodynamic behaviour. [14]
We note that this same study also functions as a critical assessment of the limits of generalization. The difference in the model’s validation (R2 = 0.84) and training (R2 = 0.99) scores demonstrates the repeated problem of overfitting during the development of complex models trained on sparse experimental data. To provide greater assurance of the ability of high-throughput screening to consistently translate into industrially valuable, highly stable sorbents for hypothetical frameworks, future work on physics-informed machine learning should move toward using a more diverse dataset by including a wider variety of chemical functions. [14]

2.2. Universal Property Prediction and Isotherm Generalization

Kirtil [15] showed that the recent trend in zeolite-based carbon capture is to move away from empirical models based on individual materials and utilize a single set of experimental data for all zeolites. Researchers have also been able to create a “universal” predictive capability far superior to that achieved by using typical Langmuir isotherm fittings, both in terms of accuracy and generalization, by building a Gradient Boosted Trees (GBT) model from greater than 5,700 sets of experimental data. This method uses the Si/Al ratio and cation composition as the key chemical parameters, allowing for model generalization across multiple framework types (e.g., FAU, ZSM-5, and 13X) and elevated operational pressures (up to 45 bar).
One important confirmation of the model’s robustness comes from external validation on datasets that have not been examined previously and effectively address the overfitting failure often encountered when training with data generated from published literature. However, a useful critique of the dataset suggests it is somewhat biased toward low data (uptake) (2 mmol/g) and that 32% of the data used to train (i.e., surface area) was not reported. In order to improve industrial scalability, future physics-enabled machine learning efforts should focus on utilizing high-data (i.e., uptake) experimental benchmarks and standard reporting of structural parameters so that the sensitivity of the model is optimized for maximum performance and scaled effectively across the industry. [15]

2.3. Property-Driven ML Applications

Kotov et al. [16] used Machine Learning to predict CO2 working capacity from Metal-Organic Frameworks and CO2/N2 selectivities using Artificial Neural Networks due to their computationally efficient capabilities compared to Graph Neural Networks for rapid screening of MOFs with high prediction accuracy (R2). The model combined industrial, field, and simulation data into a single design principle showing that pore size and surface area primarily determine CO2 working capacity. Evidence of this relationship is a weak negative correlation between CO2 working capacity and increasing chemical complexity. The CO2 working capacity predictions were accurate (MAE = 0.8 mmol/g); however, the CO2/N2 selectivity predictions were highly dispersed from each other (MAE = 25). This result could be indicative of intrinsic properties not able to account for the competitive adsorption physics that are required to accurately predict high selectivity. Moreover, because no thermal criteria (temperature, pressure) or outside testing has been done for an unknown structure, the model does not have the capacity to find new or previously uncharacterized fragments from a structural assessment perspective. Future directions for physics-informed research should deal with the uncertainty in selectivity modelling to ensure that rapid AI-related screening processes can be effectively translated into industrial carbon reductions. [16]
Even when high-throughput screening significantly reduces computation cost, predictive accuracy alone does not guarantee scientific understanding. The black-box models used to describe adsorption behaviour can mask the underlying physicochemical principles driving the process. Consequently, it is critical to be able to derive mechanistic knowledge from trained predictive models and ensure that predictions are based on the underlying physics of adsorption via interpretability frameworks and thermodynamic mapping techniques.

3. Model Interpretability and Thermodynamic Mapping

3.1. Model Interpretability and Physical Insight

Zheng et al. [17] developed quantum chemistry-based machine-learning potentials (MLPs) developed as an alternative to the rigid-lattice assumptions typically employed for gas transport molecular simulations. The researchers attained an exceptional level of predictive accuracy (R^2 = 0.9916) for system energies of MgMOF-74 via the DeepPot-SE model, which captures the dynamically flexible character of this material. Consequently, it is clear that structural flexibility plays an important role in the prediction of CO2 diffusivity; rigid models yield predictions of adsorption free-energy barriers that reflect significant overestimates of true values and produce diffusion coefficient estimates more than 10 times lower than are experienced with flexible materials. The integration of thermodynamics and machine learning highlights the importance of the “hopping” mechanism of diffusion between open-metal sites, but further rigorous evaluation is necessary in determining the generalizability of this model. Accuracy is crucially dependent on the quality of the underlying DFT-MD reference data (the MLP was only trained on one 30 ps frame of the DFT-MD trajectory), since there are no direct measurements available to use as experimental standards for characteristic properties like diffusivity of Mg-MOF-74. As a result, there is a need for further development of standardized experimental benchmarks against which the industrial scalability of such physics-informed machine learning models can be confirmed. While these smaller modelled datasets help to reveal some physical patterns, such as enhanced diffusion due to flexibility, they should be utilized with caution in the real world where there will be more complex mixtures of gas and the possibility of varying transport kinetics due to structural deficiencies. [17]

3.2. Ensemble-Averaged Thermodynamics and Potential Energy Surface (PES) Mapping

Lim et al. [18] showed that the development of transferable machine learning force fields (MLFFs) has provided a radically new way of performing high-throughput screenings of porous materials for direct air capture (DAC) and allowed researchers to achieve thermodynamic ab initio-quality calculations at speeds comparable to classical methods by optimizing a base model (MACE-MP) against the GoldDAC dataset. This is significant because it demonstrates that the UFF+DDEC models form the basis for all current literature on hybrid materials with complex chemistries and are susceptible to systemic errors in the case of lanthanide-based frameworks if the models were to be used for predicting water adsorption in these materials, potentially overestimating H2O adsorption energies by as much as 17.8%. The use of the DAC-SIM package for identifying materials has allowed for 161 candidates to be identified by moving from single-point interaction energy calculations to the use of ensemble-averaged properties. Results of this investigation demonstrate the need to sustain a high level of CO2/H2O selectivity (KH ratio 1.00) through a “selective pocket” mechanism enabled by the presence of some parallel benzene rings (PAR) with free or uncoordinated nitrogen atoms. Current assumptions based solely on rigid frameworks still create obstacles to existing Understanding renewal costs and chemical dynamics necessary for large-scale industrial use does not exist today (despite the use of real-world physics as the basis for current screening). [18]
This combination of thermodynamics and machine learning indicates that the “cooperative insertion mechanism” is more about dynamic redistribution than static binding and leads to faster kinetic rates; however, relying upon an entirely computational MLP will have significant limitations for modelling chemical transition processes occurring from carbamates/carbonates to humid streams due to difficulties associated with accurately simulating these processes despite relying upon real NMR data to confirm results. To effectively connect these transport models with the complicated stability and regeneration needed in high-humidity flue gas environments for industrial use, physics-informed ML approaches will probably need to combine molecular simulation and modelling with computational descriptions more thoroughly.

4. Molecular Transport and Multicomponent Separation Mechanisms

4.1. Molecular-Level Transport and Adsorption Mechanisms

Tayfuroglu et al. [19] utilized Fragment-based Neural Network Potentials (NNPs) to provide a potential solution for the long-standing challenge of reliably simulating flexible frameworks containing open metal sites (OMS) at ab initio accuracy. The results presented here show extraordinary agreement with X-ray findings regarding force RMSE (0.039 eV/Å) and provide excellent transferability based on a training set of only ~2000 density functional theory (DFT) conformers with only 0.54% variance in lattice parameters from the training set to empirical data (i.e., using an E(3)-equivariant architecture called NequIP). Importantly, new hybrid molecular dynamics (MD) and grand canonical Monte Carlo (GCMC) workflows that combine thermodynamic and machine-learning methods indicate that consideration of structural dynamic effects is critical for predicting accurate adsorption behaviour: framework flexibility facilitates structural relaxation/leads to delocalization of CO2 over time and allows for a more favourable overall energy landscape at lower pressures (0.1–1.0 bar).
At temperatures above 298 K, conventional GCMC models, which use rigid-lattice approximations, tend to significantly underpredict the amount of gas being absorbed, because those models assume that the structure of the host/guest will not change. While this active-learning approach for screening these fragment-based models has proven successful, a systematic evaluation of the transferability of the training set is needed to confirm that the model has not forgotten to include complex interactions from other chemical environments outside of the rigorous training environment. The high-fidelity potentials obtained from the performance of physics-based machine learning modelling methods will be essential in investigating the cooperative insertion process and regeneration cycle of these materials in order to enable large-scale commercial carbon capture. [19]

4.2. Multicomponent Gas Separation and Structural Design Strategies

Zhou et al. [20] noted that the next step to implementing a MOF-based system for carbon capture will be changing from binary/ternary gas models to real (natural) 6-component gas mixtures (N2, CO2, CH4, C2H6, C3H8, H2S). This paper found an R2 = 0.922 predictive accuracy for the renderability of materials based on 10 top-performing MOFs (from 12,020 structures) that were identified via GCMC simulations combined with a Random Forest architecture. Also, the database has identified a ‘volcanic’ structure-property relationship, where the maximum separation performance exists in a density range of 0.5 to 1.7 g/cm3; any MOFs with densities outside of this range will experience either kinetic exclusion or reduced selectivity due to large pore sizes.
Next-generation adsorbents will have a clear tripartite design strategy as follows: (i) replacing metal nodes with new metal negative nodes—e.g., replacing cadmium (Cd) with manganese (Mn) to achieve 4 times greater working capacity; (ii) using nitrogen- and hydrogen-rich organic linkers—i.e., to utilize electrostatic interactions of linker materials (e.g., using pyridine or azoles). The rigid framework assumptions successfully identified the best-performing materials (e.g., XIGWUF and ETECOX), but going forward into future physics-based machine learning applications, the static lattice models will not be adequate for modelling or developing adsorbents because the thermodynamic and kinetic effects of adsorbents will vary under real-world, high-pressure natural gas streams due to the effects of framework fit and flexibility. [20]
Chemical conditions that exist at the local level within the framework will ultimately control how selective you’ll be when adsorbing as well as when creating a catalyst. CO2 movement through porous architectures can be quantified using transport models. But more than just structural optimization, connecting chemical reactivity with thermodynamics of adsorption is achieved through engineering Lewis’s acid-base interactions, modifying the electronic structure, and creating synergistic arrangements of sites.

5. Active-Site and Electronic Structure Engineering

5.1. Lewis Acid-Base Site Engineering and Catalytic Kinetics

Bai et al. [21] showed that, with understandable machine learning being combined with low-cost descriptors, it is now reasonable to rapidly screen catalytic frameworks without the excessive costs associated with density functional theory (DFT). 372 high-yield experimental data points were used to train a Random Forest architecture to predict CO2 cycloaddition activities at an amazing 97% level of accuracy. The primary aim of this research is to expand the boundaries of “black-box” predictions and determine the “optimal Lewis’s acidity” of metal nodes via SHAP and PDP analyses. The research results indicate that the maximum epoxide substrate activation occurs with a metal charge value of between 1.2 and 2.0; metal nodes in this charge range prevent saturation mode coordination leading to deactivation of the active site.
MOF-76(Y)’s confirmed exceptional activity at a top-tier TOF of 64.72 h−1 shows that ML may be the solution to creating a bridge between CO2 utilization in real-world processes and the potential to build hypothetical structure(s). However, the thorough study completed in this work shows that, in order to successfully compare results across non-linear reaction profiles, the need for a very extensive degree of data cleaning (e.g., eliminating data points based on very low yields) from datasets reported in the literature is the threshold that must be achieved prior to conducting further testing of any physics-informed machine-learning-directed catalyst(s). The future directions for physics-informed machine learning should be focused on the goal of standardizing experimental benchmarks to enable the further commercial reproduction of screened catalysts for a variety of catalytic processes. [21]

5.2. Electronic Property Modulation and Electrocatalytic Selectivity

Xing et al. [22] showed that Two-Dimensional Conjugated Metal-Organic Frameworks (2D c-MOFs) have been identified as a superior alternative for CO2 electroreduction performance compared to traditional copper (211) reference materials through DFT-based discovery and Gradient Boosting Regression (GBR) design principles. The study creates a hierarchy of fundamental stability for c-MOFs (e.g., TMN4TMN2O2TMO4) and determines that certain MOF types, such as NiN4−HDQ MOFs, possess very low limiting potentials of -0.04 V (UL) for carbon monoxide (CO) by systematically modifying both their metal catalyst sites and the associated organic ligands (HDQ series).
Electron affinity (EA) and electronegativity (χ) are, according to an analysis of feature importance for a model. They combine to account for 34% of all feature importance in the sensitivity study performed. The main driver of the strength of intermediate adsorption in 2D c-MOFs is the ligand-mediated orbital overlap, with both type 2D c-MOFs having pristine surfaces in an aqueous environment and extremely high thermal stability at 400 K. To achieve scalability to industrial processes and realize the potential of highly conductive and tunable materials for artificial carbon cycle applications, future use of physics-informed ML approaches must link the high-fidelity C1 model to the C2+ coupling kinetics. [22]

5.3. Synergistic Site Engineering and Composite Pore Modulation

Sheng et al. [23] used convolutional neural networks (CNNs) with inception modules to create a robust multi-criteria screening approach for simulating complex composite adsorbents will avoid a computational time trap from simulating complex composite adsorbents. The CNN model was trained with 700 GCMC-validated structures, which allowed the model to successfully navigate the inherent trade-offs associated with CO2 working capacity and selectivity to achieve a high level of predictive fidelity for its ability to replicate the measured data (R2 ≈ 0.90). Furthermore, the analysis of the composite pool of 1,631 composites indicates that unusual synergistic frameworks (e.g., IL@MARJAQ) will use the IL to produce entirely new potential energy minima relative to CO2. By contrast, most ionic liquids will improve selectivity by limiting the physical transport of N2.
This investigation establishes the non-linear effect due to the amount of IL loaded, indicating that “more isn’t necessarily better”; changing the number of molecules injected into the system can give the IL@GUBKUL a range of selectivity between 614 and 7000+. Precision loading techniques will replace conventional trial-and-error methodologies that currently exist for post-synthetic modification. Thus far, the current model assumes rigid-framework behaviour while underestimating uptake in frameworks that contain open metal sites (OMS); for future physics-informed ML developments, consideration of lattice dynamic behaviour for accurate prediction of regeneration efficiencies and process economics required for large-scale separation of flue gases is necessary. [23]
The advances achieved through electrical control and creating an active site influence the evolutionary progression of this technology through an inverse design approach. Generative machine learning methods can use the learned relationships between structure and properties to generate entirely new MOF designs that target numerous goals, rather than just screen current materials. This transition represents a fundamental change in how we discover new materials, shifting from making predictions about materials using models to generating new forms of material without human input.

6. Multi-Objective Inverse Design and Generative MOF Discovery

6.1. Multi-Objective Inverse Design and Chemical Subspace Exploration

Park et al. [24] noted that for the purpose of designing new materials for sequestration of CO2 through physisorption, it was necessary to find an efficient way to explore the vast number of chemical options available to scientists. With traditional searching techniques becoming less effective as chemical spaces increase in size, a new approach to finding suitable materials involves switching from brute-force screening to Deep Reinforcement Learning (DRL). Researchers were able to successfully create physi-sorbents suitable for Direct Air Capture (DAC) with Qst values above 40 kJ/mol and CO2/H2O selectivities in excess of 1 by using DRL based on transformer-style predictors integrated into a reward-driven environment. Analysis of these findings via database-driven searches also revealed very interesting trade-offs in “genes” controlling the materials. More specifically, certain combinations of Cu and Zn clusters provide high selectivity towards CO2, and those clusters fall into a separate subspace within the overall chemical design landscape. Conversely, many open-metal sites that are responsible for providing very high affinity for CO2 (e.g., Mn-based N131 nodes) also possess the ability to attract water.
The current use of classical force field models to predict chemisorption involving charge transfer has reduced the DRL method’s prediction capacity, despite having superior ability to extrapolate and produce structures that can compete with those produced using the best experimental methods (i.e., KAUST-7 MOF). The next step in developing these physics-informed ML methods will involve incorporating charges obtained from DFT calculations and active learning loops around these to create accurate predictive models that can be scaled industrially. This will ensure that inverse-designed candidates will continue to be effective in real-world atmospheric capture conditions such as humid or highly diluted. [24]
Badrinarayanan et al. [25] showed that the shift from hand-based high-throughput screening methods to RL-assisted generative design methods is having a profound effect on the ability of scientists to create new materials for specific carbon capture applications. Here we show that when using the MOFGPT machine learning framework, an RL strategy can obtain 100% valid (i.e., satisfying chemical and MOF-specific structural constraints) MOFs that exhibit high capacity for CO2 adsorption (mean + 2σ); however, RI methods using standard supervised fine-tuning do not generate a single valid MOF (0%). By optimizing sequences of MOFIDs, i.e., strings that encode both SMILES-based chemistry and RCSR-based topology, this approach provides the means to explore an essentially infinite chemical space without the need for manually assembling building blocks.
The RL model includes correlations between underlying structures and their properties. The evidence is clear from the generated candidate database. Open Cu2+ paddle wheel units and nitrogen-rich linkers improve electrostatic interactions. These types of units are strongly correlated with high CO2 uptake. The framework demonstrated its ability to probe the under-represented areas in the chemical landscape where traditional screeners cannot identify novel candidates. The framework consistently produced over 63% structurally novel, out-of-distribution candidates with no prior pedigree for developing ACs. A key limitation of the RL framework is its use of rigid-lattice assumptions, which omit the necessary dynamic flexibility required for CO2 transport kinetics. To better match artificially intelligent string designs and the properties of robust manufactured sorbents, future work utilizing physics-guided machine learning (ML) should involve additional 3D structural builds and filters that use DFT-based measures of stability. [25]
Just because there have been substantial improvements in terms of performance predictions from generatively discovering new materials, it does not necessarily mean that these newly discovered materials are commercially viable. Ultimately, for any given material to be assessed for its adsorption capabilities, it must be put through actual industrial processes that include temperature and pressure cycling. Thus, integrating machine learning and process simulation helps bridge the gap between molecular design and the system-scale performance of the overall carbon capture system. In effect, this will enable companies involved in carbon capture to convert the predicted performance values obtained from their computer programs into actual carbon capture solutions that can be used today.

7. Multiscale Generative Design and Process-Oriented Optimization

7.1. Process-Integrated Generative Design and Material Optimization

Progressing from mere property screening toward developing a multiscale generative workflow for process-level material design represents an important milestone. The use of the MOF-NET architecture (derived from natural language processing-style word embeddings) has enabled researchers to traverse a landscape consisting of trillions of possible structural variations and identify superior representatives among candidate materials based on comparisons to two quality metrics—CALF-20 (Synthetic Homogeneous Metal-Organic Frameworks) and 13X Zeolite. The determination of two strategies for reaching optimum performance (strict size exclusion [3 – 5 Å pores] or formation of a high-density binding pocket [6 – 30 Å sized pores] where CO2 is stabilized by the presence of multiple oxygen-rich coordination nodes) is derived from database-driven analyses of the best-performing materials.
The present investigation reveals to experimentalists an explicit guide of design by defining Cu-based nodes / fluorinated short linkers as genetic markers for outstanding sorbents. In addition, there appears to be a continuous innovation gap since this (the best material manufactured by HJM + N387 + E44) provides considerable improvements in productivity versus CALF-20 because it rejects significantly more N2. The next technological leap of ML for physics-informed methods will be the combination of synthetic accessibility (SA) scores with lattice dynamics, which when combined will allow “theoretical possibilities” to successfully traverse from the computer environment into the lab, as many computationally designed materials still struggle with synthesizability and stability when in contact with water during use cases. [4]
“Despite rapid advancements in predictive accuracy and generative capabilities, several critical challenges remain, particularly in bridging the gap between computational predictions and experimentally viable materials. Addressing these issues will require the integration of physics-informed models, experimental validation, and process-level optimization frameworks.”
Table 1. Comparative overview of machine learning models discussed in this review for MOF-based CO2 capture and gas remediation, summarizing the primary algorithm, key descriptors, predictive performance, and core scientific insight reported by each cited study.
Table 1. Comparative overview of machine learning models discussed in this review for MOF-based CO2 capture and gas remediation, summarizing the primary algorithm, key descriptors, predictive performance, and core scientific insight reported by each cited study.
Study Focus Primary ML Algorithm(s) Key Descriptor(s) Predictive Performance (R2) Core Scientific Insight
Pore Energy Mapping [8] XGBoost Energy-based RDFs & Surface Histograms > 0.81 CO2 > 0.97 N2 Spatially aware energy RDFs resolve the “intermediate pressure bottleneck” in isotherms.
Composite Modulation [23] CNN (Inception) Geometric + Chemical (Ionic Liquids) ≈ 0.90 Ionic liquids can act as synergistic sites, creating new potential energy minima for CO2.
Kinetic Transport [17] DeepPot-SE (MLP) Atomic coordinates (Flexible) 0.9916 (Energy) Framework flexibility accelerates CO2 diffusivity by 10x compared to rigid models.
Generative Design [25] MOFGPT (Transformer) MOFid (NLP-based strings) 35–100% Validity Reinforcement learning effectively navigates the “extreme tail” of property distributions.
Process-Level Design [4] MOF-NET (ANN) Word Embeddings of Building Blocks Elite purity/recovery Optimal design bifurcates into small-pore exclusion vs. large-pore binding.
Mixed Matrix Membranes [13] Stacking Ensemble Polymer FFV + MOF PLD/LCD 0.96 A “10x permeability rule” exists where filler must exceed polymer permeability for gain.
Experimental Benchmarking [12] Stacking (RF/XGB/MLP) Textural (BET) + Operational (P, T) 0.9833 Identified a metal-affinity hierarchy where Mg and Cu centers provide superior binding sites.
Hybrid Optimization [11] LSSVM-GO Textural + Operational 0.9798 Growth Optimization (GO) significantly reduces prediction errors in high-uptake regimes.
Multicomponent Separation [20] Random Forest Structural + Chemical Descriptors 0.922 (R%) MOF renderability is optimized within a specific density window of 0.5–1.7 g/cm3.
Electrocatalytic Selectivity [22] Gradient Boosting (GBR) Electronic (EA, chi, d-band) 0.9998 Catalytic activity is primarily governed by electron affinity and electronegativity.
Universal Zeolite Prediction [15] GBT / RF / DL Si/Al Ratio + Cation type 0.936 Provides a universal framework without case-specific parameter fitting required by Langmuir models.

Conclusions

Since ML is being used to identify and optimize MOFs, the landscape of carbon capture research has shifted away from traditional methods of large-scale brute-force screening towards data-driven methods that integrate generative design, prediction, and interpretation into cohesive computational workflows (Figure 3 summarizes the key challenges currently limiting real-world deployment of these methods and the corresponding future directions synthesized across this review).
Findings consistently supported across multiple studies:
Framework flexibility matters: across independent studies, rigid-lattice approximations for diffusivity calculations have repeatedly resulted in underestimated diffusivity and misinterpreted adsorption behavior at open metal sites, while machine-learning interatomic potentials that retain dynamic structural detail consistently correct for this. Physics-informed, interpretable descriptors (SHAP, PDP, energy-based radial distribution functions) achieve strong and reproducible predictive performance (R2 = 0.81-0.97) across multiple independent algorithms (XGBoost, Random Forest, stacking ensembles), suggesting descriptor quality is a stronger driver of accuracy than the choice of algorithm itself. Generative and inverse-design approaches, from reinforcement learning to transformer-based frameworks such as MOFGPT, consistently shift the field from passive screening toward proactive candidate generation, producing high proportions of chemically valid, structurally novel candidates across the studies reviewed here.
Unresolved technical limitations:
The gap between computational discovery and industrial deployment remains wide: synthetic accessibility, competition from co-adsorbed gases, and process-level integration are all comparatively underexplored relative to raw predictive accuracy. Water stability, while now an active machine-learning subfield in its own right, remains poorly integrated with CO2-capture-performance modeling; the two research threads are still largely developed in parallel rather than combined. Overfitting and limited generalization to new material classes recur across multiple studies (for example, the R2 = 0.84 validation versus R2 = 0.99 training gap noted in one CATBoost-based benchmarking study [14]), suggesting that current benchmarking practices may overstate real-world readiness.
Limitations of the present review:
This review is narrative rather than systematic in design: it draws on a single database (Google Scholar) over a defined 2022-2026 window, meaning relevant earlier foundational work and non-English-language literature were not captured. Reported performance metrics (R2, RMSE, validity percentages) are drawn directly from the source studies without independent re-verification or standardized cross-study benchmarking, so comparisons across studies in this review should be read as indicative rather than strictly equivalent. Finally, the review’s coverage of experimental (as opposed to purely computational) validation is inherently limited by how few of the surveyed studies report experimental data at all.
Prioritized future directions:
1. Development of standardized, cross-study benchmarks and shared held-out test sets to enable fair comparison of ML models for MOF-based CO2 capture.
2. Deeper integration of water-stability screening and other deployment-relevant filters (synthetic viability, multi-component gas tolerance) directly into performance-prediction pipelines, rather than as separate research threads.
3. Wider adoption of physics-informed, flexible-framework simulation as the default for transport and kinetics modeling, in place of rigid-lattice assumptions.
4. Greater emphasis on experimental validation of computationally generated candidates, particularly for generative and inverse-design outputs.
5. Process-level integration connecting molecular-scale predictions to PSA/TSA cycle performance and techno-economic viability, to close the gap between computational discovery and industrial deployment identified throughout this review.

References

  1. Ozkan, M.A.; Coley, Amir-Ali; Shang, William; Ma, Ruoxu; Yi. Progress in carbon dioxide capture materials for deep decarbonization   . Chem 2022, 8, 141–173. [Google Scholar] [CrossRef]
  2. Carrascal-Hernández, D.C.; Grande-Tovar, C.D.; Mendez-Lopez, M.; Insuasty, D.; García-Freites, S.; Sanjuan, M.; Márquez, E. CO2 Capture: A Comprehensive Review and Bibliometric Analysis of Scalable Materials and Sustainable Solutions   . Molecules 2025, 30(3), 563. [Google Scholar] [CrossRef] [PubMed]
  3. Mahajan, Shreya; M.L. Recent Prog. Metal.-Org. Fram. (MOFs) CO2 Capture At. Differ. Press. J. Environ. Chem. Eng. 2022. [CrossRef]
  4. Deng, Z.; Sarkisov, L. Multi-Scale Computational Design of Metal-Organic Frameworks for Carbon Capture Using Machine Learning and Multi-Objective Optimization Chem. Mater. 2024, 36(19), 9806–9821.
  5. Hussin, F.N.; Aqilah, Siti; Mohamed Hatta, Nur Syahirah; Aroua, Mohamed; Mazari; Ali, Shaukat. A systematic review of machine learning approaches in carbon capture applications   . J. CO2 Util. 2023. [Google Scholar] [CrossRef]
  6. Coudert, François-Xavier. Recent advances in stimuli-responsive framework materials: Understanding their response and searching for materials with targeted behavior Coord. Chem. Rev. 2025, 539.
  7. Fathalian, F.A.; Sepehr; Ghaemi, Ahad; Hemmati, Alireza. Intelligent prediction models based on machine learning for CO2 capture performance by graphene oxide-based adsorbents Sci. Rep. 2022, 12, 21507. [PubMed]
  8. Deng, Z.; L.S. Engineering machine learning features to predict adsorption of carbon dioxide and nitrogen in metal-organic frameworks J. Phys. Chem. C 2024. [CrossRef]
  9. Zuo, J.; Qu, F.S.; Yang, Z.; Xie, C.; Zhang, L.; Li, Y.; Li, X.; J. Unraveling the coupling effect of micropore confinement and functional sites of carbon-based adsorbents on flue gas CO2 adsorption: A machine learning study based on multi-scale simulations Carbon Capture Sci. Technol. 2025. [CrossRef]
  10. Teng, Y.; G.S. Interpret. Mach. Learn. Mater. Discov. Predict. CO2 Adsorpt. Prop. Metal.-Org. Fram. 2024. [CrossRef]
  11. Longe, P.O.; Mehrad, S.D.; Wood, M.; D.A. Robust machine-learning model for prediction of carbon dioxide adsorption on metal-organic frameworks; J. Alloys Compd., 2024. [Google Scholar]
  12. Iyiola, Z.; Okeke, E.T.B.; Sanni, N.J.; Longe, K.; P. Carbon capture using metal-organic frameworks (MOFs): Novel custom ensemble learning models for prediction of CO2 adsorption Processes, 2025.
  13. Wan, H.; Hu, Y.F.; Guo, M.; Sui, S.; Huang, Z.; Liu, X.; Zhao, Z.; Liang, Y.; Wu, H.; Gao, Y.; Qiao, H.; Z. Interpretable machine learning and big data mining to predict the CO2 separation in polymer-MOF mixed matrix membranes   . Adv. Sci. 2024. [Google Scholar] [CrossRef] [PubMed]
  14. Achour, S.; Z.H. ML-Driven Model. Predict. CO2 Uptake Metal.-Org. Fram. (MOFs) Can. J. Chem. Eng. 2024. [CrossRef]
  15. Kirtil, E. Universal prediction of CO2 adsorption on zeolites using machine learning: A comparative analysis with Langmuir isotherm models ChemEngineering 2025. [CrossRef]
  16. Kotov, E.V.; Logabiraman, J.S.; Dhall, G.; Chandna, H.; Madan, M.; Sharma, P.; V. Carbon capture and storage optimization with machine learning using an ANN model   . E3S Web Conf. 2024, 588, 01003. [Google Scholar] [CrossRef]
  17. Zheng, B.; dos Santos, G.X.G.; Ferreira, C.; Steiner, R.N.B.; Luan, M.; B. Simulating CO2 diffusivity in rigid and flexible Mg-MOF-74 with machine-learning force fields   . APL Mach. Learn. 2024, 2, 026115. [Google Scholar] [CrossRef]
  18. Lim, Yunsung; H.P.; Walsh, Aron; Kim, Jihan. Accelerating CO2 direct air capture screening for metal-organic frameworks with a transferable machine learning force field Matter 2025. [CrossRef]
  19. Tayfuroglu, O.; K.S. Modeling CO2 adsorption in flexible MOFs with open metal sites via fragment-based neural network potentials J. Chem. Phys. 2025, 163(5), 054704. [PubMed]
  20. Zhou, Y.; He, S.J.; Fan, S.; Zan, W.; Zhou, L.; Ji, L.; He, X.; G. Machine-learning-assisted high-throughput screening of metal-organic frameworks for CO2 separation from CO2-rich natural gas Ind. Eng. Chem. Res. 2024. [CrossRef]
  21. Bai, X.; Xie, Y.L.; Chen, Y.; Zhang, Q.; Li, X.; J.-R. High-throughput screening of CO2 cycloaddition MOF catalyst with an explainable machine learning model Green Energy Environ 2024. [CrossRef]
  22. Xing, G.; Sun, S.L.; Liu, G.; J.-Y. Modification of metals and ligands in two-dimensional conjugated metal-organic frameworks for CO2 electroreduction: A combined DFT and machine learning study SSRN Electron. J 2024. [CrossRef]
  23. Sheng, Mengjia; X.Z.; Cheng, Hongye; Song, Zhen; Qi, Zhiwen. Multi-Criteria Comput. Screen. [BMIM][DCA] @MOF Compos. CO2 Capture Green Chem. Eng. 2025. [CrossRef] [PubMed]
  24. Park, H. Inverse design of metal-organic frameworks for direct air capture of CO2 via deep reinforcement learning Digital Discovery 2024, 3(4), 728–741.
  25. Badrinarayanan, Srivathsan; Antony, R.M.; Meda, Akshay; Farimani, Radheesh Sharma; Barati, Amir. MOFGPT: Generative Design of Metal-Organic Frameworks using Language Models   . J. Chem. Inf. Model. 2025, 65(17), 9049–9060. [Google Scholar] [CrossRef] [PubMed]
  26. Zhang, Z.; Pan, F.; Mohamed, S.A.; Ji, C.; Zhang, K.; Jiang, J.; Jiang, Z. Accelerating Discovery of Water Stable Metal-Organic Frameworks by Machine Learning   . Small 2024, 20(42), 2405087. [Google Scholar] [CrossRef] [PubMed]
  27. Terrones, G.G.; Huang, S.-P.; Rivera, M.P.; Yue, S.; Hernandez, A.; Kulik, H.J. Metal-Organic Framework Stability in Water and Harsh Environments from Data-Driven Models Trained on the Diverse WS24 Data Set   . J. Am. Chem. Soc. 2024, 146(29), 20333–20348. [Google Scholar] [CrossRef] [PubMed]
  28. Zhang, Z.; Jiang, Z. Discovering Ultra-Stable Metal-Organic Frameworks for CO2 Capture from A Wet Flue Gas: Integrating Machine Learning and Molecular Simulation Environ. Sci. Technol. 2025, 59(18), 9123–9133. [PubMed]
  29. Wilmer, C.E.; Leaf, M.; Lee, C.Y.; Farha, O.K.; Hauser, B.G.; Hupp, J.T.; Snurr, R.Q. Large-Scale Screening of Hypothetical Metal-Organic Frameworks   . Nat. Chem. 2012, 4(2), 83–89. [Google Scholar] [CrossRef] [PubMed]
  30. Willems, T.F.; Rycroft, C.H.; Kazi, M.; Meza, J.C.; Haranczyk, M. Algorithms and Tools for High-Throughput Geometry-Based Analysis of Crystalline Porous Materials   . Microporous Mesoporous Mater. 2012, 149, 134–141. [Google Scholar] [CrossRef]
  31. Wang, W.; Wang, D.; Song, H.; Hao, D.; Xu, B.; Ren, J.; Wang, M.; Dai, C.; Wang, Y.; Liu, W. Size Effect of Gold Nanoparticles in Bimetallic ZIF Catalysts for Enhanced Photo-Redox Reactions   . Chem. Eng. J. 2023, 455, 140909. [Google Scholar] [CrossRef]
  32. Shi, W.-J.; Chen, M.-W.; Li, Y.-Z.; Chen, H.-Y.; Lin, Y.; Wang, H.; Hou, L.; Wang, G.-D. Integrating Hydrogen-Bonding Nanotrap into a Cage-like MOF for Natural Gas Upgrade and MTO Product Separation   . Chem. Eng. J. 2025, 526, 170811. [Google Scholar] [CrossRef]
  33. Obaid, A.; Sun, T.; Babarao, R.; Wang, L.; Gao, L.; Wang, Y.; Hao, D. Integrating Machine Learning in Catalyst Design for Sustainable Hydrogen from Plastic Waste   . Energy Mater. 2026, 6(5). [Google Scholar] [CrossRef]
Figure 2. “Comparison of major machine learning models used in MOF-based CO2 capture, highlighting their strengths, limitations, and best-use scenarios. The figure emphasizes the trade-offs between interpretability, data requirements, and predictive capability.”.
Figure 2. “Comparison of major machine learning models used in MOF-based CO2 capture, highlighting their strengths, limitations, and best-use scenarios. The figure emphasizes the trade-offs between interpretability, data requirements, and predictive capability.”.
Preprints 230370 g001
Figure 3. Key challenges limiting the real-world deployment of ML-assisted MOF design (data scarcity, overfitting, bias toward well-studied MOFs, rigid-framework assumptions, poor humidity handling, limited gas-mixture modeling, and the simulation-experiment gap) and the corresponding future directions discussed throughout this review.
Figure 3. Key challenges limiting the real-world deployment of ML-assisted MOF design (data scarcity, overfitting, bias toward well-studied MOFs, rigid-framework assumptions, poor humidity handling, limited gas-mixture modeling, and the simulation-experiment gap) and the corresponding future directions discussed throughout this review.
Preprints 230370 g002
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.