Preprint
Article

This version is not peer-reviewed.

Comparative Assessment of RF and XGBoost Machine Learning Models for Urban-Scale PM2.5 Estimation in Quito, Ecuador Using Satellite and Ground-Based Observations

Submitted:

12 June 2026

Posted:

16 June 2026

You are already at the latest version

Abstract
Estimating PM₂.₅ exposure in high-altitude Andean cities is challenging due to limited monitoring coverage and the complex interactions among topography, meteorology, and atmospheric chemistry. This study presents a comparative assessment of Random Forest and Extreme Gradient Boosting (XGBoost) for urban-scale PM₂.₅ estimation in Quito, Ecuador. Ground-based observations from the Metropolitan Atmospheric Mon-itoring Network of Quito (REMMAQ) were integrated with Sentinel-5P TROPOMI products, meteorological variables, and topographic predictors processed in Google Earth Engine. Models were developed separately for wet and dry seasons to account for seasonal variability. XGBoost achieved the highest predictive accuracy during the wet season (R² = 0.73), when topographic controls dominated PM₂.₅ variability and the DEM emerged as the most influential predictor. In contrast, RF demonstrated greater robustness during the dry season (R² = 0.63), when photochemical interactions became increasingly important and the CO–SO₂ combustion index was the dominant predictor. Spatial predictions identified a persistent north–south pollution corridor within the urban core of Quito. These findings indicate that PM₂.₅ dynamics in inter-Andean val-leys are governed by seasonally shifting physical and chemical controls. The proposed framework provides a scalable approach for generating spatially continuous air-quality estimates in mountainous urban environments with limited monitoring in-frastructure.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Atmospheric pollution represents one of the most significant threats to public health worldwide, yet scientific evidence regarding the dynamics of pollutants such as particulate matter and their associated health risks remains limited in regions such as Latin America and the Caribbean [1].
Among the various air pollutants, fine particulate matter (PM2.5) is of particular concern due to its microscopic size, with an aerodynamic diameter smaller than 2.5 μm [2]. This characteristic enables PM2.5 particles to penetrate deep into the pulmonary alveoli, inducing oxidative stress and systemic inflammation, and contributing to the development of both respiratory and cardiovascular diseases [3]. PM2.5 originates from a wide range of natural and anthropogenic sources, including wildfires, dust storms, fossil fuel combustion, and biomass burning [4].
Traditionally, air quality monitoring has relied on fixed ground-based monitoring stations [5]. However, as noted by [6,7], these measurements are representative only of their immediate surroundings and lack the spatial coverage required to provide a continuous regional perspective. This limitation hinders the identification of pollution hotspots in areas not covered by monitoring networks [8], particularly in cities characterized by complex topography.
In recent years, remote sensing has emerged as an effective solution to overcome these spatial limitations. Advances in satellite technology and data accessibility have substantially improved the spatiotemporal resolution of air quality monitoring [9]. Within this context, the Sentinel-5P mission and its TROPOspheric Monitoring Instrument (TROPOMI), launched in 2017 as part of the Copernicus Programme [10], provide daily measurements of trace gases, including nitrogen dioxide (NO2), sulfur dioxide (SO2), carbon monoxide (CO), and the Aerosol Index (AI) [7].
Although satellites do not directly measure PM2.5 concentrations, studies such as [11] have demonstrated that PM2.5 can be estimated by integrating satellite-derived pollutant proxies, meteorological variables, and in situ observations through artificial intelligence algorithms. Previous research has validated the application of Machine Learning techniques, particularly Random Forest (RF), for integrating satellite observations and accurately estimating air quality conditions [12]. More recent studies [13,14] indicate that the atmospheric constituents measured by Sentinel-5P serve as effective proxies for PM2.5 because they share common emission sources and exhibit similar spatial and temporal dynamics. In this regard, [7] applied a RF model in Croatia by combining Sentinel-5P observations with meteorological and topographic variables, achieving high predictive performance and accurately identifying urban pollution hotspots.
Despite the growing use of Machine Learning approaches for PM2.5 estimation, most studies have focused on megacities located at sea level or in northern hemisphere regions with stringent regulations governing sulfur content in transportation fuels [7,15]. Consequently, the robustness of these models under the complex topographic conditions of the Andes remains insufficiently explored. In high-altitude mountainous environments, reduced atmospheric pressure and lower oxygen availability significantly decrease the efficiency of internal combustion engines, leading to substantially higher primary pollutant emissions [16,17].
This phenomenon is further exacerbated in developing countries, where lower fuel quality and aging vehicle fleets pose additional challenges to public health. Furthermore, mountain–valley circulation patterns, steep elevation gradients, and heterogeneous emission sources introduce highly nonlinear relationships that may affect model predictive performance [18].
The Metropolitan District of Quito (MDQ) exhibits a unique atmospheric environment characterized by complex wind-channeling effects and intense equatorial solar radiation, both of which substantially influence aerosol chemistry and transport processes. Evaluating Machine Learning models under such extreme environmental conditions constitutes a rigorous test of their spatial adaptability. The Metropolitan Atmospheric Monitoring Network (REMMAQ) [19] provides continuous, high-quality temporal observations [20], creating an ideal empirical foundation for training and validating high-precision Machine Learning architectures while establishing a methodological benchmark for predictive modeling under extreme orographic conditions.
Located within an inter-Andean valley at an average elevation of approximately 2,850 m above sea level, Quito consistently experiences PM2.5 exposure levels exceeding the annual guidelines established by the World Health Organization (WHO), with approximately 80% of emissions attributed to the transportation sector [21]. Additionally, the city exhibits a relatively stable equatorial meteorological regime throughout the year, with precipitation and relative humidity representing the primary sources of seasonal variability. These hydrological variables govern the transition between the dry and wet seasons.
Addressing the nonlinearities associated with complex terrain requires exploring architectural variations within ensemble-learning frameworks as a methodological necessity rather than a simple model comparison. RF constructs multiple decision trees independently and in parallel using randomly selected subsets of the training data. Through its bagging-based resampling mechanism, it effectively reduces model variance and mitigates overfitting in noisy datasets [22]. In contrast, Extreme Gradient Boosting (XGBoost) employs a sequential boosting strategy in which each new tree is trained to correct the residual errors of the preceding trees, thereby reducing model bias [23].
The simultaneous application of both approaches enables a robust assessment of how different mathematical formulations handle the severe vertical and horizontal dispersion constraints characteristic of Andean valleys, ultimately establishing a reference framework for continuous air quality mapping in metropolitan areas with complex terrain.
Therefore, the objective of this study is to assess the predictive performance of RF and XGBoost models for estimating the spatiotemporal dynamics of PM2.5 within the MDQ. Through a systematic comparison of these ensemble-learning architectures, this research seeks to identify the optimal mathematical approach for capturing the nonlinear dispersion patterns inherent to high-altitude inter-Andean valleys. Furthermore, through a rigorous variable importance analysis, the study quantifies the relative influence of Sentinel-5P-derived proxies, meteorological factors, and topographic characteristics on PM2.5 concentrations. In doing so, it provides a robust diagnostic framework for identifying pollution hotspots and supporting air quality management in urban environments with complex terrain. Ultimately, this work delivers a scalable diagnostic tool capable of overcoming the spatial limitations of fixed monitoring networks, offering valuable information for air quality management and health-risk assessment in topographically complex urban centers worldwide.

2. Materials and Methods

This study adopts a quantitative, observational, non-experimental research design with a cross-sectional and descriptive-correlational approach. PM2.5 data collected throughout 2024 across the urban parishes of the Metropolitan District of Quito (MDQ) were analyzed to characterize spatial and temporal patterns of pollutant accumulation during the dry and wet seasons. The study correlates in situ measurements from the Metropolitan Atmospheric Monitoring Network (REMMAQ) [19] with satellite-derived, meteorological, and topographic variables processed within the Google Earth Engine platform, without experimental manipulation of the variables under investigation (Figure 1).

2.1. Study Area

The MDQ, located in the northern Ecuadorian Andes (Figure 2), constitutes the study area. The analysis focused on the urban parishes of the district. In recent years, concentrations of particulate matter (PM2.5 and PM10) have frequently exceeded the air quality guidelines established by the World Health Organization (WHO), primarily due to vehicular emissions, industrial activities, and biomass burning [24].
The MDQ covers approximately 400,000 ha and is home to nearly 2.8 million inhabitants. Elevation ranges from approximately 2,500 to 2,850 m above sea level across the valley regions and the consolidated urban area [25]. The city is bordered by mountainous terrain, which increases the likelihood of thermal inversions. Combined with rapid population growth over recent decades, these conditions make Quito particularly susceptible to episodes of elevated air pollution [26].
Quito experiences a temperate high-altitude equatorial climate characterized by pronounced diurnal temperature variability. The mean annual temperature is approximately 13.4 °C. Precipitation follows a predominantly wet period extending from September to May and a relatively drier season between June and August. Furthermore, the city’s high elevation results in persistently elevated levels of solar and ultraviolet radiation throughout the year [27]. This strong solar irradiance, coupled with marked diurnal variability, contributes to the formation of photochemical pollutants such as ozone.
From a biogeographical perspective, the MDQ encompasses ecoregions associated with the Northern Andean páramo and montane forest systems. The southeastern periphery transitions toward high-Andean ecosystems, whereas the central and northern sectors are predominantly urbanized. This heterogeneous land-use configuration modulates emission patterns and population exposure to air pollutants [28].
The urban core concentrates the majority of the population, industrial activities, and vehicular traffic, while peripheral valleys such as Los Chillos continue to experience rapid urban expansion. Consequently, the MDQ represents a highly suitable case study for investigating air quality dynamics in complex mountainous environments [29].

2.2. In Situ, Satellite, and Auxiliary Data Sources

Hourly PM2.5 concentration records were obtained from the REMMAQ [19], managed by the Environmental Secretariat of the Metropolitan District of Quito. These publicly available data, validated by the Environmental Secretariat, covered the entire year of 2024 and were categorized according to the dry season (June–September) and the wet season (January–May and October–December).
The study incorporated data from the eight active monitoring stations within the network, selected to represent the diverse environmental settings of the Metropolitan District of Quito. These ranged from densely urbanized and high-traffic areas (Belisario, Centro, Cotocollao, and Guamaní) to peripheral valleys and urban expansion zones (Carapungo, San Antonio, Los Chillos, and Tumbaco). This spatial distribution enabled the characterization of key differences in air pollution patterns between the urban core and the surrounding valleys.
A complete-case analysis approach was adopted, retaining only observations with 100% data completeness and removing all null or incomplete records. This ensured that model training relied exclusively on valid and reliable measurements.
The primary source of information for PM2.5 estimation in this study consisted of Sentinel-5P TROPOMI Level-3 (L3 Offline) products processed [30] within the Google Earth Engine (GEE) platform. From these datasets, column densities of six atmospheric constituents were extracted: Aerosol Index (UVAI), carbon monoxide (CO), formaldehyde (HCHO), nitrogen dioxide (NO2), ozone (O3), and sulfur dioxide (SO2).
To characterize atmospheric conditions, meteorological data were obtained from the National Oceanic and Atmospheric Administration (NOAA) Global Forecast System (GFS) [31]. Unlike reanalysis products, GFS provides gridded forecast variables updated four times daily at 6-h intervals. Four key parameters relevant to atmospheric dispersion processes were selected: air temperature at 2 m above ground level, specific humidity (water vapor content per kilogram of air), and the zonal (u) and meridional (v) wind components at 10 m height, which were used to determine airflow direction and magnitude.
Topographic information was obtained from the ALOS World 3D–30 m (AW3D30) Digital Elevation Model [32] developed by the Japan Aerospace Exploration Agency (JAXA). Terrain slope was subsequently derived from the elevation model. A technical summary of all datasets processed in Google Earth Engine is provided in Table 1.

2.3. Data Preprocessing

A spatiotemporal matching procedure was performed to align REMMAQ observations with TROPOMI satellite data. Hourly GFS variables and Sentinel-5P observations were retrieved and processed using the Google Earth Engine Python API within a Google Colab environment and subsequently harmonized to a common spatial framework. The resulting dataset was subsequently stratified into dry-season (June–September) and wet-season (January–May and October–December) subsets.
Only observations exhibiting a temporal correspondence within ±1 h between in situ and auxiliary data sources were retained. Records containing missing or incomplete information were removed. Outlier detection was performed separately for each monitoring station, and anomalous observations were excluded from the final dataset.
To enhance the representation of atmospheric chemical processes, a set of derived TROPOMI variables (Table 2) was generated following the methodology proposed by [15]. These derived predictors were designed to capture chemical signatures and synergistic interactions among atmospheric constituents that are more strongly associated with PM2.5 concentrations than the original satellite variables alone.
Predictor variables were selected separately for each season using the Waikato Environment for Knowledge Analysis (WEKA) software package (Table 3), an open-source platform that provides a wide range of machine learning and data mining algorithms [33]. Feature selection was performed using the Correlation-based Feature Selection (CFS) Subset Evaluator, which identifies subsets of variables that exhibit high predictive relevance with respect to the target variable while maintaining low intercorrelation among predictors. This approach maximizes predictive capability while minimizing redundancy within the feature space [34].

2.4. Model Development and Validation

RF and XGBoost regression models were implemented in Google Colab using the Google Earth Engine Python API to estimate PM2.5 concentrations from the predictor variables selected during the preprocessing stage.
To reduce high-frequency variability, minimize temporal inconsistencies among satellite, meteorological, and ground-based observations, and enhance the representation of persistent pollution patterns, the matched observations were aggregated to weekly averages prior to model development.
The data were partitioned using a 70/30 train–test split and stratified by season to account for seasonal variability in atmospheric conditions and pollutant dynamics.
Model performance was evaluated using the coefficient of determination (R2), root mean square error (RMSE), and mean absolute error (MAE). Spatial prediction maps were generated at an effective spatial resolution of 1 km, enabling the identification and visualization of PM2.5 pollution hotspots across the study area.
To improve model interpretability, SHapley Additive exPlanations (SHAP) analysis was performed. SHAP decomposes model predictions into the individual contributions of each predictor variable, thereby providing a transparent and consistent framework for quantifying feature importance and interpreting model behavior [35]. This approach was used to identify the relative influence of satellite-derived, meteorological, and topographic predictors and to reveal potential nonlinear relationships between explanatory variables and PM2.5 concentrations.

3. Results and Discussion

3.1. Spatiotemporal Dynamics of PM2.5

Surface PM2.5 concentrations estimated by the RF model during the wet season (Figure 3) ranged from 11.21 to 24.78 µg m−3, with a spatial mean concentration of 19.58 µg m−3. The XGBoost model predicted slightly higher concentrations, with values ranging from 10.41 to 28.82 µg m−3 and a spatial average of 20.18 µg m−3.
During the dry season, the RF model revealed a modified spatial distribution pattern accompanied by a generalized reduction in PM2.5 concentrations relative to the wet season. Predicted concentrations ranged from 9.76 to 17.01 µg m−3, with an average concentration of 13.93 µg m−3 across the analyzed urban area. Similarly, the XGBoost model estimated a mean concentration of 12.82 µg m−3, with minimum and maximum values of 8.58 and 22.25 µg m−3, respectively.
The spatial predictions revealed a pronounced seasonal and topographic influence on PM2.5 distribution across the MDQ. Contrary to the atmospheric washout effect commonly associated with increased precipitation, the highest PM2.5 concentrations were observed during the wet season. This pattern suggests a complex interaction among meteorological conditions, atmospheric stability, and local topographic controls that may offset the pollutant removal capacity of rainfall.
Although precipitation is generally effective in removing coarse particulate matter (PM10), its scavenging efficiency decreases considerably for fine particles. This behavior is consistent with the time-lagged cross-correlation analysis reported by [36], which identified delayed responses between hydrological variables and atmospheric composition. Under stable atmospheric conditions and restricted ventilation, continuous pollutant accumulation near the surface may compensate for episodic wet deposition processes, resulting in elevated PM2.5 loads within the urban basin.
The wet season in the MDQ is also characterized by persistent cloud cover, which reduces incoming solar radiation and limits the development of surface-driven thermal convection. This behavior is consistent with the lower-atmosphere processes described by [37], who reported that reduced radiative forcing weakens thermal buoyancy and suppresses sensible heat fluxes. As a consequence, the erosion of thermal inversions may be inhibited, restricting daytime growth of the planetary boundary layer height (PBLH). Under conditions of elevated atmospheric stability and high relative humidity, a shallow boundary layer effectively reduces the available mixing volume, thereby limiting vertical dispersion and favoring the near-surface accumulation of PM2.5 emitted by transportation and industrial sources. Similar behavior has been documented in tropical Andean environments, where lower PBLH values were significantly associated with higher PM2.5 concentrations, confirming the fundamental role of boundary-layer dynamics in regulating pollutant accumulation under complex terrain conditions. [38].
At broader temporal scales, humid conditions and elevated relative humidity do not necessarily result in lower particulate matter concentrations through wet removal processes alone. Instead, as reported by [39] in a long-term analysis covering the period 1980–2024, these meteorological conditions can promote hygroscopic particle growth and enhance secondary aerosol formation through aqueous-phase and heterogeneous chemical reactions. The positive and statistically significant relationship observed between the Standardized Precipitation Index (SPI) and PM2.5 concentrations (r = 0.24, p < 0.01) for humid conditions suggests that highly humid tropical urban environments may facilitate increases in fine particulate matter mass through physicochemical transformation processes, partially offsetting the removal effects of precipitation.
Spatial mapping generated by both machine learning models and for both seasonal periods identified a continuous north–south pollution corridor extending across the central urban sector of the Metropolitan District of Quito. The highest PM2.5 concentrations were consistently observed within the densely urbanized historical center, the north-central district, and the southern urban sector around Guamaní, areas characterized by intense vehicular activity and commercial land use.
The spatial configuration of this corridor closely follows the morphology of the inter-Andean valley and may be explained by the vertical distribution dynamics of combustion-related pollutants described by [40]. Using unmanned aerial vehicle (UAV) observations, these authors demonstrated that the vertical distribution of combustion-derived pollutants is strongly dependent on both emission source characteristics and boundary-layer conditions. Whereas emissions from elevated sources, such as industrial stacks, are released at higher atmospheric levels, traffic-related emissions tend to remain concentrated within the lowest layers of the atmosphere, typically below 50–100 m above ground level. In Quito, this near-surface confinement may interact with the surrounding Andean topography, limiting both lateral and vertical dispersion and favoring pollutant accumulation along the longitudinal axis of the urban valley.
A consistent feature observed during both seasons was the pronounced spatial contrast between the urban plateau and the peripheral valleys. Eastern valleys and topographic transition zones, including Los Chillos and Tumbaco, exhibited the lowest PM2.5 concentrations throughout the study period, with values generally ranging between 9.4 and 11.2 µg m−3. This spatial gradient suggests that the basin-like topography of the MDQ favors pollutant accumulation within the consolidated urban core, whereas peripheral valleys benefit from comparatively greater atmospheric ventilation.
The transition to the dry season was associated with an overall reduction in PM2.5 concentrations and a fragmentation of the spatial pollution pattern. Mean concentrations decreased throughout the study area, and the continuous pollution corridor observed during the wet season evolved into more localized hotspots. This persistence suggests that emission intensity may exert a stronger control on PM2.5 concentrations than seasonal meteorological variability within the most densely urbanized sectors of the MDQ.
From a public health perspective, the modeled PM2.5 concentrations in the MDQ (mean values ranging from approximately 14 to 19 µg m−3, with local maxima exceeding 24 µg m−3) consistently surpassed the annual air quality guideline established by the World Health Organization (WHO). Comparison with other urban centers across the region highlights the particular challenges faced by high-altitude Andean metropolitan areas. This pattern is consistent with the findings reported by [41], who evaluated long-term PM2.5 exposure in major Colombian cities characterized by similar topographic complexity. Their results indicated average PM2.5 concentrations ranging from 8.1 µg m−3 in Bucaramanga to 18.7 µg m−3 within the densely urbanized valley environment of Medellín, values that are comparable to those estimated in the present study.

3.2. Influence of Predictor Variables: Orography Versus Atmospheric Chemistry

Model interpretability analysis based on Gini impurity (Figure 4) revealed substantial seasonal shifts in the relative importance of predictor variables. During the wet season, the physical and dynamical factors selected through the WEKA feature-selection procedure dominated model performance. The DEM emerged as the most influential predictor, accounting for 38.09% of the total importance in XGBoost and up to 60.13% in RF.
This pronounced concentration of predictive importance within the topographic variable reduced the relative contribution of the remaining predictors, relegating meteorological and atmospheric composition variables to secondary roles. This finding is consistent with the results reported by [42], who demonstrated through geographic detector analysis that topography exerts a persistent and statistically significant influence on the spatial variability of atmospheric pollutants in urban environments.
Furthermore, as reported by [43], machine-learning models frequently experience substantial reductions in the availability of satellite-derived optical observations under cloudy and rainy conditions. Cloud cover limits the capability of satellite sensors to retrieve surface-level pollution information, reducing the predictive value of optical atmospheric products. Under these circumstances, tree-based algorithms appear to rely more heavily on predictors that are less affected by cloud interference. In the present study, the wind-dynamics parameter (P14) contributed between 14.9% (RF) and 19.56% (XGBoost) of total feature importance, whereas satellite-derived chemical proxies exhibited comparatively lower influence.
Following the transition to the dry season, the predictive structure of the models became more balanced. In the absence of persistent rainfall, increased solar radiation and elevated precursor-gas concentrations favor photochemical processes associated with secondary aerosol formation [43]. Under clear-sky conditions, atmospheric trace-gas retrievals also benefit from improved observation quality because solar radiation interacts directly with atmospheric constituents without substantial cloud interference.
Feature importance patterns during the dry season (Figure 4b) indicate a transition toward a chemically driven atmospheric regime. In both models, the CO–SO2 combustion index (P18) emerged as the dominant predictor, accounting for 33.5% of total importance in RF and 22.1% in XGBoost. Similarly, the relevance of ozone-related indicators—including P05 (O3/precursor ratio, approximately 19.6–21.6%), P17 (meteorology–ozone index, approximately 18.2–18.7%), and P13 (CO–O3 ratio, approximately 18.6–21.4%)—suggests that atmospheric oxidative capacity plays a key role in regulating PM2.5 formation processes.
Unlike the wet season, where predictive information was strongly concentrated within a single topographic variable, the dry season exhibited a more balanced distribution of explanatory power across multiple predictors. This pattern suggests that PM2.5 variability during the dry season is governed by a network of interacting photochemical and meteorological processes rather than by topographic controls alone.

3.3. Performance Evaluation of Machine Learning Models

To assess the reliability of the spatial estimates, the predictive performance of the implemented machine-learning models was evaluated using independent validation data (Table 4). Both algorithms demonstrated the feasibility of integrating satellite-derived atmospheric information with meteorological and topographic variables for PM2.5 estimation in a complex mountainous environment. These findings are consistent with the methodological conclusions of [15], who reported that mathematically derived relationships among atmospheric trace gases (e.g., variables P01–P18) provide greater predictive power than raw satellite observations alone.
The results indicate that model performance was strongly influenced by seasonal conditions. During the wet season, XGBoost achieved the highest predictive accuracy (R2 = 0.73), together with the lowest MAE and RMSE values. In addition, residual distributions were comparatively homogeneous, indicating stable predictive behavior across the concentration range. Although RF achieved a similar level of performance (R2 = 0.71), it exhibited a tendency to underestimate the highest PM2.5 concentrations and showed mild signs of heteroscedasticity (Figure 5) at the upper end of the distribution. The predictive performance obtained during the wet season is comparable to that reported in PM2.5 estimation studies based on Aerosol Optical Depth (AOD) products [44], despite the absence of direct AOD observations in the present framework.
During the dry season, RF substantially outperformed XGBoost. RF maintained a moderate predictive agreement with observed concentrations (R2 = 0.63), whereas XGBoost experienced a marked decline in explanatory power (R2 = 0.44). This contrast suggests that, under conditions characterized by increased data dispersion and stronger photochemical influences, the parallel ensemble structure of RF was better able to generalize PM2.5 patterns than the sequential boosting strategy employed by XGBoost. The validation metrics obtained in this study are consistent with those reported in previous PM2.5 modeling investigations based on machine-learning approaches [7,15], indicating that satellite-derived atmospheric predictors can provide reliable estimates of fine particulate matter concentrations even in topographically complex environments.
The SHAP analysis (Figure 6) revealed important differences in model behavior and predictive stability between seasons. During the wet season (Figure 6a,b), model explainability was largely governed by relatively stable physical constraints. The SHAP distributions showed a clear dominance of the Digital Elevation Model (DEM), with higher elevation values generally associated with increased predicted PM2.5 concentrations, particularly in the RF model. This concentration of predictive influence suggests that, under rainy conditions, topography acts as the primary control on pollutant accumulation by modulating atmospheric confinement and dispersion processes [37].
Another notable finding concerns the role of the wind-dynamics variable (P14_WIND). Whereas chemically derived indices dominated the dry-season models, wind dynamics emerged as an important adjustment variable during the wet season. The broader spread of SHAP values associated with P14_WIND indicates that both models relied on wind direction and velocity to account for variability introduced by cloud cover and precipitation-related processes. In contrast, the relatively compact SHAP distributions of the chemical proxy variables (P02, P03, and P04), centered around near-zero impacts, suggest that photochemical aerosol formation played a comparatively limited role under wet-season conditions.
A comparison of the two model architectures further indicates that XGBoost exhibited greater sensitivity to elevation than RF during the wet season (Figure 6b). The wider range of SHAP values associated with DEM suggests that XGBoost relied more heavily on topographic information to discriminate between areas of high and low PM2.5 concentrations.
The dry-season SHAP results (Figure 6c,d) revealed a different explanatory structure. In the RF model, high values of the combustion-related index P18 consistently exerted a strong positive influence on predicted PM2.5 concentrations, producing a coherent and interpretable response pattern. In contrast, XGBoost displayed a more complex distribution of feature contributions. Although P05 emerged as the most important predictor, SHAP values associated with P13, P17, and P18 exhibited greater dispersion, indicating stronger interactions among chemical predictors.
This increased variability in feature contributions suggests that XGBoost may have been more sensitive to localized photochemical fluctuations during the dry season. Such behavior is consistent with the reduced predictive performance reported in Table 4, where XGBoost exhibited lower explanatory power than RF. By contrast, the bootstrap aggregation strategy employed by RF appeared to produce more stable estimates by averaging predictor contributions across multiple decision trees. These findings suggest that, under chemically complex atmospheric conditions with strong photochemical activity, RF may provide greater robustness and generalization capability than XGBoost for PM2.5 prediction in high-altitude Andean environments.
This seasonal shift highlights the importance of matching model architecture to the dominant atmospheric processes governing PM2.5 variability.

3.4. Limitations and Future Research Directions

Despite the strong predictive performance achieved by the ensemble models under different meteorological regimes, several limitations should be acknowledged when interpreting these findings.
First, the limited number of observations available during the dry season remains an important constraint. Owing to the intra-annual nature of the study, the density of valid observations was lower during the dry period than during the wet season. This imbalance in sample size likely contributed to the increased variability observed in model performance during the dry season. Expanding the temporal coverage and spatial density of monitoring observations would therefore strengthen future implementations of the proposed framework.
Second, potential uncertainty associated with the chemical signatures used during model training warrants further investigation. Although trace-gas ratios (P01–P18) have consistently been shown to improve predictive performance relative to raw satellite observations [7,15], the results suggest that some of these indicators may exhibit complex nonlinear interactions that are not fully captured by the current modeling approaches. If these chemical proxies are influenced by unresolved meteorological feedbacks, they may introduce systematic uncertainty and limit the ability of machine-learning models to represent small-scale PM2.5 formation processes. Future studies should therefore investigate the interaction between localized boundary-layer dynamics and atmospheric chemical signatures to better characterize and mitigate these sources of uncertainty.
Finally, the geometric complexity of Andean terrain continues to pose challenges for conventional machine-learning architectures. Although the integration of a high-resolution (30 m) Digital Elevation Model successfully captured large-scale topographic effects, sub-grid terrain heterogeneity and its influence on local pollutant dispersion remain difficult to parameterize. Future research should explore the incorporation of additional predictors, including high-resolution land-use and land-cover (LULC) information and local emission inventories, together with more advanced spatial modeling approaches, such as geographically weighted ensemble methods, to better represent the influence of complex orography on PM2.5 distributions.

4. Conclusions

This study assessed the predictive performance of RF and XGBoost for estimating PM2.5 concentrations in the Metropolitan District of Quito using Sentinel-5P satellite products, meteorological variables, and topographic information. The results demonstrated that model performance was strongly dependent on seasonal atmospheric conditions. XGBoost achieved the highest predictive accuracy during the wet season (R2 = 0.73; RMSE = 3.63 µg m−3), whereas RF provided more robust predictions during the dry season (R2 = 0.63; RMSE = 2.74 µg m−3), indicating that no single algorithm consistently outperformed the other under all atmospheric regimes.
The comparison of ensemble-learning architectures revealed that the optimal modeling strategy depends on the dominant environmental processes governing PM2.5 variability. XGBoost performed better under wet-season conditions characterized by strong topographic constraints, while RF demonstrated greater stability under dry-season conditions dominated by complex photochemical interactions. These findings highlight the importance of adopting seasonally adaptive approaches for PM2.5 estimation in high-altitude inter-Andean valleys.
An important atmospheric finding was the observation of higher PM2.5 concentrations during the wet season despite the presence of precipitation. Contrary to the expected washout effect, PM2.5 accumulation remained elevated within the urban core of Quito, suggesting that topographic confinement, reduced planetary boundary-layer development, and limited atmospheric ventilation can outweigh precipitation-driven removal processes in tropical inter-Andean valleys. This result highlights the complex interplay between meteorology and topography in regulating pollutant accumulation at high elevations.
Variable-importance analyses based on Gini importance and SHAP values quantified the relative influence of satellite-derived, meteorological, and topographic predictors. During the wet season, elevation emerged as the dominant explanatory variable, accounting for up to 60.13% of the total importance in the RF model, while wind dynamics represented a secondary control. In contrast, dry-season predictions were primarily driven by photochemical indicators, particularly the CO–SO2 combustion index (P18), which contributed up to 33.5% of the explanatory power. These results suggest a seasonal transition from physically controlled pollutant accumulation to chemically driven aerosol formation processes.
Finally, the proposed framework generated spatially continuous PM2.5 estimates and successfully identified persistent pollution hotspots within the urban core of Quito, including a north–south pollution corridor associated with topographic confinement and intense anthropogenic emissions. By overcoming the spatial limitations of fixed monitoring networks, the methodology provides a scalable decision-support tool for air-quality management and public-health assessment in mountainous urban environments. Future research should evaluate the transferability of this framework to other tropical mountain cities and incorporate additional boundary-layer and emissions-related variables to further improve predictive performance.

Author Contributions

Conceptualization, P.A. and J.M.; methodology, P.A.; software, P.A.; validation, P.A.; formal analysis, J.M., P.A. and L.V.; investigation, J.M.; resources, L.V. and J.E; data curation, J.M and P.A.; writing—original draft preparation, J.M. and P.A; writing—review and editing, J.M.; visualization, P.A.; supervision, L.V. and J.E; project administration, L.V.; funding acquisition, L.V and J.E. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The datasets used in this study are publicly available from the Metropolitan Atmospheric Monitoring Network of Quito (REMMAQ), and the Google Earth Engine datasets for: Sentinel-5P TROPOMI, NOAA GFS, and JAXA ALOS AW3D30.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
Random Forest RF
Extreme Gradient Boost XGBoost
TROPOspheric Monitoring Instrument TROPOMI
Metropolitan District of Quito MDQ
The Metropolitan Atmospheric Monitoring Network RMMAQ
Planetary boundary layer height PBLH

References

  1. Gouveia, N.; Rodriguez-Hernandez, J.L.; Kephart, J.L.; Ortigoza, A.; Betancourt, R.M.; Sangrador, J.L.T.; Rodriguez, D.A.; Diez Roux, A.V.; Sanchez, B.; Yamada, G. Short-Term Associations between Fine Particulate Air Pollution and Cardiovascular and Respiratory Mortality in 337 Cities in Latin America. Sci. Total Environ. 2024, 920, 171073. [Google Scholar] [CrossRef] [PubMed]
  2. Thangavel, P.; Park, D.; Lee, Y.-C. Recent Insights into Particulate Matter (PM2.5)-Mediated Toxicity in Humans: An Overview. Int. J. Environ. Res. Public. Health 2022, 19, 7511. [Google Scholar] [CrossRef] [PubMed]
  3. Lim, E.Y.; Kim, G.-D. Particulate Matter-Induced Emerging Health Effects Associated with Oxidative Stress and Inflammation. Antioxidants 2024, 13, 1256. [Google Scholar] [CrossRef] [PubMed]
  4. Garcia, A.; Santa-Helena, E.; De Falco, A.; de Paula Ribeiro, J.; Gioda, A.; Gioda, C.R. Toxicological Effects of Fine Particulate Matter (PM2.5): Health Risks and Associated Systemic Injuries—Systematic Review. Water. Air. Soil Pollut. 2023, 234, 346. [Google Scholar] [CrossRef] [PubMed]
  5. Bainomugisha, E.; Ssematimba, J.; Okedi, D.; Nsubuga, A.; Banda, M.; Settala, G.W.; Lubisia, G. AirQo Sensor Kit: A Particulate Matter Air Quality Sensing Kit Custom Designed for Low-Resource Settings. HardwareX 2023, 16, e00482. [Google Scholar] [CrossRef] [PubMed]
  6. Li, T.; Shen, H.; Yuan, Q.; Zhang, L. Geographically and Temporally Weighted Neural Networks for Satellite-Based Mapping of Ground-Level PM2.5. ISPRS J. Photogramm. Remote Sens. 2020, 167, 178–188. [Google Scholar] [CrossRef]
  7. Mamić, L.; Gašparović, M.; Kaplan, G. Developing PM2.5 and PM10 Prediction Models on a National and Regional Scale Using Open-Source Remote Sensing Data. Environ. Monit. Assess. 2023, 195, 644. [Google Scholar] [CrossRef] [PubMed]
  8. Plakotaris, D.D.; Kassandros, T.; Bagkis, E.; Karatzas, K. Estimation of Particulate Matter Levels in City Center Pedestrian Routes with the Aid of Low-Cost Sensors. Atmosphere 2024, 15, 965. [Google Scholar] [CrossRef]
  9. Xia, H.; Chen, X.; Wang, Z.; Chen, X.; Dong, F. A Multi-Modal Deep-Learning Air Quality Prediction Method Based on Multi-Station Time-Series Data and Remote-Sensing Images: Case Study of Beijing and Tianjin. Entropy 2024, 26, 91. [Google Scholar] [CrossRef] [PubMed]
  10. Tonion, F.; Pirotti, F. SENTINEL-5P NO2 DATA: CROSS-VALIDATION AND COMPARISON WITH GROUND MEASUREMENTS. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2022, XLIII-B3-2022, 749–756. [Google Scholar] [CrossRef]
  11. Ibrahim, S.; Landa, M.; Pešek, O.; Brodský, L.; Halounová, L. Machine Learning-Based Approach Using Open Data to Estimate PM2.5 over Europe. Remote Sens. 2022, 14, 3392. [Google Scholar] [CrossRef]
  12. Yu, R.; Yang, Y.; Yang, L.; Han, G.; Move, O. RAQ–A Random Forest Approach for Predicting Air Quality in Urban Sensing Systems. Sensors 2016, 16, 86. [Google Scholar] [CrossRef] [PubMed]
  13. Wang, Y.; Yuan, Q.; Li, T.; Tan, S.; Zhang, L. Full-Coverage Spatiotemporal Mapping of Ambient PM2.5 and PM10 over China from Sentinel-5P and Assimilated Datasets: Considering the Precursors and Chemical Compositions. Sci. Total Environ. 2021, 793, 148535. [Google Scholar] [CrossRef] [PubMed]
  14. Son, R.; Kim, H.C.; Yoon, J.-H.; Stratoulias, D. Estimation of Surface Pm2.5 Concentrations from Atmospheric Gas Species Retrieved from Tropomi Using Deep Learning: Impacts of Fire on Air Pollution Over Thailand. SSRN Electron. J. 2022. [Google Scholar] [CrossRef]
  15. Mamić, L.; Kaplan, G.; Gašparović, M. FUSING MULTIPLE OPEN-SOURCE REMOTE SENSING DATA TO ESTIMATE PM2.5 AND PM10MONTHLY CONCENTRATIONS IN CROATIA. GIS Odyssey J. 2022, 2, 59–77. [Google Scholar] [CrossRef]
  16. Tsunemoto, H.; Ishitani, H. The Role of Oxygen in Intake and Exhaust on NO Emission, Smoke and BMEP of a Diesel Engine with EGR System; SAE International, 1980. [Google Scholar]
  17. Zalakeviciute, R.; López-Villada, J.; Rybarczyk, Y. Contrasted Effects of Relative Humidity and Precipitation on Urban PM2.5 Pollution in High Elevation Urban Areas. Sustainability 2018, 10, 2064. [Google Scholar] [CrossRef]
  18. Eguis Cuentas, M.C. Dispersión atmosférica de material particulado en ciudades de montaña y tropicales. Tesis de Maestría, Universidad Nacional de Colombia, Medellín, Colombia, 2023. [Google Scholar]
  19. Secretaría de Ambiente del Distrito Metropolitano de Quito Red Metropolitana de Monitoreo Atmosférico de Quito (REMMAQ).
  20. Oña Proaño, D.; Tafur, P. Diseño de Un Modelo Geoestadístico de La Calidad Del Aire. AXIOMA 2025, 1, 21–30. [Google Scholar] [CrossRef]
  21. Vega, D.; Ocaña, L.; Parra Narváez, R. Inventario de emisiones atmosféricas del tráfico vehicular en el Distrito Metropolitano de Quito. Año base 2012. ACI Av. En Cienc. E Ing. 2015, 7. [Google Scholar] [CrossRef]
  22. Kumar, M.; Agrawal, Y.; Adamala, S.; Pushpanjali; Subbarao, A.V.M.; Singh, V.K.; Srivastava, A. Generalization Ability of Bagging and Boosting Type Deep Learning Models in Evapotranspiration Estimation. Water 2024, 16, 2233. [Google Scholar] [CrossRef]
  23. Jafarzadeh, H.; Mahdianpari, M.; Gill, E.; Mohammadimanesh, F.; Homayouni, S. Bagging and Boosting Ensemble Classifiers for Classification of Multispectral, Hyperspectral and PolSAR Data: A Comparative Evaluation. Remote Sens. 2021, 13, 4405. [Google Scholar] [CrossRef]
  24. Cazorla, M.; Giles, D.M.; Herrera, E.; Suárez, L.; Estevan, R.; Andrade, M.; Bastidas, Á. Latitudinal and Temporal Distribution of Aerosols and Precipitable Water Vapor in the Tropical Andes from AERONET, Sounding, and MERRA-2 Data. Sci. Rep. 2024, 14, 897. [Google Scholar] [CrossRef] [PubMed]
  25. Barrera, A.; Cabrera-Barona, P.; Velasco-Oña, P. Derechos, Calidad de Vida y División Social Del Espacio En El Distrito Metropolitano de Quito. EURE 2022, 48. [Google Scholar] [CrossRef]
  26. Zhou, J.; Gladson, L.; Díaz Suárez, V.; Cromar, K. Respiratory Health Impacts of Outdoor Air Pollution and the Efficacy of Local Risk Communication in Quito, Ecuador. Int. J. Environ. Res. Public Health 2023, 20, 6326. [Google Scholar] [CrossRef] [PubMed]
  27. Zalakeviciute, R.; Vallejo, F.; Erazo, B.; Chimborazo, O.; Bonilla-Bedoya, S.; Mejia, D.; Tapia-Flores, T.I.; Chuquimarca, G.; Rybarczyk, Y. Beyond Global Trends: Two Decades of Climate Data in the World’s Highest Equatorial City. Atmosphere 2025, 16, 1080. [Google Scholar] [CrossRef]
  28. Guarderas, P.; Smith, F.; Dufrene, M. Land Use and Land Cover Change in a Tropical Mountain Landscape of Northern Ecuador: Altitudinal Patterns and Driving Forces. PLoS ONE 2022, 17, e0260191. [Google Scholar] [CrossRef] [PubMed]
  29. Raysoni, A.; Armijos, R.; Weigel, M.; Echanique, P.; Racines, M.; Pingitore, N.; Li, W.-W. Evaluation of Sources and Patterns of Elemental Composition of PM2.5 at Three Low-Income Neighborhood Schools and Residences in Quito, Ecuador. Int. J. Environ. Res. Public. Health 2017, 14, 674. [Google Scholar] [CrossRef] [PubMed]
  30. Veefkind, J.P.; Aben, I.; McMullan, K.; Förster, H.; de Vries, J.; Otter, G.; Claas, J.; Eskes, H.J.; de Haan, J.F.; Kleipool, Q.; et al. TROPOMI on the ESA Sentinel-5 Precursor: A GMES Mission for Global Observations of the Atmospheric Composition for Climate, Air Quality and Ozone Layer Applications. Remote Sens. Environ. 2012, 120, 70–83. [Google Scholar] [CrossRef]
  31. National Centers for Environmental Prediction/National Weather Service/NOAA/U.S. Department of Commerce NCEP GFS 0.25 Degree Global Forecast Grids Historical Archive 2015.
  32. Japan Aerospace Exploration Agency ALOS World 3D 30 Meter DEM 2021.
  33. Cassales, G.W.; Liu, J.J.; Bifet, A. Accelerated Weka: GPU Machine Learning with Weka Workbench. Neurocomputing 2025, 646, 130432. [Google Scholar] [CrossRef]
  34. Frank, E.; Hall, M.A.; Witten, I.H.; Pal, C.J. The WEKA Workbench. Online Appendix for “Data Mining: Practical Machine Learning Tools and Techniques”. In Data Mining: Practical Machine Learning Tools and Techniques; Morgan Kaufmann, 2016. [Google Scholar]
  35. Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions; arXiv: Long Beach, CA, USA, 25 November 2017. [Google Scholar]
  36. Shikwambana, L.; Kganyago, M.; Zhang, X. Analyzing the Effect of the 2015/16 Catastrophic El Niño Event on Wildfire Emissions in Southern Africa Using Lagged Correlation and Interrupted Time-Series Causal Impact Technique. Earth 2026, 7, 42. [Google Scholar] [CrossRef]
  37. Mantovani, J.A.; Carneiro, R.; Borges, C.K.; Ibarra-Espinosa, S.; Aravéquia, J.A.; Fisch, G.; Herdies, D.L. Impact of PBL Schemes on the Simulation of PBL Height in the Central Amazon Basin. Geosciences 2026, 16, 134. [Google Scholar] [CrossRef]
  38. Arias-Arana, D.; Montilla-Rosero, E.; Calderón-Losada, O.; Reina, J.H. Correlating Particulate Matter and Planetary Boundary Layer Dynamics in Northwestern South America: A Case Study of Santiago de Cali. Atmos. Pollut. Res. 2025, 16, 102352. [Google Scholar] [CrossRef]
  39. Palácios, R.; Machado, A.; Franco, R. de C.; Morais, F.G.; Franco, M.A.; Oliveira, F.; Cirino, G.; Imbiriba, B.; Silva, J. de A.; Curado, L.F.A.; et al. Long-Term Links Between Precipitation Regimes and PM2.5 in an Urban Area of Eastern Amazonia (Belém, Brazil), 1980–2024. Atmosphere 2026, 17, 399. [Google Scholar] [CrossRef]
  40. Wang, H.; Huang, C. Investigating Vertical Distributions and Driving Factors of Black Carbon in the Atmospheric Boundary Layer Using Unmanned Aerial Vehicle Measurements in Shanghai, China. Atmosphere 2023, 14, 1472. [Google Scholar] [CrossRef]
  41. Marín, D.; Herrera, V.; Piñeros-Jiménez, J.G.; Rojas-Sánchez, O.A.; Mangones, S.C.; Rojas, Y.; Cáceres, J.; Agudelo-Castañeda, D.M.; Rojas, N.Y.; Belalcazar-Ceron, L.C.; et al. Long-Term Exposure to PM2.5 and Cardiorespiratory Mortality: An Ecological Small-Area Study in Five Cities in Colombia. Cad. Saúde Pública 2025, 41, e00071024. [Google Scholar] [CrossRef] [PubMed]
  42. Zhang, X.; Liu, Y.; Chen, R.; Si, M.; Zhang, C.; Tian, Y.; Shang, G. The Spatiotemporal Evolution and Driving Forces of the Urban Heat Island in Shijiazhuang. Remote Sens. 2025, 17, 781. [Google Scholar] [CrossRef]
  43. Chimla, S.; Chotamonsak, C.; Chaipimonplin, T. Integration of WRF-Chem Model-Based, Satellite-Based, and Ground-Based Observation Data to Predict PM2.5 Concentration by Machine Learning Approach. Atmosphere 2025, 16, 1304. [Google Scholar] [CrossRef]
  44. Lin, H.; Li, S.; Niu, J.; Yang, J.; Wang, Q.; Li, W.; Liu, S. Estimation of Ultrahigh Resolution PM2.5 in Urban Areas by Using 30 m Landsat-8 and Sentinel-2 AOD Retrievals. Remote Sens. 2025, 17, 2609. [Google Scholar] [CrossRef]
Figure 1. Methodology flowchart.
Figure 1. Methodology flowchart.
Preprints 218380 g001
Figure 2. Location of the study area: (a) South America; (b) Ecuador; and (c) the Metropolitan District of Quito (MDQ).
Figure 2. Location of the study area: (a) South America; (b) Ecuador; and (c) the Metropolitan District of Quito (MDQ).
Preprints 218380 g002
Figure 3. Spatial prediction of PM2.5 concentrations using machine learning models: (a) wet season—RF; (b) dry season—RF; (c) wet season—XGBoost; and (d) dry season—XGBoost.
Figure 3. Spatial prediction of PM2.5 concentrations using machine learning models: (a) wet season—RF; (b) dry season—RF; (c) wet season—XGBoost; and (d) dry season—XGBoost.
Preprints 218380 g003
Figure 4. Seasonal comparison of feature importance based on Gini impurity for the RF and XGBoost models. (a) Wet season: dominance of the DEM and wind dynamics. (b) Dry season: diversification of predictive importance among complex gas ratios (e.g., CO–SO2 and CO–O3).
Figure 4. Seasonal comparison of feature importance based on Gini impurity for the RF and XGBoost models. (a) Wet season: dominance of the DEM and wind dynamics. (b) Dry season: diversification of predictive importance among complex gas ratios (e.g., CO–SO2 and CO–O3).
Preprints 218380 g004
Figure 5. Observed versus predicted PM2.5 concentrations for the RF and XGBoost models under different seasonal conditions: (a) RF—wet season; (b) XGBoost—wet season; (c) RF—dry season; and (d) XGBoost—dry season.
Figure 5. Observed versus predicted PM2.5 concentrations for the RF and XGBoost models under different seasonal conditions: (a) RF—wet season; (b) XGBoost—wet season; (c) RF—dry season; and (d) XGBoost—dry season.
Preprints 218380 g005
Figure 6. SHAP summary plots showing feature importance and contribution to model predictions. Panels (a) and (b) correspond to RF and XGBoost during the wet season, respectively, while panels (c) and (d) present the corresponding results for the dry season.
Figure 6. SHAP summary plots showing feature importance and contribution to model predictions. Panels (a) and (b) correspond to RF and XGBoost during the wet season, respectively, while panels (c) and (d) present the corresponding results for the dry season.
Preprints 218380 g006
Table 1. Remote sensing and auxiliary datasets used in this study.
Table 1. Remote sensing and auxiliary datasets used in this study.
Mission/Source Product (Variable) Parameter Description
Sentinel-5P TROPOMI NO2 Nitrogen dioxide column density
CO Carbon monoxide column density
HCHO Formaldehyde column density
O3 Ozone column density
AI UV Aerosol Index
SO2 Sulfur dioxide column density
NOAA (GFS / ERA5) 2-m air temperature (T2M) Air temperature at 2 m above ground level
HUM Specific humidity at 2 m above ground level
U-WIND Zonal wind component at 10 m
V-WIND Meridional wind component at 10 m
JAXA ALOS DEM Digital Elevation Model
SLOPE Terrain slope derived from the DEM
Table 2. Derived variables used in this study.
Table 2. Derived variables used in this study.
Derived Variable Code
(NO2+SO2)/(NO2-SO2) P01
NO2/SO2 P02
HCHO/CO P03
(CO+HCHO)/CO-HCHO) P04
O3/(NO2+SO2+CO) P05
AI*(NO2/SO2) P06
O3/(NO2/SO2) P07
O3/((NO2+SO2)/(NO2-SO2)) P08
SQRT(1/(NO2+SO2+O3)) P09
(AI+HUM)/(AI-HUM) P10
(AI+DEM)/(AI-DEM) P11
(CO+NO2)/CO-NO2) P12
(CO+O3)/(CO-O3) P13
WIND1 WIND
(U-WIND + V-WIND)/2 P14
WHT2 P15
(WHT+AI)/(WHT-AI) P16
(WHT+O3)/(WHT-O3) P17
(CO+SO2)/(CO-SO2) P18
1 (U-WIND + V-WIND)/2, 2 (((U-WIND + V-WIND)/2)+HUM+LST)/3.
Table 3. Predictors Selected by CFS Feature Selection.
Table 3. Predictors Selected by CFS Feature Selection.
Season Instances Selected Predictor Variables
Wet season 1153 DEM, P02, P03, P04, P14_WIND
Dry season 256 P05, P13, P14_WIND, P15, P17, P18
Table 4. Performance metrics for the RF and XGBoost models in PM2.5 estimation, stratified by season.
Table 4. Performance metrics for the RF and XGBoost models in PM2.5 estimation, stratified by season.
Season Model R2 MAE (µg m-3) RMSE (µg m-3)
Wet RF 0.71 2.73 3.76
XGBoost 0.73 2.56 3.63
Dry RF 0.63 2.25 2.74
XGBoost 0.44 2.52 3.35
R2: coefficient of determination; MAE: mean absolute error; RMSE: root mean square error.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings