Preprint
Article

This version is not peer-reviewed.

A Low-Cost IoT Architecture for Micro-Zone Climate Prediction and Meteorological Forecasting

A peer-reviewed article of this preprint also exists.

Submitted:

15 June 2026

Posted:

16 June 2026

You are already at the latest version

Abstract
This article presents the design and deployment of ClimaBogotá v1.2, a climate prediction system tailored for high-altitude urban micro-zones in Bogotá, Colombia. The system combines low-cost IoT sensing, machine learning modeling, and cloud-based orchestration to enable scalable and affordable meteorological forecasting. Its architecture comprises Raspberry Pi-based weather stations, a Random Forest model trained on engineered temporal features, and an n8n-driven automation pipeline for real-time inference and dissemination via Telegram, PostgreSQL, and Grafana. With a Mean Absolute Error of 2.59°C and an R2 of 0.6286 on a 30-minute forecast horizon, the system demonstrates both predictive reliability and operational feasibility using free-tier cloud resources. Unlike traditional weather systems, ClimaBogotá emphasizes modularity, adaptability, and cost-efficiency, offering a replicable framework for decentralized climate monitoring in data-scarce urban environments. Temporal misalignment between sensor nodes was identified as the primary constraint, informing future enhancements toward distributed learning strategies.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Urban climate in high-altitude Latin American cities presents particular characteristics that hinder prediction using conventional meteorological models designed for regional or national scales. Bogotá, Colombia, located on the Cundiboyacense savanna at approximately 2,600 metres above sea level, experiences marked intra-diurnal climate variability, with temperature changes that may exceed 10°C between early morning and midday. Such variability becomes even more significant when analysed at the micro-zone scale, where local topography, land use, urban density, vegetation, and proximity to water bodies produce highly heterogeneous thermal behaviour [1].
This challenge is part of a broader transformation in digital infrastructures for smart environments, where sensing, communication, prediction, automation, and decision support increasingly operate as integrated layers rather than isolated components. Recent studies show that artificial intelligence has become a core enabler for adaptive and data-driven systems in domains such as smart buildings, e-learning, operational command, and smart environments [5,10]. In these contexts, the value of AI does not lie only in prediction accuracy, but also in the system’s capacity to support timely decisions, automate workflows, and remain usable and trustworthy for end users [6]. The challenge of processing heterogeneous and multi-scenario data is well-documented in the machine learning literature: Zhou et al. [35] demonstrate that models operating on multi-source datasets require dedicated strategies to handle variability across scenarios, a concern directly applicable to multi-station IoT deployments where each node exhibits distinct temporal coverage and environmental exposure.
In climate-sensitive urban environments, this perspective is particularly relevant. Weather station networks from official agencies such as Colombia’s Institute of Hydrology, Meteorology and Environmental Studies (IDEAM) provide reliable observations at regional scale, but their spatial density is often insufficient to capture gradients at the micro-zone level. This limitation affects practical domains such as precision agriculture, thermal comfort assessment, urban planning, local risk management, and real-time operational response. Therefore, hyperlocal climate intelligence requires not only low-cost sensing infrastructure, but also robust data engineering, predictive analytics, service orchestration, and secure digital connectivity [2]. Machine learning has proven effective for quality diagnostics and pattern recognition in resource-constrained deployments [37], reinforcing the case for data-driven approaches in low-cost IoT architectures.
The democratization of low-cost hardware platforms such as Raspberry Pi, combined with open-source software ecosystems for machine learning and workflow automation, enables the deployment of local climate intelligence systems at a fraction of the cost of traditional meteorological infrastructure [2]. At the same time, recent research in smart environments emphasizes that predictive systems must be designed with adaptability, resilience, and trustworthiness in mind, especially when they operate in operational or semi-critical contexts [8,27]. Likewise, studies on human interaction with intelligent systems highlight that usability, acceptance, and transparent delivery channels strongly influence whether predictive information is effectively incorporated into decision processes [7,9].
Recent work has established conceptual and technical foundations for Raspberry Pi-based weather station networks targeting micro-zone climate monitoring in Colombian cities, defining hardware architectures, database schemas, and integration frameworks for machine learning-based prediction. Building upon these foundations, the present article documents the complete implementation and validation of a fully operational end-to-end system named ClimaBogotá v1.2. The system integrates four Raspberry Pi weather stations, a Python data-cleaning and feature-engineering pipeline, a Random Forest forecasting model with 43 predictors, a Flask REST API for real-time inference, automated n8n orchestration executing prediction cycles every 5 minutes, persistent PostgreSQL cloud storage, Telegram-based dissemination to subscribers, and Grafana Cloud visualization.
The principal scientific contributions of this work lie in three methodological advances that address persistent challenges in hyperlocal IoT-based climate prediction. First, the session-based feature engineering scheme enables supervised learning from temporally asynchronous and heterogeneous historical datasets without requiring strict time alignment between stations, overcoming a fundamental limitation of conventional multi-sensor fusion approaches that discard non-overlapping observations. Second, the fault-tolerant orchestration architecture ensures continuous operation and partial prediction delivery even when individual sensor nodes fail, a critical requirement for operational deployment that is absent from most research prototypes. Third, the complete integration of the prediction pipeline with accessible end-user delivery channels (Telegram subscription bot, web dashboard) demonstrates that hyperlocal climate intelligence can be made operationally usable at zero marginal cost, addressing the well-documented gap between technically accurate models and systems that non-expert users can trust and incorporate into decision-making workflows. These contributions extend beyond the specific case of Bogotá micro-zones and provide reusable design patterns for low-cost IoT-based environmental monitoring in resource-constrained settings.
The remainder of this article is organized as follows. Section 2 reviews the literature on low-cost environmental sensing, short-term predictive analytics, smart-environment orchestration, and trustworthy digital infrastructures, and positions the present work within the broader research trajectory on adaptive environmental intelligence. Section 3 describes the complete methodology, including the five-layer microservices architecture, the BME280 sensor network, the session-based data processing pipeline with mathematical formulations of the feature engineering operators, the Random Forest ensemble and confidence estimation, the Flask REST API, and the n8n orchestration, Telegram messaging, and PostgreSQL persistence subsystems. Section 4 presents the experimental results on the temporal validation set, feature importance analysis, error patterns, end-to-end latency breakdown, and operational continuity metrics. Section 6 concludes with a summary of contributions, limitations, and the roadmap for future system versions.

3. Methodology

ClimaBogotá v1.2 follows a five-layer microservices architecture with clear separation of responsibilities, as illustrated in Figure 1. This design allows each layer to be independently modified, replaced, or scaled without affecting the others, which constitutes a fundamental decision for the long-term sustainability of the project. The five layers defined in the system are: (1) IoT hardware data acquisition; (2) processing and inference with Python/Flask; (3) orchestration with n8n; (4) persistence with PostgreSQL; and (5) distribution and visualisation with Telegram and Grafana. Communication between these layers is performed exclusively via HTTP REST, guaranteeing interoperability and facilitating the replacement of any component within the ecosystem.
The microservices architecture follows the principle of loose coupling and high cohesion, where each layer exposes a well-defined contract through RESTful interfaces. This architectural pattern enables independent scalability: the Flask inference layer can be horizontally scaled by deploying multiple instances behind a load balancer, while the n8n orchestration layer can be parallelised across multiple workflow executions. The separation of concerns also facilitates incremental updates: the machine learning model can be retrained and redeployed without modifying the orchestration logic, and new visualisation endpoints can be added to the Grafana dashboard without affecting the prediction pipeline. The Flask server acts as the central hub of the system and is the only component with direct access to the trained machine learning model. All other components interact with the model exclusively through the REST API endpoints, which enforces a strict separation between inference logic and orchestration logic. This design choice also simplifies future model updates: replacing the model file requires only restarting the Flask server, with no changes to the orchestration pipeline or the distribution layer.

3.1. Data Acquisition Layer

Each weather station consists of a Raspberry Pi 4B (4 GB RAM, ARM Cortex-A72 quad-core processor at 1.5 GHz) with a Bosch Sensortec BME280 sensor connected via the I2C protocol. The BME280 was selected for its precision (±1.0°C absolute accuracy for temperature, ±1 hPa for atmospheric pressure, and ±3% for relative humidity), low power consumption, and availability in the Colombian market. The use of affordable, widely available sensing hardware is consistent with recent findings showing that low-cost sensor platforms can match the accuracy of more expensive instruments when paired with appropriate data processing pipelines [36]. Figure 2 shows one of the deployed weather stations.
The sensor specifications of the BME280 are particularly well-suited for the climatic conditions of Bogotá. At 2,600 m above sea level, atmospheric pressure ranges between 730 and 760 hPa, well within the sensor’s operational range of 300–1100 hPa. The temperature range of the city, which rarely falls below 5°C or exceeds 25°C, is also comfortably within the sensor’s measurement range. The I2C protocol operates at 400 kHz (Fast Mode) on the Raspberry Pi, enabling a theoretical maximum sampling rate of approximately 30 Hz for the BME280, although the current implementation uses a conservative 0.2 Hz (5-second interval) to balance data granularity with storage overhead and sensor lifespan. The I2C address of the BME280 is configurable via the SDO pin (0x76 or 0x77), allowing multiple sensors on the same bus, although the current deployment uses one sensor per Raspberry Pi to simplify wiring and reduce the risk of address conflicts.
The data collection script on each Raspberry Pi executes a reading cycle every 5 seconds, storing each record in a local MariaDB database with fields for timestamp (ISO 8601), device identifier, temperature in degrees Celsius, pressure in hectopascals, and relative humidity in percentage. The 5-second sampling interval was chosen as a balance between data granularity and storage requirements: at this rate, each station generates approximately 17,280 records per day, or roughly 6.3 million records per year. The four stations accumulated a total of 310,448 raw records during the operating period, of which 216,949 were retained after quality filtering. A significant temporal asymmetry was identified: the Norte and Sur stations have 18 months of history (March 2023 to December 2024), while the Este and Oeste stations have only 9 days and 3 weeks of data respectively, both from December 2024. Table 3 presents the unified view of the sensor specifications and the historical data distribution across the four stations.

3.2. Data Processing Pipeline

The difference between raw and valid records is attributable to three distinct categories of anomalies. The first category consists of cross-contamination records: 7,240 readings from the Este station file that were actually generated by the Norte station device (rasberry_1), resulting from an incorrect database export that merged two tables. The second category consists of overheating records: 4,985 readings from the Oeste station with temperatures above 27°C, which correspond to sensor failures caused by direct solar radiation on the enclosure. The third category consists of integrity failures: 1,274 records with null values or invalid timestamp formats distributed across all four stations.
The first significant technical challenge encountered during data processing was the format of the CSV files exported from MariaDB. The standard exports generated files where all records were concatenated into a single line of text without line breaks, making them unreadable by conventional tools including pandas.read_csv(), Unix awk/sed commands, and Microsoft Excel. A custom parser (parsear_estaciones.py) was developed using Python’s re module to extract each record via a regular expression that detects the pattern of five quoted fields. After parsing, three cascaded quality filters were applied.
The feature engineering scheme was designed specifically to overcome the temporal asymmetry limitation between stations. The challenge of aligning heterogeneous temporal signals from multiple independent sources is a well-recognised problem in multi-modal data processing: Luo et al. [43] address a similar issue in live-streaming platforms by mapping signals from different modalities into a unified temporal semantic space, while Kim et al. [38] demonstrate that fusing local and global feature representations improves model robustness when input sources have different temporal granularities. A session Session k is defined as a maximal block of consecutive records where the gap between successive readings does not exceed 5 minutes:
Session k = { r i : t i + 1 t i 5 min } i I k
where I k denotes the index set of the k-th session. This session-based approach is a deliberate departure from the conventional time-alignment strategy, which would require all four stations to have simultaneous records and would therefore discard the majority of the Norte and Sur data when combined with the much shorter Este and Oeste series. The target variable for each sample is defined as:
y i = T i + 360 ( 30 - minute ahead prediction )
where the offset of 360 positions corresponds to 360 × 5 s = 1 , 800 s = 30  minutes at the 5-second sampling rate.
The feature engineering pipeline transforms the raw time series { T i , P i , H i } i = 1 N into a supervised learning dataset where each sample consists of a 43-dimensional feature vector x i R 43 and a scalar target y i = T i + 360 representing the temperature 30 minutes ahead. The transformation includes temporal lag operators L k [ T i ] = T i k for k { 1 , 6 , 12 , 36 , 72 , 180 , 360 } , differential operators D w [ T i ] = T i T i w for w { 12 , 72 } , and rolling window operators R w [ T i ] as defined in Equation (3). The cyclic encoding of hour-of-day and day-of-year via sine and cosine transforms (Equation (5)) ensures that the model correctly interprets the circular nature of time: hour 23 is close to hour 0, and December 31 is close to January 1.
The rolling mean, differential, and temporal encoding operators are formally defined as:
roll w ( T i ) = 1 w j = i w + 1 i T j
Δ w ( T i ) = T i T i w
h sin = sin 2 π h 24 , h cos = cos 2 π h 24
where w is the window size in number of samples and h is the hour of day. Analogous sine/cosine encodings are applied to the day-of-year variable with a period of 365.25 days.
Table 4 provides a unified summary of the quality control filters and the feature engineering categories, including the mathematical formulations applied in each case.

3.3. Random Forest Prediction Model

The ensemble prediction and its associated uncertainty are computed as follows. Given a forest of B decision trees, the predicted temperature is the average of the individual tree outputs:
y ^ = 1 B b = 1 B T b ( x )
where T b ( x ) denotes the prediction of the b-th tree for input vector x . The prediction uncertainty is estimated via the empirical standard deviation across trees:
σ ( x ) = 1 B 1 b = 1 B T b ( x ) y ^ 2
The model confidence score, reported to end users via the Telegram bot and the Grafana dashboard, is derived from σ through a logistic transformation that maps the unbounded standard deviation to the [ 0 , 1 ] interval:
Confidence ( x ) = 1 1 + σ ( x )
The Random Forest ensemble leverages the bias-variance tradeoff by combining B = 300 independently trained decision trees, each constructed on a bootstrap sample of the training set (bagging) and using a random subset of n features at each split (random subspace method). This double randomisation decorrelates the individual trees, reducing the variance of the ensemble prediction without significantly increasing bias. The standard deviation σ ( x ) across the 300 tree predictions (Equation (7)) provides a natural measure of prediction uncertainty: low σ indicates consensus among trees (high confidence), while high σ indicates disagreement (low confidence), often occurring at the boundaries of the training distribution or during rapid transient events.
The choice of Random Forest over alternative algorithms such as gradient boosting or recurrent neural networks was motivated by four considerations: (1) robustness to sensor noise; (2) ability to capture non-linear relationships; (3) interpretability through feature importance analysis; and (4) computational efficiency, with training completing in approximately 4 minutes and inference taking less than 2 seconds per prediction. The final dataset of 210,225 samples was split using a temporal 85%/15% split, preserving chronological order to prevent data leakage. Table 5 lists the complete set of hyperparameters and their mathematical justifications.

3.4. Inference API and Service Layer

The Flask v2.3 server loads the 269 MB serialised model into RAM at startup and exposes it through a REST API. The design of a lightweight inference layer capable of real-time response on standard cloud hardware is consistent with recent work on compact model architectures for resource-constrained environments: Riaz et al. [40] demonstrate that models combining local and global feature extraction can achieve real-time inference at approximately 15 FPS on edge devices, a performance target analogous to the sub-2.5-second response requirement of the ClimaBogotá inference endpoint. Table 6 describes all available API endpoints.
The REST API follows the JSON:API specification for resource naming and response formatting, with CORS headers enabled to allow cross-origin requests from the Grafana Cloud dashboard. All endpoints support HTTP OPTIONS for preflight requests, and the server implements rate limiting (100 requests/minute per IP) to prevent abuse. The /predecir endpoint accepts an array of exactly 400 timestamped readings, validates the temporal ordering, checks for missing fields, and returns a 400 Bad Request if the input is malformed. The inference latency of 1.5–2.5 seconds is dominated by the feature construction step (rolling windows and lags), which requires iterating over the 400-element input array multiple times; the actual Random Forest predict() call takes only 50–100 ms.

3.5. Orchestration, Messaging, and Persistence

The n8n orchestration pipeline executes the complete prediction cycle every 5 minutes through 7 sequential stages: reading the four CSV files, preparing the last 400 readings per station, calling POST /predecir for each of the four stations in parallel, consolidating the four predictions into a single zone report with fault tolerance, saving the consolidated result to PostgreSQL via POST /guardar, formatting the Telegram message, and sending it to all subscribers. The fault-tolerant consolidation node is a key design feature: it reads the results of the four prediction nodes by name using n8n’s $(’NodeName’).first().json syntax, wrapped in try/catch blocks, so that a failure in any individual station branch does not abort the entire pipeline. Figure 3 shows the complete n8n workflow as deployed.
The Telegram bot (@ClimaBogotaBot) supports subscription management via a persistent suscriptores.txt file. The webhook architecture uses ngrok with a fixed domain to expose the Flask /webhook endpoint to the internet, ensuring that the webhook URL remains stable across server restarts. The PostgreSQL database on Render Cloud stores one consolidated row per prediction cycle, with nullable columns for individual station temperatures that allow the system to record partial results when one or more stations are unavailable. Table 7 consolidates the three integration subsystems into a unified view.
The Grafana Cloud dashboard connects to this database via the native PostgreSQL plugin and displays four panels: time series of temperature by station, model confidence over time, current status statistics, and a tabular history of the last 20 records, all updated every 5 minutes. Table 8 provides a complete summary of all system components, their roles, and their associated costs.
The total hardware cost of 160–240 USD covers four Raspberry Pi 4B units with their respective BME280 sensors, power supplies, microSD cards, and ventilated enclosures. The zero monthly cloud cost is achieved by using the free tiers of Render (PostgreSQL, 1 GB), Grafana Cloud (10,000 metrics), Telegram Bot API, and ngrok (one fixed domain). This cost structure makes the system economically viable for academic research, community initiatives, and small organisations in middle-income countries.

3.6. Security Architecture and Q-OPSEC Alignment

The five-layer microservices architecture of ClimaBogotá v1.2 was designed with security extensibility as a first-class requirement, informed by the PRISEC–Oraculum–DN25 research trajectory [???,?,?]. The current prototype deliberately separates the validation of the prediction pipeline from the hardening of the communication layer: the system first demonstrates operational feasibility ( MAE = 2 . 59 C , R 2 = 0.6286 , zero monthly cloud cost) before introducing cryptographic overhead that could complicate performance profiling during the initial deployment phase.
All inter-layer communication in ClimaBogotá v1.2 is performed via plain HTTP REST. Each Raspberry Pi station writes sensor readings to a local MariaDB database and exports them as CSV files, which are consumed by the n8n orchestration pipeline every 5 minutes. This architecture has a measurable security deficit: a compromised station could inject a fabricated reading y ˜ i into the feature vector x k (Equation (3)), producing a corrupted prediction y ^ zone without triggering any integrity check. The magnitude of this attack is bounded by the influence of station i on the zone prediction:
Δ y ^ zone = w i · y ^ i ( x k ) y ^ i ( x ˜ k ) w i · ϵ max ,
where x ˜ k is the feature vector with the injected reading and ϵ max is the maximum injection magnitude. For the current uniform weighting w i = 0.25 , a single compromised station can shift the zone prediction by up to 0.25 · ϵ max , which for ϵ max = 10 C yields a maximum zone error of 2 . 5 C —comparable to the current MAE of 2 . 59 C and therefore sufficient to render the prediction operationally misleading.
The PRISEC benchmarks [???] establish that, for a BME280 payload of b = 64 bytes, the encryption latency satisfies τ c ( b ) < 2 ms on ARM Cortex-A72, representing an overhead fraction:
ϕ = τ c ( b ) Δ t s < 2 × 10 3 5 = 4 × 10 4 ,
where Δ t s = 5 s is the sampling interval. This confirms that AES-128-GCM or ChaCha20-Poly1305 can be applied to every sensor reading without affecting the acquisition cycle. The contextual security policy of PRISEC II [?] will govern cipher selection in v1.3, mapping the station’s operational state s i = ( ρ i , μ i , λ i ) to the optimal cipher c i * = π ( s i ) as defined in Equation (2) of the related works section, ensuring that the cryptographic overhead remains within the bound ϕ < 4 × 10 4 even under variable CPU load.
The fault-tolerant consolidation logic of the n8n pipeline is formalised as a weighted aggregation over the set of active stations A { 1 , 2 , 3 , 4 } :
y ^ zone = i A w i · y ^ i i A w i , w i = 1 1 + MAE i ,
where MAE i is the historical mean absolute error of station i computed over its validation set. This inverse-MAE weighting scheme assigns higher influence to stations with longer and more accurate operational histories (Norte and Sur, with 18 months of data) and lower influence to recently deployed stations (Este and Oeste, with fewer than 30 days at the time of writing), directly addressing the temporal asymmetry in the training corpus of 310,448 records. The Oraculum framework [?] provides the pathway for replacing this deterministic weighting with a learned policy π θ that optimises the composite objective:
J ( θ ) = E MAE zone + λ r · E n a λ s · E σ zone ,
where λ r and λ s are regularisation coefficients that penalise low station availability and high prediction uncertainty respectively, and σ zone is the inter-station standard deviation of individual predictions.
The confidence score delivered via Telegram is computed as:
C ( x ) = 1 σ ( x ) σ max , σ ( x ) = 1 T t = 1 T y ^ t ( x ) y ¯ ( x ) 2 ,
where T is the number of trees in the Random Forest ensemble, y ^ t ( x ) is the prediction of tree t, y ¯ ( x ) is the ensemble mean, and σ max is the maximum observed inter-tree standard deviation over the validation set. A confidence score C ( x ) 0.75 indicates that the ensemble is in a low-uncertainty region of the feature space, while C ( x ) < 0.50 signals that the current atmospheric conditions are underrepresented in the 310,448-record training corpus.
The integration of DN25 [?] into ClimaBogotá v1.3 will replace the current ngrok HTTP tunnel with a Q-OPSEC-secured channel. Each station e i will authenticate to the Flask server using a CRYSTALS-Kyber-768 key pair, and the shared session key K session will be derived as:
K session = KDF Kyber . Dec ( sk i , c i ) nonce i ,
where sk i is the station’s private key, c i is the ciphertext received from the server, nonce i is a session-specific random value, and KDF is a key derivation function (HKDF-SHA3-256). This authenticated channel prevents the data injection attack quantified in Equation (6), reducing Δ y ^ zone to zero for any station operating under a valid Q-OPSEC session. The security roadmap from the current unauthenticated HTTP prototype to a Q-OPSEC-secured operational service reflects the broader design philosophy of ClimaBogotá: each version advances one architectural layer while preserving the validated components of the previous version, ensuring that the transition from research prototype to urban operational service is incremental, reproducible, and grounded in empirically validated security primitives.

4. Results

The model was evaluated on the temporal validation set of 31,534 samples, corresponding to 15% of the total dataset reserved exclusively for validation and never exposed during training. The temporal split strategy, preserving chronological order, is critical in time-series forecasting to prevent data leakage, since a random split would allow the model to implicitly learn future patterns from training samples interleaved with validation samples, artificially inflating performance metrics. The Mean Absolute Error (MAE) achieved was 2.59°C, the Root Mean Square Error (RMSE) was 4.46°C, the coefficient of determination R 2 was 0.6286, the Mean Absolute Percentage Error (MAPE) was 16.2%, and the median absolute error was 1.87°C. These results indicate that the model explains approximately 62.9% of the temperature variability in the validation set, and that half of all predictions carry an error below 1.87°C. To contextualise these values, the standard deviation of temperature in the validation dataset is approximately 3.2°C, which means the model captures the majority of the temporal patterns present in the data. It is worth noting that an R 2 of 0.6286 represents a strong result for this class of prediction problem: Naskinova et al. [42] report that even well-tuned ensemble models (CatBoost, Random Forest) achieve R 2 values of only 0.19–0.27 when predicting physiological variables from routinely collected clinical data, underscoring that modest explained variance is a common outcome when the target variable is inherently complex and the available features capture only a fraction of the underlying dynamics. The gap between MAE (2.59°C) and RMSE (4.46°C) indicates the presence of a subset of high-error predictions, which are analysed in detail below. Table 9 presents the complete set of performance metrics.
The feature importance analysis reveals that the most influential predictors are recent temperature statistics, which is consistent with the physics of short-term prediction: the 6-minute rolling mean of temperature accounts for 15.3% of the model’s decision weight, followed by the 1-minute rolling mean (10.8%) and the current temperature value (9.7%). Together, these three features alone account for 35.8% of the total predictive weight, confirming that the thermal inertia of the local microclimate is the dominant signal at the 30-minute horizon. Temporal cyclic features (hour sine and cosine) appear in positions 8 and 10, confirming that the model has learned the diurnal patterns of Bogotá’s climate, particularly the characteristic morning warming between 08:00 and 13:00 and the afternoon cooling after 16:00 associated with convective activity. Pressure-based features appear from position 9 onwards, suggesting that atmospheric pressure contributes secondary but non-negligible predictive information at the 30-minute horizon. Figure 4 presents the top 10 most important features.
The analysis of prediction errors reveals a systematic pattern: the largest errors (above 5°C) are concentrated in abrupt temperature changes typical of Bogotá afternoons between 14:00 and 17:00, when direct sunlight can rapidly raise the temperature while convective clouds produce sudden drops. This phenomenon, locally known as veranillo de la tarde, is characterised by temperature gradients exceeding 4°C in under 20 minutes, which falls precisely within the 30-minute prediction horizon of the model. The Random Forest architecture, being a tree-based ensemble that interpolates between training observations rather than extrapolating, is inherently limited in its ability to predict temperature values outside the range observed during training for a given feature combination. This structural limitation explains why the RMSE (4.46°C) is disproportionately larger than the MAE (2.59°C). The end-to-end latency of the prediction pipeline was measured across 50 consecutive execution cycles during the testing period. Table 10 presents the latency breakdown by pipeline stage.
The dominant stage is the parallel Flask inference calls (8–12 s), which involve loading the last 400 readings per station from CSV, constructing the 43-feature vector, and invoking the Random Forest’s predict() method on the 269 MB serialised model. The parallelisation of the four station calls via n8n’s native HTTP node reduces this stage from a theoretical sequential maximum of 48 s to the observed 8–12 s. The PostgreSQL persistence stage (3–5 s) is bounded by the round-trip latency to Render Cloud’s free-tier infrastructure in the US-East region, which introduces a fixed network overhead of approximately 180 ms per request. The total cycle latency of 30–45 s represents only 10–15% of the 5-minute cycle period, leaving ample margin for additional processing stages in future versions of the pipeline.
Figure 5. Scatter plot of model predictions versus actual temperature values on the temporal validation set (31,534 samples).
Figure 5. Scatter plot of model predictions versus actual temperature values on the temporal validation set (31,534 samples).
Preprints 218747 g005
The n8n pipeline operated continuously during the development and testing period without manual errors, with all four stations active simultaneously. The Telegram bot responded to /predecir commands in under 5 seconds. The Grafana dashboard updated every 5 minutes with data from all four stations, and the PostgreSQL database on Render accumulated one consolidated prediction row per cycle with no data loss during the test period, resulting in a database of approximately 2,016 prediction records per week of continuous operation. The total cloud infrastructure cost was 0 USD/month, with a one-time hardware investment of 160–240 USD for the four Raspberry Pi stations.
The most significant limitation of the current model is the temporal asymmetry between stations: Norte and Sur have 18 months of history covering multiple seasons, while Este and Oeste have only 9 days and 3 weeks of data from December 2024. This means the model has not learned the seasonal variation patterns of the Este and Oeste stations, nor the spatial gradients between all four zones. It is estimated that with 6 months of simultaneous data from all four stations, the R 2 should improve to a range of 0.72–0.80. Additional current limitations include the operation with historical CSV files rather than real-time Raspberry Pi readings, the dependency on a local machine to run Flask and n8n, and the absence of external validation data from IDEAM stations to quantify the systematic bias of the model with respect to official measurements.

5. Discussion

The results presented in Section 4 confirm that the ClimaBogotá v1.2 architecture achieves its primary design objectives: operational feasibility, economic accessibility, and predictive reliability at the micro-zone scale. The MAE of 2.59°C and R 2 of 0.6286 are consistent with the expected performance of a Random Forest model trained on a temporally asymmetric dataset, where two of the four stations contribute only short-term data from a single month. As demonstrated by Naskinova et al. [42], even state-of-the-art ensemble methods face fundamental limits when the available features capture only a fraction of the target variable’s variance; the R 2 achieved here is therefore a strong result given the data maturity constraints. The session-based feature engineering scheme proved effective in enabling supervised learning from heterogeneous historical records without discarding the majority of the Norte and Sur data, which would have been the outcome of a conventional time-alignment strategy. This approach is conceptually aligned with the cross-modal temporal alignment strategies proposed by Luo et al. [43] and the local-global feature fusion demonstrated by Kim et al. [38], both of which address the challenge of integrating signals with different temporal granularities.
The fault-tolerant n8n orchestration architecture demonstrated continuous operation throughout the testing period, validating the design decision to use try/catch consolidation logic rather than strict pipeline dependencies. The Telegram-based dissemination channel proved operationally effective, with response latencies under 5 seconds for on-demand predictions, confirming that the system can serve non-expert users without requiring additional applications or proprietary infrastructure. The planned integration of a large language model for natural-language climate summaries in Version 2.0 is supported by recent evidence that context-aware LLM systems can produce more relevant and personalized outputs when environmental and user-profile signals are incorporated into the retrieval pipeline [41].
The identified limitation of temporal asymmetry between stations is not a fundamental architectural constraint but rather a data maturity issue that will be resolved as the Este and Oeste stations accumulate longer operational histories. The projected R 2 improvement to 0.72–0.80 after six months of simultaneous four-station operation is consistent with the expected reduction in model uncertainty when spatial gradient features become available across all zones.

6. Conclusions

This work demonstrated that it is possible to implement a fully functional hyperlocal climate prediction system for Colombian micro-zones using low-cost hardware, open-source software, and free cloud services. The successful implementation of the five system layers, from Raspberry Pi data collection to Grafana Cloud visualisation, validates the conceptual proposal of Guerrero Mateus and De Paz Santana (2024) and establishes a concrete experimental platform for subsequent research phases.
The main technical contributions of this work are: (1) the development of a robust parser for the non-standard data format generated by the Raspberry Pi stations, reusable in any similar deployment of the proposed architecture; (2) the design of a session-based feature engineering scheme that enables training machine learning models with temporally spaced and asynchronous data between stations, overcoming the conventional limitation of temporal alignment; (3) the integration of n8n as a low-code orchestrator for IoT climate prediction systems, with a fault-tolerant design that continues operating when one or more stations are inactive; and (4) the implementation of a Telegram-based prediction distribution system with persistent subscriptions that democratises access to predictions without requiring additional applications or proprietary notification infrastructure.
From a cost and accessibility perspective, the system demonstrates that the total implementation budget, estimated at 160–240 USD in hardware plus 0 USD/month in software infrastructure, is compatible with academic projects, community initiatives, and small organisations in middle-income countries such as Colombia. This economic accessibility is fundamental for the replicability of the system in other micro-zones across the country.
Future work is organised into four priority lines. Version 1.3 will connect n8n directly to the MariaDB databases of the four Raspberry Pi stations via the n8n MySQL plugin, eliminating the CSV file dependency, and will deploy Flask on Render with a paid plan to guarantee 24/7 availability. Version 1.4 will retrain the model after six months of simultaneous operation of all four stations in their definitive triangulated locations within the 10 km2 micro-zone, enabling the model to learn spatial gradients and complete seasonal patterns, with an expected R 2 improvement to 0.72–0.80. Version 1.5 will integrate historical and real-time data from the nearest IDEAM stations as additional model features, enabling external validation and potentially extending the prediction horizon to 1–3 hours. Version 2.0 will integrate a large language model (Gemini 2.0 Flash) to generate climate analyses in accessible colloquial language for non-technical users [41], and will develop a custom web dashboard with an interactive micro-zone map visualising real-time temperature gradients [14,15,16].

Author Contributions

Conceptualization, C.A.G.M. and J.F.D.P.; methodology, C.A.G.M. and J.M.B.S.; software, C.A.G.M. and J.M.B.S.; validation, C.A.G.M. and D.N.; formal analysis, C.A.G.M.; investigation, C.A.G.M.; resources, V.R.Q.L.; data curation, C.A.G.M. and J.M.B.S.; writing—original draft preparation, C.A.G.M.; writing—review and editing, D.N., V.R.Q.L. and J.F.D.P.; visualization, C.A.G.M. and J.M.B.S.; supervision, V.R.Q.L. and J.F.D.P.; project administration, V.R.Q.L. All authors have read and agreed to the published version of the manuscript.

Funding

The authors acknowledge financial support from CNPq (National Council for Scientific and Technological Development, grant 307137/2022-8) and CAPES (Coordination for the Improvement of Higher Education Personnel – Brazil, Finance Code 001). This research was partially funded by national funds through the FCT – Foundation for Science and Technology, I.P. within the scope of projects UIDB/04466/2025 and UIDP/04466/2025, and project 16881, LISBOA2030-FEDER-00816400, DOI:

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The data and code supporting the results of this study are available upon reasonable request to the corresponding author.

Acknowledgments

The authors thank ISCTE-IUL (University Institute of Lisbon) and the University of Salamanca for their institutional support. Colaboración Consejería de educación de la Junta de Castilla y León grupo de investigación ESAL-EXPERT SYSTEM AND APPLICATIONS LAB (ESALAB).

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Guerrero Mateus, C.A.; De Paz Santana, J.F. Proposal for a Raspberry Pi-based weather station network for micro-zone climate monitoring in Colombian cities. In Doctoral Research Proposal; Universidad de Salamanca: Salamanca, Spain, 2024. [Google Scholar] [CrossRef]
  2. Alali, Y.; Harrou, F.; Sun, Y. A proficient management of solar energy production using IoT-based monitoring system. IEEE Access 2020, 8, 48845–48857. [Google Scholar] [CrossRef]
  3. Mathon, V.; Boone, F.; Lafore, J.-P. Short-term forecasting of convective rainfall using machine learning. Q. J. R. Meteorol. Soc. 2021, 147, 2305–2323. [Google Scholar] [CrossRef]
  4. Holgado, B.; Urquiza-Aguiar, L.; Paredes-Paredes, M.C. Context-aware security and post-quantum cryptography for IoT networks. Internet Things 2024, 28, 101358. [Google Scholar] [CrossRef]
  5. Bakhouyi, A.; Dehbi, R.; Talea, M.; Benhaddou, A. A hybrid AI system for predictive analytics in smart learning environments. Array 2025, 25, 100199. [Google Scholar] [CrossRef]
  6. Emon, M.A.K.; Hasan, M.M.; Islam, T. Mediating role of attitudes in AI usage among professionals. Array 2025, 25, 100188. [Google Scholar] [CrossRef]
  7. Sasongko, P.S.; Wibowo, A.; Hartanto, D. Extended TAM analysis for e-learning acceptance in higher education. Array 2025, 25, 100192. [Google Scholar] [CrossRef]
  8. Tari, Z.; Mahmood, A.; Bertok, P. Human-centric cybersecurity for resilient digital ecosystems. Future Gener. Comput. Syst. 2025, 162, 107–121. [Google Scholar] [CrossRef]
  9. Ibrahim, R.; Zainuddin, N.; Azizan, S. Effect of VR on decision-making, presence, and task load in operational scenarios. Array 2026, 26, 100282. [Google Scholar] [CrossRef]
  10. Khan, M.A.; Alazab, M.; Xu, Q. AI and blockchain integration for efficiency and security in smart buildings: A review. Array 2026, 26, 100284. [Google Scholar] [CrossRef]
  11. Narayan, V.; Sharma, S.; Gupta, A. Detecting misinformation in LLM-generated reviews: Governance and trust challenges. Array 2026, 26, 100285. [Google Scholar] [CrossRef]
  12. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
  13. Pedregosa, F.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. Available online: https://scikit-learn.org (accessed on December 2024).
  14. n8n GmbH. n8n—Workflow Automation Tool, 2024. Available online: https://n8n.io (accessed on December 2024).
  15. Grafana Labs. Grafana: The Open Observability Platform, 2024. Available online: https://grafana.com (accessed on December 2025).
  16. Projects, Pallets. Flask: A Lightweight WSGI Web Application Framework, 2024. Available online: https://flask.palletsprojects.com (accessed on December 2025).
  17. Saraiva, D.A.F.; Savii, C.R.; Barbosa, J.L.V.; Leithardt, V.R.Q. PRISEC: Comparison of symmetric key algorithms for IoT devices. Sensors 2019, 19, 4312. [Google Scholar] [CrossRef] [PubMed]
  18. da Costa, P.; Noetzold, D.; Barbosa, J.L.V.; Leithardt, V.R.Q. PRISEC II: A comprehensive model for IoT security. In Proceedings of DiTTET 2025, AISC; Springer: Berlin, Germany, 2025; vol. 1465, pp. 147–159. [Google Scholar] [CrossRef]
  19. Sohail, H.; Noetzold, D.; Hadi, M.; Leithardt, V.R.Q.; Barbosa, J.L.V. PRISEC III: Cryptographic techniques for enhanced security in IoT environments. Proc. IEEE CSCS 2025, 2025, 354–361. [Google Scholar] [CrossRef]
  20. Noetzold, D.; Leithardt, V.R.Q.; Barbosa, J.L.V. Oraculum: A model for self-adaptive system optimization in smart environments. Expert Syst. Appl. 2026, 315, 131705. [Google Scholar] [CrossRef]
  21. Noetzold, D.; Leithardt, V.R.Q.; Barbosa, J.L.V.; da Costa, C.A. DN25: An adaptive quantum cryptography protocol for secure and efficient communication in IoT networks. Rev. I+D Tecnológico 2025, 19. [Google Scholar] [CrossRef]
  22. Suarez-Roman, M.; Tapiador, J. Attack structure matters: Causality-preserving metrics for Provenance-based Intrusion Detection Systems. Comput. Secur. 2025, 157, 104578. [Google Scholar] [CrossRef]
  23. Picazo-Sanchez, P.; Tapiador, J.; Schneider, G. After you, please: browser extensions order attacks and countermeasures. Int. J. Inf. Secur. 2020, 19, 623–638. [Google Scholar] [CrossRef]
  24. Blazquez, E.; Tapiador, J. Practical Android Software Protection in the Wild. ACM Comput. Surv. 2026, 58, 36. [Google Scholar] [CrossRef]
  25. Yuste, J.; Pardo, E.G.; Tapiador, J. Optimization of code caves in malware binaries to evade machine learning detectors. Comput. Secur. 2022, 116, 102643. [Google Scholar] [CrossRef]
  26. Marson, L.M.; de Moura Costa, H.J.; Leithardt, V.R.Q.; Crocker, P.A. Detecting Faults in Furniture Parts Using Neural Networks in a Mobile App. In Proceedings of the 25th International Conference on Control Systems and Computer Science (CSCS), Bucharest, Romania, 2025; pp. 306–313. [Google Scholar] [CrossRef]
  27. Noetzold, D.; Leithardt, V.R.Q.; de Paz Santana, J.F.; Barbosa, J.L.V. A Self-Adaptive Architecture for Predictive and Reinforcement-Based Optimization in Smart Environments. In Proceedings of the 2025 International Symposium on Networks, Computers and Communications (ISNCC), 2025; pp. 1–6. [Google Scholar] [CrossRef]
  28. Noetzold, D.; de Moraes Rossetto, A.G.; Silva, L.A.; Crocker, P.; Leithardt, V.R.Q. JVM optimization: An empirical analysis of JVM configurations for enhanced web application performance. SoftwareX 2024, 28, 101933. [Google Scholar] [CrossRef]
  29. Zou, K.; Wang, S.; Li, Y.; Zhao, H. Research on Building System of Adaptive Security Protection System of Cloud Platform Based on Localization Large Model. In Proceedings of the 2024 International Conference on Electronics and Devices, Computational Science (ICEDCS), 2024; pp. 980–985. [Google Scholar] [CrossRef]
  30. Ari, I.; Balkan, K.; Pirbhulal, S.; Abie, H. Ensuring Security Continuum from Edge to Cloud: Adaptive Security for IoT-based Critical Infrastructures using FL at the Edge. In Proceedings of the 2024 IEEE International Conference on Big Data (BigData), 2024; pp. 4921–4929. [Google Scholar] [CrossRef]
  31. Mirsadri, S.J.; Chaves, R.; Pedrosa, L. Energy-Aware Adaptive Security for Smart Farming (EAASF): A Hybrid IDS-IPS Framework with SDN-Orchestrated for Agriculture 4.0. In Proceedings of the 2025 23rd International Symposium on Network Computing and Applications (NCA), 2025; pp. 328–329. [Google Scholar] [CrossRef]
  32. Sohail, H.; Noetzold, D.; Leithardt, V.R.Q. PRISEC III: Dynamic Cryptographic Adaptation for Balancing Performance and Security. J. Internet Serv. Appl. 2026, 17, 174–191. [Google Scholar] [CrossRef]
  33. Zhang, Y.; Wang, X.; Liu, J. Context-aware security management in smart environments: A machine learning approach. J. Netw. Comput. Appl. 2021, 177, 102932. [Google Scholar] [CrossRef]
  34. Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018; ISBN 978-0-262-03924-6. [Google Scholar]
  35. Zhou, D.; Shi, L.; Wang, B.; Xu, H.; Huang, W. Are Your Comments Positive? A Self-Distillation Contrastive Learning Method for Analyzing Online Public Opinion. Electronics 2024, 13, 2509. [Google Scholar] [CrossRef]
  36. Vasta, N.; Jajo, N.; Graf, F.; Zhang, L.; Biondi, F.N. Evaluating a Camera-Based Approach to Assess Cognitive Load During Manufacturing Computer Tasks. Electronics 2025, 14, 467. [Google Scholar] [CrossRef]
  37. Firza, N.; Bakiu, A.; Monaco, A. Machine Learning for Quality Diagnostics: Insights into Consumer Electronics Evaluation. Electronics 2025, 14, 939. [Google Scholar] [CrossRef]
  38. Kim, N.; Lim, H.; Li, Q.; Li, X.; Kim, S.; Kim, J. Enhancing Review-Based Recommendations Through Local and Global Feature Fusion. Electronics 2025, 14, 2540. [Google Scholar] [CrossRef]
  39. Santamaría, A.; Alfaro, M.; Antón, C.; Sánchez-Quiñones, B.; Ibarra, N.; Gil, A.; Reinoso, O.; Payá, L. Machine Learning-Based Predictive Model for Risk Stratification of Multiple Myeloma from Monoclonal Gammopathy of Undetermined Significance. Electronics 2025, 14, 3014. [Google Scholar] [CrossRef]
  40. Riaz, W.; Ji, J.; Ullah, A. TriViT-Lite: A Compact Vision Transformer–MobileNet Model with Texture-Aware Attention for Real-Time Facial Emotion Recognition in Healthcare. Electronics 2025, 14, 3256. [Google Scholar] [CrossRef]
  41. Karlović, R.; Rovis, M.; Smajić, A.; Sever, L.; Lorencin, I. Context-Aware Tourism Recommendations Using Retrieval-Augmented Large Language Models and Semantic Re-Ranking. Electronics 2025, 14, 4448. [Google Scholar] [CrossRef]
  42. Naskinova, I.; Kolev, M.; Karova, D.; Milev, M. Machine Learning-Based Blood Pressure Prediction Using Cardiovascular Disease Data: A Comprehensive Comparative Study. Electronics 2026, 15, 312. [Google Scholar] [CrossRef]
  43. Luo, J.; Zhu, P.; Wang, Y.; Xiao, Z.; Li, J.; Kong, X.; Zhan, Y. A Data-Driven Multimodal Method for Early Detection of Coordinated Abnormal Behaviors in Live-Streaming Platforms. Electronics 2026, 15, 769. [Google Scholar] [CrossRef]
Figure 1. Overview of the ClimaBogotá v1.2 architecture. Five layers communicate exclusively via HTTP REST: (1) IoT acquisition with Raspberry Pi 4B + BME280; (2) Python processing pipeline and Random Forest inference via Flask; (3) n8n orchestration every 5 minutes; (4) PostgreSQL persistence on Render Cloud; (5) distribution via Telegram bot and Grafana Cloud dashboard.
Figure 1. Overview of the ClimaBogotá v1.2 architecture. Five layers communicate exclusively via HTTP REST: (1) IoT acquisition with Raspberry Pi 4B + BME280; (2) Python processing pipeline and Random Forest inference via Flask; (3) n8n orchestration every 5 minutes; (4) PostgreSQL persistence on Render Cloud; (5) distribution via Telegram bot and Grafana Cloud dashboard.
Preprints 218747 g001
Figure 2. One of the four ClimaBogotá weather stations: Raspberry Pi 4B with Bosch BME280 sensor (temperature, pressure, humidity) connected via I2C, housed in a ventilated enclosure for outdoor deployment.
Figure 2. One of the four ClimaBogotá weather stations: Raspberry Pi 4B with Bosch BME280 sensor (temperature, pressure, humidity) connected via I2C, housed in a ventilated enclosure for outdoor deployment.
Preprints 218747 g002
Figure 3. n8n orchestration workflow for ClimaBogotá v1.2.
Figure 3. n8n orchestration workflow for ClimaBogotá v1.2.
Preprints 218747 g003
Figure 4. Top 10 most important features of the Random Forest model.
Figure 4. Top 10 most important features of the Random Forest model.
Preprints 218747 g004
Table 1. Research evolution leading to ClimaBogotá v1.2.
Table 1. Research evolution leading to ClimaBogotá v1.2.
Work Primary Focus Contribution to ClimaBogotá
PRISEC [17] IoT Symmetric Crypto Benchmarking Performance baseline for low-cost sensor-to-server secure communication.
PRISEC II [18] Multi-Layer Security Modeling Integration of contextual signals (device, network, environment) for adaptive policies.
PRISEC III [19] Cryptographic Agility Dynamic protocol switching based on operational context and resource constraints.
Oraculum [20] RL-Based Self-Adaptive Optimization Autonomous decision engine for parameter tuning and workflow orchestration.
DN25 [21] Quantum-Safe Protocol Foundation for future secure communication between distributed weather stations.
Table 2. Comparison Between Representative Related Works and the Proposed ClimaBogotá System.
Table 2. Comparison Between Representative Related Works and the Proposed ClimaBogotá System.
Work Domain Main Contribution Relation to This Article Difference from ClimaBogotá
Alali et al. [2] IoT monitoring / smart farming Low-cost sensor network with cloud services Viability of low-cost environmental sensing Not focused on hyperlocal forecasting with ML or messaging delivery
Mathon et al. [3] Short-term env. prediction ML for short-term forecasting Supports ML for short prediction horizons No low-cost IoT, orchestration, or user-facing services
Holgado et al. [4] IoT security Context-aware security and post-quantum cryptography Motivates secure communication in climate-IoT Security-focused; no climate modelling or orchestration
Bakhouyi et al. [5] AI hybrid / big data Hybrid AI for predictive analytics in smart learning Importance of full data-processing pipelines Not focused on IoT climate sensing or forecasting
Emon et al. [6] AI adoption Mediating role of attitudes in AI usage Trust and adoption of AI-based predictive systems Adoption behaviour only; no system implementation
Noetzold et al. [27] Smart environments Self-adaptive architecture for RL-based optimization Closest to adaptive smart-system logic used here No climate sensing, micro-zone prediction, or Telegram/Grafana
Sasongko et al. [7] Technology adoption Extended TAM for e-learning acceptance Usability and user acceptance in digital services No climate component or sensor infrastructure
Tari et al. [8] Cybersecurity resilience Human-centric cybersecurity for digital ecosystems Resilience perspective for distributed smart services No environmental sensing or climate forecasting
Ibrahim et al. [9] Decision support / VR Effect of VR on decision-making and task load Usable interfaces for operational decision support VR-focused; not IoT-based climate intelligence
Khan et al. [10] AI + blockchain / smart buildings AI/blockchain for efficiency and security Smart-environment infrastructures requiring trust Conceptual review; no hyperlocal weather prediction
Narayan et al. [11] AI governance Misinformation in LLM-generated reviews Trust and governance in AI-mediated environments Authenticity detection; not environmental prediction
Naskinova et al. [42] ML comparative study ML models for complex prediction tasks Contextualises R 2 for constrained problems Physiological data; not IoT climate sensing
Luo et al. [43] Multimodal detection Cross-modal temporal alignment Parallel to session-based multi-station alignment Live-streaming domain; not environmental IoT
Karlović et al. [41] Context-aware LLM LLM with contextual signals for personalized outputs Informs LLM-based climate summaries in v2.0 Tourism domain; not climate prediction
This work Hyperlocal climate intelligence End-to-end: IoT sensing, ML, Flask API, n8n, PostgreSQL, Telegram, Grafana Integrates all layers in a single system for Bogotá micro-zones Combines low-cost hardware, ML, orchestration, and operational dissemination
Table 3. Data Acquisition Infrastructure and Historical Dataset.
Table 3. Data Acquisition Infrastructure and Historical Dataset.
(A) BME280 Sensor Specifications
Parameter Range Accuracy Resolution
Temperature −40 to +85°C ±1.0°C 0.01°C
Pressure 300–1100 hPa ±1.0 hPa 0.18 Pa
Humidity 0–100% ±3% 0.008%
Interface I2C (400 kHz) / SPI
Supply 1.71–3.6 V
(B) Historical Data Distribution by Station
Station Raw / Valid Period Mean T
Norte 104,230 / 96,990 May 23–Dec 24 16.4°C
Sur 71,450 / 68,072 Mar 23–Dec 24 15.9°C
Este 95,848 / 17,983 Dec 24 (9 d) 17.8°C
Oeste 38,920 / 33,904 Dec 24 (3 wk) 16.7°C
Total 310,448 / 216,949 Mar 2023 – Dec 2024
Table 4. Data Processing Pipeline: Quality Control and Feature Engineering.
Table 4. Data Processing Pipeline: Quality Control and Feature Engineering.
(A) Quality Control Filters
Filter Criterion Removed % Raw
Device ID device ≠ expected ID 7,240 2.33%
Temperature T < 5 °C ∨ T > 27 °C 4,985 1.61%
Integrity NULL∨ invalid t 1,274 0.41%
Total 13,499 4.35%
(B) Feature Engineering Categories ( | x | = 43 )
Category Formulation Applied to Count
Current V i T , P , H 3
Lag L k [ V i ] = V i k T , P , H 21
Delta D w [ V i ] = V i V i w T , P , H 6
Rolling Equation (3) T , P , H 6
Temporal Equation (5) + daytime + h 6
Station ID Categorical { 1 , , 4 } 1
Total 43
Table 5. Random Forest Hyperparameters and Justifications.
Table 5. Random Forest Hyperparameters and Justifications.
Parameter Value Justification
B (n_estimators) 300 Variance reduction saturates at O ( 1 / B ) ; empirical plateau at B 250
max_depth 20 Balances expressiveness vs. overfitting; limits tree complexity to 2 20 leaves
min_samples_leaf 10 Regularisation; prevents splits on <50 s of data
max_features 43 = 6 Random subspace decorrelation [12]
random_state 42 Reproducibility seed
n_jobs −1 Parallel training on all CPU cores
Train / Val split 85% / 15% Temporal; 178,691 / 31,534 samples
Model size 269 MB Serialised via joblib
Table 6. Flask REST API Endpoints.
Table 6. Flask REST API Endpoints.
Method Endpoint Input Output
POST /predecir Last 400 readings (JSON) Predicted temp + confidence
POST /guardar Prediction result (JSON) Confirmation + record ID
GET /predicciones Optional: limit, station Historical predictions list
POST /webhook Telegram update (JSON) Bot command response
GET /salud System health status
Table 7. System Integration Layer: Orchestration, Messaging, and Persistence.
Table 7. System Integration Layer: Orchestration, Messaging, and Persistence.
(A) n8n Orchestration Nodes
Node Type Function
Schedule Trigger Trigger Fires every 5 min
Read CSV ×4 File Read Loads last 400 rows/station
Predict ×4 HTTP POST Calls /predecir in parallel
Consolidate Code (JS) Merges predictions; fault-tolerant
Save to DB HTTP POST /guardar→ PostgreSQL
Send Telegram Code (JS) Delivers to all subscribers
(B) Telegram Bot Commands
Command Function
/start Subscribe to automatic 5-min predictions
/stop Unsubscribe from automatic predictions
/predecir Request immediate on-demand prediction
(C) PostgreSQL Schema (Render Cloud)
Column Type Constraint / Description
id SERIAL PK, NOT NULL
timestamp TIMESTAMPTZ NOT NULL; prediction UTC time
temp_zona FLOAT NOT NULL; zone °C
confianza FLOAT NOT NULL; Equation (8)
temp_{N,S,E,O} FLOAT NULLABLE; per-station °C
estaciones_activas INTEGER NOT NULL; count [ 0 , 4 ]
Table 8. System Components of ClimaBogotá v1.2.
Table 8. System Components of ClimaBogotá v1.2.
Layer Technology Role Cost
Sensors Raspberry Pi 4B + BME280 Data acquisition (5 s) ∼50 USD/station
Local DB MariaDB 10.6 Local reading storage Free
Processing Python + scikit-learn Cleaning, features, ML Free
API/Server Flask 2.3 Real-time inference Free
Orchestration n8n (self-hosted) 5-min pipeline Free
Cloud DB PostgreSQL (Render) Prediction history Free (1 GB)
Notifications Telegram Bot API Subscriber distribution Free
Visualisation Grafana Cloud Real-time dashboard Free
HTTP Tunnel ngrok 3.x Exposes Flask to Telegram Free
Total hardware 160–240 USD (one-time)
Total cloud 0 USD/month
Table 9. Model Performance Metrics on the Temporal Validation Set (31,534 Samples).
Table 9. Model Performance Metrics on the Temporal Validation Set (31,534 Samples).
Metric Value Interpretation
MAE 2.59°C Average prediction error
RMSE 4.46°C Penalises large errors more
R 2 0.6286 62.9% of variability explained
MAPE 16.2% Average percentage error
Median AE 1.87°C 50% of predictions below this error
Table 10. Prediction Cycle Latency Breakdown.
Table 10. Prediction Cycle Latency Breakdown.
Stage Latency
CSV reading (4 stations) 2–3 s
Flask inference calls (4 parallel) 8–12 s
Consolidation and formatting <1 s
PostgreSQL persistence 3–5 s
Telegram delivery 5–10 s
Total cycle latency 30–45 s
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings