Submitted:
04 August 2026
Posted:
04 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Base Forecasting Models
2.2. Forecast Refinement Without Replacing Base Models
2.3. Objective-Level Forecasting Enhancements
2.4. Conditional Weighting and Soft Mixing
3. Methods
3.1. Problem Formulation and Plug-In Forecast Refinement
3.2. Model Details and Component Forecast Construction
| No. | Summary | Expression |
| 1 | Full mean | |
| 2 | Full standard deviation | |
| 3 | Late-window mean | |
| 4 | Early-window mean | |
| 5 | Late-window standard deviation | |
| 6 | Difference mean | |
| 7 | Difference standard deviation | |
| 8 | Coarse residual magnitude | |
| 9 | Last-to-recent-mean deviation | |
| 10 | Curvature magnitude |
| No. | Reference forecast | Definition |
| 1 | Last-value | |
| 2 | Trend | |
| 3 | Repeated-period | |
| 4 | Smoothed-trend |
3.3. Block-Wise Soft Weighting
3.4. Training and Inference Details
4. Experimental Results
4.1. Experimental Setup
4.1.1. Datasets
4.1.2. Base Forecasting Models and Paired Comparison Protocol
4.1.3. Training and Evaluation Details
4.2. Main Forecasting Results
4.3. Component Analysis
5. Discussion
6. Conclusions
Funding
References
- Lim, B.; Zohren, S. Time-series forecasting with deep learning: a survey. Philos. Trans. R. Soc. A 2021, 379, 20200209. [Google Scholar] [CrossRef] [PubMed]
- Benidis, K.; Rangapuram, S.S.; Flunkert, V.; Wang, Y.; Maddix, D.; Turkmen, C.; Gasthaus, J.; Bohlke-Schneider, M.; Salinas, D.; Stella, L.; et al. Deep Learning for Time Series Forecasting: Tutorial and Literature Survey. ACM Comput. Surv. 2022, 55. [Google Scholar] [CrossRef]
- Lai, G.; Chang, W.C.; Yang, Y.; Liu, H. Modeling long-and short-term temporal patterns with deep neural networks. In Proceedings of the The 41st international ACM SIGIR conference on research & development in information retrieval, 2018; pp. 95–104. [Google Scholar]
- Li, Y.; Yu, R.; Shahabi, C.; Liu, Y. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In Proceedings of the International Conference on Learning Representations, 2018. [Google Scholar]
- Yu, B.; Yin, H.; Zhu, Z. Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, 2018; pp. 3634–3640. [Google Scholar] [CrossRef] [PubMed]
- Zhao, Y. Research on Data Model Fusion Driven Dual End Uncertainty Source Load Probability Prediction and Its Application in New Energy Power Systems. IEIE Trans. Smart Process. Comput. 2025, 14, 389–399. [Google Scholar] [CrossRef]
- Zhang, B.; Tao, M. Design and Implementation of a Multi-factor Intelligent Mining System for Stocks Based on GA-TGCN. IEIE Trans. Smart Process. Comput. 2025, 14, 178–190. [Google Scholar] [CrossRef]
- Zhu, X. Application of Artificial Intelligence in the Problem of Agricultural Cold Chain Logistics Center Site Selection. IEIE Trans. Smart Process. Comput. 2025, 14, 535–544. [Google Scholar] [CrossRef]
- Chen, Y.; Liang, X.; Qi, P.; Xia, S.; Tong, J. Multimodal heart failure prediction model based on graph convolutional neural network. Biomed. Eng. Lett. 2026. [Google Scholar] [CrossRef]
- Fuadah, N.; Lim, K.M. Advances in cardiovascular signal analysis with future directions: a review of machine learning and deep learning models for cardiovascular disease classification based on ECG, PCG, and PPG signals. Biomed. Eng. Lett. 2025, 15, 619–660. [Google Scholar] [CrossRef]
- Salinas, D.; Flunkert, V.; Gasthaus, J.; Januschowski, T. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks. Int. J. Forecast. 2020, 36, 1181–1191. [Google Scholar] [CrossRef]
- Lim, B.; Arik, S.O.; Loeff, N.; Pfister, T. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting. Int. J. Forecast. 2021, 37, 1748–1764. [Google Scholar] [CrossRef]
- Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. Proc. Proc. AAAI Conf. Artif. Intell. 2021, Vol. 35, 11106–11115. [Google Scholar] [CrossRef]
- Nie, Y.; Nguyen, N.H.; Sinthong, P.; Kalagnanam, J. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In Proceedings of the International Conference on Learning Representations, 2023. [Google Scholar]
- Liu, Y.; Hu, T.; Zhang, H.; Wu, H.; Wang, S.; Ma, L.; Long, M. iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. In Proceedings of the International Conference on Learning Representations, 2024. [Google Scholar]
- Kim, T.; Kim, J.; Tae, Y.; Park, C.; Choi, J.H.; Choo, J. Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift. In Proceedings of the International Conference on Learning Representations, 2022. [Google Scholar]
- Liu, Y.; Wu, H.; Wang, J.; Long, M. Non-stationary Transformers: Exploring the Stationarity in Time Series Forecasting. Proc. Adv. Neural Inf. Process. Syst. 2022, Vol. 35, 9881–9893. [Google Scholar] [CrossRef]
- Chen, S.A.; Li, C.L.; Arik, S.O.; Yoder, N.C.; Pfister, T. TSMixer: An All-MLP Architecture for Time Series Forecasting. Transactions on Machine Learning Research, 2023. [Google Scholar]
- Das, A.; Kong, W.; Leach, A.; Mathur, S.K.; Sen, R.; Yu, R. Long-term Forecasting with TiDE: Time-series Dense Encoder. Transactions on Machine Learning Research, 2023. [Google Scholar]
- Wang, S.; Wu, H.; Shi, X.; Hu, T.; Luo, H.; Ma, L.; Zhang, J.; ZHOU, J. TimeMixer: Decomposable Multiscale Mixing for Time Series Forecasting. In Proceedings of the International Conference on Learning Representations, 2024. [Google Scholar]
- Bai, S.; Kolter, J.Z.; Koltun, V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv 2018, arXiv:1803.01271. [Google Scholar]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. Proc. Adv. Neural Inf. Process. Syst. 2021, Vol. 34, 22419–22430. [Google Scholar]
- Zhou, T.; Ma, Z.; Wen, Q.; Wang, X.; Sun, L.; Jin, R. Proceedings of the Proceedings of the 39th International Conference on Machine Learning. PMLR, 17–23 Jul 2022, Vol. 162, Proceedings of Machine Learning Research. 27268–27286.
- Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In Proceedings of the International Conference on Learning Representations, 2023. [Google Scholar]
- Liu, S.; Yu, H.; Liao, C.; Li, J.; Lin, W.; Liu, A.X.; Dustdar, S. Pyraformer: Low-Complexity Pyramidal Attention for Long-Range Time Series Modeling and Forecasting. In Proceedings of the International Conference on Learning Representations, 2022. [Google Scholar]
- Zhang, Y.; Yan, J. Crossformer: Transformer Utilizing Cross-Dimension Dependency for Multivariate Time Series Forecasting. In Proceedings of the International Conference on Learning Representations, 2023. [Google Scholar]
- Zeng, A.; Chen, M.; Zhang, L.; Xu, Q. Are transformers effective for time series forecasting? Proc. Proc. AAAI Conf. Artif. Intell. 2023, Vol. 37, 11121–11128. [Google Scholar] [CrossRef]
- Oreshkin, B.N.; Carpov, D.; Chapados, N.; Bengio, Y. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. In Proceedings of the International Conference on Learning Representations, 2020. [Google Scholar]
- Liu, M.; Zeng, A.; Chen, M.; Xu, Z.; Lai, Q.; Ma, L.; Xu, Q. SCINet: Time Series Modeling and Forecasting with Sample Convolution and Interaction. Proc. Adv. Neural Inf. Process. Syst. 2022, Vol. 35, 5816–5828. [Google Scholar] [CrossRef]
- Luo, D.; Wang, X. DecompNet: Enhancing Time Series Forecasting Models with Implicit Decomposition. Proc. Adv. Neural Inf. Process. Syst. 2025, Vol. 38, 32278–32317. [Google Scholar]
- Du, Y.; Wang, J.; Feng, W.; Pan, S.; Qin, T.; Xu, R.; Wang, C. AdaRNN: Adaptive Learning and Forecasting of Time Series. Proceedings of the 30th ACM International Conference on Information and Knowledge Management. Association for Computing Machinery 2021, CIKM ’21, 402–411. [Google Scholar] [CrossRef]
- Liu, Z.; Cheng, M.; Zhao, G.; Yang, J.; Liu, Q.; Chen, E. Improving Time Series Forecasting via Instance-aware Post-hoc Revision. Proc. Adv. Neural Inf. Process. Syst. 2025, Vol. 38, 35604–35629. [Google Scholar]
- Cuturi, M.; Blondel, M. Proceedings of the Proceedings of the 34th International Conference on Machine Learning. PMLR, 2017, Vol. 70, Machine Learning Research. 894–903.
- Le Guen, V.; Thome, N. Shape and Time Distortion Loss for Training Deep Time Series Forecasting Models. In Proceedings of the Advances in Neural Information Processing Systems, 2019; Curran Associates, Inc.; Vol. 32. [Google Scholar]
- Qiu, X.; Wu, X.; Cheng, H.; Liu, X.; Guo, C.; Hu, J.; Yang, B. DBLoss: Decomposition-based Loss Function for Time Series Forecasting. Proc. Adv. Neural Inf. Process. Syst. 2025, Vol. 38, 27741–27768. [Google Scholar]
- Wang, H.; Pan, L.; Chen, Z.; Chen, X.; Dai, Q.; Wang, L.; Li, H.; Lin, Z. Time-o1: Time-Series Forecasting Needs Transformed Label Alignment. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc., 2025; Vol. 38, pp. 8660–8690. [Google Scholar]
- Park, K.; Kim, J.; Lee, S. AliO: Output Alignment Matters in Long-Term Time Series Forecasting. Proc. Adv. Neural Inf. Process. Syst. 2025, Vol. 38, 119563–119608. [Google Scholar]
- Jacobs, R.A.; Jordan, M.I.; Nowlan, S.J.; Hinton, G.E. Adaptive Mixtures of Local Experts. Neural Comput. 1991, 3, 79–87. [Google Scholar] [CrossRef] [PubMed]
- Jordan, M.I.; Jacobs, R.A. Hierarchical Mixtures of Experts and the EM Algorithm. Neural Comput. 1994, 6, 181–214. [Google Scholar] [CrossRef]
- Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.V.; Hinton, G.E.; Dean, J. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In Proceedings of the International Conference on Learning Representations, 2017. [Google Scholar]
- Chen, P.; Zhang, Y.; Cheng, Y.; Shu, Y.; Wang, Y.; Wen, Q.; Yang, B.; Guo, C. Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series Forecasting. In Proceedings of the International Conference on Learning Representations, 2024. [Google Scholar]
- Wu, X.; Qiu, X.; Cheng, H.; Li, Z.; Hu, J.; Guo, C.; Yang, B. Enhancing Time Series Forecasting through Selective Representation Spaces: A Patch Perspective. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc., 2025; Vol. 38, pp. 23328–23354. [Google Scholar]
- Ba, J.L.; Kiros, J.R.; Hinton, G.E. Layer normalization. arXiv 2016, arXiv:1607.06450. [Google Scholar]
- Hendrycks, D.; Gimpel, K. Gaussian error linear units (gelus). arXiv 2016, arXiv:1606.08415. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations, 2019. [Google Scholar]



| Module | Configuration |
|---|---|
| Gap coefficient MLP |
Input: descriptor . Architecture: . Activation: GELU after the hidden layer. Output: bounded coefficients after tanh. |
| Direct residual MLP |
Input: descriptor . Architecture: . Activation: GELU after the hidden layer. Output: reshaped to . |
| Block-wise weighting MLP |
Input: descriptor . Architecture: . Activation: GELU after the hidden layer. Output: block-wise component weights. |
| Item | Setting |
|---|---|
| Objective | MSE between and |
| Optimizer | AdamW |
| Initial learning rate | |
| Gradient clipping | |
| Trainable parameters | Base model and GARS parameters are trained jointly |
| Frozen parameters | None |
| Auxiliary losses | None |
| Residual damping | |
| Checkpoint selection | Best validation-MSE checkpoint |
| Dataset | Sampling interval | Length | Variables | Description |
|---|---|---|---|---|
| Solar | 10 min | 52,560 | 137 | solar power production records from photovoltaic plants |
| Weather | 10 min | 52,696 | 21 | meteorological measurements collected throughout 2020 |
| Electricity | 1 hour | 26,304 | 321 | electricity consumption records from multiple clients |
| Traffic | 1 hour | 17,544 | 862 | road occupancy rates measured by freeway sensors |
| Exchange | 1 day | 7,588 | 8 | daily exchange rates from multiple countries |
| Model | iTransformer [15] | PatchTST [14] | TimeMixer [20] | TimesNet [24] | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | Original | GARS | Original | GARS | Original | GARS | Original | GARS | |||||||||
| Metric | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | |
| Solar | 96 | 0.280 | 0.313 | 0.231 | 0.301 | 0.269 | 0.302 | 0.226 | 0.292 | 0.282 | 0.312 | 0.229 | 0.300 | 0.297 | 0.315 | 0.225 | 0.304 |
| 192 | 0.325 | 0.340 | 0.261 | 0.328 | 0.318 | 0.330 | 0.260 | 0.324 | 0.340 | 0.346 | 0.254 | 0.319 | 0.346 | 0.344 | 0.237 | 0.307 | |
| 336 | 0.357 | 0.356 | 0.273 | 0.341 | 0.345 | 0.341 | 0.262 | 0.331 | 0.369 | 0.355 | 0.270 | 0.334 | 0.361 | 0.350 | 0.261 | 0.333 | |
| 720 | 0.376 | 0.364 | 0.278 | 0.342 | 0.354 | 0.346 | 0.282 | 0.342 | 0.386 | 0.362 | 0.279 | 0.343 | 0.392 | 0.377 | 0.411 | 0.440 | |
| Avg | 0.334 | 0.343 | 0.260 | 0.328 | 0.322 | 0.330 | 0.258 | 0.322 | 0.344 | 0.344 | 0.258 | 0.324 | 0.349 | 0.346 | 0.284 | 0.346 | |
| Weather | 96 | 0.199 | 0.236 | 0.177 | 0.235 | 0.200 | 0.237 | 0.177 | 0.235 | 0.206 | 0.240 | 0.179 | 0.237 | 0.178 | 0.227 | 0.173 | 0.237 |
| 192 | 0.242 | 0.273 | 0.222 | 0.277 | 0.243 | 0.272 | 0.226 | 0.283 | 0.249 | 0.276 | 0.227 | 0.283 | 0.240 | 0.277 | 0.224 | 0.282 | |
| 336 | 0.262 | 0.296 | 0.249 | 0.312 | 0.258 | 0.293 | 0.252 | 0.318 | 0.262 | 0.296 | 0.255 | 0.319 | 0.255 | 0.294 | 0.246 | 0.317 | |
| 720 | 0.338 | 0.344 | 0.318 | 0.370 | 0.332 | 0.339 | 0.316 | 0.368 | 0.337 | 0.342 | 0.318 | 0.367 | 0.332 | 0.344 | 0.316 | 0.360 | |
| Avg | 0.260 | 0.287 | 0.241 | 0.298 | 0.258 | 0.286 | 0.243 | 0.301 | 0.264 | 0.289 | 0.245 | 0.302 | 0.251 | 0.286 | 0.240 | 0.299 | |
| Electricity | 96 | 0.187 | 0.274 | 0.181 | 0.275 | 0.199 | 0.283 | 0.190 | 0.281 | 0.216 | 0.301 | 0.205 | 0.296 | 0.252 | 0.333 | 0.221 | 0.315 |
| 192 | 0.190 | 0.284 | 0.186 | 0.285 | 0.197 | 0.288 | 0.190 | 0.288 | 0.210 | 0.305 | 0.202 | 0.301 | 0.249 | 0.339 | 0.222 | 0.324 | |
| 336 | 0.213 | 0.304 | 0.210 | 0.305 | 0.215 | 0.305 | 0.210 | 0.304 | 0.229 | 0.321 | 0.223 | 0.318 | 0.262 | 0.350 | 0.247 | 0.340 | |
| 720 | 0.258 | 0.338 | 0.254 | 0.337 | 0.256 | 0.335 | 0.250 | 0.332 | 0.270 | 0.349 | 0.264 | 0.347 | 0.303 | 0.380 | 0.279 | 0.362 | |
| Avg | 0.212 | 0.300 | 0.208 | 0.300 | 0.217 | 0.303 | 0.210 | 0.301 | 0.231 | 0.319 | 0.223 | 0.316 | 0.266 | 0.350 | 0.242 | 0.335 | |
| Traffic | 96 | 0.492 | 0.345 | 0.492 | 0.345 | 0.611 | 0.384 | 0.609 | 0.384 | 0.686 | 0.413 | 0.662 | 0.411 | 0.658 | 0.359 | 0.672 | 0.368 |
| 192 | 0.495 | 0.344 | 0.495 | 0.344 | 0.578 | 0.365 | 0.578 | 0.365 | 0.624 | 0.390 | 0.619 | 0.390 | 0.686 | 0.375 | 0.729 | 0.396 | |
| 336 | 0.520 | 0.357 | 0.520 | 0.357 | 0.587 | 0.369 | 0.587 | 0.369 | 0.632 | 0.395 | 0.628 | 0.395 | 0.722 | 0.394 | 0.673 | 0.368 | |
| 720 | 0.642 | 0.428 | 0.642 | 0.428 | 0.677 | 0.421 | 0.677 | 0.421 | 0.885 | 0.515 | 0.887 | 0.517 | 0.753 | 0.406 | 0.756 | 0.403 | |
| Avg | 0.537 | 0.369 | 0.537 | 0.369 | 0.613 | 0.385 | 0.613 | 0.385 | 0.707 | 0.428 | 0.699 | 0.428 | 0.705 | 0.384 | 0.707 | 0.384 | |
| Exchange | 96 | 0.097 | 0.218 | 0.095 | 0.222 | 0.092 | 0.209 | 0.088 | 0.211 | 0.087 | 0.203 | 0.084 | 0.209 | 0.118 | 0.246 | 0.090 | 0.218 |
| 192 | 0.187 | 0.310 | 0.177 | 0.303 | 0.188 | 0.305 | 0.188 | 0.310 | 0.184 | 0.302 | 0.186 | 0.309 | 0.216 | 0.335 | 0.192 | 0.316 | |
| 336 | 0.327 | 0.416 | 0.295 | 0.404 | 0.345 | 0.425 | 0.301 | 0.404 | 0.325 | 0.412 | 0.314 | 0.411 | 0.400 | 0.462 | 0.371 | 0.443 | |
| 720 | 0.863 | 0.703 | 0.742 | 0.654 | 0.966 | 0.739 | 0.838 | 0.693 | 0.876 | 0.707 | 0.826 | 0.688 | 1.019 | 0.771 | 1.031 | 0.777 | |
| Avg | 0.369 | 0.412 | 0.327 | 0.396 | 0.398 | 0.419 | 0.354 | 0.404 | 0.368 | 0.406 | 0.353 | 0.404 | 0.438 | 0.454 | 0.421 | 0.439 | |
| Dataset | n | Mean MSE | p (MSE) | Mean MAE | p (MAE) |
|---|---|---|---|---|---|
| Solar | 16 | ||||
| Weather | 16 | ||||
| Electricity | 16 | ||||
| Traffic | 16 | ||||
| Exchange | 16 | ||||
| All | 80 |
| Horizon | MSE | MAE | ||
|---|---|---|---|---|
| Mean | Median | Mean | Median | |
| 96 | ||||
| 192 | ||||
| 336 | ||||
| 720 | ||||
| Variant | Components used | Mean MSE | Mean MAE |
|---|---|---|---|
| Full GARS | Initial + gap-aware + direct | ||
| w/o gap-aware | Initial + direct | ||
| w/o direct | Initial + gap-aware |
| Dataset | Base model | Horizon | Mean MSE | Mean MAE | ||||
|---|---|---|---|---|---|---|---|---|
| Full | w/o gap-aware | w/o direct | Full | w/o gap-aware | w/o direct | |||
| Electricity | iTransformer | 96 | ||||||
| Electricity | iTransformer | 720 | ||||||
| Electricity | PatchTST | 96 | ||||||
| Electricity | PatchTST | 720 | ||||||
| Electricity | TimeMixer | 96 | ||||||
| Electricity | TimeMixer | 720 | ||||||
| Exchange | iTransformer | 96 | ||||||
| Exchange | iTransformer | 720 | ||||||
| Exchange | PatchTST | 96 | ||||||
| Exchange | PatchTST | 720 | ||||||
| Exchange | TimeMixer | 96 | ||||||
| Exchange | TimeMixer | 720 | ||||||
| Weather | iTransformer | 96 | ||||||
| Weather | iTransformer | 720 | ||||||
| Weather | PatchTST | 96 | ||||||
| Weather | PatchTST | 720 | ||||||
| Weather | TimeMixer | 96 | ||||||
| Weather | TimeMixer | 720 | ||||||
| Variant | Active forecasts | Mean MSE | Mean MAE | Dataset-wise mean MSE | |||
|---|---|---|---|---|---|---|---|
| Weather | Electricity | Traffic | Exchange | ||||
| Full | V + T + R + S | ||||||
| w/o trend | V + R + S | ||||||
| w/o repeated | V + T + S | ||||||
| w/o last-value | T + R + S | ||||||
| w/o smoothed | V + T + R | ||||||
| Mean MSE | Median MSE | Mean MAE | |
|---|---|---|---|
| 0.1 | |||
| 0.2 | |||
| 0.3 | |||
| 0.4 | |||
| 0.5 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).