Submitted:
11 July 2026
Posted:
14 July 2026
You are already at the latest version
Abstract

Keywords:
1. Introduction
- A focused review of recent ML-based Kubernetes autoscaling studies.
- Comparative analysis across HPA, VPA, CA, and KEDA.
- Identification of dominant ML techniques and architectural trends.
- Discussion of open research challenges and future directions.
- Consolidation of emerging developments recently publish.
2. Background of Kubernetes Autoscaling
2.1. Horizontal Pod Autoscaler
2.2. Vertical Pod Autoscaler
2.3. Cluster Autoscaler
2.4. Kubernetes Event-Driven Autoscaling
3. Research Methodology
4. Machine Learning-Based Horizontal Pod Autoscaler
4.1. Deep Learning and Time-Series Forecasting
4.2. Comparative Analysis for HPA
5. Machine Learning-Based Vertical Pod Autoscaler
5.1. Intelligent Vertical Scaling
5.2. Coordinated Vertical and Horizontal Scaling
5.3. Comparative Analysis for VPA
6. Machine Learning-Based Cluster Autoscaler
6.1. Intelligent Cluster Provisioning
6.2. Edge and Distributed Resource Management
6.3. Comparative Analysis for Cluster Autoscaler
7. Machine Learning-Based KEDA and Event-Driven Autoscaling
7.1. Predictive Event-Driven Scaling
7.2. GPU-Aware and AI Inference Scaling
7.3. Comparative Analysis for KEDA
8. Cross-Category Comparative Analysis
9. Open Research Challenges
9.1. Cross-Workload Generalization
9.2. Explainability and Trustworthiness
9.3. Multi-Dimensional Scaling Coordination
9.4. Edge and AI Workload Constraints
10. Conclusions
References
- Kubernetes Documentation, “Horizontal Pod Autoscaling,” 2024.
- Kubernetes Documentation, “Vertical Pod Autoscaling,” 2026.
- Kubernetes Documentation, “Autoscaling Workloads,” 2025.
- KEDA, “Kubernetes Event-Driven Autoscaling,” 2025. [CrossRef] [PubMed]
- Kubernetes Blog, “In-Place Pod Resize Graduates to Stable,” 2025. [CrossRef] [PubMed]
- T. Guruge and B. Priyadarshana, “Time Series Forecasting-Based Kubernetes Autoscaling Using Prophet and LSTM,” Frontiers in Computer Science, 2025.
- ByteDance Research, “E3Former: Online Ensemble Transformer for Accurate Cloud Workload Forecasting,” arXiv preprint, 2025.
- Zhang et al., “Cloud Resource Auto-Scaling Strategy Based on CNN-Lightweight,” 2025. [CrossRef] [PubMed]
- Manzano et al., “Mitigating Temporal Blindness in Kubernetes Autoscaling,” arXiv preprint, 2026.
- “MAS-H2: A Hierarchical Multi-Agent System for Holistic Cloud-Native Autoscaling,” arXiv preprint, 2026.
- “AGMARL-DKS: Adaptive Graph-Enhanced Multi-Agent Reinforcement Learning for Kubernetes,” arXiv preprint, 2026.
- Vestergaard et al., “Comparing Neural and Statistical Time-Series Models for Proactive Autoscaling,” 2025. [CrossRef] [PubMed]
- Punniyamoorthy et al., “An SLO-Driven and Cost-Aware Autoscaling Framework for Kubernetes,” arXiv preprint, 2025.
- “STAR: Spatial-Temporal Autoscaling for Cloud Applications with Deep Reinforcement Learning,” 2026.
- “MARLISE: Multi-Agent Reinforcement Learning-Based In-Place Scaling Engine,” arXiv preprint, 2025.
- “A Scalable Machine Learning Strategy for Resource Allocation in Cloud-Native Environments,” Nature Scientific Reports, 2025. [CrossRef] [PubMed]
- “Understanding Kubernetes-Based Adaptive Cost Optimization for Vertical Pod Autoscaling,” 2025. [CrossRef] [PubMed]
- Akamas, “Kubernetes VPA Autoscaling: Effect on Service Cost and Efficiency,” 2023.
- CNCF Blog, “Optimizing Kubernetes Vertical Pod Autoscaler Responsiveness,” 2023. [CrossRef] [PubMed]
- “Towards Intelligent Container Orchestration in Cloud Computing,” 2026.
- “HGraphScale: Hierarchical Graph Learning for Autoscaling in Kubernetes,” IEEE Transactions on Services Computing, 2025.
- “Immersive Intelligence: ST-GNN Guided and RL-Optimized Orchestration,” 2026.
- “Towards a Proactive Autoscaling Framework for Data Stream Processing at the Edge,” arXiv preprint, 2025.
- “Deep Reinforcement Learning for Resource Management in IoT-Edge-Cloud Environments,” 2026.
- “Reinforcement Learning-Driven Kubernetes Autoscaling for High-Traffic 5G Network Functions,” 2025. [CrossRef] [PubMed]
- “A Kubernetes Custom Scheduler Based on Reinforcement Learning for Compute-Intensive Pods,” 2026.
- “Enhancing Kubernetes Resilience through Anomaly Detection and Predictive Scaling,” arXiv preprint, 2025.
- “Toward Context-Aware Anomaly Detection for AIOps in Kubernetes,” 2026. [CrossRef] [PubMed]
- “Predictive Autoscaler for Kubernetes with Kubernetes Event-Driven Autoscaling,” 2024.
- “Predictive Autoscaling in Kubernetes Microservices with KEDA and Deep Neural Networks,” 2026.
- PredictKube, “AI-Based Predictive Autoscaling for KEDA,” 2023.
- “Event-Driven Scaling for Machine Learning: KEDA and KServe,” 2026.
- Flightcrew, “Scaling Kubernetes with KEDA and RabbitMQ,” 2025. [CrossRef] [PubMed]




| Ref. | Authors / Year | ML Technique | Environment | Main Contribution |
|---|---|---|---|---|
| [6] | Guruge and Priyadarshana (2025) | Prophet + LSTM | Kubernetes workloads | Hybrid forecasting improves proactive scaling accuracy. |
| [7] | ByteDance E3Former (2025) | Ensemble Transformer | Production microservices | Transformer-based workload prediction for latency-aware autoscaling. |
| [8] | Zhang et al. (2025) | CNN-Lightweight | Cloud workloads | Adaptive cooling and scaling stabilization. |
| [9] | Manzano et al. (2026) | Bi-LSTM + PPO | Azure/Kubernetes | Reinforcement learning for latency-aware HPA. |
| [10] | MAS-H2 (2026) | Hierarchical MARL | GKE testbed | Multi-agent coordination between HPA and VPA. |
| [11] | AGMARL-DKS (2026) | Graph-enhanced MARL | Large Kubernetes clusters | Decentralized autoscaling decisions. |
| [12] | Vestergaard et al. (2025) | Comparative time-series models | FinTech workloads | Comparison of neural and statistical forecasting for proactive scaling. |
| [13] | Punniyamoorthy et al. (2025) | SLO-aware AIOps | Kubernetes microservices | Cost-aware and SLO-driven autoscaling. |
| [14] | STAR (2026) | Spatial-temporal DRL | Cloud application traces | DRL optimization for workload dynamics. |
| Ref. | Authors / Year | ML Technique | Optimization Target | Main Contribution |
|---|---|---|---|---|
| [15] | MARLISE (2025) | Multi-agent RL | In-place CPU/memory resize | Stateful workload vertical scaling without restart. |
| [16] | LSTM-MARL-Ape-X (2025) | BiLSTM + MARL | Distributed resource allocation | SLA-aware and energy-efficient scaling. |
| [17] | AI-VPA Study (2025) | Deep learning forecasting | Resource right-sizing | Long-horizon workload prediction. |
| [18] | Akamas (2023) | Bayesian optimization | JVM/container tuning | Cost-aware VPA optimization. |
| [19] | CNCF VPA Optimization (2023) | Heuristic + ML tuning | Recommender responsiveness | Improved VPA responsiveness. |
| [10] | MAS-H2 (2026) | Hierarchical MARL | HPA/VPA coordination | Unified autoscaling conflict resolution. |
| [20] | IJERT ICO (2026) | LSTM + CNN | Container provisioning | Deep learning-based vertical scaling. |
| Ref. | Authors / Year | ML Technique | Cluster Target | Main Contribution |
|---|---|---|---|---|
| [21] | HGraphScale (2025) | CHGNN + DRL | Cluster and pod scaling | Hierarchical graph learning for autoscaling. |
| [22] | ST-GNN Framework (2026) | ST-GNN + RL | Intent-aware orchestration | Unified spatio-temporal orchestration. |
| [11] | AGMARL-DKS (2026) | Graph-enhanced MARL | Large Kubernetes clusters | Decentralized scalable autoscaling. |
| [14] | STAR (2026) | Spatial-temporal DRL | Cloud application clusters | Captures temporal and spatial workload patterns. |
| [23] | GRU + Transfer Learning (2025) | GRU + DTW transfer | Edge cluster provisioning | Domain adaptation for proactive scaling. |
| [24] | Melbourne IoT-Edge MARL (2026) | Multi-agent RL | Edge-cloud continuum | Region-aware distributed provisioning. |
| [25] | WJAETS 5G DRL (2025) | Deep RL | 5G node slicing | Robust autoscaling under traffic variation. |
| [26] | RL Kubernetes Scheduler (2026) | RL scheduling | Compute-intensive workloads | Improved resource utilization. |
| [27] | Kubernetes Anomaly Detection (2025) | LSTM + Autoencoder | Fault prediction | Predictive anomaly-aware autoscaling. |
| [28] | IEEE AIOps Transformer (2026) | Masked Transformer | Container anomaly detection | Reduction of false scaling events. |
| Ref. | Authors / Year | ML Technique | Trigger Source | Main Contribution |
|---|---|---|---|---|
| [29] | IIT KEDA Study (2024) | Bi-GRU, GRU, LSTM | Multi-source KEDA metrics | Bi-GRU provides strong prediction accuracy. |
| [30] | Jisem-Journal (2026) | DNN + Prophet | CPU, memory, network, storage | Multivariate predictive autoscaling. |
| [31] | PredictKube (2023) | AI time-series forecasting | Prometheus metrics | Forecast-based proactive KEDA scaling. |
| [32] | KServe + KEDA (2026) | Demand prediction | LLM inference queues | Scale-from-zero GPU inference workloads. |
| [33] | Flightcrew RabbitMQ Study (2025) | Queue-depth scaling | RabbitMQ workloads | Event-driven queue-aware scaling. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).