Submitted:
02 October 2026
Posted:
06 October 2026
You are already at the latest version
Abstract
Reinforcement learning (RL) has shown promise for automating cloud resource scheduling, yet conventional RL-based schedulers often lack interpretability, violate service-level objectives (SLOs), and introduce substantial deployment risks in Kubernetes environments. In this paper, we propose a hybrid framework that integrates large language models (LLMs), safe reinforcement learning, and a digital twin simulator for distributed cloud resource scheduling. Rather than using the LLM as a direct decision-maker, our approach leverages it to generate scheduling constraints from SLO documents and historical cluster events, adjust reward functions, and produce human-readable explanations. A Proximal Policy Optimization agent with Lagrangian relaxation (PPO-Lagrangian) executes scheduling actions—including pod placement, autoscaling, and resource quota allocation—subject to the LLM-derived constraints. Before deployment, candidate policies are pre-evaluated in a digital twin simulator synchronized with the target cluster, and only policies passing safety verification are applied. We evaluate the framework on Google Online Boutique deployed on a Kubernetes cluster simulator under periodic, burst, and fault-injection workloads. Compared with five baselines, our approach reduces SLO violation rates by 39–68% and lowers normalized resource costs by 14–34% while maintaining competitive response times. Ablation studies confirm the contribution of each component to overall performance.
Keywords:
cloud computing
; safe reinforcement learning
; large language models
; resource scheduling
; digital twin
; Kubernetes
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.