Preprint
Article

This version is not peer-reviewed.

Meta-MARL-ESA: Toward Zero Shot Generalization for Energy Efficient Sleep Scheduling in Dynamic UAV Assisted 6G HetNets

Submitted:

01 October 2026

Posted:

05 October 2026

You are already at the latest version

Abstract
UAV, mounted aerial base stations offer remarkable deployment flexibility for 6G Non Terrestrial Networks (NTNs), yet their hovering energy demands present a critical sustainability bottleneck, particularly as network topologies continuously evolve in practice. As a first step, we develop MARL-ESA (Energy-Saving Adaptation), a Multi Agent Reinforcement Learning (MARL) framework achieving 48% energy savings in static UAV assisted heterogeneous networks (HetNets) compared with always-on operation through centralized training with decentralized execution (CTDE). However, MARL-ESA remains inherently topology specific, requiring over 1500 retraining episodes whenever UAV positions or traffic patterns change. To overcome this limitation, we propose Meta-MARL-ESA, a meta reinforcement learning extension that enables zero shot generalization across unseen network configurations. The problem is formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) over a task distribution τ∼p(T), with a permutation invariant Graph Neural Network (GNN) policy trained via Model Agnostic Meta Learning (MAML). Simulation results demonstrate that Meta-MARL-ESA achieves 12× faster convergence on new topologies, preserves 82% of optimal performance in zero shot deployment, improves energy efficiency (EE) by 21.5% over converged baselines in out of distribution (OOD) scenarios, and recovers to 98% of optimal performance within 8 decision epochs during dynamic UAV repositioning events.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.