Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

2026-07-10Networking and Internet Architecture

Networking and Internet Architecture
AI summary

The authors study a system where drones help process data quickly for different types of network services that need reliable, fast, or large-scale communication. They create a smart method using multiple AI agents that predict where users will move soon, so drones can adjust their paths and computer power to keep service promises. Their method trains drones to avoid service problems and save energy at the same time. Tests show their approach works better than others and gets close to perfect if predictions are accurate.

Unmanned Aerial Vehicle (UAV)Mobile Edge Computing (MEC)Network SlicingService-Level Agreement (SLA)Reinforcement LearningMulti-Agent Proximal Policy Optimization (MAPPO)User Mobility PredictionResource AllocationLatencyEnergy Efficiency
Authors
Mohammad Farhoudi, Zeinab Sasan, Masoud Shokrnezhad, Tarik Taleb
Abstract
Unmanned Aerial Vehicle (UAV)-enabled Mobile Edge Computing (MEC) offers flexible capacity provisioning for heterogeneous network slices, including Hyper-Reliable and Low-Latency Communication (HRLLC), Enhanced Mobile Broadband (eMBB), and Massive Machine-Type Communications (mMTC). However, guaranteeing slice-level Service-Level Agreements (SLAs) under dynamic user mobility, stochastic task arrivals, and constrained onboard energy and computing resources remains a fundamental challenge. This paper proposes a predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation. A lightweight prediction module forecasts near-future user mobility, enabling UAVs to anticipate congestion and reposition before SLA violations occur. We design an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices, alongside total energy consumption. UAV agents are trained using Multi-Agent Proximal Policy Optimization (MAPPO) with centralized training and decentralized execution, enabling scalable online decision-making. Event-driven simulations with realistic mobility traces demonstrate that the proposed framework significantly improves SLA stability compared with baselines while maintaining competitive energy efficiency and delay performance, approaching oracle-level performance with sufficiently accurate predictive information.