Nested Multi-Agent Reinforcement Learning for Adaptive Resource Management in 6G Network Slicing: A Multi-Timescale Framework with Convergence Guarantees
IEEE Open Journal of the Communications Society, 2026 (ESCI, Scopus)
- Yayın Türü: Makale / Tam Makale
- Basım Tarihi: 2026
- Doi Numarası: 10.1109/ojcoms.2026.3708409
- Dergi Adı: IEEE Open Journal of the Communications Society
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, Compendex, INSPEC, Directory of Open Access Journals
- Anahtar Kelimeler: 6G networks, adaptive resource management, agentic AI, continual learning, multi-agent reinforcement learning, nested learning, network slicing, two-timescale stochastic approximation, zero-touch networks
- İstanbul Gelişim Üniversitesi Adresli: Evet
Özet
Sixth-generation (6G) networks are expected to rely on agentic artificial intelligence for zero-touch, self-managed orchestration of heterogeneous network slices serving enhanced mobile broadband (eMBB), ultra-reliable low-latency communication (URLLC), and massive machine-type communication (mMTC). A central and under-studied challenge for adaptive multi-agent resource management (AMRM) in such settings is multi-timescale non-stationarity: channel fading evolves per time-slot, user demand shifts at the window scale, and service-level agreement (SLA) regimes change at an operational scale. Single-timescale multi-agent reinforcement learning (MARL) algorithms cannot track all three signals cleanly—a learning rate fast enough for the per-slot channel destabilises the coordination structure that governs longer-timescale policies. This paper proposes Nested-MARL, an independent-learner actorcritic algorithm in which each agent’s parameters are partitioned into three groups updated at separated rates α0≪α1≪α2, with a continuum-memory exponential moving average (EMA) anchoring the slowest group. The design is grounded in the Nested Learning paradigm of Behrouz et al. (2025) and is extended here from single-model continual learning to decentralised multi-agent coordination. We establish a finite-time convergence result in the two-timescale stochastic approximation framework showing that under standard regularity and timescale-separation conditions, Nested-MARL achieves O(T−1/2) fast-group convergence vs. an Ω(T−1/3) lower bound for any single-timescale algorithm. An empirical study on a three-agent 6G slicing simulator with continuous multitime-scale drift shows Nested-MARL outperforms independent PPO (IPPO) in mean reward at every drift severity we test (κ∈{0.5, 1.0, 1.5, 2.0}) and by +8.6% in sample efficiency over the first 40 episodes at κ=1.5 (n=10 seeds, p < 0.05). A controlled ablation establishes that stripping timescale separation reduces performance below the IPPO baseline, isolating timescale separation as the causal mechanism. Nested-MARL also reduces policy switching cost by 16.6%, an operationally meaningful benefit for zero-touch orchestration. The complete simulator, agents, and 60+ per-seed training runs are released as open source.