Massive MIMO networks require power-control schemes that jointly account for energy efficiency, inter-cell interference, fairness, and time-varying QoS constraints under mobility and traffic dynamics. This paper proposes a deep reinforcement learning (DRL) framework for downlink power control in a realistic multi-cell massive MIMO setting. We develop Smart-MIMO-Cell-Env, a four-cell downlink environment with 16 antennas and 6 users per cell, correlated Rayleigh channels with 3GPP-inspired large-scale fading, user mobility, time-varying traffic load, channel aging, and optional interference coordination. A centralized agent observes an eight-dimensional state (effective transmit power, aggregate array gain, average SINR and rate, inverse-SINR penalty, mobility factor, traffic load, and channel age) and emits a continuous 88-dimensional action vector of per-user powers and per-antenna gains. A composite reward balances throughput, power consumption, link quality, fairness, and temporal stability. We train and evaluate Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) in Ray RLlib under identical environment, state/action spaces, and network architectures, and assess both final policies over 200 independent episodes. PPO attains a mean episode reward of ?6.130, mean downlink rate 0.264, mean inverse-SINR penalty ? = 1.714, power-efficiency indicator 1.973, mean effective power 1.37 × 10?¹ W, and mean array gain 1.154, whereas SAC settles at ?108.724 reward with ? = 24.032. PPO therefore learns a higher-throughput policy with substantially better link quality at a moderate power level. The results demonstrate the viability of DRL for energy-efficient, fair, and QoS-aware power control in realistic massive MIMO deployments and identify on-policy learning as the better match for the non-stationary dynamics of such networks.
Introduction
The text presents a Deep Reinforcement Learning (DRL) framework for downlink power control in multi-cell massive MIMO networks, focusing on improving performance under realistic and changing wireless conditions.
Problem: Massive MIMO improves spectral and energy efficiency, but multi-cell networks suffer from inter-cell interference, pilot contamination, user mobility, changing traffic, channel aging, and incomplete CSI. Conventional 5G NR power-control methods and static optimization techniques do not explicitly balance energy efficiency, fairness, QoS, and adaptability.
Proposed solution: The study develops Smart-MIMO-Cell-Env, a realistic four-cell massive MIMO simulation environment with:
4 cells
16 antennas per base station
6 users per cell
Correlated Rayleigh fading and 3GPP-inspired path loss
User mobility
Time-varying traffic
Channel aging
Optional interference coordination
DRL algorithms: Two continuous-control algorithms are compared:
PPO (Proximal Policy Optimization)
SAC (Soft Actor-Critic)
Both use the same environment, state/action definitions, neural-network structure, training procedure, and hardware resources to make the comparison fair.
State and actions: The primary state contains eight aggregate features, including effective transmit power, array gain, average SINR, average rate, inverse-SINR penalty, mobility, traffic load, and channel age. The agent controls per-user transmit power and per-antenna gain coefficients.
Reward function: A multi-objective reward balances:
Spectral efficiency/rate
Power consumption
Link quality and SINR
Fairness between cells
Mobility and traffic variations
Channel aging
Temporal stability of power allocation
System model: Each base station uses zero-forcing (ZF) precoding. The model calculates received signals, intra-cell and inter-cell interference, SINR, and achievable data rates for each user.
Results: Under identical training conditions, PPO substantially outperformed SAC. Over 200 evaluation episodes:
PPO achieved approximately −6.13 mean reward, 0.264 mean downlink rate, and 1.714 inverse-SINR penalty.
SAC achieved approximately −108.72 mean reward, 0.047 mean rate, and 24.03 inverse-SINR penalty.
Although SAC operated at lower effective power, its significantly poorer rate and link-quality performance resulted in a much worse overall reward.
Main conclusion:PPO is more robust than SAC for the non-stationary, partially history-dependent massive-MIMO environment studied. Its learned policy provides a better trade-off among throughput, link quality, power consumption, fairness, and stability.
Conclusion
This paper proposed a DRL framework for energy-efficient, fair, and QoS-aware power control in multi-cell massive MIMO networks. We developed Smart-MIMO-Cell-Env, an enhanced four-cell downlink environment with M=16 antennas and K=6 users per cell, correlated Rayleigh channels with 3GPP-inspired large-scale fading, user mobility, time-varying traffic load, channel aging, and optional interference coordination. On top of this environment, we formulated a centralized MDP in which a single controller jointly selects per-UE transmit powers and per-antenna gain coefficients in an 88-dimensional continuous action space. The reward combines throughput, power consumption, link quality, mobility, traffic load, channel aging, fairness, and temporal power stability, allowing a single learning agent to internalize multiple design objectives.
We implemented PPO and SAC in Ray RLlib under the same environment, state and action spaces, network architectures, and software and hardware setup, so that differences in performance primarily reflect algorithmic rather than implementation choices. Training and evaluation on Smart-MIMO-Cell-Env showed that both algorithms learn non-trivial power-control behaviours during training, but PPO consistently achieves a much higher mean episode reward, higher downlink rates, and substantially better link quality than SAC, while operating at a moderate effective power level and exhibiting smooth, stable training dynamics. SAC converges to a low-power regime with low absolute throughput and degraded link quality (R ??0.047 and I ??24.032 in our 200-episode evaluation), despite a higher rate-per-Watt (PE ?5.408).
The observed performance gap is explained by structural differences between the algorithms and their interaction with the environment. PPO is on-policy and relies on large batches of fresh trajectories together with a clipped surrogate objective and generalized advantage estimation. These properties are well matched to the slowly drifting and only approximately Markovian dynamics induced by mobility, traffic variations, and channel aging, as well as to the high-dimensional action space. SAC, in contrast, is off-policy and trains from a replay buffer that mixes samples from many past policies and operating conditions while exploring aggressively through an entropy-regularized objective. In the presence of non-stationary dynamics and history-based reward components, this time-mixing makes SAC more sensitive to critic bias and hyperparameter choices and, in our experiments, leads to a substantially worse evaluation return.
Overall, the proposed framework demonstrates that DRL can serve as a unified mechanism for multi-objective power control in massive MIMO networks, jointly accounting for energy efficiency, interference management, fairness, and robustness to realistic dynamics. Beyond the specific PPO–SAC comparison, the Smart-MIMO-Cell-Env environment, the MDP formulation, and the reward design provide a flexible testbed for future RL-based radio resource management schemes.
Future work includes extending the centralized controller to decentralized and multi-agent architectures, in which per-cell or per-cluster agents coordinate over limited backhaul signalling; incorporating explicit delay and queue-based QoS constraints, for example through Lyapunov or effective-capacity formulations; and integrating more detailed 5G and 6G procedures such as pilot contamination, scheduling, and HARQ feedback. Another important direction is to explore transfer- and meta-learning techniques that adapt trained policies across deployment scenarios, traffic regimes, and hardware constraints, moving closer to practical deployment in next-generation wireless networks.
References
[1] NR; Physical Layer Procedures for Control, 3rd Generation Partnership Project (3GPP), Std. TS 38.213, Release 17.
[2] NR; Physical Layer Procedures for Control (3GPP TS 38.213) Release 18, v18.4.0, 3GPP/ETSI Std., 2024. [Online]. Available: https://www.etsi.org/deliver/etsi_ts/138200_138299/138213/18.04.00_60/ts_138213v180400p.pdf
[3] A. Ghazanfari, H. V. Cheng, E. Björnson, and E. G. Larsson, “A fair and scalable power control scheme in multi-cell massive MIMO,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), 2019, pp. 4499–4503.
[4] Q. Huang et al., “Energy-efficient operation of adaptive massive MIMO for 5G networks,” IEEE Trans. Wireless Commun., 2024.
[5] T. V. Chien, T. C. Nguyen, E. Björnson, and E. G. Larsson, “Power control in cellular massive MIMO with varying user activity: A deep learning solution,” IEEE Trans. Wireless Commun., vol. 19, no. 9, pp. 5732–5748, 2020.
[6] E. Björnson and L. Sanguinetti, “Making cell-free massive MIMO competitive with MMSE processing and centralized implementation,” IEEE Trans. Wireless Commun., vol. 19, no. 1, pp. 77–90, 2020.
[7] R. Han et al., “Effective capacity analysis of delay-sensitive wireless communications with power/rate allocation,” IEEE Trans. Wireless Commun., vol. 23, no. 2, pp. 1534–1549, 2024.
[8] L. Dai, J. Mei, Z. Yang, Z. Tong, C. Zeng, and K. Li, “Lyapunov-guided deep reinforcement learning for delay-aware online task offloading in MEC systems,” J. Syst. Archit., vol. 153, p. 103194, 2024.
[9] Y. S. Nasir and D. Guo, “Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,” IEEE J. Sel. Areas Commun., vol. 37, no. 10, pp. 2239–2250, 2019.
[10] F. B. Mismar, J. Choi, and B. L. Evans, “A framework for automated cellular network tuning with reinforcement learning,” IEEE Trans. Commun., vol. 67, no. 10, pp. 7152–7167, 2019.
[11] Z. Liu et al., “Cell-free XL-MIMO meets multi-agent reinforcement learning,” IEEE Wireless Commun., vol. 31, no. 4, pp. 64–71, 2024.
[12] A. A. Khan et al., “Percentile optimization in wireless networks—Part I: Problem formulation and power control,” IEEE Trans. Signal Process., vol. 72, pp. 2336–2351, 2024.