Recent Advances in RL for Self-adaptive Software Systems: A Systematic Review
Asian Journal of Research in Computer Science · pp. 143–154 · Published 10 Jul 2025
10.9734/ajrcos/2025/v18i7725Abstract
In dynamic environments like cloud computing, the internet of things (IoT), and cyber-physical systems, where conventional rule-based adaptation mechanisms frequently fall short of maintaining optimal performance in the face of uncertainty and change, self-adaptive software systems (SASS) are becoming more and more important. A promising remedy that allows systems to learn and adapt on their own through trial-and-error interactions is Reinforcement Learning (RL). After a thorough screening of 1,248 papers, 68 quantitative studies were chosen for analysis in this systematic review, which examines developments in RL for dynamic optimisation of SASS from 2021 to early 2025. Value-based, policy gradient, multi-agent, and hybrid/meta-learning approaches are the main RL methodologies identified in the review, which also looks at how they are applied in fields like cybersecurity, cloud resource management, and autonomous systems. The findings indicate that Cloud systems reduced average cost by 34.7% using PPO-based solutions, cybersecurity systems improved attack detection speed by 22.1% and false positive rates by 18.3%, and autonomous systems reduced energy consumption by 40% and adaptation latency by 27.5% in IoT and swarm robotics. Policy gradient methods (41%) dominate continuous control tasks, with PPO used in 27% of studies. Value-based approaches (32%) dominate discrete action domains, with deep Q-networks (DQN) variants used in 78% of cloud resource allocation studies. Multi-agent RL accounts for 18% of studies, with Multi-agent deep deterministic policy gradient (MADDPG) - 62% and QMIX (38%) being the most used. Serverless computing cut cold-start times by 35%, data centre optimisation lowered power usage effectiveness (PUE) by 15%, and RL-driven intrusion detection systems identified zero-day threats with 92% accuracy. Reward design difficulties were found in 63% of experiments, sample inefficiency required 1.2M episodes to converge, and real-world multi-agent reinforcement learning (MARL) deployments performed 23% worse than models. Metal analysis effect resulted in 95% in cost reduction, latency improvement and adaptation speed respectively. Practical adoption is limited because so few studies use standardised benchmarks or address safety and interpretability. In order to close the gap between research and practical implementation, the article ends by outlining open research questions and promoting formal verification, transfer learning, and hybrid learning.
Cited by 1
1 citation reported by external sources — individual citing-article records aren't available to list yet.
Related research
- Detecting Dental Caries through Captured Images Using the Machine Learning Technology Teachable Machine — shares topic coverage
- Prediction of Radiotherapy Dose Distribution for Glioblastoma Using Convolutional Neural Network Model — shares topic coverage
- A Systematic Literature Review of Machine Learning Methods in Healthcare — shares topic coverage
- Diagnostic Accuracy of Artificial Intelligence for Breast Cancer Detection: A Systematic Review — shares topic coverage
- Artificial Intelligence in the Analysis of the Fetal Genome in Utero: A Critical Review of Current Paradigms, Clinical Utility and Future Horizons — shares topic coverage
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
1
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.