Abstract
Cloud computing is a widely adopted paradigm for delivering scalable, on-demand computing services over the Internet. Because cloud workloads are inherently dynamic and often unpredictable, efficient resource allocation remains a persistent challenge, and static provisioning strategies frequently lead to either resource underutilization or performance degradation. Auto-scaling mechanisms address this problem by automatically adjusting computing resources in response to changing demand. In recent years, workload-aware auto-scaling frameworks that use workload analysis and prediction to improve scaling decisions have attracted considerable research interest. This review examines the major categories of auto-scaling techniques — reactive/rule-based, proactive/predictive, machine learning-based, and reinforcement learning-based approaches — and evaluates them against key performance indicators, including response time, resource utilization, operational cost, scalability, and Service Level Agreement (SLA) compliance. Building on reviewer feedback, this revised version additionally describes the literature search and selection methodology underlying the review, provides a more critical comparison of the reviewed techniques with respect to computational complexity, scalability, and deployment suitability across cloud environments, and discusses practical implementation issues such as security, energy consumption, and provisioning overhead. A conceptual taxonomy of workload-aware auto-scaling in cloud systems is presented, current research gaps are identified, and promising directions for future research in AI-driven cloud resource management are outlined.
KEYWORDS
Cloud Computing, Auto-Scaling, Resource Allocation, Workload Prediction, Machine Learning, Reinforcement Learning.
Neelesh Kumar Shrivastava1*, Chandra Shekhar Gautam2
1Dept. of Computer Science & Engineering, AKS University Satna-485001, M.P., India
2Dept. of Computer Science & Engineering, AKS University Satna-485001, M.P., India
