The rapid expansion of the cloud computing ecosystem necessitates scalable and energy-efficient resource management. However, achieving optimal energy efficiency is severely constrained by the complex interplay between computational workloads and thermal dynamics. Traditional data center (DC) scheduling methods typically rely on computationally expensive mechanistic models, such as computational fluid dynamics (CFD), or inefficient algorithm search, which frequently yield suboptimal solutions and struggle to adapt to real-time environments. To address these challenges, we propose an ecology-conservation-optimization virtual machine scheduling (EcoVMS) framework, a data-driven and end-to-end framework for compute-energy co-optimization. At its core, EcoVMS utilizes a latent reward generation enhanced proximal policy optimization (LARGE-PPO) algorithm, which utilizes large language models (LLMs) to analyze runtime telemetry-including central processing unit (CPU) utilization, thermal boundaries, and power consumption, and then utilizes large-scale semantic priors to generate potential reward factors. This mechanism resolves the sparse reward and credit assignment problems inherent in deep reinforcement learning without requiring complex physical modeling. Extensive simulations demonstrate that EcoVMS significantly outperforms baseline methods across diverse workload scenarios. Under a typical 75% workload, EcoVMS achieved a 12.4% improvement in significant energy savings, improving computing resource utilization while reducing hotspots and lowering service latency, resulting in guaranteed quality of service (QoS).
更多
查看译文
关键词
Cloud Computing,Data Center,Virtual Machine Scheduling,QoS,Energy Efficiency,Large Language Models,Reinforcement Learning