For over a decade, the problem of distributed cloud workload management has been studied with the goal of co-optimizing operational costs with other metrics, such as energy efficiency, using multi-objective optimization algorithms. However, there is a lack of a multi-objective algorithm that can not only provide a diverse and high-quality Pareto-optimal solution set, but also scale to different cloud management scenarios. The heterogeneity of cloud workloads and geo-distributed datacenters introduces complex operational scenarios over time and geographic locations. The introduction of emerging sustainability objectives, such as carbon emissions and wastewater generation, further aggravates the cloud management challenge. Moreover, inter-datacenter network costs must also be considered during cloud workload management to prevent unrealistic workload migration scenarios. In this article, we propose a novel cloud resource management framework called SHIELD-EB to co-optimize operational costs, operational carbon emissions, and wastewater generation in a geo-distributed cloud datacenter platform. To generate a diverse solution set with different tradeoffs, SHIELD-EB integrates a customized evolutionary strategy (CES) with eXtreme Gradient Boosting (XGB). Experimental results show that SHIELD-EB can achieve improvements of up to 33% in Pareto hypervolume, 7.1% ($0.6 M) in operational costs, 11.6% (7.6 tons) in operational carbon, and 12.0% (102.4 tons) in water usage compared to the state-of-the-art over a duration of 120 hours.
Distributed Energy Resources (DERs), such as solar panels, wind turbines, batteries, and electric vehicle chargers, increase power system complexity and threat surface for cyberattacks. This study evaluates Machine Learning models, such as Multilayer Perceptron, Convolutional Neural Networks, and Temporal Convolutional Networks, for classifying DER states under normal operation, hazards, and attacks, including False Data Injection, Man in the Middle, and Replay. Synthetic datasets were used to replicate attack conditions, and models were trained with binary and multiclass schemes. Results indicate that performance varies by DER type and attack scenario, with no single model performing best across all cases. A staged pipeline combining binary detection with multiclass attribution balances recall and precision, providing practical insights for adaptive threat detection in smart grids.
Reducing the environmental impact of cloud computing requires efficient workload distribution across geographically dispersed Data Center Clusters (DCCs) and simultaneously optimizing liquid and air (HVAC) cooling with time shift of workloads within individual data centers (DC). This paper introduces Green-DCC, which proposes a Reinforcement Learning (RL) based hierarchical controller to optimize both workload and liquid cooling dynamically in a DCC. By incorporating factors such as weather, carbon intensity, and resource availability, Green-DCC addresses realistic constraints and interdependencies. We demonstrate how the system optimizes multiple data centers synchronously, enabling the scope of digital twins, and compare the performance of various RL approaches based on carbon emissions and sustainability metrics while also offering a framework and benchmark simulation for broader ML research in sustainability.
Today there has never been a more profound codependence and synergetic convergence, such as the one between the energy and IT sectors, and the industries will need to work together more closely to meet today’s growing data center power demands.
The growing energy demands of machine learning workloads require more sustainable data centers with reduced energy consumption and lower carbon footprints. Highperformance computing (HPC) and supercomputing data centers, which utilize liquid cooling for improved efficiency over traditional air cooling, stand to benefit significantly from advanced optimization strategies. In this study, we propose a novel Reinforcement Learning (RL) based approach to optimize liquid cooling systems (RL-LC). By applying RL at varying levels of granularity in cooling control, our method achieves over 20% energy savings, delivering a more efficient and sustainable solution for data center management. Additionally, we introduce a flexible analytical model that supports simulations and digital twin applications, focusing on reducing energy consumption and carbon emissions. This work provides key advancements in liquid cooling technology, offering practical insights to mitigate the environmental impact of HPC operations.
In recent years, Large Language Models (LLM) such as ChatGPT, CoPilot, and Gemini have been widely adopted in different areas. As the use of LLMs continues to grow, many efforts have focused on reducing the massive training overheads of these models. But it is the environmental impact of handling user requests to LLMs that is increasingly becoming a concern. Recent studies estimate that the costs of operating LLMs in their inference phase can exceed training costs by 25x per year. As LLMs are queried incessantly, the cumulative carbon footprint for the operational phase has been shown to far exceed the footprint during the training phase. Further, estimates indicate that 500 ml of fresh water is expended for every 20-50 requests to LLMs during inference. To address these important sustainability issues with LLMs, we propose a novel framework called SLIT to co-optimize LLM quality of service (time-to-first token), carbon emissions, water usage, and energy costs. The framework utilizes a machine learning (ML) based metaheuristic to enhance the sustainability of LLM hosting across geo-distributed cloud datacenters. Such a framework will become increasingly vital as LLMs proliferate.
Serverless computing is an emerging cloud computing paradigm that can reduce costs for cloud providers and their customers. However, serverless cloud platforms have stringent performance requirements (due to the need to execute short duration functions in a timely manner) and a growing carbon footprint. Traditional carbon-reducing techniques such as shutting down idle containers can reduce performance by increasing cold-start latencies of containers required in the future. This can cause higher violation rates of service level objectives (SLOs). Conversely, traditional latency-reduction approaches of prewarming containers or keeping them alive when not in use can improve performance but increase the associated carbon footprint of the serverless cluster platform. To strike a balance between sustainability and performance, in this paper, we propose a novel carbon- and SLO-aware framework called CASA to schedule and autoscale containers in a serverless cloud computing cluster. Experimental results indicate that CASA reduces the operational carbon footprint of a FaaS cluster by up to 2.6x while also reducing the SLO violation rate by up to 1.4x compared to the state-of-the-art.
Function-as-a-Service (FaaS) is a growing cloud computing paradigm that is expected to reduce the user cost of service over traditional serverful approaches. However, the environmental impact of FaaS has not received much attention. We investigate FaaS scheduling and scaling from a sustainability perspective in this work. We find that the service-level objectives (SLOs) of FaaS and carbon emissions conflict with each other. We also find that SLO-focused FaaS scheduling can exacerbate water use in a datacenter. We propose a novel sustainability-focused FaaS scheduling and scaling framework to co-optimize SLO performance, carbon emissions, and wastewater generation.
An advanced digital model that emulates a physical entity or process using real-time data and simulations, a digital twin, is rapidly expanding across various industries. Digital twins enable the continuous monitoring of systems and provide data for predictive maintenance and performance analysis of either part of, or the whole system. The substantial amount of complex data used by digital twins can be simplified through visualization, streamlining real-time monitoring and analysis. While augmented reality (AR) performs contextual overlay of interactive, real-time data onto the physical environment, virtual reality (VR) enables user interaction in a virtual surrounding that enables spatial awareness when using digital twins. Additionally, AR/VR facilitates remote collaboration by allowing experts to view and interact with the digital twin, thus simplifying the troubleshooting and maintenance of physical assets. In this work, we develop TwinVAR, a design tool that provides digital twin visualization with AR/VR. In particular, we describe the design and implementation of TwinVAR, demonstrate its visualization capabilities, and exemplify its practical use to understand predictive failure analysis on digital twins for supercomputers.
Editor's notes: Realizing environmentally sustainable computing is an urgent goal given the role of the computing industry in exacerbating global warming. This article explores the intersection of sustainability and ethics in computing and how companies can pursue sustainability while protecting human rights and dignity. -Sudeep Pasricha, Colorado State University, USA -Marilyn Wolf, University of Nebraska-Lincoln, USA
Reliable and uninterrupted operation is crucial in supercomputers, especially during failures or inconsistencies i.e., anomalies. In this paper, we present a federated adaptive Digital Twin (DT) framework, with a focus on enhancing anomaly detection - a critical aspect of modern data center management. Our DT continuously monitors key metrics, detects anomalies powered by AI, and dynamically adjusts its monitoring parameters to ensure optimal performance. Using a dashboard, our system provides real-time alarms and detailed visualizations of detected anomalies, along with real-time visualization and forecast for selected metrics. Through a series of experiments, we validate the effectiveness of our approach in maintaining operational reliability and promptly identifying potential anomalies within the data center.
Data centers often suffer from energy inefficiencies caused by suboptimal heat flow, cooling imbalances, and temperature disparities. Although Thermal Computational Fluid Dynamics (CFD) can model these challenges, its computational demand often impedes timely, adaptive energy-saving measures. Addressing this, we introduce an efficient machine learning CFD surrogate model, DCCFD, implemented by a 3D Convolutional Neural Network (CNN) trained on the CFD or IOT-generated data. DCCFD has an innovative input channel design that captures the data center’s 3D geometry, power usage, airflow dynamics, and AC unit set points in three dimensions, enabling 3D CNN models like U-Net to accurately predict and generate the heat map. DCCFD is orders of magnitude faster than the CFD model with high accuracy. This enables real-time control for dynamic reallocation of computational load and cooling system adjustments, yielding significant energy savings. Beyond real-time adjustments, our approach facilitates proactive optimization during the DC installation phase, minimizing potential hot spots and maximizing equipment lifespan. So, DCCFD contributes to carbon footprint reduction and sustainability by maximizing cooling efficiency, reducing energy consumption, and increasing performance while extending the life of servers.
Fueled by an unprecedented adoption of artificial intelligence, data centers are becoming the largest growing consumers of energy. This results in both an opportunity and necessity to reinvent a tighter relationship between IT and sustainable energy supplies.
Digital twins provide living digital models of physical systems that enable data-driven analysis and application of artificial intelligence to better manage selective aspects of the data center and drive efficiency for sustainability.
Megatrends are major movements at a global scale likely to have a significant impact on the global economy, society, and ecology. Mega trends are composed of many correlated, mutually dependent trends. A megatrend is both the sum of trends and a guiding force since it influences its component trends. A megatrend impacts the evolution of these multiple trends, hence the importance of understanding mega trends. Trends are hard to predict, and megatrends are hard to recognize. Evolution is easier to observe at the trend level because it is more dynamic with more immediate results. However, the impact of megatrends is broader and larger.
Today's cloud data centers are often distributed geographically to provide robust data services. But these geo-distributed data centers (GDDCs) have a significant associated environmental impact due to their increasing carbon emissions and water usage, which needs to be curtailed. Moreover, the energy costs of operating these data centers continue to rise. This paper proposes a novel framework to co-optimize carbon emissions, water footprint, and energy costs of GDDCs, using a hybrid workload management framework called SHIELD that integrates machine learning guided local search with a decomposition-based evolutionary algorithm. Our framework considers geographical factors and time-based differences in power generation/use, costs, and environmental impacts to intelligently manage workload distribution across GDDCs and data center operation. Experimental results show that SHIELD can realize 34.4x speedup and 2.1x improvement in Pareto Hypervolume while reducing the carbon footprint by up to 3.7x, water footprint by up to 1.8x, energy costs by up to 1.3x, and a cumulative improvement across all objectives (carbon, water, cost) of up to 4.8x compared to the state-of-the-art.
Understanding megatrends allows individuals and organizations to align with and benefit from the richness of trending technologies.
Sustainability has become a critical problem confronting the community. Ecological issues such as climate change and CO2 emissions, economical frictions such as energy supplies, and socio-political issues such as wars threaten growth and equity. How can technology help?
In recent years, cloud service providers have been building and hosting datacenters across multiple geographical locations to provide robust services. However, the geographical distribution of datacenters introduces growing pressure to both local and global environments, particularly when it comes to water usage and carbon emissions. Unfortunately, efforts to reduce the environmental impact of such datacenters often lead to an increase in the cost of datacenter operations. To co-optimize the energy cost, carbon emissions, and water footprint of datacenter operation from a global perspective, we propose a novel framework for multi-objective sustainable datacenter management (MOSAIC) that integrates adaptive local search with a collaborative decomposition-based evolutionary algorithm to intelligently manage geographical workload distribution and datacenter operations. Our framework sustainably allocates workloads to datacenters while taking into account multiple geography- and time-based factors including renewable energy sources, variable energy costs, power usage efficiency, carbon factors, and water intensity in energy. Our experimental results show that, compared to the best-known prior work frameworks, MOSAIC can achieve 27.45x speedup and 1.53x improvement in Pareto Hypervolume while reducing the carbon footprint by up to 1.33x, water footprint by up to 3.09x, and energy costs by up to 1.40x. In the simultaneous three-objective co-optimization scenario, MOSAIC achieves a cumulative improvement across all objectives (carbon, water, cost) of up to 4.61x compared to the state-of-the-arts.