Multiple forecast visualizations (MFVs) present curated sets of forecasts to support decision-making under uncertainty. However, the research community knows little about how people interpret and integrate competing forecasts. In this study, we investigate the strategies individuals use when predicting hypothetical future events with MFVs across five visualization types (median, 95% CIs, standard deviation intervals, density plots, and hypothetical outcome plots) and multiple probability distributions in two preregistered experiments (n = 500 each). Analysis of 18 participant strategies and open responses shows that whereas many participants attempted to visually average across forecasts, others adopted a winner-takes-all approach (e.g., selecting a single forecast as the most likely outcome), which deviates from rational agent expectations. We also observed reliance on visual artifacts, such as intersection points or end caps. These findings underscore the complexity of interpreting a range of forecasts and help explain why individuals may privilege particular predictions in real-world decision contexts.
We are excited to welcome you to IEEE VIS 2024 in sunny St. Pete Beach, Florida! The conference program is shaping up to be one of the best we have seen, and the conference venue is undoubtedly one of the most fun locations we have ever held the VIS conference.
Bayesian Neural Networks (BNNs) offer a principled approach to modeling uncertainty in addition to providing predictions, making them particularly valuable for high-stake domains where uncertainty quantification is required. However, their adoption remains low, partly due to the difficulty in tuning and interpreting these models and their results. To address this limitation, we introduce BNNVis, a visual analytics tool designed to visualize BNNs and their results. BNNVis allows the user to understand the architecture and learned posterior weight distributions of their BNN at a glance and how these distributions differ from their prior. Additionally, the system helps them understand the distribution and magnitude of the accompanying uncertainties of the model’s predictions. BNNVis provides insight into the final predictions and the model, helping practitioners tune and interpret BNNs and their results. We describe a usage scenario to demonstrate how the features of BNNVis come together to support a practitioner in using a BNN.
Despite decision-making being a vital goal of data visualization, little work has been done to differentiate decision-making tasks within the field. While visualization task taxonomies and typologies exist, they often focus on more granular analytical tasks that are too low-level to describe large complex decisions, which can make it difficult to reason about and design decision-support tools. In this paper, we contribute a typology of decision-making tasks that were iteratively refined from a list of design goals distilled from a literature review. Our typology is concise and consists of only three tasks: CHOOSE, ACTIVATE, and CREATE. Although decision types originating in other disciplines exist, we provide definitions for these tasks that are suitable for the visualization community. Our proposed typology offers two benefits. First, the ability to compose and hierarchically organize the tasks enables flexible and clear descriptions of decisions with varying levels of complexities. Second, the typology encourages productive discourse between visualization designers and domain experts by abstracting the intricacies of data, thereby promoting clarity and rigorous analysis of decision-making processes. We demonstrate the benefits of our typology through four case studies, and present an evaluation of the typology from semi-structured interviews with experienced members of the visualization community who have contributed to developing or publishing decision support systems for domain experts. Our interviewees used our typology to delineate the decision-making processes supported by their systems, demonstrating its descriptive capacity and effectiveness. Finally, we present preliminary findings on the usefulness of our typology for visualization design.
Uncertainty visualization plays a critical role in transforming ensemble simulation data into actionable insights by effectively communicating various dimensions of uncertainty within a system. The emergence of artificial intelligence-driven surrogate models trained on multirun ensemble data offers a transformative opportunity to replace computationally intensive simulations with fast estimates, enabling users to explore data spaces with unprecedented depth and interactivity. However, integrating ensemble data and surrogate models into decision-making workflows and tools introduces novel challenges for uncertainty visualization. These include reconciling and clearly communicating the unique uncertainties associated with ensembles and their surrogate model estimates, and leveraging these approximations to inform actionable decisions. This work explores these challenges in the context of high-dimensional data visualization, bridging discrete datasets with their continuous representations and addressing the complexities of systems that support iterative navigation between input and output spaces. We evaluate the role of uncertainty visualization in fostering intuitive, actionable interactions and identify critical hurdles in advancing this frontier of computational simulation.
Videos are becoming a ubiquitous means of sharing information on social media platforms. In response, data videos-short clips combining visualization with dynamic storytelling, audio descriptions, and spatial referencing-have gained popularity for communicating data. These affordances suggest that data videos might communicate data patterns, trends, and concepts more effectively than static visualizations, enhancing comprehension. However, existing research has not systematically tested this claim. To address this gap, we conducted two controlled studies to measure comprehension differences between data videos and static visualizations. Despite leveraging visual cues and audio explanations, no data video led to significantly better comprehension than an analogous static visualization. Our results suggest data videos are not categorically better and that future research should examine the tradeoffs between their engagement benefits and costs.
At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.
Scenario studies are a technique for representing a range of possible complex decisions through time, and analyzing the impact of those decisions on future outcomes of interest. It is common to use scenarios as a way to study potential pathways towards future build-out and decarbonization of energy systems. The results of these studies are often used by diverse energy system stakeholders - such as community organizations, power system utilities, and policymakers - for decision-making using data visualization. However, the role of visualization in facilitating decision-making with energy scenario data is not well understood. In this work, we review common visualization designs employed in energy scenario studies and discuss the effectiveness of some of these techniques in facilitating different types of analysis with scenario data.
Materials science has a significant impact on society and its quality of life, e.g., through the development of safer, more durable, more economical, environmentally friendly, and sustainable materials. Visual computing in materials science integrates computer science disciplines from image processing, visualization, computer graphics, pattern recognition, computer vision, virtual and augmented reality, machine learning, to human-computer interaction, to support the acquisition, analysis, and synthesis of (visual) materials science data with computer resources. Therefore, visual computing may provide fundamentally new insights into materials science problems by facilitating the understanding, discovery, design, and usage of complex material systems. This seminar is considered as a follow-up of the Dagstuhl Seminar 19151 Visual Computing in Materials Sciences, held in April 2019. Since then, the field has kept evolving and many novel challenges have emerged, with regard to more traditional topics in visual computing, such as topology analysis or image processing and analysis, to recently emerging topics, such as uncertainty and ensemble analysis, and to the integration of new research disciplines and exploratory technologies, such machine learning and immersive analytics. With the current seminar, we target to strengthen and extend the collaboration between the domains of visual computing and materials science (and across visual computing disciplines), by foreseeing challenges and identifying novel directions of interdisciplinary work. We brought visual computing and visualization experts from academia, research centers, and industry together with domain experts, to uncover the overlaps of visual computing and materials science and to discover yet-unsolved challenges, on which we can collaborate to achieve a higher societal impact.
Join us for the 14th IEEE Symposium on Large Data Analysis and Visualization (IEEE LDAV) on Sunday, October 13th 2024 collocated with IEEE VIS 2024 in St. Pete Beach, Florida, USA.
Uncertainty visualization is a key component in translating important insights from ensemble simulation data into actionable decision-making by visually conveying various aspects of uncertainty within a system. With the recent advent of fast surrogate models trained on ensemble data, we can substitute computationally expensive simulations, which allows users to interact with more aspects of data spaces than ever before. However, the use of ensemble data with surrogate models in a decision-making tool brings up new challenges for uncertainty visualization, namely how to reconcile and communicate the new and different types of uncertainties brought in by surrogates and how to utilize these new data estimates in actionable ways. In this work, we examine these issues as they relate to high-dimensional data visualization, the integration of discrete datasets and the continuous representations of those datasets, and the unique difficulties associated with systems that allow users to iterate between input and output spaces. We assess the role of uncertainty visualization in facilitating intuitive and actionable interaction with ensemble data and surrogate models, and highlight key challenges in this new frontier of computational simulation.
The continual expansion of high-performance computing (HPC) brings with it an increasing need for efficiency. Heavy investment in energy, hardware, and software infrastructure to support peta- and exascale computing requires the optimization of existing systems and, wherever possible, the discernment and adoption of best-practices towards these goals. Such is the case for runtime prediction. When a job is submitted to an HPC system, an estimate of its runtime is provided by the user in the form of “requested wallclock”. Error in this user-provided estimate can lead to jobs being prematurely killed by the scheduler, increased wait time on the queue, and decreased system utilization. More than fifteen years of research has been directed at mitigating these effects by using data-driven runtime predictions. Codified here is a set of commonalities and insights emerging from this body of work, which we present as recommendations and best practices. These practices are combined into a methodological approach described and evaluated on an 11-million-job dataset from the National Renewable Energy Laboratory’s petascale HPC system, Eagle. This dataset and the accompanying codebase have been released to the public domain for the benefit of the wider HPC research community.
Grid operators can address the inherently stochastic nature of renewables by solving a two-stage stochastic programming model that minimizes the cost of dispatch decisions while accounting for the complex grid dynamics. It is common to use a sample average approximation to estimate the expectation of the second stage costs in this model. However, the large sample count needed for numerical accuracy makes effective modeling large-scale electric grids computationally intractable. We introduce a control variate multi-fidelity estimator for the second-stage recourse that enables high quality dispatch decisions in real-time with a reduced computational burden. We obtain a hierarchy of model fidelities by linearizing the AC power flow system representation to DC power flow, and by relaxing transmission and voltage network constraints. We evaluate the performance of our proposed method on a synthetic grid with 73 buses against a deterministic baseline with persistence forecast and a high-fidelity reference. Our analysis shows a computational speed-up of 7.62x with a minimal loss in accuracy. The multi-fidelity method is well suited to fidelity combinations that use a simplified network topology in their lower fidelity model and is an attractive option for applications where accurate grid modeling needed on a limited computational budget.
As renewable energy generation deployment increases, the operation of electrical grids becomes more complex. Economic dispatch is part of a grid operator's regular decision process where the amount of energy to generate is determined based on the number of available generators and the actual level of energy demand. Renewable generators are inherently stochastic due to the chaotic nature of weather patterns, and thus, real-time decisions of economic dispatch become increasingly complex. Modeling efforts to assist in these decisions in the highest fidelity typically take hours to days to solve on leadership-class computers; too long for the 5-minute operational time-frame demanded of operators. Alternatively, multi-fidelity approximations can be used to predict generation levels quickly and with sufficient accuracy to be used for real-time operations. We have developed a visualization tool to demonstrate the utility of multi-fidelity approximations by displaying contextual results of economic dispatch approximations, comparisons across fidelity levels of generation levels and possible failures to meet demand, and meta-data on the modeling setup.
In this paper, we present the reV (Renewable Energy Potential) Dashboard, an interactive browser-based tool for uncertainty visualization and data exploration built using customized plotly dash components. With continuing development and utilization of computational models to study the power sector there is an increasing need for data-driven visualization tools which allow scientists, researchers, and engineers to interact with their data in real-time. Our principle motivation for developing this interactive uncertainty visualization was to provide domain scientists and modelers a platform which allows them to better understand and communicate scientific findings stemming from intricate information encoded in their data that is otherwise difficult to capture by conventional analysis. The development of customized dashboard components using the React programming paradigm, combined with fully pre-processed data, allows for users to select a variety of data options to update and manipulate the visualization in a straightforward computationally efficient manner.
Wind plants operate in stochastic environments characterized by complex turbulent flow dynamics and high-dimensional random variables. A key step in uncertainty quantification studies is sensitivity analysis and dimension reduction that can facilitate the development of surrogate models to be used for forward and inverse propagation or optimization under uncertainty. Prior work has shown active subspaces are an effective tool for identifying important directions in the space of stochastic inputs; however, they have only been applied to single-fidelity wind plant models. In this study, we investigate the efficacy of a multi-fidelity active subspace method for analyzing the uncertainty in wind plant power output. The multi-fidelity active subspace estimator offers the promise of increased accuracy in identifying active subspaces as compared to a single-fidelity estimator for the same computational cost, or a reduction in cost for the same accuracy. This makes the study of uncertainty in larger wind plants and with higher fidelity physics tractable. The multi-fidelity active subspace method is applied to gridded and existing wind plant layouts with single and multiple inflow conditions and its performance for surrogate modeling and uncertainty propagation is compared against a single-fidelity active subspace method. This multi-fidelity approach yields substantial computational speedups of 2x - 3.4x across the test cases along with acceptable accuracy in surrogate modeling and computing statistical moments.