Despite decision-making being a vital goal of data visualization, little work has been done to differentiate decision-making tasks within the field. While visualization task taxonomies and typologies exist, they often focus on more granular analytical tasks that are too low-level to describe large complex decisions, which can make it difficult to reason about and design decision-support tools. In this paper, we contribute a typology of decision-making tasks that were iteratively refined from a list of design goals distilled from a literature review. Our typology is concise and consists of only three tasks: CHOOSE, ACTIVATE, and CREATE. Although decision types originating in other disciplines exist, we provide definitions for these tasks that are suitable for the visualization community. Our proposed typology offers two benefits. First, the ability to compose and hierarchically organize the tasks enables flexible and clear descriptions of decisions with varying levels of complexities. Second, the typology encourages productive discourse between visualization designers and domain experts by abstracting the intricacies of data, thereby promoting clarity and rigorous analysis of decision-making processes. We demonstrate the benefits of our typology through four case studies, and present an evaluation of the typology from semi-structured interviews with experienced members of the visualization community who have contributed to developing or publishing decision support systems for domain experts. Our interviewees used our typology to delineate the decision-making processes supported by their systems, demonstrating its descriptive capacity and effectiveness. Finally, we present preliminary findings on the usefulness of our typology for visualization design.
Dimensionality reduction techniques are widely used for visualizing high-dimensional data. However, support for interpreting patterns of dimension reduction results in the context of the original data space is often insufficient. Consequently, users may struggle to extract insights from the projections. In this paper, we introduce DimBridge, a visual analytics tool that allows users to interact with visual patterns in a projection and retrieve corresponding data patterns. DimBridge supports several interactions, allowing users to perform various analyses, from contrasting multiple clusters to explaining complex latent structures. Leveraging first-order predicate logic, DimBridge identifies subspaces in the original dimensions relevant to a queried pattern and provides an interface for users to visualize and interact with them. We demonstrate how DimBridge can help users overcome the challenges associated with interpreting visual patterns in projections.
Content recommendation tasks increasingly use Graph Neural Networks, but it remains challenging for machine learning experts to assess the quality of their outputs. Visualization systems for GNNs that could support this interrogation are few. Moreover, those that do exist focus primarily on exposing GNN architectures for tuning and prediction tasks and do not address the challenges of recommendation tasks. We developed RekomGNN, a visual analytics system that supports ML experts in exploring GNN recommendations across several dimensions and making annotations about their quality. RekomGNN straddles the design space between Neural Network and recommender system visualization to arrive at a set of encoding and interaction choices for recommendation tasks. We found that RekomGNN helps experts make qualitative assessments of the GNN's results, which they can use for model refinement. Overall, our contributions and findings add to the growing understanding of visualizing GNNs for increasingly complex tasks.
This study presents insights from interviews with nineteen Knowledge Graph (KG) practitioners who work in both enterprise and academic settings on a wide variety of use cases. Through this study, we identify critical challenges experienced by KG practitioners when creating, exploring, and analyzing KGs that could be alleviated through visualization design. Our findings reveal three major personas among KG practitioners - KG Builders, Analysts, and Consumers - each of whom have their own distinct expertise and needs. We discover that KG Builders would benefit from schema enforcers, while KG Analysts need customizable query builders that provide interim query results. For KG Consumers, we identify a lack of efficacy for node-link diagrams, and the need for tailored domain-specific visualizations to promote KG adoption and comprehension. Lastly, we find that implementing KGs effectively in practice requires both technical and social solutions that are not addressed with current tools, technologies, and collaborative workflows. From the analysis of our interviews, we distill several visualization research directions to improve KG usability, including knowledge cards that balance digestibility and discoverability, timeline views to track temporal changes, interfaces that support organic discovery, and semantic explanations for AI and machine learning predictions.
Prior studies have estimated the efficiency of PGT-A using published implantation and aneuploidy rates [1]. Here we developed a more rigorous approach that incorporates a novel methodology for imputing the likelihood of failed untested transfers resulting from aneuploidy.
Abstract Study question What is the expected improvement in pregnancy rates using an artificial intelligence (AI) model for embryo ranking compared to manual grading systems? Summary answer A large-scale retrospective bootstrapped analysis shows that use of an AI model for embryo ranking can improve pregnancy rates compared to manual grading. What is known already Embryo evaluation is one of the most important steps of an in vitro fertilization (IVF) procedure. Recently, artificial intelligence (AI) models have been developed to automate embryo analysis and reduce the subjectivity of manual grading. While models are often evaluated in terms of classification accuracy or area under the curve (AUC), a more relevant metric is improvement in pregnancy rates. Here we evaluate a previously developed model using a large-scale bootstrapped analysis of virtual patient pregnancy rates and compare its performance to manual grading. Study design, size, duration Historical, de-identified images of transferred blastocyst-stage embryos and manual morphology grades were collected from 11 IVF clinics in the United States for cycles started between 2015-2020. Images were captured on day 5, 6, or 7 using the inverted microscope prior to biopsy or freeze. A total of 1,776 test set images from 3-fold cross validation were used for this analysis. Participants/materials, setting, methods Embryos were matched by age, PGT status, and race to create 16 distinct categories. Virtual patient panels were created within each category using a random selection of 3-5 embryos. Embryos were re-used across different panels, but each individual panel was unique. Three different manual ranking systems were created incorporating the morphology grade and day of image capture. The AI and one randomly chosen manual ranking system independently selected a top embryo for each panel. Main results and the role of chance On average, 105,263 unique virtual patient panels were constructed from the 1,776 embryos. Within these panels, the AI model and manual ranking system selected different top embryos from each other in 27,860 cases, or 26% of the time. The average pregnancy rate of the top-ranked embryo using manual grading was 53.1%, and the average pregnancy rate of the top-ranked embryo using the AI model was 59.4%. The average pregnancy rate improvement from using the AI model was 6.3%, with a standard deviation of 0.2% measured across 10 repetitions of the simulation with different random seeds. Limitations, reasons for caution The primary limitation is the retrospective nature of this study. Also, this bootstrapped panel study relied on recorded manual morphology grades at the time of embryo transfer or freeze rather than on the actual selection of the top embryo in each panel by an embryologist. Wider implications of the findings Our results demonstrate the potential of using an AI model for embryo ranking in terms of improved pregnancy rates. Results from this large-scale bootstrapped retrospective analysis will help inform the design of future clinical validation studies. Trial registration number not applicable
Objective: To perform a series of analyses characterizing an artificial intelligence (AI) model for ranking blastocyst-stage embryos. The primary objective was to evaluate the benefit of the model for predicting clinical pregnancy, whereas the secondary objective was to identify limitations that may impact clinical use. Design: Retrospective study. Setting: Consortium of 11 assisted reproductive technology centers in the United States. Patient(s): Static images of 5,923 transferred blastocysts and 2,614 nontransferred aneuploid blastocysts. Intervention(s): None. Main Outcome Measure(s): Prediction of clinical pregnancy (fetal heartbeat). Result(s): The area under the curve of the AI model ranged from 0.6 to 0.7 and outperformed manual morphology grading overall and on a per-site basis. A bootstrapped study predicted improved pregnancy rates between +5% and +12% per site using AI compared with manual grading using an inverted microscope. One site that used a low-magnification stereo zoom microscope did not show predicted improvement with the AI. Visualization techniques and attribution algorithms revealed that the features learned by the AI model largely overlap with the features of manual grading systems. Two sources of bias relating to the type of microscope and presence of embryo holding micropipettes were identified and mitigated. The analysis of AI scores in relation to pregnancy rates showed that score differences of >= 0.1 (10%) correspond with improved pregnancy rates, whereas score differences of <0.1 may not be clinically meaningful. Conclusion(s): This study demonstrates the potential of AI for ranking blastocyst stage embryos and highlights potential limitations related to image quality, bias, and granularity of scores. (C) 2021 by American Society for Reproductive Medicine.
Abstract Study question What is the sensitivity of an embryo-grading artificial intelligence (AI) model to different focal planes and how do we obtain consistent scores across focal planes? Summary answer Test-time augmentation and ensemble modeling reduce sensitivity of the AI model to different focal planes while maintaining performance. What is known already When prioritizing embryos for transfer, embryologists assess the 3D morphological features under a microscope, by zooming up and down, and assign a score that reflects the embryo quality. In comparison, some AI-based embryo grading models typically take one 2D focal plane of an embryo and output a score based on that focal plane. AI models such as convolutional neural networks (CNNs) are known to be sensitive to perturbations in its input. In order to reduce sensitivity and generalization error and thus improve predictive performance, techniques such as ensemble learning and test-time augmentation can be used. Study design, size, duration Historical, de-identified images of blastocyst-stage embryos were collected from 11 IVF clinics in the United States for cycles between 2015-2020. 5,100 blastocysts were matched to pregnancy outcomes as determined by fetal heartbeat. 2,900 blastocysts were matched to aneuploid PGT-A results and added to the negative training group to reduce selection bias. Data was split to 70% for training and 30% for testing. A set of 10 embryos were used for focal plane sensitivity. Participants/materials, setting, methods A single model (ResNet18), a three-model (ResNet18), and a six-model (ResNet18 and EfficientNet-b1) ensemble with and without test-time augmentation were trained to rank embryos according to their likelihood of reaching clinical pregnancy. Test-time augmentation involved taking the average scores from 4 flipped and rotated copies of the original input image. Manual grades were mapped to numeric scores for comparison. The AUC was used to evaluate the ability of the models to rank embryos. Main results and the role of chance Focal plane sensitivity was calculated as the range, or difference between the maximum and minimum score, for an embryo at different focal planes. Between 12 and 100 focal plane images were available for each of the 10 embryos. On average, the focal plane range was 0.26 for the single model, 0.22 for the single model with test-time augmentation, 0.14 for a 3-model ensemble with test-time augmentation, and 0.11 for a 6-model ensemble with test-time augmentation. Test-time augmentation on the single model reduced the range by 17%; whereas ensemble modeling with test-time augmentation reduced the range by 46% for the 3-model ensemble and 60% for the 6-model ensemble. Reduction in range did not compromise performance. The AUC for the test set for all embryos was 0.73 for the single model, 0.74 for the single model with test-time augmentation, 0.75 for the three-model ensemble with test-time augmentation and 0.74 for the six-model ensemble with test-time augmentation. All models outperformed manual grading, which was estimated to have an AUC of 0.67 for all embryos. Limitations, reasons for caution Our analysis on focal plane sensitivity was limited to a small sample size of 10 embryos, so more samples will be needed to confirm our findings. Wider implications of the findings Test-time augmentation and ensemble techniques can be used to reduce sensitivity while maintaining model performance. By reducing sensitivity to different focal planes, an AI model can produce one reliable score for a single embryo as is done currently in practice with manual grading. Trial registration number not applicable
Anomaly detection remains an open challenge in many application areas. While there are a number of available machine learning algorithms for detecting anomalies, analysts are frequently asked to take additional steps in reasoning about the root cause of the anomalies and form actionable hypotheses that can be communicated to business stakeholders. Without the appropriate tools, this reasoning process is time-consuming, tedious, and potentially error-prone. In this paper we present PIXAL, a visual analytics system developed following an iterative design process with professional analysts responsible for anomaly detection. PIXAL is designed to fill gaps in existing tools commonly used by analysts to reason with and make sense of anomalies. PIXAL consists of three components: (1) an algorithm that finds patterns by aggregating multiple anomalous data points using first-order predicates, (2) a visualization tool that allows the analyst to build trust in the algorithmically-generated predicates by performing comparative and counterfactual analyses, and (3) a visualization tool that helps the analyst generate and validate hypotheses by exploring which features in the data most explain the anomalies. Finally, we present the results of a qualitative observational study with professional analysts. These results of the study indicate that PIXAL facilitates the anomaly reasoning process, allowing analysts to make sense of anomalies and generate hypotheses that are meaningful and actionable to business stakeholders.
With the rapidly increasing resolutions of 360° cameras, head-mounted displays, and live-streaming services, streaming high-resolution panoramic videos over limited-bandwidth networks is becoming a critical challenge. Foveated video streaming can address this rising challenge in the context of eye-tracking-equipped virtual reality head-mounted displays. However, conventional log-polar foveated rendering suffers from a number of visual artifacts such as aliasing and flickering. In this paper, we introduce a new log-rectilinear transformation that incorporates summed-area table filtering and off-the-shelf video codecs to enable foveated streaming of 360° videos suitable for VR headsets with built-in eye-tracking. To validate our approach, we build a client-server system prototype for streaming 360° videos which leverages parallel algorithms over real-time video transcoding. We conduct quantitative experiments on an existing 360° video dataset and observe that the log-rectilinear transformation paired with summed-area table filtering heavily reduces flickering compared to log-polar subsampling while also yielding an additional 10% reduction in bandwidth usage.
Computational efficiency is a critical constraint for a variety of cutting-edge real-time applications. In this work, we identify an opportunity to speed up the end-to-end runtime of two such compute bound applications by incorporating approximate linear algebra techniques. Particularly, we apply approximate matrix multiplication to artificial Neural Networks (NNs) for image classification and to the robotics problem of Distributed Simultaneous Localization and Mapping (DSLAM). Expanding upon recent sampling-based Monte Carlo approximation strategies for matrix multiplication, we develop updated theoretical bounds, and an adaptive error prediction strategy. We then apply these techniques in the context of NNs and DSLAM increasing the speed of both applications by 15-20% while maintaining a 97% classification accuracy for NNs running on the MNIST dataset and keeping the average robot position error under 1 meter (vs 0.32 meters for the exact solution). However, both applications experience variance in their results. This suggests that Monte Carlo matrix multiplication may be an effective technique to reduce the memory and computational burden of certain algorithms when used carefully, but more research is needed before these techniques can be widely used in practice.