
This position paper critically examines the graphical inference framework for evaluating visualizations using the lineup task. We present a re-analysis of lineup task data using signal detection theory, applying four Bayesian non-linear models to investigate whether color ramps with more color name variation increase false discoveries. Our study utilizes data from Reda and Szafir's previous work [20], corroborating their findings while providing additional insights into sensitivity and bias differences across colormaps and individuals. We suggest improvements to lineup study designs and explore the connections between graphical inference, signal detection theory, and statistical decision theory. Our work contributes a more perceptually grounded approach for assessing visualization effectiveness and offers a path forward for better aligning graphical inference methods with human cognition. The results have implications for the development and evaluation of visualizations, particularly for exploratory data analysis scenarios. Supplementary materials are available at https://osf.io/xd5cj/.
Despite 30+ years of academic practice, visualization still lacks an explanation of how and why it functions in complex organizations performing knowledge work. This survey takes steps to bridge this knowledge gap by examining the intersection of organizational studies and visualization design, highlighting the concept of boundary objects, which visualization practitioners are adopting in both CSCW (computer-supported collaborative work) and HCI. This paper also collects the prior literature on boundary objects in visualization design studies, a methodology which maps closely to action research in organizations, and addresses the same problems of 'knowing in common'. We argue that rocess artifacts generated by visualization design studies function as boundary objects in their own right, facilitating knowledge transfer across disciplines within an organization. Currently, visualization faces the challenge of explaining how sense-making functions across domains, through visualization artifacts, and how these support decision-making. As a deeply interdisciplinary field, we advocate that visualization should adopt the theory of boundary objects in order to embrace its plurality of domains and systems, whilst empowering its practitioners with a unified process-based theory.
Empirical studies in visualisation often compare visual representations to identify the most effective visualisation for a particular visual judgement or decision making task. However, the effectiveness of a visualisation may be intrinsically related to, and difficult to distinguish from, individual-level factors such as visualisation literacy. Complicating matters further, visualisation literacy itself is not a singular intrinsic quality, but can be a result of several distinct challenges that a viewer encounters when performing a task with a visualisation. In this paper, we describe how such challenges apply to experiments that we use to evaluate visualisations, and discuss a set of considerations for designing studies in the future. Finally, we argue that aspects of the study design that are often neglected or overlooked (such as the onboarding of participants, tutorials, training, etc.) can have a big role in the results of a study and can potentially impact the conclusions that the researchers can draw from the study.
Submissions of original research that use Large Language Models (LLMs) or that study their behavior, suddenly account for a sizable portion of works submitted and accepted to visualization (VIS) conferences and similar venues in human-computer interaction (HCI). In this brief position paper, I argue that reviewers are relatively unprepared to evaluate these submissions effectively. To support this conjecture I reflect on my experience serving on four program committees for VIS and HCI conferences over the past year. I will describe common reviewer critiques that I observed and highlight how these critiques influence the review process. I also raise some concerns about these critiques that could limit applied LLM research to all but the best-resourced labs. While I conclude with suggestions for evaluating research contributions that incorporate LLMs, the ultimate goal of this position paper is to simulate a discussion on the review process and its challenges.
The generation and presentation of counterfactual explanations (CFEs) are a commonly used, model-agnostic, approach to helping end-users reason about the validity of AI/ML model outputs. By demonstrating how sensitive the model's outputs are to minor variations, CFEs are thought to improve understanding of the model's behavior, identify potential biases, and increase the transparency of 'black box models'. Here, we examine how CFEs support a diverse audience, both with and without technical expertise, to understand the results of an LLM-informed sentiment analysis. We conducted a preliminary pilot study with ten individuals with varied expertise from ranging NLP, ML, and ethics, to specific domains. All individuals were actively using or working with AI/ML technology as part of their daily jobs. Through semi-structured interviews grounded in a set of concrete examples, we examined how CFEs influence participants' perceptions of the model's correctness, fairness, and trustworthiness, and how visualization of CFEs specifically influences those perceptions. We also surface how participants wrestle with their internal definitions of 'explainability', relative to what CFEs present, their cultures, and backgrounds, in addition to the, much more widely studied phenomena, of comparing their baseline expectations of the model's performance. Compared to prior research, our findings highlight the sociotechnical frictions that CFEs surface but do not necessarily remedy. We conclude with the design implications of developing transparent AI/ML visualization systems for more general tasks.
The cognitive processes involved in understanding and misunderstanding visualizations have not yet been fully clarified, even for well-studied designs, such as bar charts. In particular, little is known about whether viewers can improve their learning processes by getting better insight into their own cognition. This paper describes a simple method to measure the role of such metacognitive understanding when learning to read bar charts. For this purpose, we conducted an experiment in which we investigated bar chart learning repeatedly, and tested how learning over trials was effected by metacognitive understanding. We integrate the findings into a model of metacognitive processing of visualizations, and discuss implications for the design of visualizations.
Visualising personal experiences is often described as a means for self-reflection, shaping one's identity, and sharing it with others. In policymaking, personal narratives are regarded as an important source of intelligence to shape public discourse and policy. Therefore, policymakers are interested in the interplay between individual-level experiences and macro-political processes that play into shaping these experiences. In this context, visualisation is regarded as a medium for advocacy, creating a power balance between individuals and the power structures that influence their health and well-being. In this paper, we offer a politically-framed reflection on how visualisation creators define lived experience data, and what design choices they make for visualising them. We identify data characteristics and design choices that enable visualisation authors and consumers to engage in a process of narrative co-construction, while navigating structural forms of inequality. Our political framing is driven by ideas of master and alternative narratives from Diversity Science, in which authors and narrators engage in a process of negotiation with power structures to either maintain or challenge the status quo.
Foundation models for vision and language are the basis of AI applications across numerous sectors of society. The success of these models stems from their ability to mimic human capabilities, namely visual perception in vision models, and analytical reasoning in large language models. As visual perception and analysis are fundamental to data visualization, in this position paper we ask: how can we harness foundation models to advance progress in visualization design? Specifically, how can multimodal foundation models (MFMs) guide visualization design through visual perception? We approach these questions by investigating the effectiveness of MFMs for perceiving visualization, and formalizing the overall visualization design and optimization space. Specifically, we think that MFMs can best be viewed as judges, equipped with the ability to criticize visualizations, and provide us with actions on how to improve a visualization. We provide a deeper characterization for text-to-image generative models, and multi-modal large language models, organized by what these models provide as output, and how to utilize the output for guiding design decisions. We hope that our perspective can inspire researchers in visualization on how to approach MFMs for visualization design.
The replication crisis has spawned a revolution in scientific methods, aimed at increasing the transparency, robustness, and reliability of scientific outcomes. In particular, the practice of preregistering study designs has shown important advantages. Preregistration can help limit questionable research practices, as well as increase the success rate of study replications. Many fields have now adopted preregistration as a default expectation for published studies. In 2022, we set up a panel “Merits and Limits of User Study Prereg-istration” with the overall goal of explaining the concept of prereg-istration to a wide VIS audience and discussing its suitability for visualization research. We report on the arguments and discussion of this panel in the hope that it can benefit the visualization com-munity at large. All materials and a copy of this paper are available on our OSF repository at https://osf.io/wes57/.
Various standardized tests exist that assess individuals' visualization literacy. Their use can help to draw conclusions from studies. However, it is not taken into account that the test itself can create a pressure situation where participants might fear being exposed and assessed negatively. This is especially problematic when testing domain experts in design studies. We conducted interviews with experts from different domains performing the Mini-VLAT test for visualization literacy to identify potential problems. Our participants reported that the time limit per question, ambiguities in the questions and visualizations, and missing steps in the test procedure mainly had an impact on their performance and content. We discuss possible changes to the test design to address these issues and how such assessment methods could be integrated into existing evaluation procedures.
Complexity is often seen as a inherent negative in information design, with the job of the designer being to reduce or eliminate complexity, and with principles like Tufte's "data-ink ratio" or "chartjunk" to operationalize mimimalism and simplicity in visualizations. However, in this position paper, we call for a more expansive view of complexity as a design material, like color or texture or shape: an element of information design that can be used in many ways, many of which are beneficial to the goals of using data to understand the world around us. We describe complexity as a phenomenon that occurs not just in visual design but in every aspect of the sensemaking process, from data collection to interpretation. For each of these stages, we present examples of ways that these various forms of complexity can be used (or abused) in visualization design. We ultimately call on the visualization community to build a more nuanced view of complexity, to look for places to usefully integrate complexity in multiple stages of the design process, and, even when the goal is to reduce complexity, to look for the non-visual forms of complexity that may have otherwise been overlooked.
This paper revisits the role of quantitative and qualitative methods in visualization research in the context of advancements in artificial intelligence (AI). The focus is on how we can bridge between the different methods in an integrated process of analyzing user study data. To this end, a process model of—potentially iterated—semantic enrichment and transformation of data is proposed. This joint perspective of data and semantics facilitates the integration of quantitative and qualitative methods. The model is motivated by examples of own prior work, especially in the area of eye tracking user studies and coding data-rich observations. Finally, there is a discussion of open issues and research opportunities in the interplay between AI, human analyst, and qualitative and quantitative methods for visualization research.
In the rapidly evolving field of information visualization, rigorous evaluation is essential for validating new techniques, understanding user interactions, and demonstrating the effectiveness and usability of visualizations. Faithful evaluations provide valuable insights into how users interact with and perceive the system, enabling designers to identify potential weaknesses and make informed decisions about design choices and improvements. However, an emerging trend of multiple evaluations within a single research raises critical questions about the sustainability, feasibility, and methodological rigor of such an approach. New researchers and students, influenced by this trend, may believe - multiple evaluations are necessary for a study, regardless of the contribution types. However, the number of evaluations in a study should depend on its contributions and merits, not on the trend of including multiple evaluations to strengthen a paper. So, how many evaluations are enough? This is a situational question and cannot be formulaically determined. Our objective is to summarize current trends and patterns to assess the distribution of evaluation methods over different paper contribution types. In this paper, we identify this trend through a non-exhaustive literature survey of evaluation patterns in 214 papers in the two most recent years' VIS issues in IEEE TVCG from 2023 and 2024. We then discuss various evaluation strategy patterns in the information visualization field to guide practical choices and how this paper will open avenues for further discussion.
Stress is among the most commonly employed quality metrics and optimization criteria for dimension reduction projections of high-dimensional data. Complex, high-dimensional data is ubiquitous across many scientific disciplines, including machine learning, biology, and the social sciences. One of the primary methods of visualizing these datasets is with two-dimensional scatter plots that visually capture some properties of the data. Because visually determining the accuracy of these plots is challenging, researchers often use quality metrics to measure the projection's accuracy or faithfulness to the full data. One of the most commonly employed metrics, normalized stress, is sensitive to uniform scaling (stretching, shrinking) of the projection, despite this act not meaningfully changing anything about the projection. We investigate the effect of scaling on stress and other distance-based quality metrics analytically and empirically by showing just how much the values change and how this affects dimension reduction technique evaluations. We introduce a simple technique to make normalized stress scale-invariant and show that it accurately captures expected behavior on a small benchmark.
I analyze the evolution of papers certified by the Graphics Replicability Stamp Initiative (GRSI) to be reproducible, with a specific focus on the subset of publications that address visualization-related topics. With this analysis I show that, while the number of papers is increasing overall and within the visualization field, we still have to improve quite a bit to escape the replication crisis. I base my analysis on the data published by the GRSI as well as publication data for the different venues in visualization and lists of journal papers that have been presented at visualization-focused conferences. I also analyze the differences between the involved journals as well as the percentage of reproducible papers in the different presentation venues. Furthermore, I look at the authors of the publications and, in particular, their affiliation countries to see where most reproducible papers come from. Finally, I discuss potential reasons for the low reproducibility numbers and suggest possible ways to overcome these obstacles. This paper is reproducible itself, with source code and data available from github.com/tobiasisenberg/Visualization-Reproducibility as well as a free paper copy and all supplemental materials at osf.io/mvnbj.
Qualitative data analysis is widely adopted for user evaluation, not only in the Visualisation community but also related communities, such as Human-Computer Interaction and Augmented and Virtual Reality. However, the data analysis process is often not clearly described and the results are often simply listed in the form of interesting quotes from or summaries of quotes that were uttered by study participants. This position paper proposes an early concept for the use of a researcher as an "Advocatus Diaboli", or devil's advocate, to try to disprove the results of the data analysis by looking for quotes that contradict the findings or leading questions and task designs. Whatever this devil's advocate finds can then be used to reiterate on the findings and the analysis process to form more suitable theories. On the other hand, researchers are enabled to clarify why they did not include this in their theory. This process could increase transparency in the qualitative data analysis process and increase trust in these findings, while being mindful of the necessary resources.
Trust is fundamental to effective visual data communication between the visualization designer and the reader. Although personal experience and preference influence readers' trust in visualizations, visualization designers can leverage design techniques to create visualizations that evoke a "calibrated trust," at which readers arrive after critically evaluating the information presented. To systematically understand what drives readers to engage in "calibrated trust," we must first equip ourselves with reliable and valid methods for measuring trust. Computer science and data visualization researchers have not yet reached a consensus on a trust definition or metric, which are essential to building a comprehensive trust model in human-data interaction. On the other hand, social scientists and behavioral economists have developed and perfected metrics that can measure generalized and interpersonal trust, which the visualization community can reference, modify, and adapt for our needs. In this paper, we gather existing methods for evaluating trust from other disciplines and discuss how we might use them to measure, define, and model trust in data visualization research. Specifically, we discuss quantitative surveys from social sciences, trust games from behavioral economics, measuring trust through measuring belief updating, and measuring trust through perceptual methods. We assess the potential issues with these methods and consider how we can systematically apply them to visualization research.
Gaining insight is considered one of the relevant purposes of visual data exploration, yet studies that categorize insights are rare. This paper reports on a study to understand if the categorization model used to describe insights and personality factors affect insight-based evaluations’ findings. Participants completed a set of tasks with three hierarchical visualizations and then reported what insights they could gather from them. Results show that the insight categorization taxonomies produce different descriptions of insights based on the same corpus of responses. In addition, our findings suggest that the openness to experience trait positively influences the number of reported insights. Both these factors may create obstacles to the design of insight-based evaluations and, consequently, should be controlled in the experimental design. We discuss the study implications, lessons learned, and future work opportunities.
Appropriate evaluation and experimental design are fundamental for empirical sciences, particularly in data-driven fields. Due to the successes in computational modeling of languages, for instance, research outcomes are having an increasingly immediate impact on end users. As the gap in adoption by end users decreases, the need increases to ensure that tools and models developed by the research communities and practitioners are reliable, trustworthy, and supportive of the users in their goals. In this position paper, we focus on the issues of evaluating visual text analytics approaches. We take an interdisciplinary perspective from the visualization and natural language processing communities, as we argue that the design and validation of visual text analytics include concerns beyond computational or visual/interactive methods on their own. We identify four key groups of challenges for evaluating visual text analytics approaches (data ambiguity, experimental design, user trust, and "big picture" concerns) and provide suggestions for research opportunities from an interdisciplinary perspective.
Population Health Management (PHM) relies on the analysis of data from several sources to account for the complex interaction of factors that contribute to the health and well-being of a population, while considering biases and inequalities across sub-populations. Visualisation is emerging as an essential tool for insight generation from data shared and linked across services including healthcare, education, housing, policing, etc. However, visualisation design is challenged by poor data connectivity and quality, high dimensionality and complexity of real-world routinely collected data, in addition to the heterogeneity of users’ backgrounds and tasks. The Creative Visualisation Opportunities (CVO) framework provides a structured approach for working with diverse communities of visualisation stakeholders and defines a set of participatory activities for the effective elicitation of requirements and visualisation design alternatives. We conducted three workshops, applying variations of the CVO framework, with over one hundred participants from the PHM domain, including clinicians, researchers, government and private sector representatives, and local communities. In this paper, we present the results of preliminary analysis of these activities and report on the perceived impact of visualisation in this domain from a stakeholders’ perspective. We report real-world successes and limitations of applying the framework in different formats (through online and in-person workshops), and reflect on lessons learned for task analysis and visualisation design in the PHM domain.