The principle of visual analytics (VA) is to provide integrated workflows where human-centric processes (e.g., visualization and interaction) and machine-centric processes (e.g., statistics and algorithms) complement each other. To implement this principle in practice, it is necessary to reason about the trade-offs among different processes and make optimal use of them in a workflow. Building on an existing ontology of the methodology for analyzing such trade-offs information-theoretically and for optimizing VA workflows systematically, we investigate ways to transform this methodology from theory to practice. In particular, we adopted the action research method. Through case studies in different application domains, VA researchers with different background knowledge and experiences offered their answers to several hypotheses about using the methodology in practice and proposed ways forward. In this paper, we present our collective analysis, the strengths and feasibility of this theory-based methodology, as well as the obstacles to its broad deployment in practice. To address these challenges, we outline a roadmap to remove such obstacles.
The rapid proliferation of social media has created new data stemming from users’ thoughts, feelings, and interests. However, this unprecedented growth has led to the widespread dissemination of misinformation—deliberately or inadvertently false content that can trigger dangerous societal ramifications. Visual analytics combines advanced data analytics and interactive visualizations to explore data and mine insights. This article introduces the Social Media Analytics and Reporting Tool (SMART) 2.0, detailing its application in tracking misinformation on social media. An updated version of its predecessor, SMART 2.0 enables analysts to conduct real-time surveillance of social media content along with complementary data streams, including weather patterns, traffic conditions, and emergency service reports. SMART 2.0 offers enhanced capabilities like map-based, interactive, and AI-powered features that enable researchers to visualize and understand situational changes by assessing public social posts and comments. As a misinformation classification and tracking case study, we collected public, geo-tagged tweets from multiple cities in the UK during the 2024 riots. We showcased the effectiveness of SMART 2.0’s misinformation detection and tracking capabilities. Our findings show that SMART 2.0 effectively tracks and classifies misinformation using a human-in-the-loop approach.
Quantitative radiological tools to assess acute ischemic stroke (AIS) survivor's functional outcomes are limited. While conventional qualitative scoring systems exist, they are limited by inter-rater variability due to subtle ischemic changes. Post-intervention estimation of final cerebral infarct volume (CIV) followed by assessment is a known objective radiological determinant of functional outcomes. Hence, this study aims to develop and evaluate a composite radiological tool using radiomic imaging markers to predict long-term functional risk in AIS survivors. The dataset consists of 50 AIS patients' clinical and radiological information. First, Alberta stroke programme early CT score (ASPECTS) and posterior circulation-ASPECTS regions were annotated on scans followed by mapping them with CIV. Multiple volumetric parameters were extracted, including TBV, CIV, and the proportion of CIV to TBV. These raw features were used to compute percentage volumes and perform summation and proportion analyses w.r.to ASPECTS regions. Premorbid and clinical features were also converted to meaningful representations. Principal Component Analysis and Recursive Feature Elimination (RFE) were employed to identify optimal feature sets. Finally, different combinations of composite features were utilized to train classification algorithms. Promising results were achieved with RFE using both feature combination approaches. Support Vector Machine (SVM) on raw features achieved the highest AUC value (0.94 +/- 0.05), while Logistic Regression (LR) attained AUC value (0.93 +/- 0.05). Additional analysis indicated models trained using transformed features perform more consistently and achieve superior overall performance. The study demonstrated the potential of composite tool using radiomic and clinical information to accurately predict AIS patient's risk of long-term functional outcomes.
This article discusses considerations on how visualization can be best positioned to help respond to future pandemics. We examine visualization, along with the corresponding and necessary enabling technologies and platforms, as a tool to facilitate a rapid and effective response to a forthcoming pandemic. We consider challenges in terms of an infrastructure supporting world-wide response, corresponding training and stakeholder engagement, integration of future technologies, and appraisal of such systems. Finally, we discuss how addressing these challenges also helps emergency response beyond infectious diseases.
A comprehensive approach to integrated one health surveillance and response Surveillance data plays a crucial role in understanding and responding to emerging infectious diseases; here, we learn why adopting a One Health surveillance approach to EIDs can help to protect human, animal, and environmental health. Over 75% of emerging infectious diseases (EIDs) affecting humans are zoonotic diseases with animal hosts, which can be transmitted by waterborne, foodborne, vector-borne, or air-borne pathways. (7) Early detection is important and allows for a rapid response through preventive and control measures. However, early detection of EIDs is hindered by several obstacles, such as climate change, which can alter habitats, leading to shifts in the distribution of disease- carrying vectors like mosquitoes and ticks. This can result in diseases such as malaria, dengue fever, and Lyme disease becoming more common in areas with established transmission or spreading to new areas entirely. (4) Environmental changes such as deforestation and urbanization disrupt ecosystems, increasing the likelihood of zoonotic disease spillover from wildlife to humans. In addition to working at the interface of these changes, detection and tracking of EIDs also requires sharing and standardization of complex data and integrating processes across different regions and health systems.
One of the most pressing public policy issues that has involved transdisciplinary research in the field of data science is rapidly detecting widespread misinformation. While data science can pose a lot of potential for solving the big-data problem of misinformation on an automated scale, it likewise requires insights from the field of communications and journalism to define quantifiable features that can assist in more accurate misinformation predictions. Currently, the preeminent tools used for misinformation detection are large language models (LLMs) as they are renowned for their ability to capture the context and meaning of textual data. However, despite advancements in developing effective data science models and tools for identifying misinformation, there are not many available options for evaluating news article content for misinformation potential. This study proposes TRUExT, an explainable, regression-based data tool that integrates multiple communication-based natural language processing (NLP) dimensions with a base LLM to holistically evaluate trustworthiness in news articles. It was found that the Hugging Face LLM RoBERTa with the added NLP dimensions as features was the most effective foundational model after testing multiple LLMs. Furthermore, TRUExT introduced a potential big-data solution to the growing problem of misinformation through research intersecting data science and communications to capture not only the technicality of misinformation data predictions but also certain communication factors in the data. In the future, this tool could likewise be deployed to be used by U.S.-based stakeholders who have an important role in the ongoing information war.
Fake news about coronavirus disease 2019 (COVID-19) can discourage people from taking preventive measures (masks, social distancing), thereby increasing infections and deaths; thus, this study tests whether attributes of users or COVID-19 tweets can distinguish tweets of true news versus fake news. We analyzed 4,165 spell-checked English tweets with a link to 1 of 20 matched COVID-19 news stories (10 true and 10 fake), across the world during 1 year, via computational linguistics and advanced statistics. Tweets with common words, negative emotional valence, higher arousal, greater dominance, first person singular pronouns, third person pronouns or by users with more followers were more likely to be true news tweets. By contrast, tweets with second person pronouns, bald starts, or hedges were more likely to be fake news tweets. Accuracy (F1 score) was 95%. While some tweet attributes for detecting fake news might be universal (pronouns, politeness, followers), others might be topic specific (common words, emotions, hedges).
The vegetation phenology of tallgrass prairie varies yearly, depending on climatic conditions, plant species composition, and location. Modeling time series of vegetation indices (VIs) using climate data can be useful for understanding and predicting how tallgrass prairie will respond to future climate scenarios and for identifying and managing areas of tallgrass prairie that are particularly susceptible to climate-induced changes. Machine or deep learning algorithms can be well-suited to model VIs for phenology studies by identifying patterns and relationships between climatic factors and VIs using historical data. This study evaluated the performance of 12 machine and deep learning algorithms, encompassing a diverse range of algorithmic families, in modeling patterns of the Moderate Resolution Imaging Spectroradiometer-derived enhanced vegetation index (EVI, greenness index) and land surface water index (LSWI) in native tallgrass prairie. The models include linear regression, Bayesian ridge, elastic net, decision tree, random forest, eXtreme Gradient Boosting (XGBoost), support vector regression (SVR), K-nearest neighbors (KNN), artificial neural network (ANN), convolutional neural network (CNN), recurrent neural network (RNN), and long short-term memory (LSTM). Air and soil temperatures showed the highest correlations with EVI (r > 0.77) and LSWI (r > 0.56). The low correlation (r <= 0.23) of EVI and LSWI with contemporaneous rainfall or soil moisture suggests vegetation's delayed response to these factors. The results indicated that ensemble methods like XGBoost and random forest performed best across all three datasets (i.e., training, testing, and validation) for modeling EVI and LSWI. Deep learning models showed varying performance across datasets, and their performance was sub-optimal compared to XGBoost and random forest. The linear regression also showed a moderate performance, while the decision tree performed the weakest overall. The strong performance of XGBoost and random forest highlights the intricate and nonlinear relationship of prairie vegetation with climatic factors. These models' strength lies in capturing such complexities. This study provides insights into the key climatic factors and underlying processes that control the vegetation dynamics of tallgrass prairie ecosystems. Our machine learning models can be a valuable tool for developing new strategies to manage tallgrass prairie ecosystems in the face of climate change.
Applying data science advances in disease surveillance and control Dr. David S. Ebert and Dr. Aaron Wendelboe explain how a cohesive, multidisciplinary, and multi-tiered approach can support a more predictive model in disease surveillance and control. Public health disease surveillance is being conducted in countless settings, including healthcare, vertebrate and invertebrate animals, wastewater, air quality, transportation, and commercial activities, but, attaining the goal of early disease detection has been somewhat elusive. For instance, one of the few key shortcoming of public health preparedness efforts is the insufficient collaboration between multidisciplinary experts, such as data scientists, computer engineers, anthropologists, social scientists, and systems engineers. To address these gaps in knowledge and preparedness, we are responding in a multi-tiered approach with a One Health perspective that will be economically feasible and sustainable. The authors have also engaged a broad set of stakeholders, created broad multidisciplinary teams, are combining relevant data sources in innovative ways that will serve as early indicators, are using advanced technologies for early diagnosis, and advancing analytic methods to maintain high specificity for true event identification.
Presents the welcome message from the VIS 2022 General Chairs.