Environmental Health Crises (EHC) are increasing rapidly in their occurrence, intensity, and global spread. Humanity is facing global warming, land-use changes, biodiversity loss, and the interconnected spread of infectious diseases and other EHC on an unprecedented scale. Public understanding plays a crucial role in instigating human action in dealing with these worldwide EHC and mitigating their destructive consequences. The example of the COVID-19 crisis has shown that the world at large faces many obstacles related to public understanding and communication of EHC. Current approaches to creating public understanding, which stem from disparate fields of study, are unable to cope with the often divisive prevalent narratives and representations of the complex multidisciplinary nature of these crises. This is the case, for instance, of media narratives surrounding the causes of global warming and the handling of pandemics. In this paper, we describe the CHRYSES project, which aims to investigate the interplay between myths and science in the ways our societies conceptualise and represent EHC, and to utilise maps to unify the corresponding perspectives and approaches of myths and science to enhance public understanding of such global crises.
Alzheimer's Disease (AD) detection from spoken language has advanced considerably with the rise of large pre-trained speech and language models, including wav2vec 2.0, BERT, generative LLMs, and emerging audio-LLMs. These models often report higher performance on shared AD detection benchmarks than handcrafted-feature pipelines, while reducing the need for manual feature design. Yet their advantages remain difficult to interpret: pre-trained representations are largely a black box, benchmark gains are not always statistically robust, and performance can vary substantially across languages, elicitation protocols, and clinical task definitions. This survey proposes a taxonomy of how large pre-trained speech and language models have been used for AD detection and progression-related tasks. We define Large Pre-trained Models (LPMs) as Transformer-based models trained with self-supervised or weakly supervised objectives, distinguishing them from earlier non-Transformer pre-trained representations. We then organize the literature into six usage strategies: direct application, fine-tuning, transfer learning, multitask learning, multilingual modeling, and generative-LLM-based assessment. Across these categories, we synthesize evidence from shared benchmarks and recent studies to answer four questions: 1) what taxonomy captures how LPMs are used for AD detection; 2) what progress and advantages do LPMs offer for AD detection; 3) what challenges and limitations do LPMs exhibit; and 4) what future opportunities and open tasks they present.
We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. Using a few minutes of voice recordings from 111 participants in the PsyVoiD database, we evaluated 12 instruction-tuned LLMs, including Llama-3 (8B, 70B), Ministral, Mistral, Gemma-2-9B, Gemma-3 (1B, 4B, 27B), Phi-4, DeepSeek (Qwen and Llama), and QwQ-Preview. A domain-informed prompt was developed in collaboration with experts in clinical psychology and linguistics. Results show that LLMs can extract semantically meaningful cues from spontaneous speech, achieving Spearman correlations of up to 0.8 on 80% of the data. Additionally, to enhance explainability, we conducted statistical analyses to characterise prediction variability and systematic biases, alongside keyword-based word cloud analyses to highlight the linguistic features driving the models' predictions.
Maps continue to play a vital role in our modern digital age. They help us explore the physical world through their virtual representations, and by doing so allow us to perform a wide range of tasks in both physical and digital environments. This is despite the fact that most people have not been trained to read and understand maps, and nor do they know that maps are ultimately subjective representations of the world around us. Although most users consider maps as an objective form of visualization of the real world without realising the subjective biases inherent in their design and representation, map-based interfaces and interactions are still seen as effective and usable by ordinary – non-expert – users. The main aim of this workshop was to provide a collaborative informal venue for its participants to showcase their latest design, research, or professional work on map-based interfaces and interactions, and to share their diverse practical expertise, best practices, learnings, and experiences with others. In particular, this workshop focused on the development of more effective maps and map-like visualizations in supporting ordinary users in performing their desired tasks. Here, we summarise the contributions made to this workshop by its participants, and provide an overview of its process, activities and outcomes.
Background:Large language models (LLMs) are increasingly demonstrating the potential to reach human-level performance in generating clinical summaries from patient-clinician conversations. LLMs are usually evaluated against clinical summaries that focus mainly on patients' biology and not on their biography (eg, preferences, values, wishes, and concerns). To achieve patient-centered care, artificial intelligence clinical summarization must incorporate patient-centered domains, implemented through patient-centered summaries (PCSs). Objective:This study aimed to develop a framework to generate PCS that capture patients' values, preferences, and wishes while ensuring clinical utility for clinicians, and assess if current open-source LLMs can achieve human-level performance in generating PCS. Methods:We developed a 4-step mixed methods process to define and evaluate PCS. First, 2 patient and public involvement and engagement groups were convened in the United Kingdom (10 patients and 8 clinicians), who participated in semistructured interviews exploring what personal and contextual information should be included in clinical summaries and how it should be structured for clinical use. Second, findings were translated into an annotation guideline, which was used by 8 clinician annotators to generate gold standard PCS from 88 transcribed patient-clinician consultations about the management of atrial fibrillation. Third, 16 consultations were used to iteratively develop and refine a prompt aligned with the annotation guideline. Finally, 5 LLMs (Llama-3.2-3B [Meta AI], Llama-3.1-8B [Meta AI], Mistral-8B [Mistral AI], Gemma-3-4B [Google DeepMind], and Qwen3-8B [Alibaba]) generated summaries from 72 consultations using zero-shot and few-shot prompting, which were evaluated against gold standard PCS using ROUGE-L (Recall-Oriented Understudy for Gisting Evaluation-Longest Common Subsequence) and BERTScore (Bidirectional Encoder Representations from Transformers Score) and assessed for correctness, completeness, conciseness, patient-centeredness, and fluency. Results:Patients emphasized that summaries should include (1) lifestyle routines and daily functioning as indicators of independence or disruption; (2) the presence and role of social support systems, especially during crises; (3) recent life events or stressors, such as trauma, loss, or caregiving demands; and (4) care preferences, values, and communication styles that provide meaning or reflect autonomy. Clinicians sought summaries that included a concise functional baseline, psychosocial context, and emotional cues, preferably in a structured, clinically digestible format. In the 72 consultations (mean age 70, SD 11 y; 32/72, 44.4% female), the best zero-shot performance was observed with Mistral-8B (ROUGE-L 0.189) and Llama-3.1-8B (BERTScore 0.673). The best few-shot prompting was found with 3 examples using Llama-3.1-8B (ROUGE-L 0.206 and BERTScore 0.683). Conclusions:The open-source LLMs we evaluated did not achieve human-level performance in generating PCSs. Without task-specific fine-tuning, current open-source LLMs cannot reach human-level performance in this task. Our framework serves as an innovative guideline for developing gold standard PCS for artificial intelligence clinical tasks.
Abstract Speech biomarkers could form a critical step in improving the accessibility, scalability, and early detection of Alzheimer's disease and related dementias. However, ethical and practical challenges remain across regulatory and cultural contexts. In this paper, we briefly review the challenges in adopting speech biomarkers, relate our experiences globally to recent advances in the neurodegenerative field, and consider how speech assessments could be integrated into clinical care. Insights from high‐ and low‐ and middle‐income countries (Taiwan, Ghana, Colombia, Brazil, Greece, UK, and US) demonstrate the potential impact of speech technology. While there are common benefits, risks, incentives, and hurdles, many aspects are specific to country or region. There is a need for more speech data sets (particularly in languages other than English), standardization in data collection and analysis, stronger collaboration between machine learning and neurodegenerative disease experts, unified privacy regulation, and finally, a consensus on the clinical interpretation of speech biomarker data.
LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability across three dimensions: intra-model consistency, ASR robustness, and evidence faithfulness. We evaluate three LLMs (Phi-4, Gemma-2-9B, and Llama-3.1-8B) on 111 English-speaking participants using ground-truth transcripts and three Whisper ASR variants (Large, Medium, Small), with three independent runs per model-condition pair. We find that (i) Phi-4 and Gemma-2-9B achieve excellent intra-model consistency (ICC > 0.89) with minimal degradation under ASR; (ii) Llama-3.1-8B shows ASR-fragile consistency, with ICC dropping from 0.82 to 0.36 at 10
Alzheimer's disease (AD) profoundly affects motor control and cognitive functions, often resulting in impaired speech characteristics such as vocal clarity, emotional expressiveness, and prosodic richness. To detect such abnormalities, we examine the role of three specific acoustic features to differentiate participants with AD from healthy controls (HC) and study the association between these acoustic features and global cognition. Speech data from 237 participants (115 HC, 110 AD) in the ADReSS-M dataset, collected during the “Cookie Theft” picture description task, were analyzed. This dataset has been matched for age and gender by propensity score to prevent bias. The HC group averaged 66.4 years (SD: 6.64), and the AD group averaged 69.4 years (SD: 6.92), with significantly lower Mini-Mental State Examination (MMSE) scores in AD (AD: 17.9; HC: 29.0). Features quantifying vocal clarity, articulatory precision (spectral contrast), vocal tone (pitch mean), and prosodic variability (pitch standard deviation) were extracted. Group differences in these features were assessed using t-tests, and Pearson correlation analyses were conducted to examine associations between acoustic measures and MMSE scores. There were significant differences between AD and HC groups for spectral contrast (t(235) = 4.26, p <0.0001), pitch mean (t(235) = 3.54, p = 0.0005), and pitch standard deviation (t(235) = 3.62, p = 0.0004). Cohen's d values for these features ranged from -0.5 to -0.6, indicating medium effect sizes, with lower values observed in the AD group. We also found a significant correlation ( p <0.01) between MMSE scores and each of the features (Pearson's r = 0.22 for pitch mean, 0.23 for pitch standard deviation, and 0.25 for spectral contrast). This preliminary study highlights the physiological basis of altered speech patterns in AD and their diagnostic relevance. Future work will focus on refining preprocessing algorithms and incorporating advanced feature extraction methods to enhance the effect sizes and correlation for AD detection and cognitive assessment.
Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram Transformer (AST) to detect depression using long-duration speech instead of short segments, along with a novel interpretation method that reveals prediction-relevant acoustic features for clinician interpretation. Our experiments show the proposed model outperforms a segment-level AST, highlighting the impact of segment-level labelling noise and the advantage of leveraging longer speech duration for more reliable depression detection. Through interpretation, we observe our model identifies reduced loudness and F0 as relevant depression signals, aligning with documented clinical findings. This interpretability supports a responsible AI approach for speech-based depression detection, rendering such tools more clinically applicable.
The resurgence of respiratory syncytial virus (RSV) epidemics after the COVID-19 pandemic highlights the critical need for high-quality surveillance dashboards to monitor RSV activity and effectively inform the public. In this study, we developed a prototype of an online dashboard, RSVdash, for visualizing RSV data using the Shiny package in R. This study demonstrates how a surveillance dashboard can effectively convey RSV-related information, raise public awareness of RSV, and enhance preparedness and management of RSV outbreaks.
Rats are gregarious rodents who naturally live in diverse social groups and communicate in part through ultrasonic vocalisations (USVs). USVs encode significant information about affective state and play an important role in social behaviour. Monitoring USVs is a non-invasive method of adding richness to data in a variety of experimental paradigms. However, manual analysis of USVs requires a significant amount of human effort. We propose a new method for automatic classification of USVs which could help automate analysis and thus reduce human input. The proposed method introduces a novel approach to USV representation called deep supervectors (DSV), which combines diverse deep embeddings feature sets extracted through our active data representation (ADR) method. The DSV method is evaluated on a multiclass recognition task involving 14 different types of rodent calls. The performance of DSV is compared to that obtained by state-of-the-art deep embeddings (alexNet, googleNet, squeezeNet and resNet). The proposed method achieves an Unweighted Average Recall (UAR) of 32.80% and outperforms both customs (8.07%) and deep embeddings (29.94%). Deep supervectors outperform pre-trained DNN in 6 out of 8 cases and fusion of the top six DSV with the top two deep embeddings improves the UAR to 37.22% in this challenging classification task, showing an improvement of 29% over the majority guess of 8%. By combining related classes, the UAR will reach 51.00%. The context of this application is the automation of the process for life-sciences laboratory technicians and scientists.
Dementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recognition of Cognitive Decline through Spontaneous Speech (PROCESS) Signal Processing Grand Challenge invites participants to focus on early-stage dementia detection. We provide a new spontaneous speech corpus for this challenge. This corpus includes answers from three prompts designed by neurologists to better capture the cognition of speakers. Our baseline models achieved an F1-score of 55.0% on the classification task and an RMSE of 2.98 on the regression task.
As part of the Intangible Cultural Heritage, Bridging the Past, Present and Future (INT-ACT) project, we investigate how Extended Reality (XR) technologies can meaningfully engage users with cultural content by monitoring their physiological and affective responses. While immersive XR systems offer new ways of exploring heritage, their impact on users’ internal states remains underexplored. In this study, we present a multimodal experimental setup using EmotiBit, a wearable biosensing platform, to monitor real-time physiological signals during cultural XR interaction. Participants were evaluated across three activities representing varying cognitive and sensory loads: immersive interaction with the INT-ACT XR Demonstrator, composing work emails (a low-stimulation control task), and passive movie watching. Our aim is to quantify how cultural XR experiences influence biosignals such as electrodermal activity and heart rate. The findings reveal distinct physiological patterns across conditions, suggesting that biosignal monitoring can inform the design of adaptive XR environments that are responsive to user states. This work contributes to INT-ACT’s broader objective of creating intelligent, inclusive, and emotionally resonant cultural heritage experiences.
People with severe mental illness have high rates of obesity, type 2 diabetes, and cardiovascular disease. Emerging evidence suggests that metabolic dysfunction may be causally linked to the risk of severe mental illness. However, more research is needed to identify reliable metabolic markers which may have an impact on mental health outcomes, and to determine the mechanisms behind their impact. In the METPSY research study, we will investigate the relationship between metabolic markers and clinical outcomes of severe mental illness in young adults. We will recruit 120 young adults aged 16–25 years living in Scotland with major depressive disorder, bipolar disorder, schizophrenia, or no severe mental illness (controls) for a prospective observational study. We will assess clinical symptoms at three in-person visits (baseline, 6 months, and 12 months) using the Structured Clinical Interview for DSM-5, and collect blood samples at each of these visits for agnostic profiling of metabolic biomarkers through an untargeted metabolomic screen, using the rapid hydrophilic interaction liquid chromatography ion mobility mass spectrometry method (RHIMMS). Participants will also complete remote assessments at 3 and 9 months after the baseline visit: Ecological Momentary Assessments to measure mental health, wrist actigraphy to measure rhythms of rest and activity, and continuous glucose monitoring to measure metabolic changes. Throughout the 12-month enrolment period, we will also measure objective markers of sleep using a radar sleep monitor (Somnofy). Using advanced statistical techniques and machine learning analysis, we will seek to better understand the mechanisms linking metabolic health with mental health in young adults with schizophrenia, bipolar disorder, and severe depression. Clinical trial number: Not applicable
Parkinson disease (PD) is a neurodegenerative disease that can impair speech production. In PD, speech impairments are typically characterised by reduced vocal loudness, monotone speech, and distorted articulation. We hypothesise that speech impairment in PD could alter energy decay patterns in speech, leading to changes in reverberation time for 60 dB of decay in sound pressure (RT60). To determine whether the reverberation characteristics can help to differentiate between healthy and PD-affected speech, RT60 was derived from speech segments extracted from recordings in Castilian Spanish and English using the Schroeder’s Integration algorithm. A t-test analysis revealed significant differences between the PD group and healthy controls in both languages. In addition, classification tasks were performed based on the descriptive statistics and temporal features of RT60 for each participant, using logistic regression, decision tree and random forest classifiers. The random forest classifier provided the best accuracy (0.77) on the Castilian Spanish data and was tested on the English data, yielding an accuracy of 0.67. This study demonstrates that RT60 can be effectively repurposed for PD detection.
Depression and Alzheimer's dementia frequently co-occur, with depression often occurring at early stages of Alzheimer's disease (AD). We investigate whether knowledge of a person's depression score gained through analysis of acoustic features of the person's speech can help detect AD. We analysed data from 239 participants, where n = 162 individuals had a clinical diagnosis of AD, to explore correlations between depression scores, cognitive test scores and AD status. First, we created multilayer perceptron models for depression assessment using speech features extracted through our active data representation method and wav2vec features. We then used outputs from the depression models as predictors in our Alzheimer's detection models. Depression scores were strongly correlated with Alzheimer's diagnosis in this data set (Spearman's correlation r = -0.610 and r = -0.536, p < 0.01 for correlations between depression scores and diagnosis and cognitive test scores, respectively). We found that depression scores could predict AD fairly accurately (Acc = 0.84), and that depression could be detected in speech with moderate accuracy (Acc = 0.79). Using these predictions and acoustic features, our dementia model reached Acc = 0.70. We also investigated classification of AD from speech in the presence and absence of depression, obtaining the same level of accuracy for this imbalanced prediction task.
A surrogate outcome is a substitute measure for a patient-final outcome; a direct measure of how a person feels, functions and/or survives. Intermediate outcomes can be defined as standardised functionality measures. Surrogate outcomes are increasingly used in clinical trials to accelerate drug approvals, yet the reporting of their use as primary endpoints in study protocols remains inconsistent. This is an issue because the primary endpoint is used to determine treatment efficacy and thus is crucial for interpreting results and making regulatory decisions. We considered various supervised learning techniques for detecting surrogate outcome usage using a three-class text classification approach. We evaluated this approach on two nervous system trial protocol datasets (EUCT-NS and NS-HRA). This study serves as a proof of concept for the potential of machine learning to automate the detection of primary surrogate endpoint use in clinical trial protocols, thus addressing gaps in reporting.
Quantitative mass spectrometry based proteomics data is high dimensional and when enriched with biological information for every measured protein, it widens the scope of analysis for end users. In silico analyses where data enrichment aids analysis goals can be further enhanced with tailored visualisations to guide differential pattern activity in proteomic data. In this study, we employed the core analysis tool of a single software platform, Qiagen’s Ingenuity Pathway Analysis (IPA), which is widely used in omic analyses, to explore comparative profiling of multi-source proteomic data. The central aim was to construct added visualisations from the wealth of exportable features available in IPA to provide further visual guides for users to utilise the built in tool metrics and domain based understanding in relation to their data.