Understanding how psychiatric patients subjectively experienced a clinical conversation is important for feedback and alliance-related process monitoring. While interviewers form post-session judgments about patient experience, these judgments do not always match patients' self-reports. Automatic approaches for predicting perceived interaction quality from conversation have been proposed, but it remains unclear whether such approaches can complement human judgment rather than simply replicate it. To address this gap, we evaluate a clinician-support framework in which post-session interviewer ratings are combined with automatic language-based predictions to estimate patient-reported interaction quality in free clinical interviews. We assess this integration across multiple standard model types, including Ridge, SVR, MLP, GRU, and BiLSTM, all trained on sentence embeddings extracted from dyadic transcripts of 107 free conversations between psychiatric patients and interviewers. Our results show that combining interviewer judgments with model predictions through simple averaging yields the strongest overall performance. The interviewer-only baseline reached a Pearson correlation of 0.365. Among fully automatic models, Ridge achieved the strongest Pearson correlation (r = 0.286), while BiLSTM achieved r = 0.270. The strongest result was obtained by BiLSTM interviewer integration (r = 0.403). Our findings suggest that automatic language analysis and interviewer judgment capture complementary aspects of patient experience and that their combination provides a more accurate approximation of the patient's own report than either source alone.
Sensorineurale Hörstörungen (SNH) sind einer der häufigsten Gründe von Schwerhörigkeit. Eine besondere Form der SNH ist die versteckte Schwerhörigkeit (HHL) bei subjektiver Normakusis. Aktuelle Forschungsergebnisse deuten darauf hin, dass bei diesen Patienten eine reduzierte Welle I im gemittelten Signal der Hirnstammaudiometrie (ABR) vorliegt. Da die Mittelungstechnik für Latenzjitter und Amplitudenhöhenvariation unempfindlich ist, wird zur Beantwortung einer weitergehenden Fragestellung eine Single-Sweep-Analyse benötigt. Insgesamt wurden 14 Mäuse mit signifikant unterschiedlichen Kalziumströmen in den IHC bei normvarianter Hörschwelle für die Analyse genutzt, um im Zeitfenster der Welle I 4 neue Parameter aus den Single Sweeps zu berechnen. Diese dienten gleichzeitig zur Beschreibung einer neuralen Aktivierungsfunktion (NAV). Alle neuen Parameter zeigen im Stimulus-abhängigen Vergleich beim Wildtyp signifikante bzw. hoch signifikante Unterschiede auf. Bei der transgenen Maus sind es signifikante bzw. nicht signifikante Unterschiede. In der neuralen Aktivität des Ruhe-EEGs zwischen der Wildtypmaus und der Mutante gibt es einen signifikanten Unterschied. Die Mittelwerte der Wellenamplituden bei der Wildtypmaus verhalten sich dabei kontrovers. Unter Ausnutzung von Single Sweeps werden innovative Ergebnisse dargestellt, die bis dato so nicht bekannt sind. Offensichtlich ist nicht die Amplitudenhöhe der Welle I für die Funktion der IHC alleinig verantwortlich, sondern besonders noch die neuen Parameter. Für die Diagnose von Hörstörungen mit normvarianter Hörschwelle scheinen die neuen Parameter hervorragend geeignet zu sein.
Estimating momentary conversational engagement is central to assistive, socially aware AI systems, yet models are typically trained and evaluated within a single domain, limiting real-world robustness. The MultiMediate '25 challenge advances engagement estimation to more challenging, cross-cultural, and multi-domain settings. Building on prior challenge editions, we expand beyond NOXI as the sole training source by introducing NOXI-J, a new multilingual corpus covering Japanese and Chinese interactions, enabling both training and evaluation in diverse linguistic contexts. Although NOXI-J conceptually extends NOXI, we treat it as a distinct domain because linguistic, cultural, capture, and annotation differences induce measurable distribution shifts. In this paper, we present new annotations, precomputed multi-modal features (visual, vocal, and verbal), baseline evaluations, and an analysis of the best performing challenge solutions. Beyond accuracy, we quantify fairness using Conditional Demographic Disparity for gender and language. Our baselines confirm strong in-domain performance (e.g., paralinguistic eGeMAPS and video-transformer features) and reveal notable cross-domain drops, underscoring the challenge of cultural, linguistic, and interactional shifts. Fairness analyses indicate generally small discrepancies for our baselines. We observe the largest disparities for the proposed challenge solutions on the Chinese language test set. All annotations, features, code, and leaderboards are made publicly available to foster sustained progress on robust and fair engagement estimation.
Introduction In daily practice, measurements of auditory brainstem responses (ABR) are used to objectively assess the hearing ability. The basic neural activity of the patient"s brain interferes with the evoked potentials resulting in a low signal-to-noise ratio. Up to 2000 stimuli are applied and the corresponding results are averaged to enable a visual annotation of the potentials. The potentials are divided into characteristic waves. At present it is not possible to automatically distinguish normal outcomes from anomalous results, as there are no clear boundaries to pathological conditions.
Accurate utterance classification in motivational interviews is crucial to automatically understand the quality and dynamics of client-therapist interaction, and it can serve as a key input for systems mediating such interactions. Motivational interviews exhibit three important characteristics. First, there are two distinct roles, namely client and therapist. Second, they are often highly emotionally charged, which can be expressed both in text and in prosody. Finally, context is of central importance to classify any given utterance. Previous works did not adequately incorporate all of these characteristics into utterance classification approaches for mental health dialogues. In contrast, we present M3TCM, a Multi-modal, Multi-task Context Model for utterance classification. Our approach for the first time employs multi-task learning to effectively model both joint and individual components of therapist and client behaviour. Furthermore, M3TCM integrates information from the text and speech modality as well as the conversation context. With our novel approach, we outperform the state of the art for utterance classification on the recently introduced AnnoMI dataset with a relative improvement of 20% for the client- and by 15% for therapist utterance classification. In extensive ablation studies, we quantify the improvement resulting from each contribution.
Brain-Computer Interface (BCI) systems represent an innovative approach to human-computer interaction, enabling users to control devices and interact with technology solely through brain activity. This study investigates the feasibility and potential of non-invasive EEG-based BCI for elevator control, addressing two primary research questions: 1) Can a person reliably control an elevator through a BCI system? and 2) What are the usability and user experience outcomes of such a system? We integrated a Muse headset with a remote-controllable elevator system using an iPhone as the interface over a local network. This setup allowed users to operate the elevator using blinking, jaw clenching, and mental focussing as triggers. Performance, accuracy, and user experience were evaluated through experiments involving 50 participants aged 12 to 60. Usability was measured with the System Usability Scale (SUS) questionnaire along with additional feedback questions. Key findings indicate that the system achieved an average SUS score of 80.3, reflecting excellent usability on the adjective rating scale. Moreover, 94
Einleitung In der täglichen Routine werden zur Feststellung des objektiven Hörvermögens BERA-Messungen durchgeführt. Über EEG Elektroden werden dabei akustisch ausgelöste Potentiale gemessen. Aufgrund hoher neuraler Grundaktivitäten werden bis zu 2.000 Reize appliziert, um im gemittelten Signal die Morphologie des Wellenmusters visuell zu annotieren. Eine Abgrenzung von auffälligen Wellenerhebungen ist aktuell automatisiert nicht möglich und es fehlen Grenzen für pathologische Zustände.
Estimating the momentary level of participant's engagement is an important prerequisite for assistive systems that support human interactions. Previous work has addressed this task in within-domain evaluation scenarios, i.e. training and testing on the same dataset. This is in contrast to real-life scenarios where domain shifts between training and testing data frequently occur. With MultiMediate'24, we present the first challenge addressing multi-domain engagement estimation. As training data, we utilise the NOXI database of dyadic novice-expert interactions. In addition to within-domain test data, we add two new test domains. First, we introduce recordings following the NOXI protocol but covering languages that are not present in the NOXI training data. Second, we collected novel engagement annotations on the MPIIGroupInteraction dataset which consists of group discussions between three to four people. In this way, MultiMediate'24 evaluates the ability of approaches to generalise across factors such as language and cultural background, group size, task, and screen-mediated vs. face-to-face interaction. This paper describes the MultiMediate'24 challenge and presents baseline results. In addition, we discuss selected challenge solutions.
AIMS:To create and assess the performance of an artificial intelligence-based image analysis tool for the measurement and quantification of the corneal neovascularisation (CoNV) area. METHODS:Slit lamp images of patients with CoNV were exported from the electronic medical records and included in the study. An experienced ophthalmologist made manual annotations of the CoNV areas, which were then used to create, train and evaluate an automated image analysis tool that uses deep learning to segment and detect CoNV areas. A pretrained neural network (U-Net) was used and fine-tuned on the annotated images. Sixfold cross-validation was used to evaluate the performance of the algorithm on each subset of 20 images. The main metric for our evaluation was intersection over union (IoU). RESULTS:The slit lamp images of 120 eyes of 120 patients with CoNV were included in the analysis. Detections of the total corneal area achieved IoU between 90.0% and 95.5% in each fold and those of the non-vascularised area achieved IoU between 76.6% and 82.2%. The specificity for the detection was between 96.4% and 98.6% for the total corneal area and 96.6% and 98.0% for the non-vascularised area. CONCLUSION:The proposed algorithm showed a high accuracy compared with the measurement made by an ophthalmologist. The study suggests that an automated tool using artificial intelligence may be used for the calculation of the CoNV area from the slit-lamp images of patients with CoNV.
Human emotions are often not expressed directly, but regulated according to internal processes and social display rules. For affective computing systems, an understanding of how users regulate their emotions can be highly useful, for example to provide feedback in job interview training, or in psychotherapeutic scenarios. However, at present no method to automatically classify different emotion regulation strategies in a cross-user scenario exists. At the same time, recent studies showed that instruction-tuned Large Language Models (LLMs) can reach impressive performance across a variety of affect recognition tasks such as categorical emotion recognition or sentiment analysis. While these results are promising, it remains unclear to what extent the representational power of LLMs can be utilized in the more subtle task of classifying users' internal emotion regulation strategy. To close this gap, we make use of the recently introduced Deep corpus for modeling the social display of the emotion shame, where each point in time is annotated with one of seven different emotion regulation classes. We fine-tune Llama2-7B as well as the recently introduced Gemma model using Low-rank Optimization on prompts generated from different sources of information on the Deep corpus. These include verbal and nonverbal behavior, person factors, as well as the results of an in-depth interview after the interaction. Our results show, that a fine-tuned Llama2-7B LLM is able to classify the utilized emotion regulation strategy with high accuracy (0.84) without needing access to data from post-interaction interviews. This represents a significant improvement over previous approaches based on Bayesian Networks and highlights the importance of modeling verbal behavior in emotion regulation.
Accurate estimation of intentions is a prerequisite in a non-verbal human-machine collaborative search task. Electroencephalography (EEG) based intent recognition promises a convenient approach for recognizing explicit and implicit human intentions based on neural activity. In search tasks, implicit intent recognition can be applied to differentiate if a human is looking at a specific scene, i.e., Navigational Intent, or is trying to search a target to complete a task, i.e., Informational Intent. However, previous research studies do not offer any robust mechanism to precisely differentiate between the intents mentioned above. Additionally, these techniques fail to generalize over several participants. Thus, making these methods unfit for real-world applications. This paper presents an end-to-end intent classification pipeline that can achieve the highest mean accuracy of $$97.89 \pm 0.74$$ (%) for a subject-specific scenario. We also extend our pipeline to support cross-subject conditions by addressing inter and intra-subject variability. The generalized cross-subject model achieves the highest mean accuracy of $$96.83 \pm 0.53$$ (%), allowing our cross-subject pipeline to transfer learning from seen subjects to an unknown subject, thus minimizing the time and effort required to acquire subject-specific training sessions. The experimental results show that our intent recognition model significantly improves the classification accuracy compared to the state-of-the-art.
Automatic analysis of human behaviour is a fundamental prerequisite for the creation of machines that can effectively interact with- and support humans in social interactions. In MultiMediate'23, we address two key human social behaviour analysis tasks for the first time in a controlled challenge: engagement estimation and bodily behaviour recognition in social interactions. This paper describes the MultiMediate'23 challenge and presents novel sets of annotations for both tasks. For engagement estimation we collected novel annotations on the NOvice eXpert Interaction (NOXI) database. For bodily behaviour recognition, we annotated test recordings of the MPIIGroupInteraction corpus with the BBSI annotation scheme. In addition, we present baseline results for both challenge tasks.
Technology readiness levels (TRL), used in technological R&D, also frame European Commission biomedical funding frameworks. Providing an adaptation to biomarker development would improve communication with industry, align stakeholders and professionals, reduce risk of development failure, costs and waste. Here, we aim to align TRLs with the Strategic Biomarker Roadmap (SBR) defined to validate biomarkers for Alzheimer’s disease (AD) and related disorders. We analyzed the NASA definition of TLR phases and a previously proposed adaptation to the biomedical field. We drafted an alignment of TRLs with the SBR Phases and Aims. We refined such alignment through discussion and consensus among coauthors and outlined discrepancies, gaps and SBR limitations relative to the TRL framework, to improve the development process for future biomarkers. The SBR covers all of the 9 TRLs. The demonstration of Analytical Validity entails TRLs_1-4. Specifically, the SBR Phase_1 includes TRLs_1-2; Phase_2 Primary Aim (PA) corresponds to TRL_3, and Phase_2 Secondary Aims (SAs) to TRL_4. Clinical Validity entails TRLs_5-8. Specifically, the SBR Phase_3 corresponds to TRL_5, while TRLs_6-8 correspond respectively to Phase_4 PA, SAs, and full completion. Demonstration of Clinical Utility corresponds to TRL_9. Despite such correspondence, the SBR entails only biomarker diagnostic performance and utility. Aspects like manufacturing and usability, fundamental in the TRL framework as well as in the previous biomedical adaptation, are still lacking. The definition of tools (e.g., Target Product Profile, introduced at TRL 3 in the biomedical adaptation) may help complement the SBR in this direction. The SBR and TRLs follow a consistent sequence despite different pacing and granularity, SBR steps corresponding to one-to-multiple TRLs. Targeting such inconsistencies and upgrading the SBR validation roadmap into a development roadmap, entailing manufacturing, usability, and an integrated procedure to assess cost-effectiveness and plan implementation requires a structured dialogue with industry partners and other stakeholders, like regulators. Such attempt to evolve the SBR from a linear to a multi-dimensional model may help boost the efficiency of translational research and our response to the global priority of dementia.
Purpose Cornea guttata (CG) prevalence post keratoplasty varies from 15 to 18%, with 1 to 2% of the cases presenting with significant negative outcomes. The purpose of this research project is to create a program based on artificial intelligence (AI) that helps with the detection of CG in the donor corneas (DC) in the eye bank. Methods Preoperative corneal endothelial images (PCEI) of patients who underwent keratoplasty were collected and classified into 2 groups according to the postoperative CG grade. Group 1 included healthy corneas and those having mild postoperative CG, while group 2 included corneas with severe postoperative CG. Using previously tested semi-quantitative morphological criteria along with other characteristics such as donor age and lens status, the PCEI were analyzed and used to create and train an AI-based tool for the detection of CG. The underlying concept of the tool compares previous cases with comparable properties to the DC in test. The postoperative CG grades of previous cases similar to the DC in test determine the prediction for its CG grade. Finally, the features and CG grade of the analyzed DC are stored in the database for future use. Results In total, 6221 PCEI belonging to 1078 patients were used to create a transparent and explainable decision support tool for the detection of CG through a hybrid approach combining 2 components. (1) Graphical analytic tools, whereby the PCEI pass multiple OpenCV-based image processing steps including the Watershed transform algorithm. In this step, cell membranes are delineated, and abnormally large cells or cell depleted areas are marked in red. Several other cell representations such as 'honeycomb' representation are created for an enhanced visualization of the endothelial layer (EL). (2) Machine learning (ML) classifiers including Case-Based Reasoning were created to detect CG. Initial experiments showed a performance comparable to humans (4-fold evaluation yielded precision: weighted F1 score:0.93). Conclusion We presented an AI-based program able to facilitate the detection of CG in the DC in the eye bank by comparing the PCEIs with relevant previous cases, using ML classifiers and offering an enhanced visualization of the EL. The evaluation and optimization of this program will follow as the next stage of our project.
This paper and the accompanying demonstration video showcase an interactive counting aid implemented with PARTAS, our personalizable, Augmented Reality-based worker assistance system. PARTAS combines contour-based instructions with a pick-by-projection approach and is particularly designed to be adaptable for people with different cognitive disabilities.
Commercially available assembly systems for industrial settings do not address the needs and requirements of people with disabilities and sheltered workshops. This is due to the complexity of the system’s setup including hardware, software and application in combination with the workers’ and supervisors’ capabilities. In the case of sheltered workshops, resources for implementing functionalities and accessibility requirements push expenses beyond available resources. In this work, we present a prototype of an intuitive, adaptive and cost-efficient worker assistance system utilizing only one RGB camera, one projector and a small single-board computer. The system implements a combination of projected, contour-based instructions and pick-by-projection functionality with the goal of maximizing the instructions’ affordance. An interdisciplinary team including workers from a sheltered workshop, their caregivers, technicians and a psychologist followed an iterative user-centered design methodology approach based on two personas with different cognitive disability profiles. Finally, we conducted a user study based on eight participants matching the main persona. All participants showed a strong learning rate and performed successfully completely new assembly tasks after a short training phase with the proposed system.
In recent years, more and more AI models and algorithms get used in previously uncharted domains. The medical domain is one of them and already shows a significant usage of AI methods, for example computer vision algorithms for the analysis of medical imagery. The field of ophthalmology studies medical conditions relating to the eye. One of those conditions, Cornea guttata, can be identified by analysing post mortem microscope images of the donor's cornea endothelium, which needs to be done manually by a skilled professional. To help facilitate this analysis, this paper proposes a hybrid Decision Support System that combines computer vision methods and AI classifiers to guide the decision of the clinicians. By conducting a UX-driven study with professionals from an eye bank, we show that our Decision Support System is able to help users with the classification of Cornea guttata in microscope images. Moreover, the system was able to boost the agreement between two professionals classifying the same cases. The implemented classifiers showed a higher performance compared to the human baseline and the combination of human expertise and AI classifiers detected most of the guttata cases.
Identifying objective and reliable markers to tailor diagnosis and treatment of psychiatric patients remains a challenge, as conditions like major depression, bipolar disorder, or schizophrenia are qualified by complex behavior observations or subjective self-reports instead of easily measurable somatic features. Recent progress in computer vision, speech processing and machine learning has enabled detailed and objective characterization of human behavior in social interactions. However, the application of these technologies to personalized psychiatry is limited due to the lack of sufficiently large corpora that combine multi-modal measurements with longitudinal assessments of patients covering more than a single disorder. To close this gap, we introduce Mephesto, a multi-centre, multi-disorder longitudinal corpus creation effort designed to develop and validate novel multi-modal markers for psychiatric conditions. Mephesto will consist of multi-modal audio-, video-, and physiological recordings as well as clinical assessments of psychiatric patients covering a six-week main study period as well as several follow-up recordings spread across twelve months. We outline the rationale and study protocol and introduce four cardinal use cases that will build the foundation of a new state of the art in personalized treatment strategies for psychiatric disorders.