Medical students increasingly encounter artificial intelligence (AI) and explainable AI (XAI) in clinical training, yet their mental models of diagnostic decision-making and expectations for AI support remain poorly understood. Understanding these expectations is crucial for designing human-centered clinical decision support systems that are pedagogically effective and safe. This study explores medical students' diagnostic processes in their final year of undergraduate training and AI support preferences across four research questions: (1) What information do medical students use to make a diagnosis? (2) In what form would they like (X)AI to support them in this process? (3) What questions do they have for such (X)AI? (4) What form of presentation would they prefer for XAI? Semi-structured interviews (N=14 medical students, ages 24-31) were analyzed using reflexive thematic analysis, following information power principles. Participants worked through a standardized interactive online CASUS™ case based on a real multimorbid patient. During the CASUS™ case, participants used think-aloud techniques, followed by semi-structured interviews and a drawing task in which they sketched desired XAI interfaces for four modalities (laboratory results, electrocardiograms, radiographs, patient photographs). Interviews were audio-recorded, transcribed, and analyzed using an iterative, mixed deductive–inductive content analysis with a collaboratively developed codebook. (1) Students prioritized patient-led symptom narratives, progression patterns, comorbidities, and physical signs during the diagnostic process. (2) They sought diagnostic-specific AI support: guideline prompts, ranked differentials (including rare diseases), test checklists, multimodal interpretation, and therapy verification. (3) Trust in AI prerequisites included data origins and specificity metrics. Fears regarding AI centered on loss of control, while benefits included time savings. (4) Preferred XAI featured step-by-step mentor guidance, counterfactuals, similar cases, and chat-based learning triggered by AI actions, surprises, or knowledge gaps. Final-year medical students conceive AI not only as a diagnostic assistant but also as a teaching partner that should scaffold reasoning, broaden differentials, and support confidence calibration. Their preferences yield concrete design requirements for trainee-oriented clinical AI/XAI: longitudinal workflow support, layered and multimodal explanations, and explicit communication of uncertainty and limitations to mitigate automation bias. While limited by its single-center, small-sample design and simulated case setting, this study offers actionable insights for the design and evaluation of human-centered AI tools in medical education and motivates prospective studies in authentic clinical learning environments.
This paper investigates in an exploratory manner small group interactions in Social VR classroom conversations using eye-tracking data (group eye-gaze configurations, blink rate and blink duration) in terms of task engagement. Task engagement has been previously studied in educational settings as it is an important construct for learning. Prior research showed that eye signals in groups can assess rapport, attention, agreement or communication engagement. However, little research focused on whether eye-tracking data can predict task engagement in educational small group Social VR conversations. We conducted an exploratory study with 52 pedagogy university students (8 dyads, 12 triads) engaging in free-flow conversations on a given topic. Results indicated that eye-tracking data could predict task engagement using multiple linear regression analysis. However, group eye-gaze configurations were the significant predictors in dyads, whereas the blink rate was the strongest predictor associated with task engagement in triads. Based on these findings, we discuss the higher complexity in triadic gaze dynamics and why the blink rate was the task engagement predictor. On the other hand, the simplicity of dyadic gaze configurations might as well explain why it was found as a predictor. Furthermore, we argue that future work can explore machine learning detectors of task engagement dependent of group size. The results from this exploratory study help understanding non-verbal group interactions in VR and contribute to further work on engagement models for more socially-abled virtual experiences in Social VR applications.
The OPTIMAL theory of motor learning, from its introduction to its subsequent refinement, has catalyzed a substantial body of research into motivational effects on motor learning with both supportive evidence and critical debate. This paper examines the effects of goal-directed practice, provided either through autonomy-supportive practice conditions—hypothesized by the OPTIMAL theory to yield motivational benefits—or a yoked group or a low-autonomy instructor, on implicit motor sequence learning of a complex, bimanual dual task. Participants practiced a motor sequence in a virtual reality serial reaction time (SRT) task and were either given or denied control over task difficulty as a task-relevant choice. While all groups successfully acquired the target sequence, differences between groups were negligibly small or absent altogether. These results suggest that the motivational effects of autonomy support do not substantially impact the motor learning of complex tasks.
This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that rely on pixel-level reconstructions, V-JEPAs learn by predicting embeddings of masked regions from the embeddings of unmasked regions. This enables the trained encoder to not capture irrelevant information about a given video like the color of a region of pixels in the background. Using a pre-trained V-JEPA video encoder, we train shallow classifiers using the RAVDESS and CREMA-D datasets, achieving state-of-the-art performance on RAVDESS and outperforming all other vision-based methods on CREMA-D (+1.48 WAR). Furthermore, cross-dataset evaluations reveal strong generalization capabilities, demonstrating the potential of purely embedding-based pre-training approaches to advance FER. We release our code at https://github.com/lennarteingunia/vjepa-for-fer.
In dyadic interactions, various human facial reactions could be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023, 2024 and 2025 challenge series, a body of generative deep learning (DL) models have been developed for the problem of multiple appropriate facial reaction generation (MAFRG). This year, we propose the REACT 2026 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can generate multiple personalised, appropriate, diverse, realistic and synchronised human-style facial reactions expressed by a specific human listener for responding to each given speaker behaviour. As a key of the challenge, we continuously provide challenge participants with MARS dataset introduced by REACT 2025 but additionally provide individual-level Big-Five personality labels and EEG recordings. This introduces a new one-to-many personalised facial reaction generation setting combining human expressive behavioural, affective and neurophysiological signals, which remains largely unexplored in current dyadic interaction modelling. This paper also presents the challenge guidelines and new baselines on the four proposed sub-challenges: Offline generic and personalised MAFRG as well as Online generic and personalised MAFRG, respectively, which are publicly available at https://github.com/reactmultimodalchallenge/baseline_react2026.
Agentic AI assistants that autonomously perform multi-step tasks raise open questions for user experience: how should such systems communicate progress and reasoning during extended operations, especially in attention-critical contexts such as driving? We investigate feedback timing and verbosity from agentic LLM-based in-car assistants through a controlled, mixed-methods study (N=45) comparing planned steps and intermediate results feedback against silent operation with final-only response. Using a dual-task paradigm with an in-car voice assistant, we found that intermediate feedback significantly improved perceived speed, trust, and user experience while reducing task load - effects that held across varying task complexities and interaction contexts. Interviews further revealed user preferences for an adaptive approach: high initial transparency to establish trust, followed by progressively reducing verbosity as systems prove reliable, with adjustments based on task stakes and situational context. We translate our empirical findings into design implications for feedback timing and verbosity in agentic in-car assistants, balancing transparency and efficiency.
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing, and assume each speaker only speaks a single language. However, in real-world applications, such assumptions often do not hold. Visual or audio information may be missing due to occlusions, camera or microphone failures, or privacy constraints. Multilingual speakers introduce additional complexity due to linguistic variability across languages. These situations constitute substantial challenges for the robustness and generalization capabilities of multimodal speaker identification systems. Aim of the POLY-SIM 2026 challenge is to address these aspects of speaker identification and to provide a standardized setup for the comparison of the proposed solutions.
Künstliche Intelligenz (KI) hat in den letzten Jahren große Fortschritte in der Medizin erzielt, insbesondere bei der Analyse medizinischer Bilddaten. Weniger im Fokus steht bislang ihr Potenzial als Analysetool für die Psychologie und Psychotherapie, obwohl aktuelle Entwicklungen zunehmend zeigen, dass KI-basierte Verfahren wertvolle Beiträge zur Diagnostik psychischer Störungen und zur Analyse psychotherapeutischer Gespräche leisten können. Ziel dieses Beitrags ist es, den internationalen Forschungsstand zu KI-basierten Werkzeugen und Techniken für Diagnostik und Gesprächsanalyse darzustellen, deren Potenziale und Grenzen kritisch zu diskutieren und zukünftige Anforderungen für eine verantwortungsvolle klinische Nutzung aufzuzeigen.
Existing benchmarks for Large Language Model (LLM) agents focus on task completion under idealistic settings but overlook reliability in real-world, user-facing applications. In domains, such as in-car voice assistants, users often issue incomplete or ambiguous requests, creating intrinsic uncertainty that agents must manage through dialogue, tool use, and policy adherence. We introduce CAR-bench, a benchmark for evaluating consistency, uncertainty handling, and capability awareness in multi-turn, tool-using LLM agents in an in-car assistant domain. The environment features an LLM-simulated user, domain policies, and 58 interconnected tools spanning navigation, productivity, charging, and vehicle control. Beyond standard task completion, CAR-bench introduces Hallucination tasks that test agents’ limit-awareness under missing tools or information, and Disambiguation tasks that require resolving uncertainty through clarification or internal information gathering. Baseline results reveal large gaps between occasional and consistent success on all task types. Even frontier reasoning LLMs achieve less than 50% pass rate on Disambiguation tasks due to premature actions, and frequently violate policies or fabricate information to satisfy user requests in Hallucination tasks, underscoring the need for more reliable and self-aware LLM agents in real-world settings.
Robots are becoming more prominent in assisting persons with disabilities (PwD). Whilst there is broad consensus that robots can assist in mitigating physical impairments, the extent to which they can facilitate social inclusion remains equivocal. In fact, the exposed status of assisted workers could likewise lead to reduced or increased perceived stigma by other workers. We present a vignette study on the perceived cognitive and behavioral stigma toward PwD in the workplace. We designed four experimental conditions depicting a coworker with an impairment in work scenarios: overburdened work, suitable work, and robot-assisted work only for the coworker, and an offer of robot-assisted work for everyone. Our results show that cognitive stigma is significantly reduced when the work task is adapted to the person's abilities or augmented by an assistive robot. In addition, offering robot-assisted work for everyone, in the sense of universal design, further reduces perceived cognitive stigma. Thus, we conclude that assistive robots reduce perceived cognitive stigma, thereby supporting the use of collaborative robots in work scenarios involving PwDs.
What happens when a virtual character speaks with the cloned voice of a person you know? Our results suggest that such a known-voice clone can subtly shape social perception, but does not reliably transfer the perceived personality of the original speaker to the virtual character. We investigate whether voice cloning can shape the attribution of personality and social traits to a virtual character when that character speaks with a cloned version of a known voice. Across two studies, we compared ratings of an original voice speaker with ratings of a virtual character speaking with either a known-voice clone or an unknown voice. The first study was a pilot in which students were exposed to the voice of their lecturer. The second study was a preregistered follow-up in which content creators’ voices were presented to their respective communities. Participants evaluated personality traits and social characteristics, and we measured how similar these ratings were to those of the original speaker. Quantitative results revealed small and selective effects. In the pilot study, differences emerged for likability and agreeableness, while in the follow-up, effects were observed for extraversion. Qualitative analyses showed a shift in focus away from perceptions of the salience of artificiality and unnaturalness toward the incongruence between voice and visual appearance. Overall, the findings indicate that known-voice cloning is not a general mechanism for personality transfer. Instead, it functions as a weak and context-dependent cue whose influence is constrained by multimodal coherence and users’ expectations. These results provide practical guidance on when known-voice cloning may be beneficial, when it may backfire, and which aspects should be prioritized in the design of embodied voice interfaces.
It has been widely recognized in educational research that, if addressed appropriately, errors can promote learning. Adaptive reactions to errors-including affective-motivational and action-related reactions-have been proposed to capture such adequate adjustments of affect and behavior to regulate learning processes. However, prior research addressing reactions to errors has left primary school students underrepresented. At the same time, findings from assessments beyond self-reported data are still sparse. Therefore, we multimodally assessed adaptive reactions to errors in an authentic learning situation within a digital mathematics learning environment. In a sample of 284 third-grade students, we found a positive relation between self-reported adaptive affective-motivational reactions to errors, as well as the corresponding behavioral indicator, and knowledge acquisition. For action-related reactions to errors, only the behavioral indicator was related to knowledge acquisition. These results highlight the relevance of adaptive reactions to error for primary school students while emphasizing the necessity of multimodal assessment. Educational relevance statement: The results show that adaptive reactions to errors, including adaptive affectivemotivational and action-related reactions to errors, positively relate to knowledge acquisition in primary school students. To fully understand their impact, multimodal assessment appears to be necessary. Therefore, teachers should take students' reports and behaviors into account while adopting a constructive approach to addressing errors in their classrooms.
Wolfgang Minker合作论文数Faculty of Engineering and Computer Science,University of Ulm
Institute of Information Technology20