
There is extensive research on measuring Cognitive Load (CL) using both behavioural and physiological indicators. However, we do not know whether existing datasets used to measure CL are transferable across contexts. We systematically reviewed non-invasive measurement methods in naturalistic settings using PRISMA guidelines, searching five databases for studies published between 2015-2025. We analysed 24 studies focusing on non-invasive measurement methods and found that multimodal approaches reported the highest accuracy (up to 95.4
Collective intelligence refers to a team’s general ability to perform well across a wide variety of tasks, a stable team-level analogue of individual intelligence. Although it fundamentally emerges from interaction, it is still largely treated as a static property or as an aggregate performance outcome. This article proposes a computational proof of concept that operationalizes a dynamic, process-oriented model of collective intelligence as a multilayer pipeline instantiated in immersive virtual reality. The model articulates three complementary views: collective intelligence as a measurable capacity (the C-factor), as an emergent state produced by three transactive subsystems, that is, shared processes through which teammates jointly manage memory (“who knows what”), attention (what the team focuses on), and reasoning (how goals are aligned), and as a performance gain over individual baselines, all embedded within a recursive Input-Mediator-Output-Input cycle. A multi-user VR collaborative task built on role interdependence and distributed information enables the synchronized capture of each participant’s multimodal behavioral traces (speech, gaze, facial expressions). These signals are transformed into windowed features and fused into subsystem-level indicators, whose recursive updates formalize collective intelligence as a temporally unfolding process rather than a fixed score. A pilot study (six triads) demonstrates the technical feasibility and user acceptability of the system, as well as meaningful variability in performance and interaction dynamics; higher-performing teams exhibit multimodal patterns consistent with more efficient attentional and regulatory coordination. This work bridges these perspectives within a unified computational framework and lays the groundwork for the longitudinal study of collective intelligence and for multimodal systems designed to support it.
Human-computer interaction exploits various input methods to improve user experience, but those who don’t have visual access suffer more than the mainstream. Hand gestures are becoming increasingly popular in the design of assistive technologies, representing a natural means of communication. In this context, this paper proposes a novel Gesture-based Computer Interaction (GCI) system to facilitate computer interaction for visually impaired people. However, a large set of gestures for a simple application introduces complexity, which poses challenges for the visually impaired to interact with computers. Therefore, an accessible assistive application (keyboard) with a minimal set of gestures is designed here. The performance of the proposed GCI is evaluated from a cognitive load perspective. Nineteen (19) participants engaged with the proposed GCI in various cognitive scenarios with the computer. Several dimensions are evaluated, including physiological measures (skin conductance), NASA TLX, and performance indicators (response time, gesture response time, false responses, forgot word counts). Among all metrics, the gesture response time revealed that the proposed GCI technique had a 39
Adapting agricultural practices to climate change requires both technical innovation and behavioural change among farmers. The Internet of Things supports this transformation by generating large, diverse data that require adaptive content and presentation. Technology adoption in agriculture is a human decision-making process shaped by psychological, social, and cultural factors. This paper presents AgriCog, a cognitive framework that leverages psychology-informed design to support behavioural change in agricultural assistance systems. It addresses key psychological barriers such as social norms, conflicts between farmers’ attitudes and recommended practices, and lack of psychological safety. Six behavioural design elements (emotion-aware design, manifest-based learning, yield challenges, risk level support, social awareness, and best practices delivery) foster engagement, trust, and gradual adoption of sustainable practices. A decision-support tool based on this framework was evaluated with farmers, agricultural experts, and social scientists. Results show that AgriCog improves both user experience and interface usability, demonstrating that cognitive psychology principles can effectively support behavioural change and sustainable agriculture.
Robotaxis are gradually taking the place of traditional taxis, removing the human driver as a source of local insights and conversation, fostering passive travel and potentially hindering passengers’ spatial learning of their surroundings. This paper introduces RideGuide, a customizable multimodal conversational tour guide system for autonomous vehicles that integrates voice, touch, and vision capabilities to provide context-aware and engaging interactions. In an exploratory lab study (n = 12), participants customized RideGuide ’s language, voice, and chatbot personality before experiencing a pre-recorded robotaxi ride (t = 10 min, 270 ^∘ car back-seat video). Results indicate a positive hedonic user experience (UEQ-S, 1.10), moderate chatbot usability (CUQ, 67.6
The design of spatial user interaction systems requires knowledge of the actions performed in the environment and the underlying system that detects these actions. While studies that utilize spatial data collected from sensors recognize these aspects as valuable, more research can be done to investigate how users with differing embodied knowledge of a task can improve the design of spatial user interaction systems. We present a user-centered design study of a prototype system that allows users to share their hand gestures through sound as they crochet to enhance the experience of this embodied creative task. We recruited novice crocheters (N = 10), experienced crocheters (N = 9), and subject matter experts (N = 5) in sonification, audio technology, and human–computer interaction to holistically uncover opportunities to improve our prototype. We present our qualitative results and a set of guidelines that can support researchers in utilizing differences in embodied knowledge to further their design process.
Virtual agents offer a powerful and underutilized methodology for studying social interaction, particularly how bias and stereotypes operate. Drawing on theoretical frameworks such as the Media Equation and supported by empirical studies, we demonstrate how virtual agents can serve as controlled proxies for human social actors to study social bias. Through two empirical studies in the negotiation domain, we show how gender cues can be studied in negotiation contexts. Study 1 examines the effects of gendered job descriptions on negotiation behavior — finding that gendered language and gender interact to affect participants’ minimally acceptable salary. Study 2 explores how agent embodiment and emotional tone influence user perceptions across gender lines — finding that agent embodiment and one’s gender interact to affect participants’ subjective feelings about the dispute. Together, these findings imply the need for greater use of virtual agents in research, as they help inform the development of more inclusive and equitable AI systems.
The growing presence of Socially Interactive Agents (SIAs) in educational, professional, and domestic contexts raises questions about their role in reflecting and shaping social biases. This overview summarises and critically reflects on a series of studies that were conducted on bias and diversity in interactions with SIAs. Building on theories of social identity, linguistic bias, and intersectionality, we explore how users perceive and evaluate SIAs based on social cues such as language, accent, ethnicity, and gender. Our findings suggest that established psychological mechanisms, such as in-group favouritism, stereotyping, and categorisation also apply to human-agent interaction. Beyond diagnosis, the paper discusses how SIAs might be used to challenge and reduce bias, drawing on approaches such as intergroup contact, and counter-stereotypical design. At the same time, we critically question the ethical implications of simulating marginalised identities in virtual environments and address risks. We conclude by outlining future research directions aimed at designing culturally sensitive and socially responsible agents that not only reflect human diversity but actively support fairness and inclusion.
To realize their full potential, Socially Interactive Agents (SIAs) must effectively engage with human users from diverse individual, social, and cultural backgrounds. However, most current SIAs are grounded in White- and Western-centric assumptions, limiting their ability to express and interpret social cues appropriately across cultures. Here, we demonstrate how the data-driven psychophysical method of reverse correlation can help address these limitations by modeling users’ perceptual expectations, preferences, and sociocultural norms and strategically integrating these insights into SIA design. Drawing on examples from our research group, we show how this method could enable SIAs to exhibit social signals that are psychologically grounded, culturally adaptive, and ethnically inclusive. By informing the design of SIA appearance and expressive behavior with empirically derived user models, our approach aims to improve user engagement and trust while contributing to broader efforts to mitigate algorithmic bias, reduce access inequality, and challenge real-world prejudice in both human-AI and human–human interaction contexts.
This paper tackles the complex issue of "race" and "diversity" in contemporary robotics. Studies conducted by Sparrow (Sci Technol Hum Values 45 3 538 560) and Bartneck et al. (Proceedings of the ACM/IEEE International Conference on Human Robot Interaction, Chicago, 2018) concluded that the (mediated) image and (material) shape of robots reveal discrimination and segregation processes similar to those experienced by ethnic minorities in contemporary societies—and they are as such, machine embodiments of human racism. It is true that there seems to exist a formal standardisation of robots' appearance, often inspired by Western cultural models. But in the global and rapid developments of robotics, robots are also increasingly dyed with the colours of ethnic diversity—they are "racialised" to better fit alternative models (ethnic black robots, Asian-styled humanoids, and so on). Yet, most anthropomorphic robots are race-free. Beyond the burning issue of apparent/alleged "racism" in robot design and, by extension, the existence of the WEIRD (Western, Educated, Industrialised, Rich, Democratic) complex in contemporary robotics and technologies, this paper examines different models of robots designed in countries from the Global South (Africa, Asia) and raises the fundamental ambivalence of robots that are both subjected to racism (as almost human) and racialised (coloured) / used as weapons of resistance against racial discrimination.
Socially Interactive Agents (SIAs) are increasingly involved in people’s daily lives, assuming roles that range from virtual tutors to conversational companions. In this context, the agent’s personality and its use of argumentation schemes can shape decision-making processes and the way it interacts. When an SIA resorts to emotional or potentially fallacious arguments, it may reinforce stereotypes and discriminatory behaviors. However, these same schemes can also be employed in a legitimate or ethical manner if integrated within a framework that anticipates bias detection and behavioral adaptation. This work explores the integration of the Five-Factor Model (FFM) of personality with argumentation schemes to predict the types of arguments an agent would use based on its personality profile. Such integration enables the simulation of both unethical or discriminatory behaviors (when the traits of the FFM suggest a more hostile interaction style) and respectful, inclusive strategies (when empathy-related traits prevail). We present a framework that combines personality traits, situational contexts, and discrimination detection mechanisms with the goal of designing SIAs that behave consciously and ethically.
Socially Interactive Agents (SIAs), including physically embodied social robots and virtual agents, are increasingly used in applications involving social interactions, including Motivational Interviewing (MI). This study investigates how physical embodiment and nonverbal behaviors, specifically facial expressions generated by our custom diffusion model, influence user’s perceptions and interaction dynamics within the MI framework. By upgrading the diffusion model from an offline configuration to a real-time architecture, we make live counseling sessions with human participants possible. An experimental design was developed to compare three facial expression generation conditions (model-based, mismatched, and control) and two embodiment modalities (Robot vs. Agent). Verbal behavior was supported by a pre-prompted Mistral Large Language Model (LLM). This design allowed us to evaluates, how aligned facial expressions and physical embodiment influence social presence, attitude perception, social rapport perception, self-disclosure, and the overall quality of MI using objective and subjective measures. Results show that aligned facial expressions or the absence of expressions were perceived more positively than mismatched expressions. We also observe that physical embodiment significantly improves social presence and attitude perception, with the social robot consistently rated higher than the virtual agent. A mediation analysis revealed that social rapport serves as a factor linking attitude perception to the MI quality.
Automated engagement detection has to date focussed exclusively on engagement annotations collected via external observation. It is therefore unknown whether model predictions align with a participant’s own perception of their engagement level. In this study, we examine self-reported and external observations of engagement in a corpus of small-group conversational interactions. We find no evidence of correlation between self-reported and externally observed engagement. We show that annotators rely heavily on speaking activity as a proxy for engagement, which we demonstrate is a poor indicator of self-reported engagement. The focus of recent literature has been to improve the engagement detection accuracy with deep learning. In contrast, we use a simple multimodal combination of low-level linguistic features (e.g. number of words spoken) and facial expression. We find that none of the linguistic features significantly predict self-reported engagement. However, we find that each linguistic feature significantly predicts annotator engagement, with the odds of a high engagement prediction increasing with each word and utterance spoken. Finally, we report the performance of a model which predicts real-time engagement levels. We find that the best model overall utilises a multimodal combination of linguistic and facial expression features. Our work raises a number of important questions concerning the ecological validity of current approaches for engagement detection and underlines the need to validate external observations against self-reported metrics.
Traditional teaching techniques, with their heavy reliance on oral instruction and written practice, create serious barriers for persons with speech, mobility, visual, or learning disabilities. This work introduces a multisensory, gesture-based learning device to provide an inclusive and interactive mathematics platform. The main input modality is through gestures. The users wear a motion sensor to perform mid-air gestures, representing digits from zero to nine and mathematical operators like addition and subtraction to answer computer-generated problems. A customized, lightweight deep learning model recognizes gestures by interpreting their spatial and temporal gesture-based features. The model is optimized for efficiency by making the system compact and minimizing delay while preserving the accuracy required for correct recognition of intended numbers and symbols. The recognized gestures are fed into a large language model that generates relevant math problems and provides real-time audio feedback. The system automatically adjusts problem difficulty based on user performance, creating a personalized learning experience. The system was evaluated with twenty persons with disabilities, representing the target population, demonstrating strong effectiveness, usability, and accessibility. The proposed gesture-based approach demonstrates strong promise in empowering persons with disabilities by closing gaps in accessibility and encouraging engagement through adaptive, multisensory learning.
Emotion recognition is crucial for enhancing human-computer interaction systems. However, the development of robust methodologies for French emotion recognition is hindered by the scarcity of labeled, interactive multimodal datasets. In this work, we outline the acquisition and annotation procedures and provide an evaluation benchmark for the Card Game-based Multimodal Emotion Recognition (CG-MER) dataset that we designed to capture spontaneous emotional expressions in French conversations. The dataset comprises approximately ten hours of video recordings featuring dyadic interactions between 20 French participants (11 males, 9 females) engaged in a card game, capturing natural expressions through facial cues, speech, and gestures. Unlike existing corpora, CG-MER provides refined annotations across all three modalities, enabling a detailed investigation of emotion dynamics and their associated gestures in a French-speaking context. Additionally, we establish baseline results using state-of-the-art models for each modality and propose a standardized evaluation protocol, facilitating future comparative studies on multimodal emotion recognition.
The research study aimed to evaluate the effectiveness of a personal device designed to assist blind users in navigating their environment. The wearable device consists of a stereo vision camera with structure illumination, a central control unit, a haptic belt, speakers, and a custom-made keyboard. The device uses haptic stimuli and auditory cues to provide environmental information and facilitate navigation. An important feature of the system is that the user can select operating modes and their settings while navigating. The procedures for processing the depth map, its segmentation and a mapping scheme for converting the segmented image into a matrix of haptic actuators are explained. The usability of the system was verified in static and mobility trials with 10 blind people. A post-trial survey provides valuable feedback from users, highlighting strengths of the system, such as its functionality and safety features, as well as areas for improvement, such as the intuitiveness of the user interface. The results of the study show that the combination of haptic and auditory modalities to present the environment can be an effective approach to building usable navigation aids for the visually impaired.
The use of semantics in immersive environments such as Virtual Reality (VR) and Augmented Reality (AR) offers new opportunities to enhance user interaction through multimodal interfaces, including gestures and voice commands. However, significant challenges persist due to the lack of formal and explicit semantic representations of both interaction contexts and virtual scene content. This paper offers a comprehensive review of current methods for integrating semantic data into virtual environments, with a focus on knowledge representation and rule-based reasoning. Using the PRISMA guidelines for systematic reviews, our analysis spans various applications, including science, education, and architecture in virtual and augmented environments. We explore the role of ontologies and knowledge representation in managing virtual environments, emphasizing how knowledge graphs and rule-based reasoning can improve the efficiency and comprehension of user interactions, particularly in generating multimodal commands. This paper identifies research gaps and proposes future directions for developing adaptable and interoperable semantic frameworks that can be reused across diverse virtual environment applications.
In this study, we design and test a smart accessory for the visually impaired that can be attached to the white cane and help the user avoid above-waist obstacles that cannot be detected by a conventional white cane. The device integrates multimodal capabilities, combining ultrasonic distance measurement and haptic and auditory feedback to enhance spatial awareness. To design the device, 103 blind Mexicans, their caregivers, and teachers were interviewed to obtain a clear understanding of their specific needs. People from different socioeconomic strata and from various states of the country were considered. Once designed, 7 units were manufactured to conduct usability tests. Results indicate a significant 29
With the increasing complexity of human-computer interaction activities in civil aircraft cockpits, the introduction of multimodal interaction to improve cockpit simplicity, efficiency, and safety is imminent. Effects of four interaction modalities (single voice commands, phased voice commands, voice, gesture, and eye-movement, and voice and eye-movement interaction) on flight performance were explored. In a 2 × 2 within-subjects flight simulation experiment, we observed that, under complex conditions, the use of multimodal interaction significantly reduced mental demand and shortened task completion time. These results suggest that the impact of complex conditions can be minimized by multimodal interaction. Furthermore, the combination of gestures and eye movements caused a shortage of visual resources, which affected pilot performance, as evidenced by faster task completion times and a better user experience for voice and eye-movement interaction compared to voice, gesture, and eye-movement interaction. The above conclusions indicated that voice and eye-movement interaction was ideal for the next-generation civil aircraft cockpit for flight tasks.
Touch sensation is valuable in mobile augmented reality games, especially for improving experience in serious games by linking the virtual and real worlds. This study develops haptic stimuli for a mobile augmented reality serious game to enhance haptic and user experience. We designed haptic stimuli based on the game’s visual components, testing two versions: fantasy and ordinary. Three multimodal combinations (visual-audio, visual-haptic, and visual-audio-haptic) were created and evaluated in a user study to assess their impact on haptic and user experience. Quantitative findings showed significant differences in the involvement aspects of haptic experience, indicating that participants felt more involved with the fantasy version than with the ordinary version when playing with the visual-audio-haptic combination. The visual-audio-haptic combination consistently provided a better user experience in terms of attractiveness, stimulation, and novelty, regardless of the version. Qualitative feedback offered insights into future design improvements for haptic and user experience in mobile augmented reality serious games.