
Much research in human–artificial intelligence (AI) interaction treats cooperation as a design problem of alignment, asking how system behavior can be coordinated with human intentions, task goals or expectations in order to support effective collaborative action. In this debate contribution, we shift the analytic focus from alignment to sequential organisation and ask how cooperation is interactionally accomplished when one source of action does not demonstrably orient to a shared sense of reality, yet its outputs are nevertheless incorporated into an unfolding sequence of action. Drawing on a sequential comic-making practice involving a generative AI system, we analyse three examples to examine how creative work can progress through misrecognition, breakdown and material effects, not necessarily requiring consensus or control. We argue that such misalignment is not a failure to be repaired but can function as a recurring interactional condition in human–AI cooperation, sustained through the sequential uptake, reinterpretation and transformation of AI-generated outputs. Based on this analysis, we suggest that human–AI interaction and design research would benefit from more explicitly interactional accounts of cooperation that treat misalignment as a productive condition rather than something to be minimised. We position artistic creative practice as a methodologically revealing site for examining how agency, authorship and coordination are practically accomplished in cooperative work with AI and we discuss implications for human-centered design.
A motion-based virtual reality (VR) bicycle simulator allows for steering and pedaling as input, while visual information and platform movements are provided as output. The simulator offers a potential means of safely experiencing mountain biking (MTB); however, the impact of complex multidirectional tilting of the bicycle on user experience remains unclear. Therefore, the aim of this study was to investigate the effects of integrating tilt and pitch on users’ psychological experience during downhill cycling in a simulator. Twenty-one participants rode a simulator course designed to replicate a real MTB course and were instructed to pass through balls placed at ten turns (i.e., banks) along the course. Measurements were taken under two conditions: the nonmotion (NM) condition, in which the platform remained stationary, and the motion-based (M) condition, in which the platform moved in tilt and pitch according to the visual environment. After each condition, participants completed the Simulator Sickness Questionnaire (SSQ), the Igroup Presence Questionnaire (IPQ), and the short version of the User Experience Questionnaire (UEQ-S). Eighteen participants completed both conditions. There was no significant difference in the SSQ and the UEQ-S between the two conditions. In the IPQ, only the subscale spatial presence was significantly higher in the M condition than in the NM condition. The platform’s tilt and pitch movements during downhill cycling in a VR bicycle simulator had only a minimal impact on simulator sickness, presence, and user experience. These results indicate that a platform without motion may be sufficient for rehearsing MTB downhill courses for individuals with no prior MTB experience.
Augmented reality (AR) introduces new dynamics for personal information sharing by embedding digital content into shared physical spaces. While prior work has examined comfort and privacy concerns in AR, less is known about why people choose to disclose personal information and how social context shapes these decisions. We conducted a mixed-methods study with 20 participants across three fictional scenarios that varied in formality and social norms. Participants disclosed information to clarify identity, build relationships, and strategically manage impressions, yet patterns varied by setting: intimate details were sometimes shared in casual contexts despite acknowledged risks, while professional contexts raised concerns about appropriateness. Based on these findings, we present an AR-adapted disclosure decision model that treats context as an active moderator of perceived risks and benefits and highlights the roles of control and reciprocity in co-present environments. We conclude with design recommendations for context-sensitive privacy mechanisms that support selective and reciprocal disclosure in AR.
Depending on the degree of disability, simple tasks of daily living can be challenging for people with physical disabilities, such as picking up and placing objects, eating, or reaching for a cup to drink independently. Pervasive technologies such as robotic arms can be used to assist with these daily tasks, allowing patients to regain independence while reducing the need for care. Specialized devices, such as assistive forks or spoons, can facilitate these tasks. Image datasets of everyday objects such as MS COCO do not contain assistive devices, which tend to look different from their non-assistive counterparts. We present the dataset WLRI-AD (Work-Life Robotics Institute–Assistive Devices) to enable a robot to interact with devices in assisted living homes. The benefits of including assistive devices are demonstrated by comparing versions of the dataset with each other and to a baseline. Initial results show an improvement in the detection of assistive devices by training a YOLOv8 model on the assistive devices.
In Ambient Intelligence (AmI), seamless interaction between humans and AI systems is crucial for the effective utilization of smart environments. This paper investigates human–AI interaction in AmI, focusing on the management of uncertainty—situations where the AI is unsure how to choose between multiple possible outputs—within the context of Opportunistic Software Composition. The Opportunistic Composition Engine (OCE) dynamically constructs applications or assemblies from available software components, learning from human feedback to tailor applications to user preferences and context. We design, implement, and evaluate three approaches to uncertainty handling: explicitly asking the user to resolve uncertainty, offering a choice between multiple assemblies, and using a visual gradient to indicate the AI’s confidence. Each approach is implemented in a distinct version of OCE and assessed through a user study (N = 121) measuring usability and learning performance in a 2D online simulated ambient environment. Our results show that providing clear visual or informational cues about uncertainty improves user satisfaction, while forcing users to resolve uncertainty themselves reduces usability and can hinder learning. Importantly, the other interaction strategies do not impair the system’s learning, suggesting that lightweight, system-guided handling of uncertainty can effectively support both user experience and AI adaptation. Beyond these conclusions, this work demonstrates the utility and general applicability of simulation-based technologies for prototyping and evaluating AmI applications.
The number of surveilling and recording devices increases steadily, contributing to raising concerns on individuals’ privacy. Previous research has addressed this issue by investigating end-users’ reactions to various types of video equipment, mainly with surveys and interviews. However, less is known regarding the impact of such technology on bystanders’ affective state and actual behavior. In the present explorative study, we investigated the effect of different video-surveilling conditions (i.e., traditional camera, smart glasses, drone, human observer, and no observer) on participants’ performance and anxiety levels. To this end, a controlled between-participant laboratory experiment was run. We collected both objective metrics, namely, the performance in a memory task, and self-reported data, i.e., the level of state anxiety. Our results showed that in all of the groups in which a surveilling technology was present, i.e., a traditional camera, a drone, and smart glasses, the levels of state anxiety increased. However, no effect emerged on the participants’ performance in any group. The present paper contributes by showing that besides raising privacy concerns, monitoring technology induces an increase in the level of state anxiety. Nevertheless, participants’ performance was not affected by the surveillance taking place, endorsing the view that individuals are becoming increasingly accustomed to it.
Privacy research demonstrates a dichotomy between what individuals describe as their level of concern and the protective behaviors they exhibit. We posit this dichotomy or “privacy paradox” is impacted by a feeling of hopelessness regarding an individual’s ability to protect their personal information. Younger generations have experienced all of life on social networks and much of their private information has been online since before they made individual informed privacy decisions. We evaluate the proposed research model using a survey method targeted primarily at a younger population. This research demonstrates the initial impact of hopelessness on privacy concern and also exhibits the impact of privacy awareness and perceived control. Our findings indicate that a strong antecedent to privacy concern is the feeling there is nothing we can do to protect our privacy. We contribute to the ongoing privacy research stream by incorporating this concept from psychology into a privacy concern model.
Human cognitive performance affects a wide range of aspects of our daily lives. Numerous factors influence our cognitive performance, and cognitive performance in turn impacts our capabilities. Partial sleep deprivation in particular negatively affects vigilance, a key factor in many work tasks. Sleep in general plays a large role in physiological recovery and our capability to perform mental tasks. In this work, we focus on two research questions. First, we investigate how fluctuations in sleep quality influence cognitive vigilance. Second, we study how smartphone typing can be leveraged as a continuous measurement for cognitive vigilance and can thus be an indicator of decline in cognitive capabilities and sleep quality. We report on a 2-month field study in which we collected cognitive performance data using the Psychomotor Vigilance Task (PVT), mobile keyboard typing metrics from participants’ personal smartphones, and sleep quality metrics through a wearable sleep-tracking ring. Our findings highlight that individual sleep metrics such as night-time heart rate, sleep latency, sleep timing, sleep restfulness, and overall sleep quantity significantly influence vigilance. Long sleep latencies can reduce reaction times up to 30 ms, abnormal sleep durations up to 20 ms, and night-time awake time up to 10 ms. Heart rate is a well-known indicator of recovery quality, and improvements in both heart rate and heart rate variability (HRV) show positive variations of 15–20 ms in reaction test performance. To expand the current research on cognitive computing, we introduce smartphone typing metrics as a proxy or a complementary method for continuous passive measurement of cognitive vigilance and report on statistically significant correlations in PVT performance and typing speed and error rates. Together, our findings contribute to ubiquitous computing via a longitudinal case study with a novel wearable device, the resulting findings on the association between sleep and cognitive function, and the introduction of smartphone keyboard typing as a proxy of cognitive function.
Many Coronavirus disease 2019 (COVID-19) and post-COVID-19 patients experience muscle fatigues. Early detection of muscle fatigue and muscular paralysis helps in the diagnosis, prediction, and prevention of COVID-19 and post-COVID-19 patients. Nowadays, the biomedical and clinical domains widely used the electromyography (EMG) signal due to its ability to differentiate various neuromuscular diseases. In general, nerves or muscles and the spinal cord influence numerous neuromuscular disorders. The clinical examination plays a major role in early finding and diagnosis of these diseases; this research study focused on the prediction of muscular paralysis using EMG signals. Machine learning-based diagnosis of the diseases has been widely used due to its efficiency and the hybrid feature extraction (FE) methods with deep learning classifier are used for the muscular paralysis disease prediction. The discrete wavelet transform (DWT) method is applied to decompose the EMG signal and reduce feature degradation. The proposed hybrid FE method consists of Yule-Walker, Burg's method, Renyi entropy, mean absolute value, min-max voltage FE, and other 17 conventional features for prediction of muscular paralysis disease. The hybrid FE method has the advantage of extract the relevant features from the signals and the Relief-F feature selection (FS) method is applied to select the optimal relevant feature for the deep learning classifier. The University of California, Irvine (UCI), EMG-Lower Limb Dataset is used to determine the performance of the proposed classifier. The evaluation shows that the proposed hybrid FE method achieved 88% of precision, while the existing neural network (NN) achieved 65% of precision and the support vector machine (SVM) achieved 35% of precision on whole EMG signal.
Recently, the leading cause of preventable blindness is diabetic retinopathy (DR). Although there are several undiagnosed and non-treated cases of DR, accurate and adequate retinal screening could facilitate the early detection and treatment of DR. The goal of this research is to develop a reliable DR screening and detection model to reduce the risk of DR-related blindness. DR-infected eyes describe ophthalmologist for further examination and diagnosis might reduce the risk of vision loss and provide timely and accurate diagnostic information. Hence, this paper proposes a hybrid inductive machine learning algorithm (HIMLA) as an automated DR detection diagnostic tool. HIMLA processes and classifies colored fundus images as healthy (no retinopathy) or unhealthy (presence of DR) by identifying the appropriate medical DR cases. The proposed algorithm comprises four stages: pre-processing, segmentation, feature extraction, and classification. At the pre-processing stage, colored fundus images are normalized to a specific brightness level to enhance the quality of the images. In the segmentation stage, the processed image is encoded and decoded to segment the images for improving image quality. Furthermore, feature extraction and classification are performed using multiple instance learning (MIL). The proposed method was evaluated on CHASE datasets for the detection of DR. The accuracy, sensitivity, and specificity of the proposed approach are 96.62%, 95.31%, and 96.88%, respectively. These results indicates that HIMLA outperforms other DR models, such as ML-based neovascularization detection in the optic disc (MLB-NVD), genetic algorithm–based diabetic retinopathy (GAB-DR), DL algorithm diabetic retinopathy (DLA-DR), and diagnostic assessment–based DL for diabetic retinopathy (DAD-DR ), which reduces the risk of vision loss.
The novel human coronavirus disease COVID-19 has become the fifth documented pandemic since the 1918 flu pandemic. COVID-19 was first reported in Wuhan, China, and subsequently spread worldwide. Almost all of the countries of the world are facing this natural challenge. We present forecasting models to estimate and predict COVID-19 outbreak in Asia Pacific countries, particularly Pakistan, Afghanistan, India, and Bangladesh. We have utilized the latest deep learning techniques such as Long Short Term Memory networks (LSTM), Recurrent Neural Network (RNN), and Gated Recurrent Units (GRU) to quantify the intensity of pandemic for the near future. We consider the time variable and data non-linearity when employing neural networks. Each model’s salient features have been evaluated to foresee the number of COVID-19 cases in the next 10 days. The forecasting performance of employed deep learning models shown up to July 01, 2020, is more than 90% accurate, which shows the reliability of the proposed study. We hope that the present comparative analysis will provide an accurate picture of pandemic spread to the government officials so that they can take appropriate mitigation measures.
Besides the enhancement of the Internet of Things (IoT) distributed environment, anomalous activities are also escalating rapidly. Therefore, improving the trustworthiness of distributed networks is required for the extensive adoption of IoT infrastructure. Establishing a security mechanism in IoT networks is a challenging task as communication links are lossy and connected devices are resource-dependent. Conventional security techniques such as intrusion detection systems (IDS) are insufficient to shelter the IoT-distributed environment due to less computational capacity, restricted upgraded devices, and mismatched protocols. This paper proposes a novel machine learning-based trustworthy model for IoT attack detection. The proposed system combines the capability of Ada-boost and Gradient-boost to classify anomalous activities with low computational capacity proficiently and within a minimum time frame. Experiments were conducted on Distributed Smart Space Orchestration System (DS2OS) IoT dataset to assess the significance of the novel attack detection model. The demonstration shows that the proposed model obtains 98.28
Intelligent systems, such as chatbots, are likely to strike new qualities of UX that are not covered by instruments validated for legacy human–computer interaction systems. A new validated tool to evaluate the interaction quality of chatbots is the chatBot Usability Scale (BUS) composed of 11 items in five subscales. The BUS-11 was developed mainly from a psychometric perspective, focusing on ranking people by their responses and also by comparing designs’ properties (designometric). In this article, 3186 observations (BUS-11) on 44 chatbots are used to re-evaluate the inventory looking at its factorial structure, and reliability from the psychometric and designometric perspectives. We were able to identify a simpler factor structure of the scale, as previously thought. With the new structure, the psychometric and the designometric perspectives coincide, with good to excellent reliability. Moreover, we provided standardized scores to interpret the outcomes of the scale. We conclude that BUS-11 is a reliable and universal scale, meaning that it can be used to rank people and designs, whatever the purpose of the research.
Telemedicine has become increasingly popular due to its convenience and accessibility. However, one of its drawbacks is the difficulty of performing lung and heart auscultation remotely. As a solution, we propose smartphone-based tele-auscultation for capturing lung and heart sounds. However, one of the obstacles to this approach is that low-frequency sounds, such as heart sounds, cannot be accurately reproduced by a smartphone speaker, due to their frequency response. To address this issue, we suggest using pitch-shifted versions of these sounds, specifically designed for hearing through smartphones. We conducted an initial evaluation of these processed sounds by obtaining 10 heart sounds and 20 lung sounds from open-source databases, which were pitch-shifted using algorithms based on Paul’s Stretch and SoundStretch libraries, respectively. These processed audios were validated by 40 final-year medical students using a web survey and conventional headphones and were compared against the original versions. The results showed that 72% and 80% of responses indicated that clinical information was preserved in the samples of respiratory and heart sounds, respectively. These findings suggest that pitch-shifted sounds could be potentially used in tele-auscultation devices such as smartphones. However, further research is needed regarding the recording and playback capabilities of smartphones.
Free-text keystroke dynamics, the unique typing patterns of an individual, have been applied for the security of mobile devices by providing the non-intrusive and continuous user authentication. Existing authentication approaches mainly concentrate on the keystroke dynamics when operating a specific device, and overlook the generality of keystroke dynamics for cross-device user authentication. To tackle this problem, in this paper, we propose an efficient federated free-text keystroke dynamics mechanism to mitigate the difference in keyboards for cross-device authentication. Specifically, we explore and analyze the keystroke features of various keyboards and extract cross-device keystroke features. To protect user privacy, their type of rhythm information must be kept locally. We utilize federated learning based on the auxiliary model to train the authentication model. Our proposed solution was evaluated on a large-scale data set with 168,000 users. The experimental results show that our proposed solution performs well with great robustness across different types of keyboards.
Conversational technologies have become increasingly prevalent, with interactions becoming more and more personalized. This paper reviews literature on the phenomenon of self-disclosure to conversational technologies. Five types of factors emerge as influencing self-disclosure: interface modality, conversational factors, user characteristics, mediating mechanisms, and contextual factors. We describe each type of factor, cover findings from the literature, present the framework of factors influencing self-disclosure that thus emerges, and put forth pertinent questions for future research on self-disclosure to conversational technologies.