Social robots are coming into our daily life. Existing conversational robots are mostly reactive in that the interactions are usually initiated by the users. With the knowledge of the environmental context such as people's daily activities, robots can be more intelligent and proactive. In this paper, we proposed a context-aware conversation adaptation system (CACAS) for human-robot interaction (HRI). First, a context recognition module and a language processing module are developed to obtain the context information, user intent and slots, which become part of the system state. Second, a reinforcement learning algorithm is utilized to train an initial policy in a simulated HRI environment. User feedback data is collected through HRI using the initial policy. Third, a new policy that combines the reinforcement learning-based policy and a supervised learning-based policy is adapted based on the user feedback. We conducted both simulated user tests and real human subject tests to evaluate the proposed CACAS. The results show that the CACAS achieved a success rate of 85% in the real human subject test and 87.5% of participants were satisfied with the adaptation results. For the simulation test, the CACAS had the highest success rate compared with the baseline methods.
Multi-modal large language models (LLMs) are expected to significantly enhance the intelligence of home service robots. However, reliance on cloud processing of raw visual data poses critical privacy risks. To address this problem, we propose a novel two-stage cloud-edge hybrid architecture for robots in domestic environments. This architecture employs a lightweight local LLM to perform sensitive content screening and semantic abstraction before transmitting the data to a more powerful cloud-based LLM for high-level planning and reasoning. Experiments with our end-to-end system demonstrate that it effectively protects a wide range of private data with minimal impact on task success rates. Without modifying cloud models, our approach offers a deployable performanceprivacy trade-off for home robots, advancing safe and socially acceptable autonomy.
To recognize continuous hand gestures from real-time video streams quickly and enable natural human-robot interaction, this paper proposes continuous hand gesture-based human-robot interaction for indoor mobile robot based on Multi-scale Local and Global skeleton Spatiotemporal feature extraction Network (MLGSNet). Firstly, we propose spatiotemporal feature extraction method of hand skeletal joint flow based on uniform sampling to address the challenge of extracting spatiotemporal features from videos. Subsequently, we propose multi-scale local and global spatiotemporal graph convolution network for hand gesture recognition to effectively capture both local and global dynamic spatiotemporal features during gesture execution, which is characterized by combining multi-scale attention, graph and temporal convolution. Furthermore, we design hand gesture activation based on mean filtering of dual confidence to accurately activate continuous gesture streams and recognize gesture command units under uneven gesture stream distribution. Experimental results on two public datasets demonstrate that the MLGSNet achieves state-of-the-art performance in both isolated and continuous gesture recognition. Finally, case study on the robot platform shows an overall human-robot interaction performance of 93.05
In a mixed-traffic environment, where autonomous vehicles (AVs) and human-driven vehicles are operating side by side, cooperative driving can become a useful tool to facilitate safe coordination. Unlike AVs, human-driven vehicles are difficult to model and control. Existing literature is largely focused on collaboration between AVs considering human vehicles as obstacles or unconnected and uncontrollable entities. In this paper, we present a cooperative driving framework that actively influences human-driven vehicles via advising commands and enables them to cooperate with AVs to achieve coordinated driving behaviors. We formulate cooperative driving as a stochastic model predictive control (sMPC) problem and consider human drivers' various stochastic aspects, such as attentiveness and tendency to follow an advisory action. The solutions to the sMPC provide advisory for longitudinal actions and lane change to human-driven vehicles and control commands to autonomous vehicles. With simulation and human-in-the-loop experimental results, we examine human drivers' reactions in complex driving scenarios and demonstrate the effectiveness of the developed method.
Most existing simultaneous localization and mapping (SLAM) algorithms demonstrate strong performance in static environments but encounter notable challenges in complex dynamic scenarios, particularly with camera rotation and objects moving along approximate epipolar lines. Moreover, dynamic environments with low texture further exacerbate positioning challenges. To address these issues, we propose TS-VINS, a novel visual-inertial SLAM (VINS) framework. TS-VINS enhances feature point extraction and tracking through a well-designed mesh-based method and a feature point group (FPG)-based mistracking detector, ensuring stability even in low-texture scenes. Additionally, this method incorporates a two-stage dynamic feature recognition (TDFR) algorithm that uses geometric constraints to distinguish between static and dynamic objects. Extensive experiments on public datasets, including VIODE and KITTI, and real-world scenarios demonstrate that TS-VINS notably improves dynamic object recognition, system robustness, and trajectory accuracy, outperforming the state-of-the-art (SOTA) SLAM systems in dynamic environments with low texture.
In this paper, a collaborative activity monitoring system (CAMS) is developed by combining a smart watch and a robot for elderly care. To improve the performance of ADL monitoring, an optimization problem on sensor selection is formulated and solved. First, we presented an overview of the CAMS. Second, to balance activity recognition accuracy, power consumption on the watch, and privacy preferences, we developed a Deep Q-Learning (DQL) model that enables the robot to learn optimal sensor selection strategies, ensuring adaptive and efficient monitoring. Third, the proposed method was evaluated using both offline and real time data, which were collected in a smart home testbed and a real apartment, respectively. The results showed that, compared with the baseline methods, the proposed method could recognize ADLs with higher accuracy while saving energy and respecting users' privacy preferences.
Natural language understanding is crucial for home robots to help people in their daily lives. However, existing home robots mainly rely on keyword matching to understand explicit commands, while struggling with understanding human instructions expressed more implicitly and naturally. Recently, large language models (LLMs) have demonstrated great potential in human language understanding. In this paper, we explore the application of LLM in home robots with a focus on navigation, one of the core capabilities in mobile robots. First, we designed and implemented an LLM-assisted robot navigation framework which adopts a modular architecture to integrate human-robot interaction, semantic mapping, and motion planning, thereby enhancing scalability and deployment flexibility. Second, we established a systematic method and benchmarks to evaluate the performance of LLMs in robot navigation applications, quantifying the performance differences between cloud-based and local LLM models. Finally, we conducted experiments on a custom-designed mobile robot in a real apartment setting. Our findings provide practical insight in selecting and optimizing LLMs for robotic applications.
For mobile robots to autonomously traverse the scene and perform interactions with the environment, perceiving and understanding the surroundings is essential. Many current incremental panoptic mapping frameworks rely on depth map segmentation algorithm to discover 3D objects. However, depth maps cannot reflect the information of objects with vanishing relative depths, and the depth map acquired by a consumer-grade RGB-D camera degrades with increasing distance, adversely affecting algorithm performance when relying solely on this input. This paper presents a novel incremental panoptic mapping framework that deals with the above problems by fusing multi-source information. We first propose a panoptic mask fusion algorithm to discover 3D objects by fusing information from color images and depth maps, which not only notably reduces the impact of long-distance depth map degradation, but also additionally discovers objects with relatively low depth. To further improve the mapping accuracy, a 3D instance fusion algorithm that incorporates semantic information, geometric information and priori knowledge is designed to achieve map regularization. Extensive experiments demonstrate that compared to the state-of-the-art incremental panoptic mapping framework, our method improves the average panoptic quality of thing classes by 4.1% and 3% on the SceneNN and ScanNet v2 dataset, respectively. Our approach even discovers many objects that are not annotated in the datasets, but are truly present. Furthermore, evaluations were conducted on a robot platform in different real-world scenarios characterized by low-depth objects and significant depth map degradation, demonstrating the reliability of our approach for real robot environment perception.
The ability to search for objects is a fundamental prerequisite for mobile robots when addressing a wide range of automation tasks. However, how to effectively estimate the positions of unobserved objects in a continuously changing environment remains an open challenge. Previous works have utilized probabilistic models to estimate the co-occurrence property between the target object and the observed landmark objects in a scene. However, few approaches can predict the precise spatial relations between objects based on a specific scene configuration. In this letter, we propose a novel unobserved object localization framework that achieves context-specific relation prediction based on the particular configuration of a scene. First, we leverage a 3D scene graph as a compact representation of the environment and propose a relation prediction model based on graph neural networks. This model can effectively interpret the information provided by the 3D scene graph and make accurate relation predictions. Second, to address the challenge of a high number of non-existent links between objects in the scene graph, we introduce a novel loss function that can better address imbalanced training data. Additionally, we propose an evaluation framework to comprehensively assess whether the relation prediction model benefits object search tasks. Comprehensive evaluation results obtained on public datasets and real-world scenes reveal the superiority of our method over competing approaches.
Cracks are the most common damage type on the pavement surface. Usually, pavement cracks, especially small cracks, are difficult to be accurately identified due to background interference. Accurate and fast automatic road crack detection play a vital role in assessing pavement conditions. Thus, this paper proposes an efficient lightweight encoder–decoder network for automatically detecting pavement cracks at the pixel level. Taking advantage of a novel encoder–decoder architecture integrating a new type of hybrid attention blocks and residual blocks (RBs), the proposed network can achieve an extremely lightweight model with more accurate detection of pavement crack pixels. An image dataset consisting of 789 images of pavement cracks acquired by a self‐designed mobile robot is built and utilized to train and evaluate the proposed network. Comprehensive experiments demonstrate that the proposed network performs better than the state‐of‐the‐art methods on the self‐built dataset as well as three other public datasets (CamCrack789, Crack500, CFD, and DeepCrack237), achieving F1 scores of 94.94%, 82.95%, 95.74%, and 92.51%, respectively. Additionally, ablation studies validate the effectiveness of integrating the RBs and the proposed hybrid attention mechanisms. By introducing depth‐wise separable convolutions, an even more lightweight version of the proposed network is created, which has a comparable performance and achieves the fastest inference speed with a model parameter size of only 0.57 M. The developed mobile robot system can effectively detect pavement cracks in real scenarios at a speed of 25 frames per second.
To ensure safe cooperative driving in mixed traffic with both manned and unmanned vehicles, it is crucial to understand and model the driving styles of human drivers. This paper explores how to develop accurate recognition of driving style and use that for the prediction of vehicle motion, which enables better performance in cooperative driving. A simulation testbed that consists of a driving simulator and a copilot is first introduced for the purpose of data collection and testing. A Long Short-Term Memory (LSTM)-based network that models human driving styles and predicts driving acceleration is developed. Standalone tests are conducted to examine the model performance in the simulation testbed. Finally, the model is evaluated in a series of merging experiments that involves 5 vehicles.
This paper presents a cooperative driving testbed based on vehicle-to-vehicle (V2V) communication, which can be used for research in intelligent transportation systems, such as collision avoidance in mixed traffic of both human-driven vehicles and autonomous vehicles. To achieve the goal, an intelligent copilot is developed. The copilot can share the data regarding vehicle status, intention, etc, with other nearby vehicles through V2V communication. Several case studies are conducted to validate the proposed testbed and evaluate the performances of cooperative driving. When dangerous situations occur, the copilot solves the collision avoidance problem using Mixed Integer Programming (MIP), which either provides control commands to the autonomous vehicle, or advises the human driver to take action. Experimental results show that the safety and stability of the involved vehicles have been significantly enhanced. This cooperative driving testbed can be used by researchers to develop and test cooperative driving algorithms before they are deployed in real vehicles.
Different people have different preferences when it comes to human-robot interaction. Therefore, it is desirable for the robot to adapt its actions to fit users’ preferences. Human feedback is essential to facilitating robot adaptation. However, when the task is complex or the robot action space is large, it requires a large amount of user feedback. ChatGPT is a powerful generative AI tool based on large language models (LLMs), which possesses a significant corpus of information obtained from human society, and exhibits robust proficiency in the comprehension and acquisition of natural language. Therefore, in this paper, we proposed a ChatGPT-powered adaptation system (ChatAdp) for human-robot interaction which requires less user feedback to achieve a good adaptation result. In the proposed ChatAdp, we use ChatGPT as a user simulator to provide feedback. We evaluated ChatAdp in a case study for context-aware conversation adaptation. The results are very promising. Our proposed method can achieve a mean success rate of 92% on the user’s natural language-described preferences after receiving 33 rounds of feedback from a user on average, which is only 2% of the number of states covered by the user preferences and outperforms the two baseline methods.
Automatic pavement crack detection is an important task to ensure the functional performances of pavements during their service life. Inspired by deep learning (DL), the encoder-decoder framework is a powerful tool for crack detection. However, these models are usually open-loop (OL) systems that tend to treat thin cracks as the background. Meanwhile, these models can not automatically correct errors in the prediction, nor can it adapt to the changes of the environment to automatically extract and detect thin cracks. To tackle this problem, we embed closed-loop feedback (CLF) into the neural network so that the model could learn to correct errors on its own, based on generative adversarial networks (GAN). The resulting model is called CrackCLF and includes the front and back ends, i.e. segmentation and adversarial network. The front end with U-shape framework is employed to generate crack maps, and the back end with a multi-scale loss function is used to correct higher-order inconsistencies between labels and crack maps (generated by the front end) to address open-loop system issues. Empirical results show that the proposed CrackCLF outperforms others methods on three public datasets. Moreover, the proposed CLF can be defined as a plug and play module, which can be embedded into different neural network models to improve their performances.
Object goal navigation tasks are critical for robots operating in unfamiliar environments, where they must locate specific objects using visual cues. The ability to leverage prior knowledge significantly enhances a robot’s associative capabilities, leading to improved navigation performance. However, existing methods struggle with the generalization challenge when transferring navigation models to new environments, a key issue addressed in this paper. To overcome this challenge, on the one hand, a time-varying knowledge graph is proposed to update the prior knowledge graph with context vectors derived from co-occurrence objects in the current observation. This approach prioritizes local graphs centered around the target and co-occurring objects, allowing for efficient and accurate target localization. Furthermore, the dynamic updating mechanism facilitates efficient exploration in new scenarios. On the other hand, to embed prior knowledge more rationally in the reinforcement learning-based navigation strategy, a time-varying knowledge graph inference network (TVGN) is presented. The TVGN utilizes context vectors and global spatial semantic information to perceive and understand the environment in real-time. It formulates navigation strategies based on the precise goal information encoded within the graph, thereby enhancing the robot’s efficiency in reaching the target. Based on the widely applied dataset AI2-THOR, extensive comparative experiments are conducted to illustrate the effectiveness of the proposed method. Experimental results indicate that our model outperforms state-of-the-art competitors, demonstrating notable advantages in navigation effectiveness and efficiency in previously unseen environments.
In this article, we presented a multimodal approach to monitor older adults' activities of daily living (ADLs) using the combination of a wearable device and a companion robot. A dynamic Bayesian network (DBN) model was developed for activity recognition, which fuses different data, including location, object, sound event, body action, and time. The walking action is detected as the transition between consecutive activities, which helps capture the inception of activities and save energy on the wearable device. Three tests were conducted to evaluate the proposed approach. First, multiple daily activities were simulated and evaluated the approach based on a public ADL dataset. Second, the proposed approach was tested based on an offline dataset collected in our smart home testbed, which contains images, sound events, motion, and time data. Third, the proposed approach was tested in real time and a web-based interface was developed, which helps caregivers better monitor the ADLs of older adults and provide further assistance. In the offline test and the real-time test, the results show that the system achieved 91% and 93% activity detection ratio, respectively, which significantly outperformed the baseline periodic sampling methods. In addition, the camera and microphone sensor trigger times were reduced from 1537 to 140 and 78, leading to energy reduction of 36.0% and 37.6% on the wearable device, respectively.
In this paper, we developed an activities of daily living (ADLs) monitoring system using a smart watch and a robot for elderly care. In order to balance the activity recognition accuracy, privacy concerns, robot resource cost and power consumption on the watch during monitoring, we proposed a Deep Q Learning model to solve a sensor selection problem. Based on the above four criteria, the robot runs the Deep Q Learning algorithm to decide whether to activate the robot sensors or the watch sensors for data collection. First, we presented the overview of the ADL monitoring system. Second, we developed the Deep Q Learning algorithm by considering the transition motions as part of the environment states. Third, we created a smart watch application for data collection and communication between the robot and the watch. Finally, the proposed model was trained and evaluated based on both offline data and real time data collected in our smart home testbed. The results showed that the proposed method could recognize ADLs with high accuracy while saving about 14.6% energy compared with the baseline periodic methods.
Background: As the older adult population increases there is a great need of developing smart healthcare technologies to assist older adults. Robot-based homecare systems are a promising solution to achieving this goal. This study aims to summarize the recent research in homecare robots, understand user needs and identify the future research directions. Methods: First, we present an overview of the state-of-the-art in homecare robots, including the design and functions of our previously developed ASCC Companion Robot (ASCCBot). Second, we conducted a user study to understand the stake- holders’ opinions and needs regarding homecare robots. Finally, we proposed the future research directions in this exciting research area in response to the existing problems. Results: Our user study shows that most of the interviewees emphasized the importance of medication reminder and fall detection functions. The stakeholders also emphasized the functions to enhance the connection between older adults and their families and friends, as well as the functions to improve the efficiency and productivity of the caregivers. We also identified three major future directions in this exciting research area: human-machine interface, learning and adaptation, and privacy protection. Conclusions: The user study discovered some new useful functions that the stakeholders want to have and also validated the developed functions of the ASCCBot. The three major future directions in the homecare robot research area were identified.
3D Panoptic perception is essential for the understanding of real-world environment and plays an increasingly important role in the field of robotics. However, most existing methods heavily rely on image panoptic segmentation networks to acquire panoptic information of the environment, which is time-consuming and susceptible to interference. In this paper, we propose a novel and efficient panoptic mapping method based on multi-source information. Specifically, to improve the real-time performance of the system, we first apply lightweight object detection and semantic segmentation to extract 2D semantic and instance information from images. Second, a panoptic inference algorithm is designed that fully utilizes multi-source information, including geometry-based and learning-based information, to simultaneously reason about background and foreground objects in the environment. Finally, we take advantage of the scalability of the framework by introducing a multi-object tracking algorithm into the framework, thus providing the temporal information among consecutive frames to the data association module. Based on two popular datasets, extensive comparison experiments are conducted to illustrate the effectiveness of the proposed method. Experimental results show that compared with state-of-the-art panoptic mapping methods, the proposed method achieves superior performance in accuracy, real-timeness and stability. Furthermore, we also evaluate our method in real-world scenarios and CPU-only device to demonstrate the feasibility of its practical deployment.
Existing conversational robots are mostly reactive in that the interactions are usually initiated by the users. With the knowledge of the environmental context such as people's daily activities, robots can be more intelligent and proactive. In this paper, we proposed a context-aware conversation adaptation system (CACAS) for human-robot interaction (HRI). First, a context recognition module and a language processing module are developed to obtain the context information, user intent and slots, which become part of the state. Second, a reinforcement learning algorithm is developed to train an initial policy with a simulated user. User feedback data is collected through HRI using the initial policy. Third, a policy combining the reinforcement learning-based policy with the neural network-based policy is adapted based on the user feedback. We conducted both simulated user tests and real human subject tests to evaluate the proposed system. The results show that CACAS achieved a success rate of 85% in the real human subject test and 87.5% of participants were satisfied with the adaptation results. For the simulation test, CACAS had the highest success rate compared with the baseline methods.