We present V-HPOT, a novel approach for improving the cross-domain performance of 3D hand pose estimation from egocentric images across diverse, unseen domains. State-of-the-art methods demonstrate strong performance when trained and tested within the same domain. However, they struggle to generalise to new environments due to limited training data and depth perception – overfitting to specific camera intrinsics. Our method addresses this by estimating keypoint z-coordinates in a virtual camera space, normalised by focal length and image size, enabling camera-agnostic depth prediction. We further leverage this invariance to camera intrinsics to propose a self-supervised test-time optimisation strategy that refines the model's depth perception during inference. This is achieved by applying a 3D consistency loss between predicted and in-space scale-transformed hand poses, allowing the model to adapt to target domain characteristics without requiring ground truth annotations. V-HPOT significantly improves 3D hand pose estimation performance in cross-domain scenarios, achieving a 71
Stroke represents the third cause of death and disability worldwide, and is recognised as a significant global health problem. A major challenge for stroke survivors is persistent hand dysfunction, which severely affects the ability to perform daily activities and the overall quality of life. In order to regain their functional hand ability, stroke survivors need rehabilitation therapy. However, traditional rehabilitation requires continuous medical support, creating dependency on an overburdened healthcare system. In this paper, we explore the use of egocentric recordings from commercially available smart glasses, specifically RayBan Stories, for remote hand rehabilitation. Our approach includes offline experiments to evaluate the potential of smart glasses for automatic exercise recognition, exercise form evaluation and repetition counting. We present REST-HANDS, the first dataset of egocentric hand exercise videos. Using state-of-the-art methods, we establish benchmarks with high accuracy rates for exercise recognition (98.55 https://github.com/wiktormucha/rest-hands .
The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by partial vision impairment, which leads to discomfort and dependency in daily activities. To address the limitations of this part of society, we propose an intelligent reading assistant based on smart glasses with embedded RGB cameras and a Large Language Model (LLM), whose functionality goes beyond corrective lenses. The video recorded from the egocentric perspective of a person wearing the glasses is processed to localise text information using object detection and optical character recognition methods. The LLM processes the data and allows the user to interact with the text and responds to a given query, thus extending the functionality of corrective lenses with the ability to find and summarize knowledge from the text. To evaluate our method, we create a chat-based application that allows the user to interact with the system. The evaluation is conducted in a real-world setting, such as reading menus in a restaurant, and involves four participants. The results show robust accuracy in text retrieval. The system not only provides accurate meal suggestions but also achieves high user satisfaction, highlighting the potential of smart glasses and LLMs in assisting people with special needs.
Hand pose represents key information for action recognition in the egocentric perspective, where the user is interacting with objects. We propose to improve egocentric 3D hand pose estimation based on RGB frames only by using pseudo-depth images. Incorporating state-of-the-art single RGB image depth estimation techniques, we generate pseudo-depth representations of the frames and use distance knowledge to segment irrelevant parts of the scene. The resulting depth maps are then used as segmentation masks for the RGB frames. Experimental results on H2O Dataset confirm the high accuracy of the estimated pose with our method in an action recognition task. The 3D hand pose, together with information from object detection, is processed by a transformer-based action recognition network, resulting in an accuracy of 91.73 https://github.com/wiktormucha/SHARP .
Action recognition is essential for egocentric video understanding, allowing automatic and continuous monitoring of Activities of Daily Living (ADLs) without user effort. Existing literature focuses on 3D hand pose input, which requires computationally intensive depth estimation networks or wearing an uncomfortable depth sensor. In contrast, there has been insufficient research in understanding 2D hand pose for egocentric action recognition, despite the availability of user-friendly smart glasses in the market capable of capturing a single RGB image. Our study aims to fill this research gap by exploring the field of 2D hand pose estimation for egocentric action recognition, making two contributions. Firstly, we introduce two novel approaches for 2D hand pose estimation, namely EffHandNet for single-hand estimation and EffHandEgoNet, tailored for an egocentric perspective, capturing interactions between hands and objects. Both methods outperform state-of-the-art models on H2O and FPHA public benchmarks. Secondly, we present a robust action recognition architecture from 2D hand and object poses. This method incorporates EffHandEgoNet, and a transformer-based action recognition method. Evaluated on H2O and FPHA datasets, our architecture has a faster inference time and achieves an accuracy of 91.32% and 94.43%, respectively, surpassing state of the art, including 3D-based methods. Our work demonstrates that using 2D skeletal data is a robust approach for egocentric action understanding. Extensive evaluation and ablation studies show the impact of the hand pose estimation approach, and how each input affects the overall performance. The code is available at https://github.com/wiktormucha/effhandegonet.
Egocentric action recognition is essential for healthcare and assistive technology that relies on egocentric cameras because it allows for the automatic and continuous monitoring of activities of daily living (ADLs) without requiring any conscious effort from the user. This study explores the feasibility of using 2D hand and object pose information for egocentric action recognition. While current literature focuses on 3D hand pose information, our work shows that using 2D skeleton data is a promising approach for hand-based action classification, might offer privacy enhancement, and could be less computationally demanding. The study uses a state-of-the-art transformer-based method to classify sequences and achieves validation results of 94%, outperforming other existing solutions. The accuracy of the test subset drops to 76%, indicating the need for further generalization improvement. This research highlights the potential of 2D hand and object pose information for action recognition tasks and offers a promising alternative to 3D-based methods.
Action recognition is at the core of egocentric camera-based assistive technologies, as it enables automatic and continuous monitoring of Activities of Daily Living (ADLs) without any conscious effort on the part of the user. This study explores the feasibility of using 2D hand and object pose information for egocentric action recognition. While current literature focuses on 3D hand pose information, our work shows that using 2D skeleton data is a promising approach for hand-based action classification and potentially allows for reduced computational power. The study implements a state-of-the-art transformer-based method to recognise actions. Our approach achieves an accuracy of 95% in validation and 88% in test subsets on the publicly available benchmark, outperforming other existing solutions by 9% and proving that the presented technique offers a successful alternative to 3D-based approaches. Finally, the ablation study shows the significance of each network input and explores potential ways to improve the presented methodology in future research.
Active and Assisted Leaving (AAL) devices that use cameras and images raise concerns about the privacy of the monitored individuals. These devices capture images that include personal and behavioral data during the day. Most authors decide to switch from RGB to depth sensors to maintain privacy. Nevertheless, not all available works agree that depth image is private, which creates an open legal problem for AAL applications. In this paper, privacy is discussed in vision-based systems using depth sensors. Various factors of depth and RGB images that might affect privacy are presented to define the privacy level of depth devices. One of the main issues that make an image non-private is that the subjects' faces are visible and can be identified. In the experimental part, a state-of-the-art Face Recognition (FR) model in depth images is developed. It is used to establish boundary conditions allowing correct recognition of a person's face. A comparison between FR in RGB and depth images is performed, including the ability to learn the model by training the two modalities from scratch on identical data. This study answers under which conditions depth cameras protect the privacy and how much privacy is disclosed by them.
Image-based assistive solutions raise concerns about the privacy of the individuals being monitored. The issue involves the situation when such technology is used in medical institutions to protect patients' health and support the personnel. These devices are installed in facilities and process images that include personal and behavioral data during the day. Other types of images than RGB are used to maintain privacy in this type of application, like depth images. Usage of depth cameras in the majority of publications is considered private protective. This paper discusses the issue of privacy in vision-based applications using depth modality. The factors affecting privacy in depth images are presented. The main problem that makes an image non-private is that the subjects' faces allow identification. This paper compares the Face Recognition (FR) technique between RGB and depth images. In the experimental part, a state-of-the-art model for FR in depth images is developed, which is used to establish boundary conditions when a person is recognized. The performance of FR between these two modalities is compared on two existing datasets containing images in both versions, including the training process. The study aims to determine under which conditions depth cameras preserve privacy and how much privacy they reveal.
TMFD Dataset addresses the lack of available depth and thermal face detection benchmarks. The corresponding depth and thermal data frames accompany each RGB image in the dataset. The dataset encompasses a wide range of variations. The samples include different numbers of people in the scene, different individuals, various types of backgrounds, varying distances of measurement, and wearable accessories worn by the individuals, such as hoodies, headphones, hats, glasses, and face masks. Additionally, the dataset includes different types of illumination in the scene. The dataset is categorized into three separate groups based on the complexity and difficulty level of face detection. The first subset represents the easiest detection conditions, with images containing a single person against a simple background. The sensor is positioned close to the target. This subset consists of 781 images for each modality. The second group represents medium detection conditions. The sensor placement remains the same as in the first subset, but additional variety is introduced. Individuals are captured wearing accessories that serve as obstacles to make detection more challenging. These accessories include hoodies, headphones, hats, and glasses. The second group also includes images with multiple persons in the scene to test the performance of multiple face detection. There are 965 images for each modality in this group. The last group represents the most challenging images for detection. In this case, the sensor is placed further away from the target. The scenes feature various obstacles, such as computer screens, plants, desks, and chairs, with individuals interacting with them. The presence of singular and plural individuals is also included. A distinct subset within this collection consists of images from a previous, easier scene, but with the addition of face mask usage to increase the difficulty of the prediction task. This particular group contains 1062 images for each modality. This database may be used for non-commercial research purpose only. If you publish material based on this database, we request you to include a reference to: Mucha W., Kampel M. (2022) “Depth and Thermal Images in Face Detection – A Detailed Comparison Between Image Modalities”, The 5th International Conference on Machine Vision and Applications (ICMVA 2022), February 18-20, 2022, Singapore https://doi.org/10.1145/3523111.3523114
The report illustrates the state of the art of the most successful AAL applications and functions based on audio and video data, namely (i) lifelogging and self-monitoring, (ii) remote monitoring of vital signs, (iii) emotional state recognition, (iv) food intake monitoring, activity and behaviour recognition, (v) activity and personal assistance, (vi) gesture recognition, (vii) fall detection and prevention, (viii) mobility assessment and frailty recognition, and (ix) cognitive and motor rehabilitation. For these application scenarios, the report illustrates the state of play in terms of scientific advances, available products and research project. The open challenges are also highlighted.
Face detection is a well-known issue in image processing, and numerous studies are present in this field. A prominent part of the work is devoted to RGB images, leaving depth and thermal data with less interest. However, in some conditions like low-light areas where face detection is needed, non-RGB sensors might perform better. Also, mounting an additional RGB camera could be challenging or not possible, considering privacy concerns. In this work, current deep learning methodologies are employed to train depth and thermal detection models. The training is done using combined publicly available data that is processed by us for this purpose in order to create necessary annotations for a learning process. The resulting models are validated on a new trimodal dataset collected for this experiments purpose. It contains images captured with RGB, depth, and thermal sensors. Various scenes with single and multiple faces appearances can be found. The results show that non-RGB solutions can be applied in practice with highly robust accuracy and their efficiency is close to RGB detectors. However, their performance depends on the environment and that circumstances are described later in this article.