
The widespread adoption of depth cameras has been driven by their cost-effectiveness and depth measurement reliability across diverse applications. These sensors, however, face common challenges inherent to optical systems, particularly motion blur artifacts when deployed on moving platforms. This paper introduces a novel end-to-end approach for motion blur removal in Amplitude Modulated Continuous Wave (AMCW) depth cameras, specifically designed to handle High Dynamic Range (HDR) depth imagery. Our method leverages an optical flow network trained on Differential Correlation Sample (DCS) images to achieve proper alignment prior to phase unwrapping. We present a new real-HDR DCS dataset and develop an unsupervised training framework utilizing custom loss functions with adaptive scheduling, eliminating the need for ground truth data. While our implementation did not achieve fully successful optical flow generation for phase unwrapping, our experiments demonstrate the potential viability of optical flow-based approaches for depth image deblurring.
The emergency evacuation systems are often installed in crowded places such as airports, train stations, and shopping malls to evacuate and protect people when incidents occur quickly. With the development of surveillance cameras, vision-based emergency evacuation systems have demonstrated their ability to observe and promptly warn flexibly. This paper proposes a human behavior detector by fine-tuning the YOLOv11n detection network with the Global Attention Mechanism (GAM) to enhance the individual human action recognition. Extensive experiments are trained and evaluated on the Human Behavior Detection Dataset (HBDset) using a NVIDIA Tesla V100 32GB GPU. The proposed network achieves 62.2% of mAP and an inference speed of 1.3 milliseconds (ms), and outperforms other networks of the same scale.
The paper presents preliminary analyses of pulse transit time (PTT) derived from photoplethysmography (PPG) measurements taken at the elbow, wrist, and finger. Data were collected from ten participants under the approval of a bioethics committee. The study examined the correlation coefficients of PTT between these locations using four fiducial points of the PPG signal. The results show strong dependence on the individual participant, the choice of fiducial point, and the sensor placement (the correlation coefficient varies from -0.69 to 1.00). These findings highlight the need for further analysis on a larger and more diverse population, including individuals with varying health conditions and broader ranges of blood pressure fluctuations.
Currently, monocular 3D object detection has attracted increasing attention due to its cost-effectiveness and scalability compared to LiDAR-based solutions or stereo-based solutions. However, the inherent absence of depth information in a single RGB image introduces significant uncertainty in 3D localization and pose estimation of objects. To address the ambiguity in estimating object dimensions and depth, many existing approaches rely on complex and heavy neural network architectures, which often hinder their practicality in real-world applications. This paper proposes a novel and lightweight multi-task framework for monocular 3D object detection, which leverages keypoint estimation and 2D object detection to infer 3D spatial attributes. Extensive experiments conducted on the KITTI benchmark demonstrate the efficiency and effectiveness of the proposed method. Notably, our model achieves an inference speed of up to 40 FPS, making it highly suitable for real-time deployment in intelligent driving scenarios.
The advancement of robotics has been driven by the integration of artificial intelligence, machine learning, and sophisticated sensing technologies, enabling more seamless Human-Robot Interaction (HRI). Facial Attribute Classifier (FAC) plays a crucial role in HRI by helping robots understand human emotions, intentions, and social cues, fostering personalized and intuitive interactions. However, while existing methods achieve high accuracy, their computational complexity limits real-time applications on low-cost or CPU-based devices, highlighting the need for lightweight models that balance accuracy and efficiency. This work proposes an Efficient Network (ENet) designed to achieve an optimal trade-off between efficiency and accuracy of FAC. ENet introduces an Enhanced Sequential Efficient Attention Module (ESEAM) to improve the quality of feature maps while maintaining high efficiency. Accordingly, ENet demonstrates a compromise between efficiency and accuracy on the CelebA and LFWA datasets. The proposed ENet is computationally efficient, generating a few parameters, making it well-suited for CPU-based applications. When combined with a face detector, the optimized FAC achieves a processing speed of 25.88 frames per second (FPS) on an Intel Core i7-9750H CPU, demonstrating its suitability for real-time use.
Peripheral artery disease (PAD), condition in which narrowed arteries reduce blood flow to the arms or legs It is estimated that 13% of the population over 50 years old suffers either from symptomatic or asymptomatic PAD. Ankle-brachial index (ABI) due to relatively low sensitivity and specificity is not recommended for screening. In this article we evaluate use of Active Dynamic Thermography as the alternative technique for the ischemia classification. By initially cooling the subject's tissue and then observing with a thermal camera, we were able to visualize vascularization and measure the tissue's ability to return to its initial state. Thermographic parameters obtained correspond well with age of the patient.
As vehicle systems become increasingly connected and intelligent, insurance providers are turning to machine learning techniques to personalize billing based on individual driving behavior. This shift raises important questions about how to balance predictive performance with user privacy. In this paper, we present PrivFedProfiling, a decentralized privacy-preserving learning framework designed for use-based insurance (UBI) systems. Our method leverages Federated Learning (FL) to collaboratively train behavior models across distributed driver devices without transferring raw data. To further strengthen privacy, we integrate Differential Privacy (DP) and Homomorphic Encryption (HE) within the training process, protecting sensitive patterns in shared model updates. The proposed approach uses a Multilayer Perceptron (MLP) architecture and is validated using synthetic driving behavior data generated from the SUMO simulator. It offers a realistic yet controllable environment for testing. Results indicate that our method maintains high model accuracy while ensuring strong privacy guarantees, making it suitable for real-world deployment.
Monkeypox was declared a public health emergency of international concern by WHO in July 2022 due to the unprecedented global spread of the disease outside of previously endemic countries in Africa. Previously proposed classification models based mainly on pre-trained CNNs are ineffective in detecting Monkeypox. In this paper, a novel model is presented for the Monkeypox virus detection hybrid window attention and depthwise asymmetric dilation convolution. This study's novelty lies in the proposed lightweight and robust multi-stage model, utilizing depthwise dilation convolution and pointwise convolution in the two first stages with a large spatial dimension. Window attention combination with depthwise asymmetric dilation convolution is applied at the two last stages to reduce model complexity. We have evaluated the proposed approach with ViT, Swin Transformer, and MaxViT, DenseNet201, and ResNet50 on the public dataset of “Monkeypox Skin Images Dataset” (MSID) on Mendeley. The model performance was evaluated using metrics such as accuracy, recall, precision, specificity, and F1-score. The proposed method achieves the best results with an average of 99.30%, 98.65%, 98.60%, 99.56%, and 98.60% in accuracy, precision, recall, and specificity, F1-score., respectively. Additionally, the proposed model has only 19M parameters, which is smaller than MaxViT and Swin Transformer.
In this study, we investigate the robustness of various deep learning models to inpainted lung CT scans. We train several image classification models (EfficientNet, MobileNet, DenseNet, and a custom CNN) as well as one object detection model (YOLO) to determine whether a scan contains a tumor or not. We then apply a pretrained GAN to inpaint tumor regions in test images, effectively hiding cancer presence. The resulting set of modified images is then evaluated by the previously trained models. We observe that the classifiers test accuracy remains almost unchanged, whereas the YOLO model adjusts its predictions, failing to detect tumors that were previously visible. This suggests that classification models may have learned non-obvious patterns (or biases) associated with cancer presence, possibly integrating features from larger image regions rather than focusing solely on localized tumor areas. In contrast, the YOLO model depends more on local visual features and is more vulnerable to inpainting-based adversarial attacks. These findings highlight the need to consider model architecture and susceptibility to image manipulation when deploying AI in human-system interaction tasks.
Occupancy detection and activity recognition are crucial for building automation and patient monitoring. Traditional motion sensors have limitations, such as binary output and the inability to detect stationary individuals, compromising detection accuracy. This study explores environmental sensors - such as ambient light, temperature, humidity, and sound levels - as alternatives or complements to motion sensors for occupancy detection. Using the “Smart Home environment data across 4 European countries”, compromising various sensor configurations in individual rooms, we utilize event-based sensors (primarily motion sensors) as proxy labels for presence. We train both generalized and personalized XGBoost models: the generalized model predicts occupancy in new environments, while the personalized model is tailored to individuals. Comparing model performance, we find that personalized models generally outperform generalized ones, though the best configurations' F1-scores are close (0.77 vs. 0.76). The sound level features are the most influential in both models, indicating that environmental sensors, particularly with sound data, can enhance occupancy detection and potentially reduce the dependence on motion sensors. This study improves our understanding of the importance and configurations of sensors for occupancy detection and highlights the advantages of personalized models in smart home applications.
On waste-sorting conveyor lines, illumination often fluctuates sharply or drops to very low levels, causing conventional object detectors to fail. To achieve lighting-robust performance without extra lamps, vision rooms, or site-specific retraining, we propose an Illumination-Invariant Convolution (IIC) block that can be inserted into any backbone network. Working in log-intensity space under a Lambertian model, IIC applies learnable zero-mean cross-channel filters that mute lighting artefacts and boost material cues, and then merges the resulting maps with base features to produce lighting-robust representations. We integrate IIC into the “nano” versions of YOLOv5/8/10/11 and lightweight RT-DETR, training on roughly 200 k conveyor-belt images (29 classes) from the AI-Hub waste dataset. IIC raises YOLO mAP50 by 2.5-4.5 pp and mAP50-95 by up to 4.6 pp; even the data-hungry, transformer-based RT-DETR gains up to 1.5 pp. In low-light video tests, the IIC-augmented models successfully detected objects the original networks missed, raising recall. This demonstrates that inserting the IIC module at a network's input provides a straightforward path to illumination-robust, field-ready waste-sorting systems.
Froth flotation is a crucial separation technique in the metal beneficiation process. Among the various control variables, acid dosage plays a significant role in adjusting the slurry pH and directly impacts flotation efficiency. However, traditional acid dosage prediction methods struggle to achieve accurate modeling due to the nonlinear, time-varying nature of the flotation process and the complex coupling among variables, thereby hindering process optimization and stable control. This study proposes a deep learning-based time series prediction model, Probformer, to address the acid dosage forecasting task. Probformer incorporates a time series decomposition mechanism and a probabilistic autocorrelation attention mechanism to effectively capture long-term trends and periodic variations during prediction. Experiments conducted on real-world industrial data demonstrate that Probformer outperforms conventional models such as Autoformer, Transformer, Informer, and Reformer in terms of prediction accuracy, achieving lower prediction errors and higher fitting accuracy. The results highlight strong performance of Probformer in acid dosage forecasting for froth flotation, indicating its promising potential for practical applications.
Object detection and tracking are two critical tasks in computer vision, widely applied in security surveillance, autonomous vehicles, and behavioral analysis. Strong performance in object detection has been demonstrated by recent Transformer-based models, such as RT-DETR (Real-Time Detection Transformer), due to their global context modeling and high accuracy. However, an inherent tracking mechanism is lacking in RT-DETR, which requires additional components to maintain identity consistency across frames. To address this limitation, an integration of RT-DETR with DeepSORT is proposed, leveraging the strengths of both models to enhance real-time object detection and tracking. A comparative evaluation with YOLOv8, a widely used real-time detector, is conducted to highlight the advantages of the proposed approach in tracking accuracy and robustness. Experiments show that effective performance is achieved in challenging scenarios such as object occlusion and intersection. Specifically, an IDF1 score of 60.0%, a MOTA of 42.4%, and a MOTP of 43.3% are obtained by RTl+DeepSORT on the MOT17-02-DPM dataset, outperforming YOLOv8x+DeepSORT. These results indicate that significant improvements in tracking accuracy are attained while real-time efficiency is maintained, making the proposed approach well-suited for applications such as intelligent surveillance and autonomous navigation.
Service robots need to be able to avoid pedestrians safely and reliably in environments where they coexist with pedestrians. However, it is difficult for robots to detect and avoid pedestrians from a safe distance because pedestrians at the end of a corridor or around a corner are in blind spots that the robot cannot detect with its external sensors. Therefore, this research proposes a robot navigation method that detects pedestrians in the robot's blind spot by receiving Bluetooth low energy (BLE) Beacon signals emitted by Bluetooth devices such as smartphones and wireless earphones that pedestrians generally have, and safely avoids them. By using the proposed method, our mobile robot can detect pedestrians that suddenly appear from a corner of a corridor earlier than when using a laser rangefinder, and avoid them safely in our preliminary experiment.
The auricular concha is a crucial component of the auricle and holds significant importance in areas such as auricular acupuncture point localization in traditional Chinese medicine, auricle recognition, and the design of ear-worn devices. In this study, a segmentation model was first employed to perform eight-region segmentation of auricle images, from which the auricular concha subregion images were extracted. Subsequently, considering the characteristics of ear-worn devices, key points within the auricular concha region were selected and detected using a keypoint detection model. A method for constructing geo-metric features of the auricular concha region was then designed. Finally, three clustering methods-K-means, Deep Clustering, and Autoencoder combined with K-means-were applied to the geometric features of the auricular concha region to generate multiple representative auricular concha templates, followed by comparative analysis. Experimental results demonstrate that the Autoencoder + K-means clustering method achieves superior performance in generating personalized design templates for ear-worn devices. This study provides robust technical support for applications in personalized intelligent wearable device design, auricular acupuncture point localization, and auricle recognition.
Generating descriptive annotations for person re-identification (Re-ID) images is essential for bridging vision and language domain, improving both interpretability and cross-modal retrieval performance. However, large vision-language models (LVLMs), which trained on broad web-scale corpora, often struggle to generate accurate, context-relevant descriptions for ReID samples due to inherent challenges such as occlusions, low resolution, varying illumination, and diverse viewpoints. In this paper, we propose to apply Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) to tune Qwen2-VL for ReID-specific captioning tasks. Leveraging existing Re-ID datasets with paired image-text annotations, our fine-tuned model generates domain-aligned and discriminative captions. Experiments show significant improvements in caption relevance and identity descriptiveness, highlighting the potential of PEFT-tuned LVLMs for real-world ReID applications.
Thermal image super-resolution remains a challenging task due to the limited spatial detail captured by infrared sensors. Although RGB-guided methods and domain adaptation techniques have shown promise, they often introduce architectural complexity or require multimodal inputs at inference time. In this study, we investigate the integration of Fourier Domain Adaptation (FDA) as a lightweight preprocessing strategy to improve the performance of the state-of-the-art Dense-Residual-Connected Transformer (DRCT) model for image super-resolution. FDA transfers low-frequency information from grayscale versions of RGB images to thermal images during training, enabling the model to incorporate additional structural information without introducing any additional architectural complexity or adding overhead during inference. Experimental results on the Tufts Face Database demonstrate that the FDA-enhanced model improves PSNR by up to 0.17 dB and SSIM by up to 0.0004 compared to thermal-only training at a x2 scaling factor. These findings highlight the effectiveness of simple frequency-based adaptation techniques in improving the generalization of thermal SR models while preserving the simplicity of the inference pipeline.
Biomedical engineering relies heavily on simulators as essential tools for evaluating and assessing various design options for Wireless Sensor Networks (WSN). In medical contexts, diverse wireless biosensors are deployed on the human body to monitor patients and deliver precise treatments. The topology and mobility of Wireless Body Sensor Networks (WBSNs) within medical healthcare systems can dynamically shift due to changes in posture and movement patterns. The research provides a comparative evaluation of diverse simulators designed for body sensor network protocols. The suggested framework is flexible and can be utilized with various simulators like OMNeT++, Castalia, and NS3. The results indicate that this comparative analysis plays a crucial role in improving biomedical engineering applications by assisting in the choice of suitable simulators.
This study presents a specialized system for detecting textual data in medical images, aimed at enhancing the anonymization of sensitive patient information. Conventional OCR tools, such as EasyOCR and docTR, exhibit notable limitations in medical settings - particularly in handling complex imagery such as lung CT scans - resulting in high false positive rates and suboptimal recall. To overcome these challenges, we adapted the YOLOv11n object detection model for the task of text localization in medical images. The model was trained on a hybrid dataset comprising real CT images with annotated labels and synthetically generated samples, improving robustness under various imaging conditions. In a comparative evaluation, YOLOv11n achieved a precision of 100.00%, recall of 92.09% and F1 score of 95.86%, significantly outperforming EasyOCR (F1 = 66.9%) and docTR (F1 = 89.58%). In particular, YOLOv11n did not produce false positives in the test set, which makes it highly effective in preserving relevant diagnostic information while accurately detecting sensitive text regions. These results underscore the potential of adapting state-of-the-art object detection frameworks such as YOLO to specialized tasks in medical image processing and privacy preservation.
The aim of this study was to develop a system capable of recognizing textual data in medical images and identifying which labels contain private information. We designed an intuitive, modular pipeline with extensibility in mind, allowing for the seamless integration or replacement of individual components as technologies evolve. This architecture ensures long-term adaptability to future advancements in OCR and machine learning. For text detection on images, we used a YOLO model fine-tuned on a specialized dataset of lung CT images. To improve text analysis, we integrated a large language model (LLM) into the workflow. Although effective, LLM introduces noticeable latency when processing numerous detected text regions, which may impact the overall image processing time. Nevertheless, with an accuracy of approximately 80% achieved on the test set, the system demonstrates robustness and scalability for the detection and interpretation of textual data in medical imaging contexts.