This study systematically evaluates the impact of sensor configuration, body location, classification granularity, and model choice on inertial-based human activity recognition in a laboratory dataset aligned with the Spanish IMPaCT cohort design. Data were collected from 85 participants instrumented with thigh-, wrist-, and hip-mounted inertial measurement units over a structured protocol of 13 semi-structured daily activities, a resting phase and a structured activity. After manual correction of timestamp drift, signals were segmented into overlapping 10-s windows and analyzed using convolutional neural networks, Random Forest, and XGBoost classifiers.Two classification targets were defined: fine-grained recognition of 15 laboratory-controlled activities and coarse-grained classification into four MET-based intensity levels. Results showed that classification granularity is the primary determinant of performance (F=224.85, p-value = 2.304×10−13 through the analysis of variance of the F1-score), with intensity-level classification substantially outperforming fine-grained activity recognition. Sensor configuration, model type, and body location also significantly influenced classification outcomes. Wrist-mounted sensors achieved the highest overall F1-scores. Incorporating gyroscope-derived features consistently improved performance across configurations, and feature importance analysis confirmed their substantial contribution. These findings, derived from models developed under controlled laboratory conditions, provide practical guidance for the design of wearable sensing protocols and modeling strategies in large-scale population-based studies, and support their extension to everyday physical activity, laying the foundation for future real-world applications.
Federated domain generalization has shown promising progress in image classification by enabling collaborative training across multiple clients without sharing raw data. However, its potential in the semantic segmentation of autonomous driving remains underexplored. In this paper, we propose FedS2R, the first one-shot federated domain generalization framework for synthetic-to-real semantic segmentation in autonomous driving. FedS2R comprises two components: an inconsistency-driven data augmentation strategy that generates images for unstable classes, and a multi-client knowledge distillation scheme with feature fusion that distills a global model from multiple client models. Experiments on five real-world datasets, Cityscapes, BDD100K, Mapillary, IDD, and ACDC, show that the global model significantly outperforms individual client models and is only 2 mIoU points behind the model trained with simultaneous access to all client data. These results demonstrate the effectiveness of FedS2R in synthetic-to-real semantic segmentation for autonomous driving under federated learning
This study introduces a magnetometer-free, custom-built inertial measurement unit (IMU) combined with an open-source sensor fusion algorithm and provides a systematic experimental comparison against commercial IMUs that integrate proprietary hardware and closed-source filtering software. The main innovation lies in demonstrating that an open and fully reproducible hardware and software architecture can achieve orientation estimation accuracy comparable to that of state-of-the-art commercial solutions under both static and dynamic conditions. Experimental results show that the proposed system achieves a mean orientation error of 2.10% (0.11 degrees) while consistently outperforming proprietary filtering approaches in static scenarios and exhibiting competitive variability during dynamic motion. Furthermore, the study presents a comprehensive benchmarking framework that analyzes the influence of sensor noise characteristics and filtering strategy on estimation performance across multiple hardware and algorithm configurations. These findings establish that custom-built IMUs paired with open-source sensor fusion algorithms constitute a reliable, cost-effective, and transparent alternative to commercial hardware and software solutions for biomedical and engineering applications.
Rendering 3D virtual scenarios has become a popular alternative for generating per-pixel-labeled image datasets, especially in fields like autonomous driving. The approach is valuable for training neural perception models, such as semantic segmentation models, particularly when data might be scarce, expensive, or difficult to collect. However, fundamental questions persist within the research community regarding the generation and processing of these synthetic images, particularly a better understanding of the key factors influencing the performance of deep learning models trained with such synthetic images. In response, we conducted a series of experiments to elucidate the impact that common aspects involved in the generation of rendered synthetic images may have on the performance of neural semantic segmentation tasks. Our study used a recent autonomous driving synthetic dataset as our main testbed, allowing us to investigate the effect of different approaches when modeling their geometric, material, and lighting details. We also studied the impact of rendering noise, typically produced by path-tracing algorithms, as well as the impact of using different color transformations and tonemapping algorithms.
3D multi-object tracking is crucial for enhancing the understanding of the environment in autonomous driving and robotics. Low-quality detections and less robust associations are two challenges in the point-aware tracking-by-detection paradigm. Conventional approaches suffer from inadequate pre-processing of detected outliers, and poor appearance-based associations during occlusion. To address these issues, this paper proposes a real-time and robust 3D multi-object tracking framework based on the fusion of camera and LiDAR data. Firstly, a two-level association strategy is introduced, whereby high-confidence tracks and detections are initially linked through a straightforward 3D IoU cost, followed by the association of remaining entities using discriminative deep appearance features, emphasizing the similarity between the recently updated track appearance and reemerging targets within dynamically constrained search boundaries. Secondly, a track drift compensation method is presented to refine the low-quality detections using their historically matched tracks, facilitating accurate updates accordingly. Experiments show that the proposed method achieved 79.36 % HOTA and 74 % AMOTA in KITTI and nuScenes benchmarks, respectively. This result surpasses many advanced solutions, particularly exhibiting robust performance in occluded environments.
Instance segmentation is crucial for autonomous driving, but is hindered by the lack of annotated real-world data due to expensive labeling costs. Unsupervised Domain Adaptation (UDA) offers a solution by transferring knowledge from labeled synthetic data to unlabeled real-world data. While UDA methods for synthetic to real-world domains (synth-to-real) excel in tasks such as semantic segmentation and object detection, their application to instance segmentation for autonomous driving remains underexplored and often relies on suboptimal baselines. We introduce UDA4Inst, a powerful framework for synth-to-real UDA in instance segmentation. Our framework enhances instance segmentation through Semantic Category Training and Bidirectional Mixing Training. Semantic Category Training groups semantically related classes for separate training, improving pseudo-label quality and segmentation accuracy. Bidirectional Mixing Training combines instance-wise and patch-wise data mixing, creating coherent composites that enhance generalization across domains. Extensive experiments show UDA4Inst sets a new state-of-the-art on the SYNTHIA-> Cityscapes benchmark (mAP 31.3) and introduces results on novel datasets, using UrbanSyn and Synscapes as sources and Cityscapes and KITTI360 as targets. Code and models are available at https://github.com/gyc-code/UDA4Inst.
This study presents a simple and accurate method for estimating joint angles with one degree of freedom using two inertial measurement units. The approach leverages quaternion-based computations and eigenvalue decomposition to determine the rotation axes, requiring only a single calibration movement, and was validated through simulation and real-case experiments. In simulations, the proposed approach achieved a mean squared error of 1.1838 degrees, demonstrating robustness to sensor misalignment and varying noise levels. A real-world validation was conducted on a healthy subject performing elbow flexion-extension, using Xsens DOT units at 60 Hz and an OptiTrack motion capture system at 100 Hz, with sensor signals interpolated for direct comparison. The results showed strong agreement between both methods, with a root mean squared error of 4.8922 degrees, ranging from 2.6995 to 5.9759 degrees during resting phases. The method’s efficiency and minimal calibration requirements make it ideal for biomechanics, rehabilitation, and robotics. Its calibration-by-motion approach is particularly useful for applications such as cycling-based knee analysis, where the motion of interest inherently serves as calibration, although elbow joint studies may require a separate calibration step to avoid interference from shoulder movement.
Connected autonomous driving and vehicle-to-everything (V2X) technology brings new opportunities for precise perception in occluded and sight-limited environments. The emergence of cooperative perception via V2X interaction positively impacts the safe and efficient driving of connected autonomous vehicles (CAVs). However, given the diversity and heterogeneity of V2X data, effectively processing such data to enhance cooperative perception remains a key and challenging task. This article introduces a heterogeneous multiscale coopera-tive perception (HM-CoPept) framework with bird's eye view features. It pays more attention to crucial cooperative data that affects driving to avoid blocked areas and extend the sensing range. For the heterogeneity of V2X data, a bidirectional cross-attention is innovatively proposed to fuse LiDAR and camera data complementarily. Furthermore, the multiscale cooperation of V2X interaction data is proposed to break perception occlusion and limitation, considering the benefit of multiscale features with different spatial importance. Cooperative perception is enhanced by learnable spatial confidence weight of safety distance constra-int and foreground estimation. Test and validation are conducted on standard benchmarks, simulated (OPV2V) and real (DAIR-V2X). HM-CoPept results show that performance increases by more than 16% compared to single-vehicle perception. Through extensive experiments and critical analysis, we demonstrate that our approach advances competitive methods and state-of-the-art in average precision. HM-CoPept enables precise and broad perception in complex driving environments and promotes the intelligent and autonomous development of CAVs.
We introduce UrbanSyn, a photorealistic dataset acquired through semi-procedurally generated synthetic urban driving scenarios. Developed using high-quality geometry and materials, UrbanSyn provides pixel-level ground truth, including depth, semantic segmentation, and instance segmentation with object bounding boxes and occlusion degree. It complements GTAV and Synscapes datasets to form what we coin as the 'Three Musketeers'. We demonstrate the value of the Three Musketeers in unsupervised domain adaptation for image semantic segmentation. Results on real-world datasets, Cityscapes, Mapillary Vistas, and BDD100K, establish new benchmarks, largely attributed to UrbanSyn. We make UrbanSyn openly and freely accessible (www.urbansyn.org).
Understanding and predicting pedestrian crossing behavioral intention is crucial for the driving safety of autonomous vehicles. Nonetheless, challenges emerge when using promising images or environmental context masks to extract various factors for time-series network modeling, causing pre-processing errors or a loss of efficiency. Typically, pedestrian positions captured by onboard cameras are often distorted and do not accurately reflect their actual movements. To address these issues, GTransPDM - a Graph-embedded Transformer with a Position Decoupling Module - was developed for pedestrian crossing intention prediction by leveraging multi-modal features. First, a positional decoupling module was proposed to decompose pedestrian lateral motion and encode depth cues in the image view. Then, a graph-embedded Transformer was designed to capture the spatio-temporal dynamics of human pose skeletons, integrating essential factors such as position, skeleton, and ego-vehicle motion. Experimental results indicate that the proposed method achieves 92% accuracy on the PIE dataset and 87% accuracy on the JAAD dataset, with a processing speed of 0.05 ms. It outperforms the state-of-the-art in comparison.
Accurately predicting whether pedestrians will cross in front of an autonomous vehicle is essential for ensuring safe and comfortable maneuvers. However, developing models for this task remains challenging due to the limited availability of diverse datasets containing both crossing (C) and non-crossing (NC) scenarios. Therefore, we propose a procedure that leverages synthetic videos with C/NC labels and an untrained model whose architecture is designed for C/NC prediction to automatically produce C/NC labels for a set of real-world videos. Thus, this procedure performs a synth-to-real unsupervised domain adaptation for C/NC prediction, so we term it S2R-UDA-CP. To assess the effectiveness of S2R-UDA-CP in self-labeling, we utilize two state-of-the-art models, PedGNN and ST-CrossingPose, and we rely on the publicly-available PedSynth dataset, which consists of synthetic videos with C/NC labels. Notably, once the real-world videos are self-labeled, they can be used to train models different from those used in S2R-UDA-CP. These models are designed to operate onboard a vehicle, whereas S2R-UDA-CP is an offline procedure. To evaluate the quality of the C/NC labels generated by S2R-UDA-CP, we also employ PedGraph+ (another literature referent) as it is not used in S2R-UDA-CP. Overall, the results show that training models to predict C/NC using videos labeled by S2R-UDA-CP achieves performance even better than models trained on human-labeled data. Our study also highlights different discrepancies between automatic and human labeling. To the best of our knowledge, this is the first study to evaluate synth-to-real self-labeling for C/NC prediction.
Accurately predicting pedestrian crossings in front of ego-vehicles is essential for intelligent transportation systems (ITS) to enhance road safety. Many existing approaches rely on multiple input modalities, such as scene images, segmentation maps, and trajectory data, which introduce complexity and inefficiencies, thus limiting real-time applicability. To address these challenges, we propose PedGT, a graph-based transformer model that integrates a graph convolutional network (GCN) for spatial feature extraction and a transformer encoder for temporal modeling. Unlike multi-modal methods, PedGT simplifies the pipeline by utilizing only pedestrian pose keypoints and bounding box center points, achieving superior performance on two benchmark datasets. On PIE, it achieves an F1 score of 91% and a recall of 93%, surpassing the previous best of 89%, and 88% by PCPNet. On JAAD, PedGT improves F1 and recall to 70%, outperforming PedFormer's 54% and 48%. Ablation studies highlight the impact of data normalization on accuracy, while frame importance analysis identifies keyframes influencing predictions. This work demonstrates that selecting optimal inputs and leveraging an efficient spatial-temporal model enable PedGT to outperform multi-modal solutions, providing a more streamlined and effective approach for pedestrian intention prediction.
This study evaluates the use of non-commercial, custom-built inertial measurement units (IMUs) for static orientation estimation, comparing them with a commercial solution. Three IMUs were analyzed: the custom-built IMU (s-III), Xsens DOT (s-I), and MATRIX IMU (s-II), using two orientation estimation methods: Proprietary Filter (PF) and Complementary Filter (CF). The results show that the custom-built IMU exhibited the lowest electronic noise, while the Xsens and MATRIX IMUs had higher noise levels due to aging and hardware characteristics. All IMUs presented noise below 3 mg for accelerometers and 0.2 degrees/s for gyroscopes. In terms of orientation accuracy, the S-C system (custom IMU with CF) had a mean error of 2.10%, comparable to the commercial solution, which achieved 0.61%. The S-A system (Xsens with CF) showed a mean error of 0.45%+/- 0.15%. The CF algorithm outperformed PF, particularly in static conditions, due to faster stabilization. However, the MATRIX IMU (S-B) had the highest error (2.88%) due to its higher noise levels. The findings suggest that non-commercial IMUs, combined with open-source algorithms, can offer accurate orientation estimates, providing a cost-effective alternative for biomedical and engineering applications.
The development of wearable devices capable of capturing high-frequency human motion has been a focus of research in recent years. In this work, we propose a new approach to improve the performance of a flexible PCB for IMU sensor mounting, by integrating a 0.3 mm steel stiffener in the PCB. A LDO circuit is used to externalize the battery management system (BMS) for 3V voltage regulation, reducing energy losses and overheating, making the design more efficient and reliable in the long term, while the internal microcontroller can handle battery monitoring and status control during operation. The STNS01PUR provides a more compact solution but at the cost of higher component density and potential thermal challenges, especially during charging. The first results show an improvement in all intended areas from previous reported designs.
The last mile of unsupervised domain adaptation (UDA) for semantic segmentation is the challenge of solving the syn-to-real domain gap. Recent UDA methods have progressed significantly, yet they often rely on strategies customized for synthetic single-source datasets (e.g., GTA5), which limits their generalisation to multi-source datasets. Conversely, synthetic multi-source datasets hold promise for advancing the last mile of UDA but remain underutilized in current research. Thus, we propose DEC, a flexible UDA framework for multi-source datasets. Following a divide-and-conquer strategy, DEC simplifies the task by categorizing semantic classes, training models for each category, and fusing their outputs by an ensemble model trained exclusively on synthetic datasets to obtain the final segmentation mask. DEC can integrate with existing UDA methods, achieving state-of-the-art performance on Cityscapes, BDD100K, and Mapillary Vistas, significantly narrowing the syn-to-real domain gap.
Semantic segmentation models need a large number of images to be effectively trained but manual annotation of such images has a high cost. Active domain adaptation addresses this problem by pretraining the model with a synthetically generated dataset and then fine-tuning it with a few selected label annotations (the "budget") on real images to account for the domain shift. Previous works annotate a percentage of either individual pixels or whole target images. We argue that the first is infeasible in practice, and the second spends part of the budget on classes that the pretrained model may have already learned well. We propose a method based on the annotation of regions computed by Segment Anything, a recently introduced foundation model for class-agnostic image segmentation. The key idea is to assign a ground truth label to each of a tiny subset of regions, those for which the model is more uncertain. In order to increase the number of annotated regions we propagate the ground truth labels to most similar regions according to a hierarchical clustering algorithm that uses the features learned by the pretrained model. Our method outperforms the state-of-the-art on the GTA5 to Cityscapes benchmark by using fewer annotations, almost closing the gap between the synthetically pre-trained model and that obtained with full supervision of the real images. Furthermore, we present competitive results for budgets less than 1% of samples and also for a larger and more challenging target dataset, Mapillary Vistas.
Accurate monocular depth estimation is a fundamental component of vision-based perception systems in intelligent transportation applications. Despite recent progress, unsupervised monocular approaches still suffer from significant performance degradation in real-world traffic scenes due to synthetic-to-real domain gaps and the presence of dynamic, non-rigid objects such as vehicles and pedestrians. In this paper, we propose Back2Color, a robust unsupervised monocular depth estimation framework that addresses these challenges through domain adaptation and uncertainty-aware fusion. Specifically, Back2Color proposes a bidirectional depth-to-color transformation strategy that learns appearance mappings from real-world driving data and applies them to synthetic depth maps, thereby constructing training samples with realistic color appearance and paired synthetic depth. In this way, the proposed approach effectively reduces the domain gap between simulated and real traffic scenes, enabling the depth prediction network to learn more stable and generalizable priors. To further improve robustness under dynamic environments, we propose an auto-learning uncertainty temporal-spatial fusion (Auto-UTSF) module, which adaptively fuses complementary temporal and spatial cues by estimating pixel-wise uncertainty, enabling reliable depth prediction in the presence of moving objects and occlusions. Extensive experiments on challenging urban driving benchmarks, including KITTI and Cityscapes, demonstrate that the proposed method consistently outperforms existing unsupervised monocular depth estimation approaches, particularly in dynamic traffic scenarios, while maintaining high computational efficiency.
The development of Autonomous Driving (AD) systems in simulated environments like CARLA is crucial for advancing real-world automotive technologies. To drive innovation, CARLA introduced Leaderboard 2.0, significantly more challenging than its predecessor. However, current AD methods have struggled to achieve satisfactory outcomes due to a lack of sufficient ground truth data. Human driving logs provided by CARLA are insufficient, and previously successful expert agents like Autopilot and Roach, used for collecting datasets, have seen reduced effectiveness under these more demanding conditions. To overcome these data limitations, we introduce PRIBOOT, an expert agent that leverages limited human logs with privileged information. We have developed a novel BEV representation specifically tailored to meet the demands of this new benchmark and processed it as an RGB image to facilitate the application of transfer learning techniques, instead of using a set of masks. Additionally, we propose the Infraction Rate Score (IRS), a new evaluation metric designed to provide a more balanced assessment of driving performance over extended routes. PRIBOOT is the first model to achieve a Route Completion (RC) of 75 and 45 datasets, potentially solving the data availability issues that have hindered progress in this benchmark.
The utilization of inertial measurement units as wearable sensors is proliferating across various domains, such as health care, sports, and rehabilitation. This expansion has produced a market of devices tailored to accommodate very specific ranges of operational demands. Simultaneously, this growth is creating opportunities for the development of a new class of devices more oriented towards general-purpose use and capable of capturing both high-frequency signals for short-term, event-driven motion analysis and low-frequency signals for extended monitoring. For such a design, which combines flexibility and low cost, a rigorous evaluation of the device in terms of deviation, noise levels, and precision is essential. This evaluation is crucial for identifying potential improvements and refining the design accordingly, yet it is rarely addressed in the literature. This paper presents the development process of such a device. The results of the design process demonstrate acceptable performance in optimizing energy consumption and storage capacity while highlighting the most critical optimizations needed to advance the device towards the goal of a smart, general-purpose unit for human motion monitoring.
Fadi Dornaika合作论文数Departamento de Ciencias de la Computacion e Inteligencia Artificial, Universidad del Pais Vasco8