We introduce MR-Compare, a mixed reality framework for spatially grounded visual comparison between 3D Gaussian splatting and mesh reconstructions with live video see-through (VST). Implemented on a PC-tethered Meta Quest 3, it combines a two-stage registration pipeline with a 3D Slider for cross-media comparison. We evaluated five representative desktop and mobile reconstruction workflows through a real-world benchmark with an exploratory user study (n=30) in two static indoor rooms. MR-Compare achieved centimetre-level translation error across all workflows. The two desktop 3DGS workflows showed the strongest overall pattern, with 3DGS-MCMC yielding the lowest registration error and strongest VST-referenced visual consistency. Room-session measures indicated high perceived usability and low workload. We further propose an anisotropy filter, a zero-shot module that leverages Gaussian anisotropies to improve 3DGS registration in MR-Compare. A controlled Replica threshold sweep shows that moderate pruning can improve robustness and reduce residual errors. These results establish system-level feasibility in the tested setting rather than task-level effectiveness or standalone deployment. The project is available at https://github.com/changruizhu96/MR-Compare.
Change detection is a cognitively challenging process that involves three stages: spotting (becoming aware of a change); localising (establishing the specific location of the change); and identifying (recognising the nature of the change). Each of these stages has the potential to be influenced by both the way the data is presented (e.g., display type) and the fidelity of that data. To explore these issues, we conducted two studies, both of which looked at the effects of display type (immersive virtual reality (VR) or desktop monitor (DM)), and the semantic availability of the scene (low or high realism). Study 1 (N = 38) explored the VR-DM differences in a broad scope, which examined six change types spanning both spatial and non-spatial changes-disappear, appear, translation, rotation, replacement, and colour. However, there were no significant differences between VR and DM in spotting, localising, and identifying at either level of (semantic) realism. Study 2 (N = 20) followed this up by exploring only two types of spatial change (translation and rotation) at a much finer degree of granularity while retaining the same experimental paradigm with necessary refinement. Study 2 showed a significant VR advantage over DM, with different patterns across realism conditions: In low-realism scenes, VR significantly outperformed DM on localisation and change-type identification overall, with the largest VR-DM contrasts observed for the smallest translations. In high-realism scenes, the only significant effect was a display-by-magnitude interaction for change-type identification at the smallest translations. Taken both studies together, VR benefits are most likely for subtle spatial changes, particularly small translations, when the semantic availability is limited. Questionnaire ratings also suggested that reliance on visual features varies with semantic availability. Semantic cues were rated significantly higher than other features in high realism scenes only. Finally, there is no significant difference between VR and DM in terms of workload, motion sickness and self-confidence, suggesting that the perceptual advantages of VR come with no additional physical or cognitive costs for change detection.
High-fidelity simulation is essential for robotics research, enabling safe and efficient testing of perception, control, and navigation algorithms. However, achieving both photorealistic rendering and accurate physics modeling remains a challenge. This paper presents a novel simulation framework, the Unreal Robotics Lab (URL), that integrates the advanced rendering capabilities of the Unreal Engine with MuJoCos high-precision physics simulation. Our approach enables realistic robotic perception while maintaining accurate physical interactions, facilitating benchmarking and dataset generation for vision-based robotics applications. The system supports complex environmental effects, such as smoke, fire, and water dynamics, which are critical to evaluating robotic performance under adverse conditions. We benchmark visual navigation and SLAM methods within our framework, demonstrating its utility for testing real-world robustness in controlled yet diverse scenarios. By bridging the gap between physics accuracy and photorealistic rendering, our framework provides a powerful tool for advancing robotics research and sim-to-real transfer. Our open-source framework is available at https://unrealroboticslab.github.io.
Autonomous and semi-autonomous systems are using deep learning models to improve decision-making. However, deep classifiers can be overly confident in their incorrect predictions, a major issue especially in safety-critical domains. The present study introduces three foundational desiderata for developing real-world risk-aware classification systems. Expanding upon the previously proposed Evidential Deep Learning (EDL), we demonstrate the unity between these principles and EDL’s operational attributes. We then augment EDL empowering autonomous agents to exercise discretion during structured decision-making when uncertainty and risks are inherent. We rigorously examine empirical scenarios to substantiate these theoretical innovations. In contrast to existing risk-aware classifiers, our proposed methodologies consistently exhibit superior performance, underscoring their transformative potential in risk-conscious classification strategies.
Change detection (CD) is critical in everyday tasks. While current algorithmic approaches for CD are improving, they remain imprecise, often requiring human intervention. Cognitive science research focuses on understanding CD mechanisms, especially through change blindness studies. However, these do not address the primary requirement in real-life CD - detecting changes as effectively as possible. Such a requirement is directly relevant to the visual comparison field - studying visualisation techniques to compare data and identify differences or changes effectively. Recent studies have used Virtual Reality (VR) to improve visual comparison by providing an immersive platform where users can interact with 3D data at a real-life scale, enhancing spatial reasoning. We believe VR could also improve CD performance accordingly. Particularly, VR offers stereoscopic depth perception over traditional displays, potentially enhancing the detection of spatial change. In this paper, we develop and analyse three 3D visual comparison techniques for CD in VR: Sliding Window, 3D Slider, and Switch Back. These techniques are evaluated under synthetic but realistic environments and frequently occurring Perceptual Challenges, including different Changed Object Size, Lighting Variation, and Scene Drift conditions. Experimental results reveal significant differences between the techniques in detection time measures and subjective user experience.
Collaborative use of mixed reality (MR) devices is blurring the line between virtual and physical worlds. A remote virtual reality (VR) user immersed in a virtual replica of a real-world environment can interact in real-time with an augmented reality (AR) user who is physically present in that location. One challenge with such a setting is that the virtual world experienced by the remote users often lacks the richness of the real world, particularly in outdoor settings where dynamic elements, such as pedestrians, are missing. The first contribution of this paper is to report findings from focus group sessions on an example collaborative outdoor mixed reality system. Participants noted that lack of synchronisation between the AR and VR worlds diminishes the VR user’s sense of having visited the real-world location together with the AR user. To address this, our second contribution is a system that brings live dynamics into a collaborative MR experience using pedestrians as an example. We conducted a user study using a tour-guide scenario, where an in-situ guide using AR interacts with a remote participant in VR. Results showed that in this scenario, most participants perceived the virtual avatars they saw as representations of real humans in situ.
As drone use has become more widespread, there is a critical need to ensure safety and security. A key element of this is robust and accurate drone detection and localization. While cameras and other optical sensors like LiDAR are commonly used for object detection, their performance degrades under adverse lighting and environmental conditions. Therefore, this has generated interest in finding more reliable alternatives, such as millimeter-wave (mmWave) radar. Recent research on mmWave radar object detection has predominantly focused on 2D detection of road users. Although these systems demonstrate excellent performance for 2D problems, they lack the sensing capability to measure elevation, which is essential for 3D drone detection. To address this gap, we propose CubeDN, a single-stage end-to-end radar object detection network specifically designed for flying drones. CubeDN overcomes challenges such as poor elevation resolution by utilizing a dual radar configuration and a novel deep learning pipeline. It simultaneously detects, localizes, and classifies drones of two sizes, achieving decimeter-level tracking accuracy at closer ranges with overall 95% average precision (AP) and 85% average recall (AR). Furthermore, CubeDN completes data processing and inference at 10Hz, making it highly suitable for practical applications.
Ensuring sufficiently accurate models is crucial in target tracking systems. If the assumed models deviate too much from the truth, the tracking performance might be severely degraded. While the models are usually defined using multivariate conditions, the measures used to validate them are most often scalar-valued. In this paper, we propose matrix-valued measures for both offline and online assessment of target tracking systems. Recent results from Wishart statistics, and approximations thereof, are adapted and it is shown how these can be incorporated to infer statistical properties for the eigenvalues of the proposed measures. In addition, we relate these results to the statistics of the baseline measures. Finally, the applicability of the proposed measures are demonstrated using two important problems in target tracking: (i) distributed track fusion design; and (ii) filter model mismatch detection.
Authoring site-specific outdoor augmented reality (AR) experiences requires a nuanced understanding of real-world context to create immersive and relevant content. Existing ex-situ authoring tools typically rely on static 3D models to represent spatial information. However, in our formative study (n=25), we identified key limitations of this approach: models are often outdated, incomplete, or insufficient for capturing critical factors such as safety considerations, user flow, and dynamic environmental changes. These issues necessitate frequent on-site visits and additional iterations, making the authoring process more time-consuming and resource-intensive. To mitigate these challenges, we introduce CoCreatAR, an asymmetric collaborative mixed reality authoring system that integrates the flexibility of ex-situ workflows with the immediate contextual awareness of in-situ authoring. We conducted an exploratory study (n=32) comparing CoCreatAR to an asynchronous workflow baseline, finding that it enhances engagement, creativity, and confidence in the authored output while also providing preliminary insights into its impact on task load. We conclude by discussing the implications of our findings for integrating real-world context into site-specific AR authoring systems.
Augmented reality displays are becoming more powerful and simultaneously more mobile. Although mobile AR is gaining popularity, it remains difficult to get an insight into users' cognitive load, despite its relevance for many mobile-based tasks. Usually, cognitive load is measured via subjective, task-disruptive self-reports such as NASA TLX. While biosensors such as galvanic skin response, heart rate variability, or pulse have been used to obtain more objective measures, these are highly susceptible to motion-induced noise. More robust techniques like EEG offer higher reliability but are impractical for mobile, real-world use. In this paper, we report on a non-contact multi-sensor approach to assess cognitive load in mobile AR. Our approach combines pupillometry, facial expression tracking, and thermal imaging for respiratory rate analysis. Within the frame of our study, we analysed the aptness of the methods, comparing load assessment for low and high cognitive load tasks under both stationary and mobile conditions. Using an XGBoost classifier, our model achieved 86.11% accuracy for binary cognitive load assessment (low vs. high cognitive load) and 84.24% accuracy for four-way classification (cognitive load $\times$ mobility). Feature importance analysis revealed that robust predictors included gaze dynamics (e.g., fixation, pursuit, and saccade durations), pupil diameter metrics (such as FFT band power and variability measures), and facial and respiratory features (including brow lowering and nostril temperature quantiles) for assessing cognitive load in mobile AR.
Motion tracking systems based on optical sensors typically often suffer from issues, such as poor lighting conditions, occlusion, limited coverage, and may raise privacy concerns. More recently, radio frequency (RF)-based approaches using commercial WiFi devices have emerged which offer low-cost ubiquitous sensing whilst preserving privacy. However, the output of an RF sensing system, such as Range-Doppler spectrograms, cannot represent human motion intuitively and usually requires further processing. In this study, MDPose, a novel framework for human skeletal motion reconstruction based on WiFi micro-Doppler signatures, is proposed. It provides an effective solution to track human activities by reconstructing a skeleton model with 17 key points, which can assist with the interpretation of conventional RF sensing outputs in a more understandable way. Specifically, MDPose has various incremental stages to gradually address a series of challenges: First, a denoising algorithm is implemented to remove any unwanted noise that may affect the feature extraction and enhance weak Doppler signatures. Secondly, the convolutional neural network (CNN)-recurrent neural network (RNN) architecture is applied to learn temporal-spatial dependency from clean micro-Doppler signatures and restore key points' velocity information. Finally, a pose optimising mechanism is employed to estimate the initial state of the skeleton and to limit the increase of error. We have conducted comprehensive tests in a variety of environments using numerous subjects with a single receiver radar system to demonstrate the performance of MDPose, and report 29.4mm mean absolute error over all key points positions, which outperforms state-of-the-art RF-based pose estimation systems.
The nonlinear and stochastic relationship between noise covariance parameter values and state estimator performance makes optimal filter tuning a very challenging problem. Popular optimization-based tuning approaches can easily get trapped in local minima, leading to poor noise parameter identification and suboptimal state estimation. Recently, black box techniques based on Bayesian optimization with Gaussian processes (GPBO) have been shown to overcome many of these issues, using normalized estimation error squared and normalized innovation error statistics to derive cost functions for Kalman filter auto-tuning. While reliable noise parameter estimates are obtained in many cases, GPBO solutions obtained with these conventional cost functions do not always converge to optimal filter noise parameters and lack robustness to parameter ambiguities in time-discretized system models. This article addresses these issues by making two main contributions. First, new cost functions are developed to determine if an estimator has been tuned correctly. It is shown that traditional chi-square tests are inadequate for correct auto-tuning because they do not accurately model the distribution of innovations when the estimator is incorrectly tuned. Second, the new metrics (formulated over multiple time discretization intervals) is combined with a student-t processes Bayesian optimization to achieve robust estimator performance for time discretized state space models. The robustness, accuracy, and reliability of our approach are illustrated on classical state estimation problems.
Most research on collaborative mixed reality (CMR) has focused on indoor spaces. In this paper, we present our ongoing work aimed at investigating the potential of CMR in outdoor spaces. These spaces present unique challenges due to their larger and more com-plex nature, particularly in terms of reconstruction, tracking, and interaction. Our prototype system utilises a photorealistic model to facilitate collaboration between remote virtual reality (VR) users and a local augmented reality (AR) user. We discuss our design considerations, lessons learnt, and areas for future work.
Normal Distribution Transformation (NDT) registration is a fast, learning-free point cloud registration algorithm that works well in diverse environments. It uses the compact NDT representation to represent point clouds or maps as a spatial probability function that models the occupancy likelihood in an environment. However, because of the grid discretization in NDT maps, the global minima of the registration cost function do not always correlate to ground truth, particularly for rotational alignment. In this study, we examined the NDT registration cost function in-depth. We evaluated three modifications (Student-t likelihood function, inflated covariance/heavily broadened likelihood curve, and overlapping grid cells) that aim to reduce the negative impact of discretization in classical NDT registration. The first NDT modification improves likelihood estimates for matching the distributions of small population sizes; the second modification reduces discretization artifacts by broadening the likelihood tails through covariance inflation; and the third modification achieves continuity by creating the NDT representations with overlapping grid cells (without increasing the total number of cells). We used the Pomerleau Dataset evaluation protocol for our experiments and found significant improvements compared to the classic NDT D2D registration approach (27.7% success rate) using the registration cost functions “heavily broadened likelihood NDT” (HBL- NDT) (34.7% success rate) and “over-lapping grid cells NDT” (OGC-NDT) (33.5% success rate). However, we could not observe a consistent improvement using the Student-t likelihood-based registration cost function (22.2% success rate) over the NDT P2D registration cost function (23.7% success rate). A comparative analysis with other state-of-art registration algorithms is also presented in this work. We found that HBL-NDT worked best for easy initial pose difficulties scenarios making it suitable for consecutive point cloud registration in SLAM application.
Augmented Reality (AR) and Virtual Reality (VR) users have distinct capabilities and experiences during Extended Reality (XR) collaborations: while AR users benefit from real-time contextual information due to physical presence, VR users enjoy the flexibility to transition between locations rapidly, unconstrained by physical space.Our research aims to utilize these spatial differences to facilitate engaging, shared XR experiences. Using Google Geospatial Creator, we enable large-scale outdoor authoring and precise localization to create a unified environment. We integrated Ubiq to allow simultaneous voice communication, avatar-based interaction and shared object manipulation across platforms.We apply AR and VR technologies in cultural heritage exploration. We selected the Euston Arch as our case study due to its dramatic architectural transformations over time. We enriched the co-exploration experience by integrating historical photos, a 3D model of the Euston Arch, and immersive audio narratives into the shared AR/VR environment.
The nonlinear and stochastic relationship between noise covariance parameter values and state estimator performance makes optimal filter tuning a very challenging problem. Popular optimization-based tuning approaches can easily get trapped in local minima, leading to poor noise parameter identification and suboptimal state estimation. Recently, black box techniques based on Bayesian optimization with Gaussian processes (GPBO) have been shown to overcome many of these issues, using normalized estimation error squared (NEES) and normalized innovation error (NIS) statistics to derive cost functions for Kalman filter auto-tuning. While reliable noise parameter estimates are obtained in many cases, GPBO solutions obtained with these conventional cost functions do not always converge to optimal filter noise parameters and lack robustness to parameter ambiguities in time-discretized system models. This paper addresses these issues by making two main contributions. First, we show that NIS and NEES errors are only chi-squared distributed for tuned estimators. As a result, chi-square tests are not sufficient to ensure that an estimator has been correctly tuned. We use this to extend the familiar consistency tests for NIS and NEES to penalize if the distribution is not chi-squared distributed. Second, this cost measure is applied within a Student-t processes Bayesian Optimization (TPBO) to achieve robust estimator performance for time discretized state space models. The robustness, accuracy, and reliability of our approach are illustrated on classical state estimation problems.
One of the most common misconceptions made about the Kalman filter when applied to linear systems is that it requires an assumption that all error and noise processes are Gaussian. This misconception has frequently led to the Kalman filter being dismissed in favor of complicated and/or purely heuristic approaches that are supposedly ``more general'' in that they can be applied to problems involving non-Gaussian noise. The fact is that the Kalman filter provides rigorous and optimal performance guarantees that do not rely on any distribution assumptions beyond mean and error covariance information. These guarantees even apply to use of the Kalman update formula when applied with nonlinear models, as long as its other required assumptions are satisfied. Here we discuss misconceptions about its generality that are often found and reinforced in the literature, especially outside the traditional fields of estimation and control.
Micro-Doppler signatures contain considerable information about target dynamics. However, the radar sensing systems are easily affected by noisy surroundings, resulting in uninterpretable motion patterns on the micro-Doppler spectrogram ( $\mu $ -DS). Meanwhile, radar returns often suffer from multipath, clutter, and interference. These issues lead to difficulty in, for example, motion feature extraction and activity classification using micro-Doppler signatures. In this article, we propose a latent feature-wise mapping strategy, called feature mapping network (FMNet), to transform measured spectrograms so that they more closely resemble the output from a simulation under the same conditions. Based on measured spectrogram and the matched simulated data, our framework contains three parts: an encoder which is used to extract latent representations/features, a decoder outputs reconstructed spectrogram according to the latent features, and a discriminator minimizes the distance of latent features of measured and simulated data. We demonstrate the FMNet with six activities data and two experimental scenarios, and final results show strong enhanced patterns and can keep actual motion information to the greatest extent. On the other hand, we also propose a novel idea which trains a classifier with only simulated data and predicts new measured samples after cleaning them up with the FMNet. From final classification results, we can see significant improvements.
Distributed virtual environments (DVEs) are challenging to create as the goals of consistency and responsiveness become contradictory under increasing latency. DVEs have been considered as both distributed transactional databases and force-reflection systems. Both are good approaches, but they do have drawbacks. Transactional systems do not support Level 3 (L3) collaboration: manipulating the same degree-of-freedom at the same time. Force-reflection requires a client-server architecture and stabilisation techniques. With Consensus Based Networking (CBN), we suggest DVEs be considered as a distributed data-fusion problem. Many simulations run in parallel and exchange their states, with remote states integrated with continous authority. Over time the exchanges average out local differences, performing a distribued-average of a consistent, shared state. CBN aims to build simulations that are highly responsive, but consistent enough for use cases such as the piano-movers problem. CBN’s support for heterogeneous nodes can transparently couple different input methods, avoid the requirement of determinism, and provide more options for personal control over the shared experience. Our work is early, however we demonstrate many successes, including L3 collaboration in room-scale VR, 1000’s of interacting objects, complex configurations such as stacking, and transparent coupling of haptic devices. These have been shown before, but each with a different technique; CBN supports them all within a single, unified system.
Binarized Neural Networks (BNNs) have the potential to revolutionize the way that deep learning is carried out in edge computing platforms. However, the effectiveness of interpretability methods on these networks has not been assessed. In this paper, we compare the performance of several widely used saliency map-based interpretabilty techniques (Gradient, SmoothGrad and GradCAM), when applied to Binarized or Full Precision Neural Networks (FPNNs). We found that the basic Gradient method produces very similar-looking maps for both types of network. However, SmoothGrad produces significantly noisier maps for BNNs. GradCAM also produces saliency maps which differ between network types, with some of the BNNs having seemingly nonsensical explanations. We comment on possible reasons for these differences in explanations and present it as an example of why interpretability techniques should be tested on a wider range of network types.
Marco Lanzagorta合作论文数the Quantum Information Group at ITT Corporation10