Human activity recognition has been a prominent research focus for several decades. While computer vision-based passive sensing is desirable for many applications, traditional RGB-based methods face significant limitations, including high computational cost, sensitivity to ambient lighting conditions, and privacy concerns. Single-photon LiDAR (light detection and ranging) is emerging as a robust alternative, offering efficient and high-resolution 3D imaging over long ranges and in difficult conditions while preserving privacy. In this study, we combine an eye-safe single-photon LiDAR system with a deep learning pipeline to achieve fast, long-range human activity recognition. To address the issue of limited available data for training, we generate synthetic datasets by combining real motion capture data with virtual models in a 3D modeling environment. A state-of-the-art recurrent neural network is trained on short-duration depth-image sequences of six activities. We also contribute two long-range, video frame rate single-photon LiDAR datasets for human activity recognition, recorded at distances of 325 m and 1.4 km, which exhibit markedly different noise levels. When evaluated on these data, the network maintains more than 80% accuracy even in the most challenging scenario, while supporting continuous and fast inference.
Three-dimensional (3D) imaging underpins applications ranging from autonomous navigation to defense and biomedicine, with single-photon avalanche diode (SPAD) light detection and ranging (LiDAR) enabling fast, long-range, and photon-efficient depth sensing. In practice, reconstruction quality is constrained by limited sensor resolution, particularly in the short- and medium-wave infrared, as well as sensor dark counts, background illumination, atmospheric effects, and motion. We introduce a unified framework for continuous-surface 3D scene representation that integrates multimodal sensing with score-based priors on latent variables. The proposed approach models scenes using a parametric continuous surface, enabling robust rendering at arbitrary spatial resolutions, even in the presence of multiple depth layers allowing imaging through camouflage. We demonstrate capabilities including data compression, targeted high-resolution rendering, and guided super-resolution of dynamic 3D videos. Validated across multiple sensing scenarios using diverse off-the-shelf priors, this framework enables compressed, high-fidelity 3D imaging in real-world environments.
Transformers have shown significant success in hyperspectral unmixing (HU). However, challenges remain. While multi-scale and long-range spatial correlations are essential in unmixing tasks, current Transformer-based unmixing networks, built on Vision Transformer (ViT) or Swin-Transformer, struggle to capture them effectively. Additionally, current Transformer-based unmixing networks rely on the linear mixing model, which lacks the flexibility to accommodate scenarios where nonlinear effects are significant. To address these limitations, we propose a multi-scale Dilated Transformer-based unmixing network for nonlinear HU (DTU-Net). The encoder employs two branches. The first one performs multi-scale spatial feature extraction using Multi-Scale Dilated Attention (MSDA) in the Dilated Transformer, which varies dilation rates across attention heads to capture long-range and multi-scale spatial correlations. The second one performs spectral feature extraction utilizing 3D-CNNs with channel attention. The outputs from both branches are then fused to integrate multi-scale spatial and spectral information, which is subsequently transformed to estimate the abundances. The decoder is designed to accommodate both linear and nonlinear mixing scenarios. Its interpretability is enhanced by explicitly modeling the relationships between endmembers, abundances, and nonlinear coefficients in accordance with the polynomial post-nonlinear mixing model (PPNMM). Experiments on synthetic and real datasets validate the effectiveness of the proposed DTU-Net compared to PPNMM-derived methods and several advanced unmixing networks.
Single-photon Lidar imaging offers a significant advantage in 3D imaging due to its high resolution and long-range capabilities, however it is challenging to apply in noisy environments with multiple targets per pixel. To tackle these challenges, several methods have been proposed. Statistical methods demonstrate interpretability on the inferred parameters, but they are often limited in their ability to handle complex scenes. Deep learning-based methods have shown superior performance in terms of accuracy and robustness, but they lack interpretability or they are limited to a single-peak per pixel. In this paper, we propose a deep unrolling algorithm for dual-peak single-photon Lidar imaging. We introduce a hierarchical Bayesian model for multiple targets and propose a neural network that unrolls the underlying statistical method. To support multiple targets, we adopt a dual depth maps representation and exploit geometric deep learning to extract features from the point cloud. The proposed method takes advantages of statistical methods and learning-based methods in terms of accuracy and quantifying uncertainty. The experimental results on synthetic and real data demonstrate the competitive performance when compared to existing methods, while also providing uncertainty information.
Time-of-flight (ToF) imaging is widely used in consumer electronics for depth perception, with compact ToF sensors often representing their data as histograms of photon arrival times for each pixel. These histograms capture detailed temporal information that enables advanced computational techniques, such as super-resolution, to reconstruct high-resolution depth images even from low-resolution sensors by leveraging the full temporal structure of the data. However, transferring full histogram data is impractical for compact systems due to the large amount of data. To address this, microcontrollers extract a few key parameters-such as peak position, signal intensity, and noise level-greatly reducing data volume. While this approach performs well for low-resolution tasks like autofocus and obstacle detection, its potential for high-resolution depth imaging has not been fully explored. In this work, we demonstrate that these few extracted parameters are sufficient to reconstruct full high-resolution depth images. We propose a compact and data-efficient neural network that enhances the spatial resolution of a basic ToF sensor from 4 × 4 pixels to 32 × 32 pixels. By focusing on only 3 key parameters per pixel, compared to the original 144 histogram bins (range ToF sensor provides), representing a 48× reduction in data, our approach significantly reduces the data requirements while maintaining performance similar to methods that rely on full histogram data. Despite this drastic reduction in data, our method achieves high-resolution depth imaging with minimal performance loss, demonstrating the feasibility of efficient and high-quality depth reconstruction using only key extracted parameters.
Wind turbine blade (WTB) surface defect detection often suffers from severe class imbalance and limited annotated data, making conventional deep learning approaches impractical. Few-shot object detection (FSOD) addresses this challenge by enabling models to detect novel defect types from only a few labelled examples. In this study, FSOD is applied to the WTB surface defect detection task using the DTU dataset. We adopt the Two-Stage Fine-Tuning Approach and utilised Contrastive Proposal Encoding (CPE) loss to improve proposal discrimination and feature representation under low-data conditions. A progressive experimental setup is designed by partitioning defect categories into base and novel classes based on sample scarcity. Our results show that integrating CPE loss leads to up to 22% improvement in novel class detection and overall gains of around 2-3% in mAP under 10-shot scenarios, while highlighting performance tradeoffs under extreme class imbalance. These findings validate the effectiveness of contrastive objectives in FSOD and underscore the importance of strategic dataset construction for robust generalisation.
Energy dispersive X-ray (EDX) spectrum imaging yields compositional information with a spatial resolution down to the atomic level. However, experimental limitations often produce extremely sparse and noisy EDX spectra. Under such conditions, every detected X-ray must be leveraged to obtain the maximum possible amount of information about the sample. To this end, we introduce a robust multiscale Bayesian approach that accounts for the Poisson statistics in the EDX data and leverages their underlying spatial correlations. This is combined with EDX spectral simulation (elemental contributions and Bremsstrahlung background) into a Bayesian estimation strategy. When tested using simulated datasets, the chemical maps obtained with this approach are more accurate and preserve a higher spatial resolution than those obtained by standard methods. These properties translate to experimental datasets, where the method enhances the atomic resolution chemical maps of a canonical tetragonal ferroelectric PbTiO3 sample, such that ferroelectric domains are mapped with unit-cell resolution.
Multifractal analysis (MFA) provides a framework for the global characterization of image textures by describing the spatial fluctuations of their local regularity based on the multifractal spectrum. Several works have shown the interest of using MFA for the description of homogeneous textures in images. Nevertheless, natural images can be composed of several textures and, in turn, multifractal properties associated with those textures. This paper introduces an unsupervised Bayesian multifractal segmentation method to model and segment multifractal textures by jointly estimating the multifractal parameters and labels on images, at the pixel-level. For this, a computationally and statistically efficient multifractal parameter estimation model for wavelet leaders is firstly developed, defining different multifractality parameters for different regions of an image. Then, a multiscale Potts Markov random field is introduced as a prior to model the inherent spatial and scale correlations (referred to as cross-scale correlations) between the labels of the wavelet leaders. A Gibbs sampling methodology is finally used to draw samples from the posterior distribution of the unknown model parameters. Numerical experiments are conducted on synthetic multifractal images to evaluate the performance of the proposed segmentation approach. The proposed method achieves superior performance compared to traditional unsupervised segmentation techniques as well as modern deep learning-based approaches, showing its effectiveness for multifractal image segmentation.
This paper introduces a Bayesian algorithm for the robust reconstruction and super-resolution of 3D video single-photon LiDAR data. The focus is on challenging scenarios with low-resolution LiDAR data, sparse photon returns or high background noise as observed in real-world applications. The proposed hierarchical Bayesian approach leverages multiscale histogram information and a high-resolution reflectivity guidance to provide high-resolution depth estimates along with corresponding uncertainty measures, aiding in better decision-making. Correlations between variables are enforced through a weighted scheme, enabling the integration of guidance from other sensors or advanced algorithms. Results on synthetic data demonstrate improved scene reconstruction in extreme conditions compared to existing methods.
Single-photon avalanche diode (SPAD) detectors offer exceptional temporal resolution and sensitivity, making them a powerful technology for depth sensing. However, reconstructing high-resolution depth maps from SPAD data is challenging due to its sparse and noisy nature, particularly in low-light or scattering conditions. This paper presents a novel SPAD-based depth map super-resolution approach that combines a SPAD’s multiscale compressive representation for robustness in noisy scenarios, with high-resolution reflectivity guidance to enhance structural details. It also leverages the generative capabilities of Denoising Diffusion Probabilistic Models to quantify uncertainty. Experimental results on simulated data demonstrate the method’s effectiveness and robustness across varying noise levels and upscaling factors.
Maintenance is a critical aspect of wind power generation as it not only ensures the efficient operation of wind turbines, but also their continuous availability and functionality. For wind turbine blades, regular maintenance is essential to optimise power output and minimise operational downtime. While various maintenance strategies are well-documented, such as predictive approaches using Machine Learning and traditional visual inspections, there is limited research on leveraging aerial imagery for detecting defects on turbine blades. The objective of this review paper is to address this by focusing on the challenges and requirements for effective surface defect detection in wind turbine blades through aerial imagery. The task of inspecting surface defects on wind turbine blades is particularly difficult due to data scarcity, substantial computational requirements, and the geometric difficulties in accurately localising defects. By addressing these issues, we aim to identify and propose promising future directions to overcome these challenges at hand, thereby ensuring a progression of research and development in this field.
Single-photon avalanche diodes (SPADs) are advanced sensors capable of detecting individual photons and recording their arrival times with picosecond resolution using time-correlated single-photon counting (TCSPC) detection techniques. They are used in various applications, such as LiDAR and low-light imaging. These single-photon cameras can capture high-speed sequences of binary single-photon images, offering great potential for reconstructing 3D environments with high motion dynamics. To complement single-photon data, these cameras are often paired with conventional passive cameras, which capture high-resolution intensity images at a lower frame rate. However, 3D reconstruction from SPAD data faces challenges. Aggregating multiple binary measurements improves precision and reduces noise but can cause motion blur in dynamic scenes. Additionally, SPAD arrays often have lower resolution than passive cameras. To address these issues, we propose a novel computational imaging algorithm to improve the 3D reconstruction of moving scenes from SPAD data by addressing the motion blur and increasing the native spatial resolution. The goal is to turn the high-speed SPAD events, recorded at a high frame rate, into non-blurred high-resolution depth images at the frame rate of the passive sensor. We adopt a plug-and-play approach within an optimization scheme alternating between guided video super-resolution of the 3D scene, and precise image realignment using optical flow. Experiments on synthetic data show that our method significantly improves image resolution across various signal-to-noise ratios and photon levels. We validate our method using real-world SPAD measurements in three practical situations with dynamic objects. First on fast-moving scenes (i.e. fan) in laboratory conditions at a short range (3 meters); second very low-resolution imaging of people with a consumer-grade SPAD sensor from STMicroelectronics; and finally, high-resolution imaging of people walking outdoors in daylight at a range of 325 meters under eye-safe illumination conditions using a short-wave infrared SPAD camera. These results demonstrate the robustness and versatility of our approach.
Single-photon Lidar is a promising 3D imaging technique, but it is challenging to deploy in real-world applications due to high noise levels and the presence of multiple surfaces per pixel. Existing statistical methods are interpretable, but limited by the assumed model. Data-driven approaches show excellent performance, but with limited interpretability, preventing their use in critical applications. In this paper, we propose an interpretable deep learning architecture with graph attention networks for the reconstruction of dual peaks per pixel in single photon Lidar. Instead of the conventional image-based representation, we represent the solution as point clouds, allowing reconstruction of more than one surface per pixel. The proposed architecture is based on a statistical Bayesian algorithm, whose iterative steps are converted into neural network layers. This approach combines the advantages of both statistical and learning-based frameworks, providing good estimates with improved network interpretability. Experimental results demonstrate the effectiveness of the proposed method.
Detecting surface defects on Wind Turbine Blades (WTBs) from remotely sensed images is a crucial step toward automated visual inspection. Typical object detection algorithms use standard bounding boxes to locate defects on WTBs. However, Oriented Bounding Boxes (OBBs) have been shown in cases of satellite imagery, to provide more precise localization of object regions and actual orientation. Existing WTB datasets do not depict defects using OBBs and this causes the lack of useful orientational information. In this paper, we consider OBBs for WTB surface defect detection through two publicly available datasets, introducing new annotations to the community. Base-lines were constructed on state-of-the-art rotated object detectors, demonstrating considerable promise and known gaps that can be addressed in the future. We present a comprehensive analysis of their performances including ablation study and discussions on the importance of angular disparity between OBBs.
This paper studies a new fusion method designed for magnetic resonance (MR) and ultrasound (US) images, with a specific focus on endometriosis diagnosis. The proposed method is based on guided filtering, leveraging the advantages of this technique to enhance the quality of fused images. The fused image is a weighted average of base and detail images from the MR and US images. The weights assigned to the US image account for the presence of speckle noise, a common challenge in US imaging whereas the weights assigned to the MR image allow the contrast of the fused image to be enhanced. The effectiveness of the method is evaluated using synthetic and phantom data, showing promising results. The image provided by the proposed fusion method holds potential for enhancing visualization and aiding decision-making in endometriosis surgery, offering a valuable contribution to the field of medical image fusion.
Automatic detection of defects from wind turbine blade images has shown tremendous progress in recent years. However, there are not many annotated datasets feasible for benchmarking purposes, and a lack of consistency in annotation procedures across existing works. In this paper, we investigate the data annotation process for wind turbine blade images to reduce inaccuracies in defect detection and to benchmark the performance of the patch-based detection framework on recent deep learning architectures. In this study, we identify challenges in the detection task that are incurred by the presence of extreme bounding box aspect ratios among the annotations. Experiments on two additional annotation sets show that the sets with altered box aspect ratios are able to improve the overall defect detection accuracy, particularly for classes containing boxes with very small aspect ratios. We also provide extensive class-wise results with visual examples of the highlighted problem.
Time-correlated single-photon technology is emerging as an important approach to 3D Imaging. This paper presents a reconstruction algorithm that exploits data statistics and multi-scale information to deliver clean depth and reflectivity images together with associated uncertainty maps. The statistical method has been implemented to run on graphics processing units (GPUs) that enable real-time reconstruction of moving scenes at more than 1000 depth frames per second on the 32 × 64 pixels real Quantic4x4 SPAD sensor array data. Comparisons with state-of-the-art algorithms on simulated and real data demonstrate the robust and efficient performance of the proposed method.
Single-photon methods are emerging as a key approach to 3D Imaging. This paper introduces a two step statistical based approach for real-time image reconstruction applicable to a transmission medium with extreme light scattering conditions. The first step is an optional target detection method to select informative pixels which have photons reflected from the target, hence allowing data compression. The second is a reconstruction algorithm that exploits data statistics and multiscale information to deliver clean depth and reflectivity images together with associated uncertainty maps. Both methods involve independent operations that are implemented in parallel on graphics processing units (GPUs), which enables real-time data processing of moving scenes at more than 50 depth frames per second for an image of $128 \times 128$ pixels. Comparisons with state-of-the-art algorithms on simulated and real underwater data demonstrate the benefit of the proposed framework for target detection, and for fast and robust depth estimation at multiple frames per second.