
We present a novel application of 3D Gaussian Splatting (3DGS) for the large-scale archival of thousands of temples in Bali. Our project, the Bali Digital Heritage Initiative, adopts a participatory approach in which local communities collect video footage of these vulnerable temples, which we transform into 3DGS scenes to serve as digital archives. Unlike existing approaches that require uploading curated sets of images, video contribution simplifies the process for the community. We expect this to encourage a high participation rate and increase the number of archived temples. However, these community-sourced videos are often of variable quality and usually aren’t ideal for 3D reconstruction, introducing technical challenges. This paper first presents several key challenges in the development of the 3DGS pipeline in this context, including managing video quality, combining multiple videos, ensuring reliable results from Structure from Motion, and maintaining the visual quality of 3DGS scenes. We then propose and test solutions to many of these challenges, addressing all parts of the reconstruction pipeline. We also introduce a postprocessing stage that cleans up the scene through a series of purpose-built 3DGS filters, which consider attributes of Gaussian primitives, including local density, geometry, and colors. Finally, we discuss future research directions to address unsolved problems and improve the performance of the pipeline.
The safety, performance, and lifespan of lithium-ion (Li-ion) batteries heavily depend on understanding their internal chemo-mechanical state. Conventional battery management systems, which depend on terminal electrical measurements, offer limited insight into the dynamic stress and strain that develop during operation. This paper introduces a non-destructive, in-situ method for monitoring the internal stress state of a commercial Li-ion pouch cell using single-frequency ultrasonic testing. Initially, a relation between stress-strain buildup within the battery and the amplitude of the transmitted ultrasonic signal is established. Then, a 120 kHz ultrasonic signal is transmitted through the cell in a through-transmission setup, and the amplitude of the received signal is analyzed as the main indicator for stress-strain and degradation monitoring. Experimental results over four charge-discharge cycles show a strong correlation between signal amplitudes and the State of Charge, with amplitude generally increasing during charging due to intercalation-induced compressive stress. The technique demonstrates high sensitivity, capturing non-monotonic behaviors attributed to electrode phase transitions. Additionally, a gradual, cycle-by-cycle decrease in overall signal amplitude is observed, providing a direct measure for tracking cumulative mechanical degradation and State of Health. These results demonstrate that ultrasonic signal amplitude is a comprehensive, versatile feature capable of offering real-time insights into the complex acousto-mechanical behavior of Li-ion batteries, offering a promising enhancement for next-generation battery diagnostics.
This paper presents a novel alternative approach to perception for aerial vehicles operating in confined spaces such as ballast tanks. It focuses on designing and implementing a system of small Time-of-Flight (ToF) sensors arranged in an array that could replace conventional, expensive, and bulky 2D or 3D LiDAR systems. Confined spaces often feature narrow, dark, and dusty environments. Here, conventional vision-based solutions fail for autonomous navigation and even basic flight tasks such as obstacle avoidance. Furthermore, carrying conventional LiDAR sensors significantly reduces flight time, especially for small drones(< 2.5 kg). Therefore, this paper proposes a system using VL53L8CX sensors as an optimal alternative, which significantly reduces the weight and computational power. A novel approach is developed to combine 12 of these ToF sensors across two SPI buses. This setup provides a 360 degrees radial field of view within a 4m radius around the aerial vehicle, which is strategically designed to accommodate these sensors. The proposed design is validated and compared against a 2D lidar system for collision prevention in a PX4 SITL simulation of a ballast tank environment. Additionally, the sparsity and accuracy of the generated point cloud are compared with those of 3D LiDAR while flying within a mock ballast tank. This design and study aim to establish a baseline for lightweight, compact, and safe navigation for small drones in confined and featureless environments.
This paper addresses the critical engineering challenge of achieving both high classification accuracy and rapid training convergence in neuromorphic vision systems. We introduce a hybrid Quantum Variational Reservoir-enhanced Multi-Layer Perceptron (QVR-MLP) and benchmark its performance against a classical MLP baseline. Both architectures are trained on an established event-based dataset from a Polarimetric Dynamic Vision Sensor (pDVS) capturing four distinct motion patterns modulated at varying speeds. Our methodology employs a computationally efficient Fast Fourier Transform (FFT) based feature extraction pipeline that processes asynchronous event streams without requiring full-frame reconstruction. The experimental results establish that quantum-enhanced neuromorphic computing is a robust and viable engineering solution for improving performance in resource-constrained neuromorphic vision applications.
Tunable Diode Laser Absorption Spectroscopy (TDLAS) is widely adopted for non-intrusive temperature and gas molar concentration imaging. Classical imaging methods in TDLAS tomography follow a two-stage approach, i.e., reconstructions of absorbance at two transitions and two-line thermometry for imaging of temperature and gas molar concentration. However, the separation of two stages lead to neglect of the physical constraint among these four variables, absorbance at two transitions, temperature and gas molar concentration, which leads to significant error propagation. An imaging model is proposed to mitigate the issue. In this model, a mapping from temperature and gas molar concentration to absorbance at two transitions is pre-established and incorporated into the reconstruction as a constraint. An optimization objective is constructed with projection constraint, mapping induced constraint and local smoothness priors. The four variables are jointly reconstructed in iteration. Numerical simulations and comparison to the typical SART with two-line thermometry method were conducted. In noise-free cases, the proposed method achieved lower image errors, with average reductions of 0.068 in temperature and 0.182 in gas molar concentration. In noisy cases, average image error decreases of 0.064 in temperature and 0.163 in gas molar concentration were yielded by the proposed method across all SNR levels. These demonstrate that the proposed method significantly reduces image errors and artifacts, improves reconstruction accuracy, and enhances robustness to noise.
Multi-perforated combustors are promising for reducing carbon emissions by improving combustion efficiency and reducing pollutant formation. These systems produce dense flame fields composed of multiple spatially confined and often overlapping flame structures, posing significant challenges for accurate flame image reconstruction, which is essential for combustion analysis and optimization. Existing methods primarily focus on reconstructing multiple depth sections of single flames, which limits their applicability to multi-flame configurations. This study proposes a novel imaging reconstruction approach that integrates sectioning tomography, light field imaging, and digital refocusing to achieve accurate depth-resolved flame images in multi-perforated burners. This method prioritizes accurate reconstruction of individual flame images by reconstructing one representative section per flame. Experiments were carried out on both non-overlapping and overlapping flames to demonstrate the method’s ability to reconstruct depth-specific images of complex flame structures. Results show that the proposed method accurately reconstructed flame contours and intensity distributions in nonoverlapping flames, while effectively isolating radiative contributions from depth-specific sections in overlapping flames. This strategy captures the overall combustion characteristics while allowing for a detailed analysis of individual flame states, providing a comprehensive understanding of both global and local combustion behaviors. Furthermore, it reduces computational complexity while preserving high-fidelity flame representation, offering an efficient tool for advanced diagnostics in multi-perforated burners.
This study investigates the use of Gaussian Splatting for 3D reconstruction of medieval castle ruins in the Rhine Valley, comparing its effectiveness with conventional photogrammetry. 3D Gaussian Splatting (3DGS), a recent real-time rendering technique, models scenes using point-based representations with anisotropic Gaussians, enabling high-fidelity reconstructions from sparse image inputs. Unlike traditional photogrammetry, which depends on dense point clouds and mesh generation, 3DGS offers smoother rendering, better handling of fine details, and improved performance in visually complex or degraded areas. 3DGS excels in producing visually compelling, immersive models with fewer artifacts in occluded or texture-deficient regions. In this paper, three analyses were performed by focusing on three individual aspects: processing time, geometric accuracy, and visual quality. Our study’s results highlight the advantages of 3DGS in this regard, which include quick data processing time and excellent visualisation capability specifically with finely detailed objects. However, this method still faces challenges when employed for a metric archiving of built heritage, where traditional photogrammetry still provides a higher quality metric result using the same input data. 3DGS suffers particularly from noisy data, especially when converted into point clouds. Nevertheless, this approach presents a promising tool for digital heritage preservation, enabling efficient and realistic visualisation of fragile historical sites.
In recent years, the number of hypertensive patients has been increasing, and routine blood pressure monitoring is essential for early detection and prevention.Our laboratory focuses on facial skin temperature, a cardiovascular indicator that can be remotely measured using infrared thermography. We aim to estimate blood pressure non-contact using spatial features of captured thermal face images (TFIs).In previous studies, resting blood pressure has been estimated based on spatial features extracted by applying Independent Component Analysis (ICA) to measured TFI. However, ICA has the property of extracting the same number of independent components as the number of input images to be applied. Therefore, when ICA is applied to a large number of thermal images, the features of the extracted independent components are widely distributed over the entire face, and the features of each region tend to become ambiguous. Thus, features that are distributed evenly across the entire face make it difficult to capture local changes related to blood pressure fluctuations.Therefore, in this study, we reduced the number of input images applied to ICA using two types of image selection methods. We then attempted to improve the accuracy of the blood pressure estimation model by enhancing the locality and discriminability of the extracted independent components. We constructed three models by applying ICA to images without selection and images selected using the two methods. A comparison of the accuracy of each model suggested that image selection affects estimation accuracy.
A histogram tomography for tunable diode laser absorption spectroscopy (TDLAS) is proposed via using optimized pairs of temperature and gas concentration to reconstruct their distributions. In the pair selection, the temperature and gas concentration are usually distributed over several equal intervals in the feasible ranges. The temperature value with the highest probability density in each interval is selected as the typical value at this interval, and a linear map is used to select the related gas concentration and form the pairs. Projection histograms along laser paths are achieved at different projection angles, optimized selected pairs are used to retrieve the distributions of temperature and concentration in the imaging region. The discrete pair selection at each pixel forms an element of a binary-valued matrix, and the twodimensional images are derived from the reconstructed pair values from the optimally selected pairs. Numerical simulations were carried out to make comparisons with the typical two-line thermometry method. The proposed method has better accuracy, especially for complex distributions, image errors of temperature and concentration images were reduced by more than 27.86% and 24.31%, respectively. The flame cross-section at the outlet of Mckenna burner was imaged to verify that the proposed method remained effective in noisy cases.
Autonomous inspection of maritime vessels using Unmanned Aerial Vehicles can prove difficult due to low light conditions, minimum space and battery limitations. This paper proposes a strategy for optimized manhole traversal in the scenario of maritime vessel ballast tank inspection. A dense point cloud, collected during the exploration of the ballast tank, is used to localize the manhole that the UAV has to traverse through. The prediction of the manhole is performed on synthesized 360. panoramic views of the point cloud, originating from positions inside the mapped area. This manhole prediction enables a Model Predictive Control framework to find the optimal path through the manhole, taking into account the capabilities of the UAV and the topological constraints of the confined space. The manhole localization method is able to predict the center of the manhole with an average success rate of 89.6%, which can be raised to 94% after the ensembling of multiple predictions from different assumed camera positions. Simulated results show the robustness of the Model Predictive Control method in a wide array of different scenarios. We demonstrate the ability of the method to predict accurately the manhole and efficiently pass through it using a UAV, in a real-world experimental setup.
Towards a more digitalized and automated future, recent years have witnessed a significant increase in the deployment of Unmanned Aerial Vehicles (UAVs) across various domains, particularly in surveillance and logistics. However, object detection in obscure environments such as parking lots, particularly for trucks, remains challenging due to the scarcity of real-world datasets. Previous datasets, such as Common Objects in Context (COCO), PASCALVisual Object Classes (VOC), and PKLot, are valuable for general object and parking occupancy detection but lack aerial-view truck images, thus limiting their applicability for real-time monitoring of truck parking availability using UAV imagery. To that end, We introduce SynthPark, a photorealistic digital-twin environment that enables scalable generation of fully labelled aerial imagery for parking occupancy studies. Built in an open-source game engine, SynthPark procedurally varies scene layout, illumination, weather and camera trajectory, yielding diverse viewpoints that closely mimic field deployments without the cost of real flights. We demonstrate the utility of the environment by exporting a synthetic image set and fine-tuning a YOLOv11 detector, which attains a mean average precision of 76.1% (mAP(50:95)) on held-out synthetic frames and translates favourably to real UAV footage. Beyond dataset creation, SynthPark constitutes a reusable test-bed for perception algorithms, contributing an environmental benchmark that complements existing collections such as COCO, PASCAL-VOC and PKLot.
This work presents an automated pipeline for statistical shape modeling and evaluation of 3D anatomical landmarks in the left and right humeri from Computed Tomography (CT) data. The process consists of training and testing stages, designed to ensure reproducibility, robustness, and anatomical accuracy. The model is designed to reduce bias related to template selection, human selection variability, and differences in laterality between left and right humeri. Anatomical landmarks play a pivotal role in semi- or fully automated implant placement. This model is capable of predicting the landmarks in a fully automated way with the possibility to adjust the results manually, if needed, in a mixed reality environment. The pipeline involves segmentation from 3D medical images (CT scans), manual landmarks selection for training images, meshing, rigid and non-rigid registration, and cross-validation of the model by performing 10 predictions with different template selection for each of the 20 test cases (10 left and 10 right humeri) for a total of 200 predictions. An interobserver variability study is conducted to validate the model, utilizing annotations from 5 operators (for a total number of 50 annotations). This model shows comparable results in comparison with human annotation, having the advantage of being automatic once the model is trained.
Unmanned aerial systems (UAS) offer substantial benefits in maritime vessel inspections by significantly reducing operational costs, since traditional inspection methods involve manual labor and extended downtime for vessels. To fully leverage these advantages, important barriers such as flight time must be addressed. Introducing a tether connected to a ground station enables constant power delivery, effectively overcoming this limitation; however, planning complexity is greatly increased. In our work, we developed a motion planning strategy for a system of four tethered drones connected in a string formation. This strategy allows the first drone in the chain to follow a predefined path in a confined space while repositioning the others to avoid collisions. To achieve this, the UAS was modeled as a serial kinematic chain, by treating the UAVs as joints and the tethers as links. MoveIt2, a popular robotic manipulator framework, is employed to provide access to sampling-based algorithms and inverse kinematic solvers, while NVIDIA’s Isaac Sim was used for visualization. A custom line-of-sight constraint was created and integrated into MoveIt2 using Bullet to enable collision checking for the tethers. Different sampling-based algorithms were compared using our heuristic under varying planning time parameters. The proposed strategy improves the performance of these algorithms across various metrics.
This study aims to develop an intelligent diagnostic system for ischemic heart disease (IHD) by integrating Spin Exchange Relaxation Free magnetocardiography (SERF-MCG) with advanced machine learning techniques. We analyzed cardiac magnetic signals from 565 patients (336 cases clinically diagnosed as IHD and 229 non-IHD cases). After data preprocessing including noise suppression and data format conversion, 2,317 features were extracted from ST segments and T-waves across time domain, frequency domain, isomagnetic maps, and current density maps to ensure the comprehensive mining of cardiac magnetic signals. Feature screening was conducted using univariate tests and four predictors: LASSO regression, random forest, Minimum Redundancy, Maximum relevance, and Light Gradient Boosting machine. Ultimately, 18 of the most distinctive features were retained for modeling. We compared the performance of different machine learning classifiers, including Logistic Regression, Support Vector Machine, K-Nearest Neighbor, Naive Bayes, Decision Tree, Random Forest, XGBoost, AdaBoost and Neural Network. The results showed that for the detection of IHD, random forest achieved the best performance with an AUC of 0.87, sensitivity of 91%, specificity of 74% and accuracy of 84%, with an F1 score of 87%. The model interpretation by the SHAP algorithm indicated roundness of positive and negative poles of T-wave peaks from isomagnetic map made major contributions to IHD classification. It provides clinicians with a rapid and accurate diagnostic tool to process and interpret MCG data, enhancing the acceptance and applicability of MCG in clinical practice. The proposed system provides a rapid and accurate diagnostic tool for clinicians, demonstrating the potential of SERF-MCG in optimizing IHD diagnosis and enhancing clinical applicability.
Compton cameras are electronically collimated cameras with higher efficiency and sensitivity when compared to PET (Positron Emission Tomography) and SPECT (Single Photon Emission Computed Tomography). One of the key challenges in realizing high resolution and accurate reconstruction in Compton cameras is the complexity in accurately modelling the detector response. Since the source is probabilistically averaged on the surface of a Compton cone, the detector response is more complex in Compton cameras compared to PET and SPECT. The detector response (system matrix) is often described by calculating the contribution of the surface of the measured Compton cones in the image space. In this paper a geometrical voxel based algorithm is proposed to compute the contributions of the Compton cone in the assumed image space. The geometrical structure of the cone and the voxel space are used to compute the contributions in the image space based on the measured events. The LM-MLEM (List Mode-Maximum Likelihood Expectation Maximization) algorithm is used to validate the modelled system matrix and compared with existing methods like RTM (Ray Tracing Method) and MC (Monte Carlo) based system matrix. These methods are evaluated on monte carlo simulated ideal data for both single and multiple point sources. The proposed method showed similar reconstruction accuracy and clear advantages in terms of modelling complexity and faster reconstruction.
Rice is a foundational staple for over half the world's population, making its protection vital for global food security. However, diseases caused by bacteria and fungi can lead to devastating yield losses, sometimes as high as 50%, posing a significant threat to economic stability. Traditional methods for detecting these diseases rely on manual inspection, a process that is not only labor-intensive and slow but also prone to human error, rendering it inefficient for large-scale monitoring. To address these challenges, this study proposes and validates a sophisticated deep learning framework for the rapid and accurate classification of rice leaf diseases. The framework employs an ensemble model that intelligently combines two powerful, pre-trained Convolutional Neural Network (CNN) architectures: EfficientNetB7 and ResNet50. These models were trained on a specialized dataset of images classified into four distinct categories: Bacterial Leaf Blight, Brown Spot, Leaf Smut, and Healthy. To enhance robustness and ensure high performance, the methodology incorporated extensive data augmentation, class weight balancing, and a five-pass TestTime Augmentation (TTA) strategy. The resulting ensemble model achieved an impressive validation accuracy of 95.31%. Crucially, it maintained high precision, recall, and F1-scores at or above 0.95 across all classes, demonstrating its reliability. These findings confirm that ensemble CNNs are a highly effective, scalable, and accurate diagnostic tool, marking a significant step forward for precision agriculture and plant pathology research.
A light-field (LF) camera with refocusing algorithms enables the reconstruction of images at different depths. Using images of various depths, the intensity distribution of each image section can then be calculated, which is particularly useful for analyzing complex scenes such as flames. However, LF images cannot be refocused in real time using conventional algorithms due to high computational demands. In this paper, we present a deep learning (DL)-based method for real-time refocusing of LF images of a burner flame, leveraging depth information obtained from the captured LF images. A Convolutional Neural Network (CNN) model is developed using a transfer learning approach, where a pretrained ResNet-50 model is extended with custom convolutional layers to perform the refocusing task. Synthetic datasets are generated using a ray tracing simulation based on the ray transfer matrix (RTM) method to train the model. The trained model produces four refocused outputs corresponding to distinct depth planes. The proposed method eliminates the need for the computationally intensive shift-and-sum algorithm traditionally applied to LF images. Simulation results show that the method accurately refocuses both geometric and flame structures, preserving depth-aware detail across planes. This approach has strong potential for enabling real-time, nonintrusive diagnostics in combustion systems and other dynamic, depth-varying environments.
The segmentation of anatomical structures in medical images and particularly in MRI scans, is essential for clinical diagnosis and monitoring disease progression. While Deep Learning (DL) architectures, such as U-Net and its extensions are very effective in medical image segmentation tasks, they often struggle with preserving fine-grained details and global contextual information. This is especially challenging for MRI data segmentation, where anatomical structures are characterized by irregular boundaries and variations in shape, contrast, and scale. To address this challenge, we propose a novel DL architecture for MRI segmentation across different anatomical structures. Specifically, the architecture introduces a module, named Multi-Resolution Feature Fusion (MRFF), that can be easily integrated into any U-Net-like architecture. The MRFF is integrated in all levels of an encode-decoder structure, along with attention mechanisms and skip connections to extract features at multiple resolutions, enabling the model to capture both fine-grained details and global contextual information. We evaluate the MRFFU-Net on two publicly available benchmark MRI datasets of different anatomical targets; one for Cerebrospinal Fluid (CSF) segmentation in spinal MR scans, and one for left atrium cardiac segmentation, from the Medical Segmentation Decathlon (MSD) challenge. Experimental results indicate that MRFFU-Net outperforms state-of-the-art models across multiple evaluation metrics, demonstrating its effectiveness in MRI segmentation.
The advent of satellite imagery has profoundly transformed the field of Earth observation and remote monitoring. This paper advances the state-of-the-art by introducing, for the first time, a novel design framework that enables the generation of a range of synthetic multispectral (MS) satellite images from true-colour RGB imagery and its associated metadata. The novel deep learning approach, introduced in this paper, leverages the generative architecture of latent diffusion models such as stable diffusion to generate multispectral bands conditioned on metadata and textual captions. A proof-of-concept case study is provided, demonstrating the feasibility of generating a single multispectral band from true-colour input images and their corresponding metadata. The work addresses a key issue in the application of deep learning to remotely sensed multispectral imagery, which is the scarce availability of high resolution, open-source imagery. Since a suitable band that overlaps with the generated band is not always available and given the goal of ensuring generalization to unseen data, the pansharpening-based superresolution techniques are not a viable solution. Instead, we adopt a metadata-conditioning approach, utilizing separate data sources from different satellites to address this challenge. The proposed deep learning model builds on the architecture of stable diffusion with two key modifications: the decoder is adapted to produce a single-channel output corresponding to a specific MS band, and a sinusoidal encoder is introduced to transform the metadata into a conditioning vector. This conditioning vector, along with textual captions, is used to guide a Denoising U-Net CNN during the image generation process. To reduce the inclusion of artifacts in the generated image, a ControlNet is also introduced conditioning the based U-Net model on the RGB image edge maps. The deep learning algorithms are trained using a dataset comprised of the Functional Map of the World (FMoW) dataset and its Sentinel counterpart, incorporating both textual captions and associated metadata. The paper presents the design methodology to training the separate components of the developed Synthetic Satellite Image Generation Framework, the dataset preparation, and the steps taken to train the models. The experimental results for the presented case study of generating NIR images from RGB imagery are introduced. The generated image clearly shows the framework’s ability to recreate the features of the NIR image from the RGB image. The results produced so far serve as an early indication that this novel framework does work and can be applied effectively to the RGB-to-NIR image generation task. The case study focuses on the generation of NIR images, however the framework can readily be extended to train on the rest of the joint FMoW dataset, enabling the generation of multispectral images of a variety of wavelengths.
Predicting the properties of cellular materials is important in engineering design as they are used ubiquitously in a broad range of applications. Here, we propose a deep learningbased approach to predict materials properties directly from 2D images, offering a fast and scalable alternative to conventional simulation techniques. Our framework is to train Convolutional Neural Networks (CNNs) by high-resolution image datasets generated using Finite Element Analysis (FEA) simulation outputs. By using semantic segmentation, class labels are assigned to individual pixels based on the stress and strain associated with that pixel. This enables precise identification of stress in a material, thus enabling the learning of spatial and hierarchical patterns in cellular materials subjected to loads. By using deep neural networks to automate the extraction of complex features from raw data, the model accurately predicts cellular material responses, aiding in the virtual characterization of cellular materials. We apply transfer learning techniques to enhance the performance of our CNN model by using the DeepLabv3 architecture with backbones such as ResNet-50, ResNet-101, and Inception-ResNet-v2, fine-tuning them for the task of predicting the properties of cellular material. This not only accelerates training, but also improves the model’s ability to generalize from limited data. The results show that the deep learning model outperforms traditional methods in terms of speed, and scalability, and can be used to predict cellular material properties applicable to a diverse range of engineering applications.