
Radiance field (RF) models stand out as powerful 3D scene descriptors capable of rendering realistic views from arbitrary directions. In the context of moving-camera scene change detection (SCD), RFs can be used to overcome common challenges such as parallax and occlusion, allowing different paths to be taken in reference and target scans. In this work, both neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) RF models are shown to be effective tools in solving the reference-target frame pairing problem. Employing modern structure-from-motion methods for localization of reference and query views, RF renders are used as references in an image-space comparison SCD, with good results. Using a new moving-camera SCD dataset, tailored towards the pairing problem challenges, a simple detector using RF renders as references achieved average precision scores of $87{\%}$ and $95{\%}$ with NeRF and 3DGS models, respectively, without any previous frame pairing.
Contrast-Enhanced Computed Tomography (CECT) is extensively utilized in cardiac imaging due to its superior anatomical detail, enabling accurate segmentation of cardiac structures. However, the use of contrast agents poses risks for patients with pre-existing medical conditions such as allergies. Non-Contrast Computed Tomography (NCCT) offers a safer alternative but presents significant challenges in segmentation, particularly for small and low-contrast structures such as the left ventricle, due to inherently lower image quality and complex three-dimensional anatomical variability. To address these limitations, this study proposes an enhanced 3D U-Net architecture integrated with dual attention mechanisms, specifically Convolutional Block Attention Modules (CBAM), designed to improve feature representation and enhance segmentation accuracy of fine-grained cardiac structures in NCCT images. Benchmark results demonstrate improved NCCT segmentation performance, indicating reduced dependency on contrast agents and enhancing the clinical viability of non-contrast cardiac imaging.
This paper describes a real-time Extended Reality (XR) Heart Twin based on ECG IoT measurements from wearable devices. Recent research has highlighted the importance of non-Euclidean features, such as wavelet graphs and heart meshes, in conveying information about a user's condition and enabling effective visualization in XR interfaces. We propose an end-to-end architecture to generate an XR Heart Twin from signals acquired by wearable devices. The XR Heart Twin architecture includes IoT communication protocols and middleware for data processing, extending to a web-based XR visualization tool that presents up-to-date features characterizing the user's condition. The proposed architecture enables a scalable monitoring solution suitable for cloud deployment. Examples of outputs presented via the XR interface are shown using ECG signals from publicly available datasets, processed in real time to extract XRrelevant heart features. This approach enables continuous user monitoring through a responsive XR dashboard. Experimental results demonstrate how an XR Heart Twin based on IoT measurements can be developed for healthcare systems using automated deployment capabilities on modern cloud platforms.
Satellite data is useful for many applications in environmental monitoring, such as vegetation cover and plant health determined by calculating vegetation indices. When using satellite data, the question often arises of how this data can be evaluated. This paper presents a method that can be used to evaluate aquatic plants detection on satellite data using an onboard camera and remote sensing methods. To this end, the Aquatic Plants and Algae (APA) composite index is used to compare estimates of the amount of seaweed in a lake in Northern Germany with satellite data. Sensor data from the mowing boat, on which the amount of seaweed is recorded with the respective mowing position, is used as a comparative value. It was shown that in the areas where the mowing boat harvested seaweed, the amount of seaweed identified by the APA index decreased compared to the amount identified before the mowing process. However, further data is required for quantitative and general results.
This work addresses the challenge of efficient Extended Reality (XR) communication in 6 G networks, where signals acquired by IoT sensors are transmitted for remote XR content reconstruction. Given the energy constraints of IoT devices and the low-latency demands of XR services, lightweight and fast encoding is essential. We propose a Semantically Weighted Principal Component Analysis (SWPCA), which selects components based on their relevance to XR rendering objectives-such as preserving object shape or motion. SWPCA applies data-dependent weights that reflect the saliency of sensor signals for reconstructing XR objects from the measurements. We analytically introduce SWPCA and apply it to XR gesture representation using wearable IoT measurements. We evaluate SWPCA performances within both an ideal and a realistic 6G end-to-end communication framework. SWPCA significantly reduces transmission bandwidth while preserving critical information for accurate XR gesture reconstruction. SWPCA provides a robust and fast encoding solution for real-time XR applications in energy/bandwidth constrained environments.
Vision language models (VLMs) are multimodal AI systems that combine a vision encoder with a large language model, enabling the AI to “see” and interpret the world. This capability has numerous applications, such as document understanding, image comparison, and visual analytics. However, it can also be leveraged for security purposes. In this paper, we propose a potential use case for VLMs in analyzing open source intelligence image data. This type of data is often publicly available on social media and other platforms and can be analyzed to extract valuable insights. Traditionally, human analysts are required to examine such visual data, but with recent advancements in VLMs and their ability to process large volumes of information efficiently, there is strong potential for their application in this domain.
This study presents a novel machine learning-based approach to mitigating digital replay attacks in face verification systems by leveraging compression artifacts to differentiate between compressed and uncompressed video frames. While traditional methods rely on active liveness detection, which can be inconvenient and negatively affect user experience, this work addresses the lack of automated solutions for detecting replay attacks. Using raw, uncompressed video datasets and widelyused video compression algorithms, the proposed method trains a classifier to identify compression artifacts as distinguishing features. Experimental results validate the model's effectiveness in detecting injected content, highlighting the critical role of compression artifacts in enhancing the robustness of video authentication systems. This contribution represents a significant step toward advancing anti-spoofing techniques by exploring a previously underutilized aspect of video integrity.
Liver radioembolization is a minimally invasive treatment for liver malignancies, requiring precise imaging to ensure accurate delivery of therapeutic agents. This paper addresses the key challenges in image registration for liver radioembolization, highlighting recent advancements with a focus on the integration of 3D Slicer software and the ellipsoid method. The oriented bounding boxes (OBBs) technique streamlines the representation of complex liver segments, which represent the object's contours. It simplifies them into simple oriented geometric shapes, making them easier to align. These shapes are then mathematically approximated as ellipsoids (a stretched sphere), enabling precise computations such as alignments, which form a central focus of this paper. This study evaluates the impact of using the OBBs as references for multi-modal medical imaging. Validation is performed with a liver phantom imaged via SPECT-CT (single-photon emission computed tomography) and PET-CT (positron emission tomography), confirming the accuracy and reliability of the proposed approach. Additionally, promising visual results on a patient's liver are also presented.
Positron Emission Tomography (PET) is a key imaging technique in clinical practice offering unique functional insights. However, its broader applicability is limited by radiation exposure and lengthy scan times, and reducing either degrades image quality by increasing noise. Over recent years, U-Net-based convolutional neural networks have become dominant approaches for PET denoising. These methods show good performance in denoising, but they offer limited uncertainty handling due to their deterministic nature and are prone to oversmoothing. This paper critically evaluates the applicability of denoising diffusion probabilistic models (DDPMs) for low-dose PET image denoising, particularly when training data is limited. We adapted a conditional DDPM architecture to the PET context and compared its performance to a U-Net baseline on clinical PET data. The conditional DDPM was both deployed as a standalone method and in a hybrid-ensemble format where it was placed in series with the baseline U-Net. The conditional DDPM, while promising in theory, introduced unrealistic features in some outputs and was computationally intensive. The hybrid model, intended to combine both strengths, underperformed due to high sample variability in DDPM outputs. Contrary to recent studies suggesting DDPM superiority, our experiments demonstrate that DDPMs underperform relative to U-Net, showing inferior PSNR and SSIM and introducing notable artifacts. Our findings highlight the importance of training dataset size and quality for DDPM effectiveness and provide practical guidelines regarding the trustworthiness of diffusion models for clinical PET denoising.
Tracheal intubation is an essential procedure for maintaining airway patency in patients, but is associated with significant risks, such as improper placement of the tube. Recent advances in video laryngoscopy have improved intubation safety by allowing precise tracheal segmentation in endoscopic images. However, many existing segmentation methods are either computationally intensive or fail to deliver robust results. In this work, we propose a novel multilevel thresholding approach that combines the statistical power of Kapur entropy with the Crayfish Optimization Algorithm (COA) to derive optimal thresholds from a two-dimensional histogram constructed using nonlocal mean filtering. Our method, termed 2DNLM-COA, achieves effective pixel separation and noise reduction while preserving critical edge features, thus facilitating improved identification of the tracheal region. Extensive experiments on video endoscopy data sets demonstrate that our approach outperforms several state-of-the-art methods. In particular, our method achieves a PSNR of 27.87, an SSIM of 0.9194, and an FSIM of 0.9080, indicating its strong potential to improve clinical intubation procedures.
This paper presents a deep learning-based framework for the automatic segmentation of kidneys, tumors, and cysts from CT images, built upon an enhanced DeepLabV3+ architecture. The model employs EfficientNet-B5 as the encoder for robust feature extraction, integrates a CBAM attention module to highlight key spatial and channel features, and uses ASPP for multi-scale context aggregation. Deep supervision is applied to improve learning stability and accuracy. Trained on 100 cases from the KiTS23 dataset, comprising 20,567 pre-processed 2D slices, the proposed method achieves a mean Dice score of 0.900, with class-wise scores of 0.955 (kidney), 0.884 (tumor), and 0.862 (cyst). It consistently outperforms existing approaches on KiTS23 dataset, demonstrating precise segmentation and accurate tumor area estimation. The study highlights the impact of combining attention mechanisms and deep supervision to advance segmentation performance on kidney CT images.
The increasing global demand for food, coupled with dwindling agricultural resources and the rising incidence of plant diseases, presents a significant challenge to global food security and sustainable agriculture. Timely and precise detection of plant diseases is crucial to minimizing crop losses, enabling early intervention, and maximizing yield. Given that leaves are often the most visibly affected organs, they serve as primary indicators for disease diagnosis. In this study, we propose a novel framework for plant disease classification that leverages both self-supervised representation learning and supervised finetuning. In the self-supervised phase, a lightweight encoderdecoder architecture learns meaningful representations by reconstructing input images, thus effectively utilizing unlabeled data. The encoder is built using group convolutions and a hybrid activation mechanism to balance computational efficiency and feature expressiveness. In the supervised phase, the pretrained encoder is fine-tuned for plant disease classification using labeled data, followed by adaptive average pooling and a lightweight fully connected classifier. Experimental results demonstrate that our model achieves a classification accuracy of $\mathbf{9 6. 3 6 \%}$ while maintaining a highly compact footprint of approximately 1.35 MB, making it well-suited for real-time, resource-constrained agricultural applications.
Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage guarantees. Full conditional guarantee is not an attainable goal, a well known fact in conformal predictions literature. As a result, several approaches have tried to approximate this behavior by adapting the conformal sets of test-time samples according to their similarity to calibration examples. Although the latter has gained traction and shown impressive performances for regression problems, its application to image classification remains under-explored. We conduct an extensive benchmarking on natural image classification tasks with vision-language models (VLMs), using our open source implementation of a recent localized conformal prediction algorithm. We show that straightforward usage of the cosine similarity between test-time and calibration visual features, an intuitive choice for VLMs, is not sufficient to improve over the non-local baselines. In response, we propose a simple non-linear transformation of the cosine similarities, which conserves marginal coverage guarantees and achieves statistically significant mean set sizes reduction. Code is available at https://github.com/cfuchs2023/lcp-vlm/.
In order to remain format compliant, cryptocompression of 3D objects is a very smart solution to guarantee both a high compression rate and a very high level of security. In this article, we present the analysis of a crypto-compression method based on the Draco format proposed by Google. The proposed analysis consists to analyze the distribution of edge lengths of crypto-compressed 3D objects as a function of the number of bits encrypted for each vertex. Experimental results, based on the Kolmogorov-Smirnov test, show that 5-bit encryption per vertex coordinate achieves an optimal trade-off between security and compression efficiency.
Semi-supervised subspace clustering has achieved promising results in hyperspectral image (HSI) analysis by leveraging spectral structure with limited labels. However, most existing methods suffer from two key limitations. First, they separate feature learning and clustering, preventing joint optimization. Second, they rely on explicit self-representation matrices with $\mathcal{O}\left(n^{2}\right)$ complexity, leading to significant memory and computation overhead. To address these issues, we propose an end-to-end semi-supervised deep subspace clustering framework tailored for HSI data. It unifies representation learning and clustering into a single trainable architecture, preserving contextual and subspace structures without computing large affinity matrices. The framework integrates cross-entropy loss, supervised contrastive learning, label-guided subspace modeling, and pseudo-label-based consistency regularization, enabling effective use of scarce labeled and abundant unlabeled data. Experiments on standard HSI datasets demonstrate that our method consistently outperforms existing unsupervised and semi-supervised clustering approaches under low-label regimes, illustrating the practical potential of semi-supervised deep subspace clustering for real-world HSI interpretation.
The rise of devices equipped with cameras resulted in a remarkable surge in user-generated content (UGC) videos to be shared on various platforms and applications. UGC videos suffers from shakiness, due to unsTable cameras hold, and this has boosted the development of video stabilization algorithms. The development of evaluation metrics has received comparatively less attention, and the existing metrics do not address the challenges posed by shakiness distortions. To fill this gap, we propose a novel no-reference metric based on deep neural network to evaluate the perceptual quality of shakiness in UGC videos. We extract spatial and temporal features using SlowFast network reinforce them by integrating optical flow features obtained by the Inflated 3D ConvNet network. The attributes of the two networks were integrated and passed through regression layers to generate the final shakiness quality score, referred to as the SlowFast-Flow metric. Comprehensive experiments confirm that the SlowFast-Flow metric outperforms state-of-the-art. The codes of the SlowFast-Flow are available at https://github.com/dborhen/SlowFast_Flow_VQA.git
Anomaly detection in industrial vision systems faces significant challenges due to the domain mismatch with large-scale datasets like ImageNet and the limited availability of labelled anomalies. This paper presents a reconstruction-based approach to detect defects in silk-screen printed glass, a domain characterized by structured patterns and scarce training data. We propose a two-stage approach: first, we use a zero-shot, class-agnostic segmentation model (Segment Anything) to generate candidate regions in the test image, then identify the region of interest by matching its features to those of a template region using feature similarity. Secondly, from the segmented areas, we extract patches to train and evaluate three generative models: a variational autoencoder (VAE), a generative adversarial network (GAN), and a DRAEM-inspired dual-network architecture with synthetically generated noise. We validate our methods on two custom datasets - S10 and M12 - collected from grayscale scans of glass plates. Results show that the GAN model consistently outperforms others, achieving high recall and precision even with limited training data, while DRAEM's reliance on synthetic noise impairs its generalization to real defects. Our study underscores the effectiveness of reconstruction-based methods in specialized industrial scenarios where conventional pre-trained features and synthetic anomalies may fall short.
Time series data often suffers from resolution limitations due to hardware constraints, sampling frequency restrictions, or economic considerations. While super-resolution techniques have seen significant advancements in computer vision, their application to spatio-temporal data presents unique challenges that remain under-explored. We argue that pure generative or auto-regressive approaches are subpar for the multi-modal super-resolution task. Hence, we introduce ChronoFusion, a novel hybrid model that simultaneously enhances both spatial and temporal resolution of time series data. Our approach leverages a graph variational autoencoder combined with adaptive attention mechanisms to generate high-resolution time series from low-resolution inputs. Unlike previous methods that handle spatial and temporal super-resolution separately, ChronoFusion integrates both dimensions through a proxy subspace. Extensive evaluation on traffic datasets in various locations demonstrates that ChronoFusion outperforms state-of-the-art methods by 10% on average in interpolation fidelity on unseen nodes while maintaining temporal consistency. Furthermore, our model demonstrates strong capabilities in handling missing data. The method's versatility across diverse spatio-temporal traffic applications makes it a valuable contribution to time series analysis and modeling. Github Repo
An increasing amount of multimedia data, including images, is being transmitted over networks and stored in the Cloud. This data requires both compression and security measures. In this paper, we present a new method based on Paillier's cryptosystem applied to images in order to perform homomorphic encryption, and compress them without loss of information. This allows us to reduce the size expansion between plain and encrypted messages, which by default is greater than 2 bits. Our method is based on an analysis of the random values of the parameter $r$ of the cryptosystem, in order to find encrypted values whose least significant bits (LSB) are equal to 0. Then we exploit this by removing k LSB par block to compress the encrypted image. This method is fully reversible, as we just need to add the previously removed LSB during the decryption phase, thus recovering the original image without any loss to quality. Experimental results show that our method can be applied to color images with large pixel block sizes.
We present a novel approach for zero-shot segmentation of electron microscopy (EM) images, where the masks obtained by the Segment Anything 2 (SAM2) model are clustered in order to retrieve just the objects of interest and filter out other structures in the image. Additionally, we propose a novel clustering algorithm inspired by the agglomerative bottom-up clustering approach that shows promising performance on two EM datasets. The results indicate that by using the proposed approach, we can significantly improve the performance of the vanilla SAM2 model without requiring any additional training data.