
In learning-based image compression methods, entropy modelling remains crucial for the achievement of higher compression performance with low computational cost. In this study, we introduce a framework for learned image compression that combines channel-wise autoregressive entropy modelling, LSTM-based context adaptation, and latent residual prediction (LRP) to improve entropy estimation and reduce reconstruction errors. The model employs channel-wise conditioning to progressively refine latent distributions, while an autoregressive prior captures spatial dependencies, allowing for more accurate probability predictions. Furthermore, latent residual prediction reduces quantization errors, resulting in more accurate reconstructions. Experimental results indicate that the proposed method surpasses other deep learning-based compression techniques, achieving a 12.7% BD-rate (Bjontegaard Delta-Rate) reduction compared to a baseline model that utilises channel-conditioning and LRP. These findings highlight the effectiveness of combining hierarchical entropy modelling with spatial and channel-wise context adaptation to achieve state-of-the-art compression efficiency.
As the security of public spaces remains a critical issue in today's world, Digital Twin technologies have emerged in recent years as a promising solution for detecting and predicting potential future threats. The applied methodology leverages a Digital Twin of a metro station in Athens, Greece, using the FlexSim simulation software. The model encompasses points of interest and passenger flows, and sets their corresponding parameters. These elements influence and allow the model to provide reasonable predictions on the security management of the station under various scenarios. Experimental tests are conducted with different configurations of surveillance cameras and optimizations of camera angles to evaluate the effectiveness of the space surveillance setup. The results show that the strategic positioning of surveillance cameras and the adjustment of their angles significantly improves the detection of suspicious behaviors and with the use of the DT it is possible to evaluate different scenarios and find the optimal camera setup for each case. In summary, this study highlights the value of Digital Twins in real-time simulation and data-driven security management. The proposed approach contributes to the ongoing development of smart security solutions for public spaces and provides an innovative framework for threat detection and prevention.
We study a recently proposed model for opinion dynamics that takes into account social pressure, that is, the fact that agents are reluctant to reveal their actual opinions in an environment where other agents may have different views. In the examined model, each agent has an inherent opinion that represents their actual views, while agents may choose to declare an opinion that is different from their inherent one when they sense that they are in an environment where neighboring agents do not agree with their inherent opinion. At any point in time, agents communicate their stated views within the network which depend on the social pressure they receive within the network as well as their inherent views. Nodes tend to conform to the views of other nodes in the network. We perform numerical experiments using this model and several types of graphs that are commonly used to represent social networks. Our scope is to measure the strength of social pressure in various scenarios, by demonstrating the honesty percentage, i.e., the percentage of agents that declare their true inherent opinion, as a function of time. The numerical results produced show that the type of network significantly affects the characteristics of the social pressure.
Objective: To compare three deep-learning architectures-Convolutional Neural Networks (CNN), Long Short-Term Memory networks (LSTM) and Transformers-for binary gait classification from a single inertial sensor. Methods: Sixty-eight participants (25 healthy controls, 43 patients with neurological gait impairment) walked for 60 s while a tri-axial IMU placed at C7 recorded linear acceleration, angular velocity and magnetic field at 50 Hz. Each trial was segmented into 2.56 s windows (128 samples) with 50 % overlap, producing 25 040 labelled samples (74 % normal, 26 % pathological). Architectures were optimised with random search and evaluated in subject-wise 5-fold stratified cross-validation. Balanced accuracy, macro-F1 and inference time were the primary metrics. Results: The bidirectional LSTM achieved the best balanced accuracy (63.5 %) and macro-F1 (0.34), marginally ahead of the CNN (62.9 %, 0.30). The Transformer reached only 50.5 % and 0.03, failing on the minority class. Conclusions: Even with a single sensor and a modest cohort, sequence-aware models capture discriminative temporal cues. LSTMs offer the best trade-off between performance and computational load, but class imbalance remains a limiting factor that future work must address.
Accurate detection of epileptic seizure events is an extremely challenging problem due to the complex and highly variable nature of seizure activity, which can manifest differently in each individual. This study aims to obtain a meaningful representation of EEG data by utilizing a) Poincare plots, which can effectively capture the underlying temporal characteristics of EEG signal; and b) hybrid autoencoder-classifier model. The bottleneck layer of autoencoder helps to provide a compressed latent representation, while classification head is used for accurate class predictions. The proposed approach is validated on four benchmark datasets, which are publicly available. The mean accuracy and false detection rate across all the subjects in each of these datasets show its superior performance, not only in identifying seizure and non-seizure events, but also in distinguishing across different types of seizures. The method is able to distinguish healthy controls and epileptic patients from BONN dataset with 100% accuracy. It provides 90% accuracy in classifying ictal, interictal, and preictal segments from New Delhi dataset. This classification task becomes comparatively more challenging, as it requires the model to differentiate between subtle variations in brain activity that occur during different phases of the seizure cycle. For the Epileptic EEG dataset, it provides 97% accuracy in the identification of different types of seizures and healthy data. In case of Turkish EEG dataset which comprises of 121 subjects, the model achieves 88% accuracy in classifying healthy and epileptic segments in a subject-independent manner. Overall, the proposed method provides results that are on par with, or even surpass, those reported in existing works.
This study presents a supervised deep-learning approach for muscle identification using surface Electromyography (sEMG). We propose an optimized hybrid Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) model to identify seven lower-limb muscles from only sEMG data, while addressing the inter-subject variability challenge. Using the publicly available ENABL3S dataset, we preprocess raw sEMG signals by detecting muscle bursts, removing overlaps, and zero-centering samples. Our model underwent iterative optimisation for both intrapersonal and interpersonal validation, achieving 97% and 93% accuracy, respectively. The results demonstrate improved feature extraction and generalization, advancing sEMG-based rehabilitation technologies.
Traditionally, dataset partitioning for training, testing, and validation purposes in Convolutional Neural Networks (CNNs) has used either random selection or cross-validation techniques. However, in complex datasets such as the VNPlant-200 medicinal plant collection of images in their natural habitat, high intra-class and low inter-class variations between captured images make it challenging to reliably achieve accurate classification. Mutual Information (MI) has recently been applied as a similarity measure to guide the CNN training set selection process by synthesising more representative images. While the MI-Guided Training (MIGT) algorithm improved accuracy performance, misclassification rates for certain species were still high due to the recurring influence of extraneous background clutter, blurring and illumination effects in many of the images. This paper introduces a novel Gaussian Mixture Model-Mutual Information Guided Training (GMM-MIGT) dataset partitioning algorithm that uniquely incorporates a combination of Expectation-Maximisation (EM), Gaussian Mixture Model (GMM) clustering, and Principal Component Analysis (PCA) as a two-stage preprocessing pipeline before MI-based partitioning. GMM-MIGT lowers colour complexity, removes less useful dimensions, and enhances similarity estimation, ensuring more representative partitions. Experimental results validate the GMM-MIGT algorithm's consistently superior performance over existing partitioning methods, achieving classification accuracies between 83% and 99%, allied with significantly lower misclassification rates, especially in complex datasets with high visual variability.
Despite significant advancements in denoising techniques for low-dose Cone-Beam Computed Tomography (CBCT) images, developing robust Artificial Intelligence (AI)-based approaches remains a critical challenge due to the unavailability of real paired CBCT datasets. While various methods have been proposed to enhance low-dose CBCT images, their effectiveness is limited by the absence of paired datasets and the use of CT images as substitutes for high-dose CBCTs, leaving room for further improvement. In this paper, we first introduce a novel CBCT dataset in which high-dose target images are synthetically generated from low-dose scans using an advanced enhancement method. Additionally, we propose a modified Selective Frequency Network (SFNet) that better captures CBCT-specific features by integrating Convolutional Block Attention Modules (CBAMs) after Residual Blocks (ResBlocks) and features an improved hierarchical shallow layer for enhanced feature extraction. Objective evaluations using full-reference and no-reference metrics demonstrate that our modified SFNet surpasses the state-of-the-art CBCT denoising approach. More importantly, our model produces CBCT images of even higher quality than the realistic high-dose targets generated from low-dose CBCT for training, as it learns a more refined mapping that reduces residual artifacts present in generated and real high-dose images. This advancement significantly improves CBCT image quality while minimizing radiation exposure, enhancing diagnostic reliability, and broadening clinical applications.
The offline Multiple Appropriate Facial Reaction Generation (OMAFRG) task is designed to model interactive communication in dyadic conversations using audio and video inputs from a speaker to generate multiple appropriate listener motions. This capability is crucial in enhancing user interactions with mental health care robots. To date, most existing research on OMAFRG has employed audio-visual modalities, utilizing video and audio encoders to form a latent behavioral representation for subsequent reaction prediction. However, this approach is computationally demanding and overlooks the latent semantic information in facial expressions, which are vital for robust appropriate reaction prediction. Besides, the research on non-deterministic motion estimation always lacks sufficient facial affect labels further complicating the task and the video encoder always demands higher computational cost than the audio encoder. From these insights, we design a novel video encoder that uses the deep facial affect prior knowledge sequences as the input feature rather than raw RGB video frames. The numerical analysis and experimental results not only demonstrate that the proposed method achieves considerable improvements compared to the existing methods but also offer the additional advantage of being more computationally efficient.
Inpainting is an image processing technique used to remove alterations and restore the original content of a modified image. The effectiveness of inpainting can be evaluated through various methods, increasingly based on criteria related to human perception. In this work, we propose a methodology to assess the effectiveness of different inpainting techniques from the perspective of an image's emotional value. We apply Visual Sentiment Analysis (VSA) algorithms to determine which techniques are most effective in restoring the original emotional value of modified images. The goal is to identify the most suitable methods for removing common image modifications, specifically manipulated elements, while preserving the image's original sentiment. This analysis is particularly relevant for assessing the communicative impact of modified images and understanding the emotional value of such alterations.
In recent years, researchers constantly attempt to derive Ultrasound-based (US-based) hand gesture recognition (HGR) solutions suitable for edge applications. This process involves improving several design aspects of US-based HGR systems such as the transducers, the wearable US acquisition systems and the algorithms employed in terms of energy consumption, computational complexity and robustness. The subject of this paper is the latter. In this paper, we present a spiking framework for US-based HGR. The proposed approach leverages a single-layer Spiking Neural Network (SNN) equipped with Spike-Timing-Dependent Plasticity (STDP) as a feature descriptor for Rate-based (RB) coded A-line US signals coupled with a lightweight linear support vector machine (SVM) classifier. According to our findings, our proposed approach achieves performance comparable to that of the state-of-the-art in the ultrasound-based adaptive prosthetic control (Ultra-Pro) dataset. Furthermore, we demonstrate that our feature descriptor exhibits inter-session generalization capabilities, i.e. re-training is not required between within-day sessions and thus reduces the burden of periodic extensive data collection from the user.
3D image classification plays a crucial role in fields such as computer vision, archaeology, and medical imaging, where accurate object recognition is essential. However, 3D classifiers alone may struggle to extract sufficient discriminative features, limiting their accuracy. To address this, we propose a hybrid framework that integrates both 3D and 2D classification techniques. Our approach extracts multi-view 2D projections from 3D objects and leverages them alongside 3D structural features to enhance classification performance. By combining the outputs of both modalities, our framework produces more robust and accurate predictions. We evaluate our framework using three state-of-the-art 3D classifiers (PointNet, PointNet++, and Mamba3D) and three 2D classifiers (ConvNeXt, EfficientNet, and ResNet). We also evaluate different methodologies to combine the predictions of the classifiers. Experimental results show that our hybrid approach consistently outperforms 3D-only classification. On the ModelNet10 dataset, PointNet++ accuracy improved from 88.93% to 94.38%, while on ModelNet40, it increased from 87.56% to 92.67%. These findings highlight the effectiveness of integrating multi-view and 3D classification for improved object recognition.
Accurate intraoperative localization of heart catheters is crucial for the success of time-critical cardiac procedures. Traditional imaging techniques, such as fluoroscopy and conventional coronary angiography, expose patients and clinicians to ionizing radiation and require contrast agents, increasing procedural complexity and risk. To circumvent some of these challenges, there is growing interest in utilizing magnetic sensor arrays to track catheters equipped with miniature magnets. However, existing approaches face critical challenges including high computational complexity, susceptibility to external magnetic noise, restricted tracking range, poor localization performance outside the physical boundary of the array. This study presents a new magnetic localization system using a hall-effect sensor array. It employs a hybrid localization strategy; integrating Weighted Centroid-Based Localization for precise tracking within the physical boundary of the sensor array and a two-dimensional Gaussian Function for accurate extrapolation outside the sensor array. A dynamic switching mechanism between two localization methods ensures accurate magnet localization across the entire operational space. Experimental validation was performed at a data acquisition and processing rate of 1 Hz. Experiments with a robotic arm moving a permanent magnet along a predefined 2D-trajectory demonstrate accurate localization within (<2 mm) and outside(<5 mm) the sensor array up to 80 mm elevations. These findings suggest that the detection range of magnetic sensor arrays can be extended without increasing the number of sensors. Future work will focus on reducing localization error below 1 mm, extending localization to three dimensions and integrating adaptive noise filtering techniques to make the system more robust in operation rooms with background noise.
Point Cloud Registration (PCR) is a fundamental task in 3D perception, with applications in Simultaneous Localization and Mapping (SLAM), 3D reconstruction, object recognition, and autonomous navigation. It involves aligning two or more point clouds into a common coordinate system, a process that is often challenged by noise, outliers, and variations in point density. Traditional registration methods, such as Iterative Closest Point (ICP) and its variants, rely on minimizing point-to-point or point-to-plane distances, but they may struggle in the presence of anisotropic errors and misaligned correspondences. In this work, we propose a novel point cloud registration method that incorporates an anisotropic error model to better account for the underlying geometric structure of the data. Utilizing key results from Linear Algebra, we derive theoretical guarantees on the robustness of our approach. Our method ensures position and orientation consistency of matched pairs by analyzing the covariance structure of registration errors, reducing the impact of outliers and enhancing the stability of the alignment process. Experimental results of this study provide valuable insights into the role of covariance modeling in point cloud registration and open new possibilities for improving the reliability of 3D perception in robotics, autonomous systems, and augmented reality applications.
Azure Kinect is a popular low-cost markerless Motion Capture (MoCap) system, showing promising results in clinical applications. However, during concurrent validation studies with a marker-based gold standard, reflective markers produce passive infrared (IR) noise, which significantly interferes with its tracking accuracy. In this study, we collected motion data from 15 healthy participants performing upper and lower limb exercises, concurrently recorded by Azure Kinect and the Vicon system. We found that Kinect's skeletal tracking primarily relies on IR images rather than depth images. Therefore, we developed a simple yet effective algorithm to mitigate noise in IR images. Our method significantly improved Kinect's skeletal tracking reliability, reducing missed poses from 10% to negligible levels and decreasing bone length variability across frames. Additionally, joint angle measurements improved, with lower Mean Absolute Error (MAE) in Range of Motion (ROM) and higher Intraclass Correlation Coefficient (ICC) of ROM. The code developed for this study is available at https://github.com/spongebobbe/pyKinectAzureImageManipulation.
This paper investigates decentralized online economic dispatch for smart grids in the presence of Byzantine attacks. In smart grids, some generation stations are subject to external manipulations or internal damages, such that they behave maliciously and send wrong messages to neighboring generation stations, thereby disrupting the online economic dispatch optimization process. We utilize the classical Byzantine attack model to characterize these malicious behaviors and propose a class of Byzantine-resilient decentralized online economic dispatch algorithms. The proposed algorithms employ a variety of existing robust aggregation rules to effectively filter out wrong messages, and are proved to achieve linear dynamic regret and accumulative constraint violation. Numerical experiments are conducted to validate our theoretical results.
Cell segmentation is a fundamental task in biomedical image analysis, particularly in pathology, cancer research, and diagnostics. Deep learning-based approaches, especially convolutional neural networks (CNNs), have significantly improved the accuracy of cell segmentation. In this study, we present a web-based platform that integrates StarDist, a state-of-the-art deep learning model specifically designed for nucleus segmentation. The system provides an intuitive interface for researchers and clinicians to upload images, apply pre-trained and custom-trained StarDist models, visualize segmentation results, and download statistical data for further analysis. This framework enables real-time nucleus segmentation without requiring specialized computational resources, making advanced deep learning-based segmentation more accessible to non-expert users. Finally, we evaluate the performance of different StarDist models on the SIPaKMeD dataset, assessing metrics such as precision, recall, accuracy, Intersection over Union (IoU), to validate their effectiveness and to provide insights into the trade-off between model complexity and segmentation accuracy.
In this paper, we present a novel approach for picture painting that extends the Intelligent Painter framework by integrating a Latent Diffusion Model (LDM) with text embedding for text-to-image generation. Similar to the original Intelligent Painter, which focuses on painting people's imagination, our method operates in the latent space using DDPM sampling and incorporates a mean-conditioned masking strategy. This new strategy allows the model to effectively utilize both object inputs and textual descriptions to guide the composition process. Experimental results demonstrate that our approach produces pictures with improved harmonization and semantic consistency compared to previous methods. The proposed framework offers a controlled and flexible solution for generating new pictures that combine user-specified spatial cues with rich textual information without the need of retraining. This further enhances people's imagination. The codes for the research is available at https://github.com/Hammond65/Advance-Intelligent-Painter
The safety of urban intersections is a critical concern for city planners. Technological advancements, such as LiDAR sensors, enable better risk assessment for road users. This study proposes a hybrid model that combines Post-Encroachment Time (PET) data with unsupervised machine learning techniques, specifically DBSCAN clustering, to detect traffic anomalies. A generalized Pareto distribution (GPD) is then applied to estimate a risk index. Finally, categorical safety risk classification is performed using an optimizable neural network (ONN), support vector machine (OSVM), efficient logistic regression (ELR), and Gaussian Naive Bayes (GNB). The impact of these methods is evaluated in real-time for urban traffic management in Trois-Rivieres, Quebec, Canada. This work aims to assist decision-makers in urban traffic planning and accident prevention.
Emotion recognition is a challenging issue in brain-computer interface (BCI) research and in online evolutive learning. The high interindividual variability in physiological signals presents challenges for the generalization of deep learning models. With the limitations of small data sets and insufficiently labeled data, enhancing the cross-domain learning capability of models remains a critical issue. Although it is possible to make judgments based on a single data source, such an approach is prone to mis-judgments because of its lack of comprehensiveness. Multimodal learning methods improve perceptual learning by constructing a multi-modal system that integrates input from multiple sources simultaneously. To address this challenge, we propose DuMoNet, a dual-modality cross-subject emotion recognition network based on semi-supervised learning, which leverages both electroencephalography (EEG) and electrocardiography (ECG) data. The framework consists of two independent networks: a temporal multiscale convolutional network (TMSCNN) for EEG data and a ResNet18-based model for ECG data, both designed to learn feature representations in the time domain. Knowledge is exchanged between models through reliability comparison, while consistency constraints mitigate the influence of noise, thereby improving the robustness of the model. Furthermore, DuMoNet can adapt to datasets from other source domains, enabling online evolutive learning.