In medical imaging, patient motion and equipment malfunctions often introduce artifacts and degrade image quality, posing significant challenges during acquisition. Existing multi-modal MRI synthesis methods typically rely on static feature fusion, which often leads to redundant information and suboptimal anatomical consistency. To address these issues, this work proposes the Selective State-Space Module (SelSSM), a plug-and-play framework that reformulates multi-modal synthesis as a dynamic selective inference process. SelSSM maintains an evolving latent state that is iteratively refined through modality-aware updates guided by feature reliability and structural consistency. An input control mechanism initializes the state using global modality priors, while the subsequent state evolution integrates local structural details and long-range contextual dependencies. To further preserve anatomical integrity, a gradient-constrained adaptation enhances spatial alignment, and a frequency-domain refinement recovers high-frequency components. A knowledge-guided contrastive inference loss regularizes the evolving representation, improving feature discrimination and robustness under missing-modality conditions. Experimental results on two public datasets demonstrate that the proposed SelSSM achieves superior fidelity, structural coherence, and generalization performance across both synthesis and downstream segmentation tasks.
Emotion significantly impacts human cognition, influencing learning, memory, and decision-making. Although software-based methods have advanced emotion recognition, they lack biological plausibility and fail to capture the interplay between emotional valence and arousal. In this work, a dual-path memristive circuit is proposed to represent emotional states along two dimensions: valence and arousal. EEG β -band signals from the F3 and F4 channels are extracted and normalized to 0–1 V as circuit inputs. The valence pathway is designed based on the Brain Emotional Learning (BEL) framework, while the arousal pathway is constructed with reference to the Long Short-Term Memory (LSTM) model. Volatile and non-volatile memristors are employed to emulate short-term and long-term synaptic dynamics, respectively. The generated valence and arousal voltages are mapped onto a 2-D emotional state space, enabling the identification of four affective states: Happy, Sad, Calm, and Terrible. A classification accuracy of 83.3
This paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent Gaussian Splatting (GS) pipelines that treat all primitives uniformly during optimization, our framework explicitly manages the attribute activation, representation and pruning of Gaussian. Specifically, our framework first introduces GauSep, a novel densification strategy that selectively activates Gaussian color or normal attributes to alleviate destructive gradient conflicts arising from dual supervision. We further propose GauRep, an adaptive Gaussian representation that dynamically adjusts spherical harmonics (SHs) orders and performs task-decoupled pruning to reduce redundancy at both the individual and global levels. To provide reliable geometric supervision for above mangement process, we additionally introduce CoRe, an regularized surface reconstruction module that distills robust normal fields from an SDF branch to the Gaussian representation through a confidence mechanism. Notably, the proposed Gaussian management is compatible with various reconstruction architectures and can be seamlessly integrated to improve performance while reducing size of the model. Extensive experiments demonstrate that our approach achieves superior or comparable performance in appearance and geometry reconstruction compared with state-of-the-art methods, while using significantly fewer parameters.
Fluctuations in human blood glucose (BG) levels trigger changes in various physiological signals, creating opportunities for non-invasive BG measurement. This study proposes a novel BG measurement method for wearable devices, comprising signal preprocessing techniques and a BG measurement network. During preprocessing, a signal quality screening algorithm based on peak detection is introduced to effectively eliminate noisy signal waveforms. To reduce reliance on expert knowledge, the proposed BG network utilizes not only photoplethysmography (PPG), electrodermal activity (EDA), and skin temperature (ST) as inputs but also incorporates derived signals from PPG. Following feature extraction, we developed a multi-modal channel interaction mechanism that considers physiological relationships among various signals. This mechanism explores information related to BG fluctuations through physiological signals by leveraging dual coupling between the cardiovascular system and the autonomic nervous system. The experimental results demonstrate that the proposed network performs effectively, achieving a root mean square error (RMSE) of 14.67 mg/dL and an overall accuracy (ACC) of 84.38%. In summary, this paper presents an exploratory methodological framework for non-invasive BG measurement using multi-modal physiological signals, with results indicating its potential applicability.
The widespread application of batch color images in real-world scenarios such as multi-camera surveillance systems, cloud-based photo albums, and social media platforms has made their secure transmission a pressing concern. Existing encryption schemes often face practicality issues by mandating uniform image sizes and suffer from efficiency bottlenecks due to serial processing, especially with large-volume data. In this paper, we fully utilized the data structure of color images and proposed a parallel encryption scheme for protecting batch color images. Here, the color channels from different images are randomly grouped and encrypted in parallel. Furthermore, in each group, a simultaneous cross channel encryption strategy is leveraged to spread the encryption influence to other channels, which further improves the efficiency. Extensive experimental results demonstrate that our scheme achieves a significant speedup, processing encryption tasks up to 36.3
Parkinson's disease (PD) remains one of the most prevalent neurodegenerative disorders, where delays in diagnosis compromise therapeutic outcomes and increase healthcare costs. Conventional unimodal approaches, based on voice, sensors, or imaging, face critical limitations, including small datasets, lack of reproducibility, and high infrastructure demands. To address these challenges, the proposed multimodal agent-based architecture integrates medical language models, audio signals, and neuroimaging, and is supported by data-machine learning pipelines and an edge-cloud infrastructure. The system leverages ensemble learning, large and vision language models, and Retrieval-Augmented Generation (RAG) to enhance clinical decision support. The transparency of the model was supported by explainability techniques (SHapley Additive exPlanations, permutation importance, partial dependence, and individual conditional expectation), which highlighted the main audio and sensor variables responsible for the predictions. Experimental evaluation confirmed the effectiveness of multimodal fusion. When integrated, the architecture achieved robust performance, with an accuracy of 0.86, an F1-score above 0.88, ROC-AUC greater than 0.93, and both sensitivity and specificity above 0.89. Calibration and hypothesis tests were validated by a low Brier score of 0.205 and an Expected Calibration Error of 0.151, while Decision Curve Analysis confirmed clinical relevance by minimizing false negatives, critical for early screening, and reducing redundant interventions. Multimodal fusion produced accurate, well-calibrated, and interpretable risk estimates for PD screening; larger prospective studies and cost-effectiveness analyses are needed to consolidate clinical applicability.
In recent years, traumatic brain injury (TBI) has emerged as a critical yet under-recognized global health concern, contributing to significant mortality and long-term neurological impairment across all demographics. Computed tomography (CT) remains the gold standard for prompt detection of intracranial injuries post-TBI, and artificial intelligence (AI) is also exploited for empowering CT based TBI diagnosis. This survey reviews AI-driven approaches for TBI detection and prognosis on CT scans, highlights limitations obstructing their adoption in clinical workflows. Meanwhile, it surveys the publicly available datasets in this domain, encompassing aspects such as image resolution, the diversity of lesion types, and the application of state of the art (SOTA) approaches. Finally, we provide targeted recommendations to enhance secondary-injury modeling, expand dataset availability, and address prevalent challenges in primary injury assessment.
Up to date, the application of digital twin (DT) in the Industrial Internet of Things (IIoT) has been continuously promoted and deepened and has become the focus of the industry. IIoT serves as the foundational infrastructure that enables pervasive connectivity, real-time data acquisition, and intelligent control within industrial environments. DTs provide enterprises with an empathetic, virtual environment that enables them to manage and operate their production facilities in a more efficient and intelligent manner. However, there is not a special summary and analysis of the combinability and combination mode of the two. Therefore, this article first sorted out the professional definitions, characteristics, and frameworks of IIoT and DT, and deeply analyzed the semantic context of data flow. Second, this article discusses the combinability and combination mode of IIoT and DT and summarizes the enabling technologies and tools at each layer. Finally, the applications status of DT empowered IIoT in different fields was summarized, and the challenges of the combined application of the two were analyzed.
Nowadays, massive amounts of facial images have been tampered with and then widely spread through social networks. Many studies have developed algorithms for frame-level DeepFake detection. However, they have low robustness due to their focus on tamper-independent features during training. To this end, we propose a framework, namely MIF-Net, based on multi-information fusion for robust frame-level DeepFake detection. Specifically, key landmarks and the facial area are first detected in the original frame. Then, the graph convolutional network constructs biometric information from these landmarks. Meanwhile, the facial region is processed into multi-view inputs by noise and edge enhancement algorithms. Finally, these products are encoded as high-level features and classified as real or fake. Five benchmark datasets are utilized for testing our model through within-dataset and cross-dataset validations. Extensive experiment results demonstrate that our proposed MIF-Net is robust and has advantages over peer algorithms.
Recent years have witnessed the increasing applications of artificial intelligence for tooth treatment, among which tooth instance segmentation and disease detection are two important research directions. Advanced algorithms have been proposed, however, two challenging issues remain unsolved, i.e., unclear prediction boundaries for adjacent teeth, and high parameters of the model. To this end, our work proposes a lightweight framework, namely UCL-Net, for efficient tooth instance segmentation and disease detection. Specifically, uncertainty-aware contrastive learning is first employed for tooth segmentation. It is based on a multivariate Gaussian distribution to model the boundary pixel and is able to highlight inter-class differences, thereby refining the segmentation boundary. In addition, a lightweight segmentation model which has only 34.9 M parameters is further developed. Benefiting from the cross-scale attention, it is able to efficiently fuse different scale features, and therefore yields accurate tooth disease detection with a lightweight load. Four benchmark datasets are employed for performance validation. Both the qualitative and quantitative results demonstrate that the proposed UCL-Net is lightweight, effective, and advantageous over peer state-of-the-art (SOTA) methods.
With the rapid growth of the elderly population, fall accidents have received increasing attention due to their serious health hazards. Pre-impact fall detection (PIFD) based on wearable sensors emerges as a promising approach for proactive fall prevention in healthcare monitoring. In this research, based on Inertial Measurement Units (IMUs), we construct and publicly provide a large-scale motion dataset named FallTL, which includes falls and activities of daily living (ADLs) collected from multiple body segments. Furthermore, we develop STA-Net, a novel Spatial-Temporal Attention Network to perform PIFD based on IMU data from a single body segment. STA-Net incorporates a dual-branch architecture: a temporal attention branch that models temporal signal dependencies and a spatial attention branch that captures cross-modality feature interactions, enabling robust representation learning from sensor data. We evaluate STA-Net across three datasets and it achieves advantageous performance and comparable lead time under cross-subject validation, outperforming state-of-the-art baselines. In addition, our analysis further investigates the influence of sensor placement and data modality on detection performance. These results indicate that accurate and robust PIFD is feasible with minimally obtrusive, single-location sensor setups, offering practical implications for wearable fall monitoring systems.
To address the high computational costs of full-frame encryption and the risk of exposing sensitive locations in partial encryption, this paper proposes a video selective encryption and steganography scheme based on object detection and image inpainting. First, YOLOv8 is employed to achieve real-time and accurate detection of human targets in video frames. Then, the LIS-HMC hyperchaotic map and a new chaotic-driven interframe chain modulation (CDICM) strategy, combined with a designed row-column interchange and Roller confusion algorithm, are applied to selectively encrypt the target regions. Next, the globally and locally consistent image completion (GLCIC) algorithm is used to restore the background panoramically, eliminating visual discontinuities. Meanwhile, based on the Walsh-Hadamard transform (WHT), a multi-round embedding (MRE) steganography strategy is developed to hide the encrypted information within the restored background. Experimental results show that the encrypted data achieve an information entropy of 7.9925, a steganographic capacity of 0.75 bpp, and a PSNR above 44.91 dB after data embedding, demonstrating that the proposed method provides a new solution for video privacy protection that balances security, real-time performance, and visual naturalness.
With the rapid development of the electronic information industry, massive private data are continuously uploaded to the Internet, posing severe challenges to data security. Thus, a high-capacity private data information protection library based on correlation generator is proposed. For private data uploaded by multiple individuals or organizations, the proposed library enables efficient bulk protection. Through the biometric images held by individuals or organizations, correlation values are generated by the correlation generator, which are combined with such initial values of the four dimensional hyperchaotic system (4DHS) to generate private keys and master keys. The chaotic sequences generated by the iteration of the system are combined with the encryption scheme to provide effective protection of private data information. Afterwards, the scheme is tested for simulation and security, which verify the feasibility and security of the proposed scheme.
Semi-supervised medical image segmentation (SSMIS) has shown great potential in alleviating the scarcity of labeled data. However, its performance is frequently hindered by domain shift arising from variations in scanners, patient cohorts, and notably disease severity, where labeled and unlabeled samples follow different distributions. To address this challenge, we investigate the more practical Mixed-Domain Semi-Supervised Segmentation (MiDSS) scenario, where limited labeled data originate from a single domain, while abundant unlabeled data are drawn from multiple heterogeneous domains. We propose a novel Multi-Domain Collaborative Segmentation (MDCS) framework, where U-Net and MedSAM are jointly trained to combine precise local segmentation and generalized global representation, achieving complementary benefits. The Domain Feature Alignment and Distillation (DFAD) module ensures cross-domain feature alignment and facilitates the transfer of domain-invariant knowledge. Meanwhile, the Dual Confidence Consistency (DCC) mechanism dynamically balances collaborative training and refines pseudo-labels through confidence-driven weighting. Furthermore, a Region Discrepancy Regularization (RDR) term is employed to constrain inter-model prediction variance, thereby improving pseudo-label stability and segmentation accuracy. Extensive experiments on four public multi-domain medical image datasets demonstrate that MDCS achieves superior performance compared to other state-of-the-art SSMIS approaches, validating its effectiveness, robustness, and adaptability in real-world mixed-domain scenarios. Code is publicly available at https://github.com/TSZ-UU/MDCS.
Whole slide images (WSIs) analysis plays a critical role in computer-aided diagnosis. Recently, weakly supervised multiple instance learning (MIL) has become a widely adopted approach for WSI processing, as it enables effective learning from slide-level labels without the need for exhaustive pixel-level annotations. However, many existing MIL methods do not fully leverage the inherent pyramidal structure of WSIs and struggle to capture dependencies among instances and local contextual information, which may restrict their ability to capture rich hierarchical information. To address these issues, this paper proposes FIA-MIL, a dual-branch network designed for weakly supervised WSIs analysis that incorporates cross-scale feature interaction and alignment. Specifically, a dual-scale feature interaction module models semantic relationships across magnifications using a pyramid-aligned two-branch architecture, with Transformer encoders capturing instance-level dependencies within each scale. Furthermore, a dual-scale feature aggregation module integrates multi-scale features and introduces alignment constraints at both the bag and instance levels to ensure semantic consistency across scales. Classification and survival analysis experiments conducted on publicly available WSIs datasets demonstrate that the proposed method exhibits promising performance.
The diversification of attractor structures and their programmable regulation represent a significant direction in neural network dynamics. However, conventional memristive Hopfield neural networks (MHNNs) are often limited by symmetry, which restricts structural diversity and adaptive control. To overcome this, a programmable memristive Hopfield neural network (MP-HNN) is proposed. The programmability is realized by introducing a memristor switchable among three modes (odd, even, and increasing functions), through which the controllable expansion of the number of equilibrium points and the flexible configuration of the attractor structure are enabled. Under the increasing function mode, an asymmetric distribution of equilibrium points is generated, effectively breaking the traditional symmetric constraint. This paper systematically analyzes the equilibrium distribution, stability, and bifurcation behavior of the MP-HNN, and examines attractor evolution under different functional modes. Results show that coordinated parameter adjustment allows the network to produce double-scroll attractors and simulate neural firing patterns such as periodic and chaotic bursting. Significant synchronization is also observed between system state variables and memristor flux. Finally, FPGA-based hardware implementation closely matches numerical simulations, validating the model and theoretical analysis.
ABSTRACT As a fundamental task in restoring continuous visual signals, image inpainting plays a critical role in autonomous driving perception, medical imaging, video editing and digital heritage preservation. Driven by deep learning and large‐scale generative models, the field has transitioned from low‐level texture synthesis to high‐level semantic generation, yielding major breakthroughs in structural fidelity and visual realism. Centring on the generative paradigm as the architectural trajectory, this survey systematically categorizes the 30‐year evolution of image inpainting into three distinct technological generations: traditional prior‐driven synthesis, deep learning data‐driven reconstruction and modern foundation model‐driven generation. Despite this progress, highly competitive methods still struggle with large‐scale missing regions, global consistency in complex scenes, fine‐grained micro‐details and alignment with human visual perception. To address these gaps, we critically evaluate the technical paradigms and main bottlenecks within each of these evolutionary stages. We categorize and compare mainstream breakthroughs across high‐resolution restoration, text‐guided synthesis and complex scene generation. Furthermore, we compile standard benchmarks, evaluation metrics and quantitative performance comparisons of representative algorithms. Finally, we dissect open challenges—focusing on cross‐scene generalization and evaluation metric alignment—and outline future trajectories, particularly the integration of inpainting with text‐guided foundation models, providing a definitive reference for future theoretical and engineering advancements.
Self-supervised learning for time-series data has broad application potential in smartphone-based early disease detection. However, time-series data often exhibit complex dynamic patterns and spatiotemporal correlations. These characteristics make it difficult to capture discriminative features and reconstruct local features. Therefore, this paper proposes a cross-patch guided reconstruction with a graph contrastive learning framework (CGR-GCL) for smartphone-based early disease detection. CGR-GCL combines the advantages of mask reconstruction to capture complex dynamic patterns and contrastive learning to extract discriminative features. For the mask reconstruction part, we design an RGB-guided multi-scale masking module. Time-series data is first converted into RGB images. These images are then used to generate attention maps and multi-scale masks. The masks apply larger occlusions to important regions of the time series and smaller ones to less critical regions. This structural prior and semantic guidance help the model focus on key information and filter out redundant information. In parallel, we develop a dual-granularity contrastive module based on Graph Neural Networks. This module is able to capture spatial-temporal dependencies and enhance local consistency modeling. Compared with six baseline algorithms, CGR-GCL generally outperforms them (e.g., achieving +6.00% accuracy and +6.20% MF1-score on the mPowerTapping dataset with 10% labels). Meanwhile, a series of ablation results demonstrate the positive contributions of different modules to the performance improvement and the rationality of the overall framework. Our code is publicly available at https://github.com/heyiyia/CGR-GCL.git.
The new generation of artificial intelligence (AI) is evolving from perception-oriented processing toward cognition-oriented intelligence, emphasizing learning, memory, decision-making, and adaptive behavior inspired by the human brain. In this paper, a memristor-based perceptual decision-making neural network circuit (PDNNC) is proposed to integrate multiple classical conditioning mechanisms and fear generalization within a unified hardware framework. Unlike existing studies that focus on isolated learning behaviors, the proposed circuit architecture simultaneously realizes associative learning, forgetting, latent inhibition, blocking effects, secondary conditioning, and emotion-modulated decision-making at the circuit level. A voltage-controlled threshold-type memristor model is employed to emulate synaptic plasticity, while modular perceptual and decision-making subcircuits are designed to ensure scalability and functional consistency. The complete circuit implementation is verified using PSPICE circuit simulations. Simulation results demonstrate the correct realization of diverse classical conditioning behaviors, controllable fear generalization, and adaptive decision responses under different perceptual stimuli. The proposed bio-inspired circuit has vast potential applications in realizing cognitive intelligence and robotics.
Electrocardiogram (ECG) analysis represents a promising field for deep learning applications in clinical diagnostics. However, practical use of current methods is still constrained by their heavy reliance on large amounts of labeled data, as well as limitations in processing efficiency and signal quality. To address these challenges, we present BioFPT (Biosignal Feature Pyramid Transformer), a novel self-supervised learning framework designed for ECG signal. The proposed framework incorporates a Split Mask-Join (SMJ) transformation for a Pre-training strategy, complemented by an overlapping embedding mechanism that eliminates positional encoding requirements. The efficiency of the architectural design is enhanced through a Spatial Reduction Attention (SRA) transformer, which achieves a reduction in computational complexity without performance degradation. Comprehensive evaluation of seven public ECG datasets comprising over 94,000 sub jects demonstrates BioFPT's effectiveness, an accuracy improvement of 4.2% and a parameter reduction of 14.8% compared to state-of-the-art models. Furthermore, it maintains robust performance across diverse pathological conditions and signal qualities. The proposed architecture rep resents a significant advancement in self-supervised ECG analysis, particularly suitable for scenarios with limited labeled data availability. Moreover, its versatile architecture shows promise for broader applications across various biosignals.