The existing methods for fake news videos detection may not be generalized, because there is a distribution shift between short video news of different events, and the performance of such techniques greatly drops if news records are coming from emergencies. We propose a new fake news videos detection framework (T^3SVFND) using Test-Time Training (TTT) to alleviate this limitation, enhancing the robustness of fake news videos detection. Specifically, we design a self-supervised auxiliary task based on Mask Language Modeling (MLM) that masks a certain percentage of words in text and predicts these masked words by combining contextual information from different modalities (audio and video). In the test-time training phase, the model adapts to the distribution of test data through auxiliary tasks. Extensive experiments on the public benchmark demonstrate the effectiveness of the proposed model, especially for the detection of emergency news.
The precise segmentation of vascular structures is vital for diagnosing retinal and coronary artery diseases. However, the complex morphology and large structural variability of blood vessels make manual annotation time-consuming and finite which in turn limits the scalability of supervised segmentation methods. We propose a semi-supervised segmentation framework named geometric orientational fusion attention network (GOFA-Net) that integrates differentiable geometric augmentation and orientation-aware attention to effectively leverage knowledge from limited annotations. GOFA-Net comprises three key complementary components: 1) a differentiable geometric augmentation strategy (DGAS) employs quaternion-based representations to diversify training samples while preserving prediction consistency between teacher and student models; 2) a multi-view fusion module (MVFM) orchestrates collaborative feature learning between quaternion and conventional convolutional streams to capture comprehensive spatial dependencies; and 3) a global orientational attention module (GOAM) enhances structural awareness through direction-sensitive geometric embeddings, specifically reinforcing the perception of vascular topology along horizontal and vertical orientations. Extensive validation on multiple retinal vessel datasets (DRIVE, STARE, CHASE_DB1, and HRF) and coronary angiography datasets (DCA1 and CHUAC) show that GOFA-Net consistently outperforms state-of-the-art semi-supervised methods, achieving particularly notable gains in scenarios with limited annotations.
Objective Accurate segmentation of ultrasound images is crucial for medical diagnosis and treatment. However, achieving precise lesion segmentation is challenging owing to inter-class indistinction caused by low contrast, high speckle noise and blurred boundaries, as well as intra-class inconsistency resulting from variations in lesion size, shape and location. Methods To address these challenges, we propose a frequency-aware vision mamba network with deformable windowed selective scan. Specifically, we introduce the frequency-aware selective state space layer, which uses fast Fourier transform to extract frequency-domain information and applies frequency-aware weighting to enhance discriminative features, effectively mitigating intra-class inconsistency. Moreover, we developed a deformable windowed selective scan module that dynamically adjusts scanning paths by a learnable offset field to focus on ambiguous boundaries, significantly reducing inter-class indistinction. Results Extensive experiments on 4 cross-domain ultrasound datasets demonstrate that the frequency-aware vision mamba network outperforms other state-of-the-art methods, achieving Dice coefficient scores of 0.84, 0.83,0.84 and 0.87 on breast ultrasound images, thyroid gland segmentation, thyroid nodule segmentation and classification in ultrasound images and multi-modal ovarian tumor ultrasound datasets, respectively, with corresponding intersection over union scores of 0.76,0.76,0.76 and 0.80, respectively. Conclusion Our results show that the proposed method provides a robust solution for clinical ultrasound image analysis, addressing inter-class indistinction and intra-class inconsistency challenges.
Synthesis pathway planning is a critical step in the design and discovery of novel inorganic compounds with desirable properties, but existing methods are often constrained by limited generalizability. This study develops a specialized language model, SynPathLM, based on tens of thousands of inorganic synthesis reactions mined from scientific literature using natural language processing (NLP) techniques, which are converted into a structured language. The model achieves end-to-end retrosynthetic planning, enabling precursor recommendation for target compounds, prediction of reaction types, and forecasting of key step temperatures (e.g., calcination and sintering). To address the inherent limitations of large language models (LLMs) in numerical prediction, we further propose SynPathLM-Reg, which integrates MAPP physicochemical prior fusion, attention pooling, a residual regression head, and temperature distribution oversampling design on top of the language model encoder. This significantly improves temperature regression accuracy, reducing the mean absolute error (MAE) for calcination temperature prediction by 54.0
Molecular property prediction faces a persistent trade-off between predictive accuracy and model interpretability. Graph neural networks achieve high accuracy through end-to-end learning from molecular graphs, but their predictions are difficult to interpret. Traditional QSPR methods produce interpretable linear equations, yet cannot capture nonlinear structure-property relationships. This paper proposes a QSPR framework based on the Kolmogorov-Arnold Network (KAN) that uses a small set of explicit molecular descriptors as input, models nonlinear relationships through learnable B-spline activation functions, and extracts symbolic equations from the trained network. Experiments on eight benchmark datasets show that KAN achieves competitive performance with graph neural networks while using only approximately six explicit descriptors on average. The extracted symbolic equations substantially outperform traditional linear QSPR formulas in fitting ability, and the learned functional forms (e.g., logarithmic, exponential) are consistent with established chemical knowledge. These results demonstrate that KAN offers a viable middle ground between purely linear QSPR models and black-box deep learning.
Multi-view subspace clustering has demonstrated superior performance over the single-view clustering by leveraging the complementary and consistent information across multiple views. However, existing clustering methods primarily focus on learning linear feature representations and deriving consensus latent representations, which overlook the nonlinear structures inherent in the data, and thus lead to the representation degradation and suboptimal clustering performance. To address these challenges, we propose a deep multi-view subspace adversarial clustering network via non-negative matrix factorization feature enhancement, termed NFEMAC. Specifically, the original data is first projected into a high-dimensional space by leveraging the separability, then the fully connected layer and the encoder network are employed to preserve the local manifold structure of the data. Furthermore, we introduce a novel Cross-view Feature Enhancement Fusion (CFEF) module, which utilizes the Non-negative Matrix Factorization (NMF) to extract the distributional structure of the data from the local manifold information, thereby enhancing the latent representations of individual views. An attention mechanism is subsequently applied to obtain a consistent latent representation. Finally, a Generative Adversarial Network (GAN) is incorporated to further enhance the robustness of the shared latent representation. Extensive experiments are implemented on seven publicly available datasets to evaluate the performance of the proposed method. Experimental results demonstrate the superiority of NFEMAC compared with state-of-the-art methods.
The rotating machinery system consists of several key components such as bearings and gears. The operating condition of the bearings directly affects equipment safety and production efficiency. However, traditional bearing fault diagnosis methods face challenges in complex operating conditions, including insufficient local feature extraction, severe noise interference, and difficulty in integrating global information due to the heterogeneity of multi-sensor data. To address these issues, this paper proposes a multi-sensor and multi-task fault diagnosis method based on the multi-scale hidden state interaction network (MHSNet). In terms of feature extraction, MHSNet integrates deep separable convolutions with hidden state-space models. By introducing multi-scale convolution units, it captures local details under different receptive fields. Additionally, the selective hidden state modeling mechanism of the Mamba module overcomes the limitations of conventional convolution networks’ local receptive fields, enabling the modeling of periodic impulses and long-range dependencies in signals. In the data fusion layer, a dynamic state space fusion module is designed to achieve parameterized interaction and adaptive alignment of multi-sensor data within the hidden state space, effectively alleviating the distribution differences and redundancy issues between multi-source information. Through the collaborative extraction of complementary features between tasks, the model further enhances robustness and discriminative accuracy under conditions of data imbalance and noise interference. Extensive experiments conducted on real bearing data and multi-condition testing platforms demonstrate that MHSNet consistently achieves high diagnostic accuracy and condition classification performance. It outperforms traditional single-modal and heterogeneous multi-sensor signal-based diagnostic networks, highlighting its significant advantages in multi-sensor collaborative representation, global and local feature fusion, and noise suppression.
The escalation of antimicrobial resistance (AMR) has become a major global public health threat, creating an urgent need for novel antimicrobial therapeutics. Antimicrobial peptides (AMPs), owing to their unique mechanisms of action and low propensity for inducing resistance, are regarded as key candidates for combating AMR. However, existing generative models struggle to balance generation quality, efficiency, and controllability. This study proposes the CFlowAMP framework, which integrates the protein language model ESM-2 with property-controlled Conditional Flow Matching to model a direct mapping trajectory from random noise to the latent representations of peptide sequences. Experimental results demonstrate that CFlowAMP substantially enhances generative performance: it increases the AMP generation success score by 39.8%, achieves an 18-fold speedup over diffusion-based models, and enables on-demand generation with controllable physicochemical properties. This approach provides a generalizable computational framework for therapeutic peptide design.
Traffic forecasting is pivotal but challenging due to intricate spatio-temporal dynamics. Existing models often apply a uniform spatial mechanism across distinct temporal scales and rely on static feature embeddings. Consequently, they are inadequate in capturing scale-specific spatial heterogeneity and dynamic feature interdependencies. To address these limitations, we propose the Multi-Scale Spatio-Temporal Attention Network (MSSTAN) with a novel dual-branch architecture: (1) A Global-Local Feature Attention Network (GLFAN) that explicitly decouples spatial interactions across decomposed temporal components to capture multi-scale spatial patterns; and (2) A Spatio-Temporal Feature Attention Network (STFAN) that dynamically recalibrates feature importance based on specific spatio-temporal contexts. A dynamic branch fusion mechanism integrates these branches to optimally aggregate their complementary views. Extensive experiments on five real-world datasets demonstrate that MSSTAN achieves state-of-the-art or highly competitive performance, validating its efficacy for traffic forecasting.
Balancing target-specific biological affinity with drug-likeness remains a central challenge in de novo molecular design, where existing generative models often exhibit limited controllability or reduced structural diversity. Here, we present DF-S4, a conditional molecular generation framework based on Structured State Space Models (S4), which addresses this limitation through a disentangled latent representation and hierarchical feature-wise linear modulation (FiLM). By decoupling structural and property variables and injecting conditional signals across multiple representation levels, DF-S4 enables fine-grained and stable multi-objective control beyond conventional input-level conditioning. Evaluated via a rigorous progressive multi-denominator auditing framework across three kinase targets (EGFR, BRAF, and FGFR1), DF-S4 exhibits robust target-steering performance, yielding favorable intradomain active ratios (74.2-83.9%) while maintaining high novelty (>95%) and competitive internal diversity (∼0.85). Furthermore, DF-S4 shifts the multi-objective Pareto Frontier toward regions of simultaneous high apparent affinity and drug-likeness, mitigating the distributional-trapping trade-offs commonly observed in prior approaches. Molecular docking analyses confirm the physical plausibility of generated candidates, yielding stronger binding affinities than reference inhibitors. Finally, comprehensive ablation studies and latent space dependence analyses mathematically validate that both explicit latent disentanglement and hierarchical FiLM modulation are critical for robust feature isolation, tighter property alignment, and generative stability.
Zero-shot image denoising has gained prominence in recent years, as it inherently relies on the intrinsic priors of images rather than learning from external data. Nevertheless, most existing methods either fail to fully exploit global priors, or do not properly preserve the fine-grained details governed by local priors. In this work, we propose a novel framework of pseudo sample generation for zero-shot denoising guided by local and global image priors. Specifically, we propose a well-crafted down-sampler based on gradient merging and grouping within a small window to generate down-sampled samples by exploiting spatial locality. Meanwhile, a global random sampler conditioned on a Gaussian distribution is devised to incorporate the nonlocal self-similarity of natural images. These two samplers build a new paradigm of pseudo sample generation powered by both local and global priors, which is termed as Zero-Shot Hybrid Prior-guided Denoising (ZS-HPD). Considering that noise is more likely to affect high-frequency details, we also present a simple yet effective loss that works in the Fourier domain and applies discriminative weights to distinct spectral bands. Numerous experiments on benchmark datasets have demonstrated the superiority of our ZS-HPD over existing advanced methods.
Diffusion models have recently emerged as a promising paradigm for time-series anomaly detection (TSAD). However, their effectiveness is fundamentally limited by unreliable conditional guidance under anomaly-contaminated and non-stationary inputs, which are common in real-world scenarios. To address this issue, we propose ReCDiff, a robust conditional encoding guided diffusion model for TSAD. The core of ReCDiff is a pre-trained robust conditional encoder (PRCE), which is optimized with simulated diffusion noise and pseudo-anomalies to learn stable and semantically consistent conditional priors of normal patterns. Once pre-trained, PRCE is frozen during downstream diffusion training and inference, providing robust conditional representations for sequence recovery. To make such guidance effective for anomaly removal, we further design an anomaly-aware conditional diffusion training strategy that explicitly incorporates structured anomalies into the corruption space by predicting a total noise composed of Gaussian perturbations and anomalous patterns. In the reverse process, anomalous sequences are treated as corrupted initial states, and the PRCE-guided denoising procedure progressively restores them toward normal sequences. Anomalies are then detected according to the discrepancy between the observed and recovered sequences. Extensive experiments demonstrate that ReCDiff achieves state-of-the-art performance on four public TSAD benchmarks.
Time series forecasting has historically concentrated on numerical data processing, and recent multimodal approaches have predominantly been limited to simple numerical-to-textual conversions. However, in reality, time series data from different domains are characteristically accompanied by abundant contextual information that remains underutilized. This necessitates a universal time series forecasting framework. Existing methods are limited by two issues: first, direct signal-to-text conversion results in information loss across different domains; second, simple feature concatenation lacks the ability to learn sophisticated cross-modal interaction features. To address these challenges, this study proposes a novel dual-stage framework, called the LLM-Enhanced Cross-Modal Fusion Framework (LCMF). Notably, LCMF achieves universal temporal forecasting by integrating semantic reasoning with numerical modeling using multilevel cross-modal attention fusion networks. In the initial stage, LCMF performs structured semantic reasoning on contextual information using LLMs through parallel modality preprocessing branches while simultaneously preserving the original numerical characteristics by processing time series from different domains. The second stage introduces a multilevel cross-modal attention fusion network (MLCA-Net) to dynamically integrate multisource information through adaptive weight allocation and hierarchical cross-modal interactions. Extensive experiments on benchmark datasets spanning multiple domains demonstrate that the LCMF framework substantially outperforms state-of-the-art methods, thereby validating its generalizability and effectiveness across different temporal forecasting scenarios.
Audio-Visual Question Answering (AVQA) is an emerging task that requires integrating question-relevant cues from both video sequences and audio signals to provide answers. To perform intricate audio-visual spatio-temporal reasoning, AVQA is asked to effectively align audio-visual targets, model temporal dynamics, and localize question-relevant segments in long videos. To address these challenges, we propose an Adaptive Proposal on Spatio-Temporal Graph (APSTG) network, which leverages a spatio-temporal graph to capture the spatial interac tions and temporal dynamics between audio and visual elements, and introduces an adaptive temporal proposal method to select query-relevant segments for question answering. Specifically, an audio-visual spatial graph is constructed for each video segment to establish the spatial relationships between audio and visual. To further capture temporal dynamics, node consistency and saliency are integrated into the spatio-temporal graph to model the complex spatio-temporal interactions. Next, a threshold based adaptive temporal proposal method is introduced to dynamically select query-relevant segments for long videos. To further optimize the proposal process, two contrastive losses are introduced—positive temporal proposal loss and negative query contrastive loss, where the former guides the model to focus on the most informative video segments, while the latter ensures that the extracted cues meet the specific reasoning requirements of different question types. Extensive experiments and ablation studies demonstrate the effectiveness of the proposed method, which significantly improves AVQA performance.
As Machine Learning (ML) and Artificial Intelligence (AI) progress rapidly, the issue of ML model generalization has emerged as a critical concern for academics and practitioners alike. In practical scenarios, it is essential for models to sustain high performance when encountering varied and novel data distributions. Nevertheless, current domain generalization techniques have their shortcomings in tackling this challenge. The objective of this paper is to introduce a novel Meta-Learning approach, incorporating Fourier transform-based Data augmentation, called MLFD, for the purpose of domain generalization. Utilizing both data augmentation and a meta-learning architecture, this proposed technique empowers models to extend their generalization to multiple unseen target domains using just a single training domain. In contrast to other domain generalization methods, the method presented in this paper achieves comparable accuracy on the Digits-DG datasets, and demonstrates substantial improvements in terms of reducing model training time.
Accurate traffic flow prediction is essential for the advancement of Intelligent Transportation Systems. However, traffic data exhibit significant nonlinearity and intricate spatiotemporal dependencies, as correlations between road traffic sensors evolve dynamically over time. Traditional deep learning prediction approaches based on static graph are often insufficient for capturing these instantaneous and time-varying spatial relationships. To address this challenge, we propose a Time-Varying Graph Convolutional Recurrent Network (TGCRN). A time-varying graph generation mechanism is designed that utilizes a combination of a static base and dynamic offset to construct evolving adjacency matrices. Specifically, a static relational topology is learned via trainable node embeddings, while dynamic offsets are generated by jointly modeling current input features and historical hidden states. This approach enables the model to adaptively capture real-time traffic variations. Furthermore, a graph fusion mechanism integrates the learned dynamic graph with the physical road network in a weighted manner to enhance model robustness. Experimental results on two real-world traffic benchmarks, METR-LA and PEMS-BAY, demonstrate that TGCRN significantly improves long-term prediction performance.
With the popularity of social media and online news, the spread of fake news has become a serious problem. Therefore, automatically detecting fake news has been a key research direction. However, how to fully utilize the consistency and inconsistency information between images and texts in multimodal news, and fully fuse multimodal features to extract effective information are still open questions. To address these issues, this paper proposes a novel Cross-modal Consistency Learning Dynamic Multimodal Fusion Framework(CCLFND). The framework combines the fine-grained fusion network with features extracted by the CLIP encoder to explore cross-modal consistency. In addition, we introduce a dynamic multimodal fusion module to adaptively aggregate the features. Extensive experiments on typical fake news datasets show that CCLFND outperforms state-of-the-art methods.
Guang-Bin Huang合作论文数School of Electrical and Electronics Engineering, Nanyang Technological University;Mind PointEye Pte Ltd9