
Micro-expression recognition (MER) is a challenging task due to the transient nature of facial movements and their susceptibility to interference from rigid head motion and high-frequency sensor noise. Current approaches, which often rely on raw optical flow or purely data-driven architectures, struggle to decouple subtle muscle deformations from this global noise. To address these issues, we propose a Spectral-Spatial-Strain Network (S3Net), a unified physics-aware framework. First, we introduce an early fusion strategy utilizing Dual TV-L1 optical flow and optical strain. By calculating the strain magnitude, we provide the network with a physical prior that helps emphasize non-rigid muscle movements and reduce the influence of head rotations. Second, we design a Spectral Gating Block (SGB) integrated into the ResNet backbone, which leverages the frequency domain to filter sensor noise. Third, a Strain Excitation Block (SEB) is used to recalibrate high-level features in the strain-fused feature stream. Extensive experiments under the official Micro-Expression Grand Challenge 2019 (MEGC2019) 442-sample composite protocol (SAMM, CASME II, and SMIC) demonstrate the effectiveness of our method. S3Net achieves an Unweighted F1-score (UF1) of 85.38
Precise traffic flow forecasting is crucial for the development of intelligent transportation systems, enabling proactive traffic management, congestion alleviation, and efficient resource allocation. Traditional predictive models frequently fail to accurately capture the intricate nonlinear dynamics and multiscale temporal correlations present in traffic data. Multiscale Temporal- Spatial Diffusion Model informed by Wavelet Transform (MTSDM/WT) is proposed to overcome these constraints. The MTSDM/WT architecture begins with wavelet decomposition to decompose traffic time series into high-frequency components (capturing short-term fluctuations such as sudden congestion) and low-frequency components (reflecting long-term trends like rush-hour patterns). Specifically, the High-Frequency Refinement Module (HFRM) leverages multiscale convolution and attention mechanisms to precisely model localized temporal variations in the high-frequency domain, enhancing the capture of transient dynamics. Meanwhile, the low-frequency components undergo a forward diffusion process that incrementally adds noise, followed by a reverse denoising diffusion process to learn and reconstruct their temporal distribution, enabling accurate modeling of long-term trends. Finally, the refined high-frequency components and reconstructed low-frequency components are fused via inverse wavelet transform to generate comprehensive and robust traffic flow predictions. Extensive experiments were conducted on real-world traffic flow datasets from key urban intersections in Linyi City, Shandong Province, China. The MTSDM/WT framework exhibits notable improvements across critical evaluation metrics-Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), Root Mean Squared Error (RMSE), and Continuous Ranked Probability Score (CRPS)-with a 3
Electric vehicle drive system (EVDS) operates under complex multi-physical field coupling load conditions, posing significant challenges to modeling its performance degradation and assessing its remaining useful life (RUL). Accurate RUL prediction is essential for implementing predictive maintenance and enhancing driving safety. This study proposes an innovative physics-informed machine learning framework for RUL prediction and uncertainty quantification in EVDS. Initially, multidimensional time-domain and frequency-domain features are constructed based on actual operational load data, while multiscale cumulative damage features are developed by integrating the failure mechanisms of core components, followed by dimensionality reduction using the sparse autoencoder models. Subsequently, a RUL prediction model is developed by integrating Bayesian optimization with the convolutional neural network and bidirectional long short-term memory architecture, enhanced by an attention mechanism. To capture realistic degradation patterns of EVDS, accelerated durability bench tests are conducted, with the root mean square of vibration signals used to characterize degradation trajectories. By accounting for differences in degradation rates between user data and the accelerated test spectrum, nonlinear degradation trajectories for various users are generated. Finally, model training and parameter optimization are performed to predict RUL. The results demonstrate that RUL prediction errors remain within 5
In recent years, outlier-detection approaches that integrate density estimation with clustering have received increasing attention, among which methods based on density-peak clustering (DPC) are the most prominent. Nevertheless, conventional DPC neglects local structural information when estimating density and still requires manual parameter tuning. Moreover, existing DPC-based outlier detectors lack an effective relative-distance mechanism for reliably discriminating normal instances from outliers. To address these limitations, this paper proposes an outlier detection algorithm based on affinity density peak clustering (ADPOD). First, an affinity density measure is defined based on both the k-nearest neighbor intersection size between samples and their corresponding distances. Second, cluster centers are automatically determined using the product of affinity density and decision distance. Third, a cluster relative distance is employed as the distance metric for data objects. Finally, a novel outlier factor is devised to quantify the outlier degree. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed algorithm significantly outperforms state-of-the-art competitors.
Diffusion has recently attracted attention in time series imputation for its potential to model the uncertainty that deterministic methods often fail to capture. However, diffusion-based approaches often exhibit suboptimal imputation performance, as they overlook the distribution shift introduced by padding and rely on noise-driven optimization that ignores direct supervision signals from observed values. This causes the models to train on biased samples and struggle to approximate the observations, ultimately leading to inconsistent imputations. To address these limitations, we propose the Implicit Trajectory-Constrained Diffusion network (ITCD), which employs a two-stage diffusion architecture to mitigate the distribution shift caused by padding, thereby enabling better adaptation to realistic data. It also implicitly guides the diffusion sampling along a coherent denoising trajectory through intermediate and terminal constraints that consider both distribution and numerical accuracy, thereby achieving more consistent imputations. Extensive experiments illustrate that ITCD outperforms baselines across various scenarios with an average improvement of 13
Accurately modeling learners’ evolving knowledge states is a long-standing challenge in intelligent education systems, especially under fine-grained temporal dynamics and inherently noisy learning interactions. Most existing knowledge tracing methods rely on deterministic graph-based or sequential modeling, which often struggle to disentangle genuine knowledge mastery from interaction noise such as guessing and careless errors. To address these challenges, this paper proposes DiffKT, a diffusion-based framework for fine-grained knowledge tracing that models learners’ knowledge states as probabilistic distributions rather than fixed point estimates. DiffKT integrates a dual-graph representation to capture student–question interactions and question–skill associations, together with a state-space sequence model that efficiently encodes long-range learning dependencies with linear complexity. Building upon these representations, a conditional diffusion model with an adaptive noise scheduling strategy is introduced to explicitly distinguish different types of interaction noise, enabling robust denoising and more accurate estimation of latent knowledge states. Extensive experiments on three real-world educational datasets demonstrate that DiffKT consistently outperforms advanced knowledge tracking methods in terms of prediction accuracy and stability, highlighting its effectiveness in modeling noisy and fine-grained learning behaviors.
Data-driven fault diagnosis methods have demonstrated remarkable success by leveraging neural networks to automatically learn discriminative features from raw data. However, existing approaches often exhibit limited feature extraction capabilities, leading to suboptimal performance. To address this limitation, this paper proposes a novel fault diagnosis framework based on recurrent attentional reinforcement learning, which is designed to adaptively search optimal temporal features and identify fault types. The framework operates in three stages: firstly, a one-dimensional convolutional neural network (1D-CNN) extracts preliminary local temporal features. Secondly, a recurrent attentional module (RAM) iteratively localizes and samples the most informative temporal fragments. Finally, the attended fragments are fused to predict the health state of the machinery. A key advantage of the framework is its ability to search for informative fragments and explicitly model long-range dependencies across the selected temporal fragments, effectively capturing historical fault patterns. A reinforcement learning paradigm is introduced to optimize the entire feature extraction process in an end-to-end manner. The proposed method is evaluated on two mechanical systems, a rolling-bearing system and a hydraulic system, achieving average accuracies of 76.4
Power transformers are critical components in electrical grids, and their operational reliability directly affects grid stability and power quality. Dissolved Gas Analysis (DGA) is a widely adopted technique for transformer fault diagnosis by analyzing decomposition gases dissolved in insulating oil. Existing data-driven DGA methods typically treat fault categories independently, neglecting the inherent hierarchical relationships among different fault types and severity levels, which limits model generalization capability. To address this issue, we propose MeFD, a multi-grained power transformer fault diagnosis framework based on enhanced dissolved gas features. Multi-grained learning refers to learning representations at multiple semantic granularity levels simultaneously, enabling the model to exploit hierarchical relationships among labels. Specifically, MeFD organizes transformer faults into a two-level hierarchy consisting of coarse-grained fault categories (normal, overheating, and discharging) and fine-grained severity levels. The diagnosis task is accordingly decomposed into coarse-grained fault classification and fine-grained severity regression. Furthermore, two discriminative feature enhancement strategies are introduced: (1) relative concentration ratios to characterize inter-gas relationships for fault classification, and (2) gas concentration deviations to quantify abnormality magnitude for severity regression. Experimental results on real-world datasets demonstrate that MeFD consistently improves the performance of multiple baseline machine learning algorithms and achieves superior accuracy in both fault recognition and severity regression.
Variational autoencoders (VAE) construct latent space by optimizing the prior distribution and posterior distribution of the model. Existing methods exhibit limited interpretability during the construction of the latent space, which hinders their capacity to effectively capture the disentangled representation of concepts. To construct an interpretable latent space, we propose the Multi-Decoder Concept Embedding Variational Autoencoder (MD-VAE), which enhances latent space interpretability by learning distinct latent variables through multiple decoders. Firstly, the MD-VAE model learns prior concept by training on generated data that represent this concept, thereby embedding the prior into the latent space. Subsequently, we propose a variational inference framework utilizing multiple decoders. In this framework, encoders map multiple latent variables into the latent space, and each corresponding set of latent variables is reconstructed by its dedicated decoder. On the basis of this, a theoretical derivation of variational lower bound of multiple decodes is combined with variation method to obtain the optimal model parameter estimates. Finally, experiments on MNIST, FashionMNIST, COIL20, and USPS datasets show that MD-VAE can improve the prediction performance of VAE while discovering differences between different concepts.
The manufacturing of complex products, prevalent in sectors such as electronics and automotive assembly, involves heterogeneous units which are connected in series and operate with interdependent scheduling decisions. This shift toward highly integrated production frameworks has rendered traditional single flowshop scheduling models inadequate. These models fail to meet the demands for global optimization across interconnected stages, making the coordinated scheduling of such multi-workshop systems a critical yet under-addressed challenge. To address this gap, we study the cascaded flowshop joint scheduling problem (CFJSP), which integrates a distributed permutation flowshop with a hybrid flowshop. Our proposed Adaptive Population-Based Iterated Greedy (APIG) algorithm begins with a collaborative initialization mechanism that blends diverse solution generation strategies. During the construction phase, an experience-driven skipping mechanism learns to evaluate operators adaptively. It intelligently prioritizes high-performance operations with a probabilistic set, effectively directing computational budget towards regions with higher payoff potential. To exploit neighborhood complementarity, the local search phase employs a hybrid strategy that alternates between insertion and swap operations. The efficacy of APIG is computationally confirmed by a 46.40
Object detection has been widely adopted in applications such as autonomous driving, security surveillance, industrial inspection, and UAV-based monitoring. However, the computational, memory, and power costs of modern detectors remain major bottlenecks for real-time deployment. This survey focuses on how complexity reductions induced by model pruning can be reliably translated into end-to-end latency gains in practical deployments. We provide a systematic review of deployment-oriented pruning for object detection and establish a unified comparative framework across three layers—algorithm, system, and evaluation. At the algorithm level, we summarize pruning granularity, importance criteria, pruning workflows, and accuracy recovery strategies. At the system level, we analyze static-graph export, operator coverage and fusion, tensor-shape constraints, and their alignment with hardware parallelism granularity. At the evaluation level, we consolidate reproducible reporting elements, including timing scope, inference configurations, and percentile latency metrics. By examining the end-to-end deployment pipeline, we highlight that reductions in complexity proxies such as FLOPs and parameter count do not necessarily lead to proportional decreases in inference time. Instead, system-level performance is often dominated by the optimization capability of inference engines, the regularity of tensor shapes after pruning, and runtime overhead introduced by dynamic shapes. Finally, we discuss open limitations in existing methods, including budget allocation mechanisms, hardware/software support for sparse acceleration, and worst-case latency constraints for dynamic strategies, and we outline future directions toward toolchain-aware pruning, reproducible deployment evaluation, and hardware-aware closed-loop optimization.
Predictive Maintenance (PdM) remains an important component in maintaining the long-term reliability and performance of industrial systems, especially in Internet of Things (IoT) environments. Existing PdM methods often fail when dealing with high-dimensional, noisy, and heterogeneous sensor data due to poor feature selection and the inability to model multi-scale temporal dependencies. Thus, the Predictive Maintenance Fault Network (PdM-FaultNet) is designed to enhance fault prediction accuracy. PdM-FaultNet is the combination of the Enhanced Wombat Optimization Algorithm (EWOA) for the feature selection and Dual Quantum-inspired Denoising Autoencoder Transformer (DQDAT) for the predictive modeling. The EWOA enhances the search diversity, convergence speed, and identifies informative feature subsets. The DQDAT employs a dual-path decomposition strategy to capture long-term and short-term temporal patterns. It integrates Quantum-inspired Long Short-Term Memory (LSTM) into a denoising autoencoder to derive nonlinear representation and a transformer to derive long-range dependencies. It attained a precision of 95.12
Extreme multi-label classification (XMLC) aims to assign the most relevant labels from an extremely large label space to each document. Existing graph-based XMLC methods model document-label interactions, but their graph construction often relies on random sampling, which may introduce noisy or weakly relevant neighbors. To address this issue, we propose SIAGXC (Code is publicly available at https://github.com/bjtu-yzh/siagxc .), a prediction-augmented graph framework for XMLC that leverages auxiliary relational signals derived from an upstream XMLC model. Specifically, upstream predictions are transformed into an auxiliary document-label graph, which is jointly encoded with the original document-label graph through a dual-branch graph encoder and an adaptive fusion module. This design enriches document and label representations and improves candidate label ranking. During inference, the prediction-derived auxiliary graph is further used to support graph-based refinement for unseen documents. Experiments on three public XMLC datasets show that SIAGXC consistently outperforms strong baselines, including graph-based and neural XMLC methods. These results demonstrate that prediction-derived auxiliary relations provide an effective way to enhance graph-based XMLC.
Multimodal recommendation models leverage item content and user–item interaction data to predict user preferences. However, observed interactions may reflect not only users’ intrinsic interest in multimodal content but also conformity-driven behavior influenced by popularity and social feedback, leading to the interest–conformity confusion problem. To address this issue, this paper proposes Dual-Channel Graph Embedding for Multimodal Recommendation with a Causal Perspective (DCGCE), a representation learning framework for multimodal recommendation inspired by a causal perspective on user behavior. The model introduces two complementary embedding channels to capture intrinsic interest and conformity-related preference signals, respectively. To mitigate semantic noise, multimodal deep features and semantic entities are extracted from item content and incorporated into graph-based representations. The framework further models collaborative signals, semantic-level preferences, and disentangled multimodal representations, which are jointly integrated for final prediction under a unified optimization objective. Extensive experiments on three real-world datasets demonstrate that DCGCE consistently outperforms state-of-the-art baselines across multiple evaluation metrics, while providing improved interpretability in user preference modeling.
Many cyber-threat knowledge databases possess complementary information. For example, the popular Common Vulnerabilities and Exposures (CVE) repository catalogs known software weaknesses. Meanwhile the MITRE ATT CK framework documents the tactics and techniques that adversaries use to exploit these vulnerabilities. Despite a natural correspondence, linking vulnerability descriptions with attacker techniques remains largely manual. This is due to differences in abstraction levels and linguistic representation. Existing automated approaches have made progress by relying on multi-stage pipelines or surface textual similarity. However, they often fail to capture the underlying attack events. Moreover, these approaches limit generalization for newly disclosed vulnerabilities. This paper introduces SemanticLink, an event-centric semantic representation framework that enables automated alignment between CVE vulnerability descriptions and ATT CK techniques through a single integrated learning pipeline. The approach uses semantic role labeling to extract structured attack events, capturing actions, targets, and exploitation context. These event representations are combined with a Siamese transformer architecture and cross-sentence attention to enable semantically grounded similarity learning. A contrastive learning strategy further improves discrimination among closely related attack techniques. We evaluate our approach on a dataset of 6,602 curated CVE-ATT CK mappings. SemanticLink achieves a mean accuracy of 93.43
Temporal Knowledge Graph Question Answering (TKGQA) aims to retrieve entities or timestamps from temporal knowledge graphs in response to temporally constrained queries; however, existing methods primarily rely on single-granularity temporal representations and lack the ability to effectively handle multi-granularity temporal reasoning, leading to suboptimal performance on complex queries. To address this issue, we propose a Multi-Granularity Implicit Temporal (MGIT) framework that enhances temporal representation and reasoning by modeling implicit temporal dependencies across different granularities. Specifically, MGIT integrates a graph attention network with multi-head attention to capture structural and implicit temporal relationships, while a combination of convolutional neural networks and gated recurrent units is employed to jointly model local temporal patterns and long-range dependencies. In addition, a self-learning encoding strategy is introduced to dynamically adapt to heterogeneous temporal granularities, and an ensemble learning mechanism is adopted to aggregate representations from different temporal segments for improved robustness. Extensive experiments on both MultiTQ dataset and CronQuestions dataset demonstrate that MGIT consistently outperforms state-of-the-art baselines, highlighting its effectiveness in capturing implicit temporal information and enhancing multi-granularity temporal reasoning.
In drug combination therapy, Drug-Drug interaction (DDI) prediction can effectively improve efficacy or reduce adverse reactions. However, existing artificial intelligence (AI) approaches have limitations in fully considering and effectively extracting rich drug features. To overcome the limitations of single-view feature extraction, this study presents a novel multiview feature fusion-based graph representation model (MFF-GRM) for predicting DDI. This architecture integrates drug molecular graphs, SMILES sequences, DDI information networks, and drug biological features (target, enzyme, transport) to learn drug features more comprehensively. First, we use a Graph Convolutional Network (GCN) to update atomic features and extract drug molecular graph features. Secondly, we propose a cascade graph representation model combined with a multi-head self-attention mechanism and graph convolution to capture the topological features of the DDI network. This design allows the model to capture the heterogeneity of neighborhood interactions and further fuse and constrain the representation through graph structure propagation. After encoding the SMILES (Simplified Molecular Input Line Entry System) sequences, SMILES representations are obtained using the Gated Recurrent Unit (GRU) model. Furthermore, we employ a hybrid similarity calculation strategy to reconstruct varying biological feature matrices for drug representations. Finally, a multi-view feature fusion and prediction module is proposed to fuse multiple drug representations for DDI prediction. We conducted comparative experiments on the ZhangDDI and DrugBank datasets. Additionally, parameter sensitivity tests, ablation parameters, and imbalance data testing were performed. The experimental results demonstrate that MFF-GRM can effectively learn drug representations and outperform the baselines for DDI prediction, achieving improvements across all evaluation metrics, with gains of 0.07-11.3% and 0.92-17.9% on the two benchmark datasets. The MFF-GRM model shows great application potential in assisting drug repositioning and clinical medication decision-making.
Intelligent Decision Support Systems (IDSS) increasingly integrate complex advances in machine intelligence, often at the cost of transparency and intelligibility. This manuscript proposes a novel trustworthy approach for intelligent support in decision-making, that addresses these limitations through class-association rule induction. The proposed approach enables transparent decision pathways and delivers actionable, intelligible rationales at the decision level. This is accomplished through simple yet expressive if-then rules that capture and explain knowledge patterns within complex data, thereby enabling transparent, fully-interpretable, and explainable decision-making support across challenging application scenarios. A tailored instance of class-association rule induction is specifically designed to ensure transparency, interpretability, and explainability, while improving computational efficiency. The computational process incorporates advanced techniques for vertical mining and pruning. The former improves efficiency, while the latter ensures compactness, expressiveness, and classification effectiveness through a combination of strategies, including incremental rule covering, mutual exclusivity, and exhaustive coverage. Furthermore, a dynamic, class-specific adaptation of support and confidence thresholds reduces parameterization complexity to only one input parameter, strengthens robustness against skewed class distributions, and ensures positive correlations between rule antecedents and consequents. The above technical peculiarities inherently foster transparency by openly revealing the decision-making mechanics, bolster interpretability via explicit clarification of input-output relationships, and enhance explainability through concise and actionable rationales for decision support. Various applications of the proposed approach are considered as testbeds to carry out an extensive experimental validation on real-world benchmark datasets from different domains in comparison with established competitors. The empirical results demonstrate the effectiveness, robustness, and scalability of our approach, alongside its transparency, interpretability, and explainability. These findings highlight the potential of our approach to enhance the trustworthiness of IDSS, thereby fostering stakeholder trust and facilitating domain-specific insights that inform strategic decision making.
Automated radiology report generation aims to alleviate documentation burdens but faces a critical trade-off between the computational overhead of Transformers and the inability of efficient State Space Models (SSMs) to capture two-dimensional spatial relationships. This study proposes MG-Hybrid to harmonize computational efficiency with anatomical fidelity while bridging the semantic gap between visual features and clinical text. We introduce a Mask-Guided Visual Disentanglement strategy that extracts structured organ-specific prototypes to preserve the spatial integrity of regions such as the heart and lungs, rather than treating images as flattened sequences. These features are processed by a Hybrid Anatomy Encoder which integrates a Mamba-based state-space module for organ-level representation refinement with a Transformer layer for global inter-organ reasoning. Additionally, a Visual-Textual Semantic Alignment mechanism is incorporated to anchor visual representations to clinical semantics explicitly. The framework was evaluated on the IU X-ray and MIMIC-CXR benchmarks. Extensive experiments demonstrate that the proposed approach achieves superior performance in clinical efficacy and precision-oriented metrics. The results indicate that the model generates reports that are more consistent with the reference reports and achieves improved clinical efficacy metrics while maintaining robust generation quality compared to existing baselines. By harmonizing the computational advantages of State Space Models (SSMs) with advanced spatial and semantic modeling, MG-Hybrid addresses the limitations of current architectures. This approach provides a viable pathway toward more accurate and computationally efficient automated diagnostic support.
Positive and unlabeled learning (PU learning) addresses classification scenarios where only positive and unlabeled samples are available, the latter comprising both hidden positive and negative instances. While most existing PU methods focus on identifying reliable negative samples, they often underutilize the remaining unlabeled data and operate primarily within a single-view framework, limiting their expressive power. To overcome these limitations, this paper proposes SMVPU, a novel similarity-based multi-view PU learning method that effectively integrates multi-view learning principles. SMVPU first extracts reliable negative samples from the unlabeled set and assigns similarity-weighted values to the remaining unlabeled instances. It then incorporates multi-view representations to enhance feature compatibility and discriminability. By leveraging both consistency and complementarity principles across views, the method constructs a robust PU classifier formulated within a large-margin learning framework. The optimization problem is efficiently solved via the Lagrangian dual method. Extensive experiments demonstrate that SMVPU achieves superior performance compared to existing PU learning methods in terms of classification accuracy and stability.