
Accurate precipitation forecasting is essential for disaster preparedness, water resource management, and climate modeling. Numerical Weather Prediction (NWP) models, while extensively utilized, suffer from inherent limitations such as systematic biases, coarse spatial resolution, and difficulty in accurately predicting extreme rainfall events. Deep learning-based post-processing has emerged as a promising approach to refine NWP outputs, but existing methods struggle with effectively capturing multi-scale meteorological features and addressing severe class imbalances in extreme precipitation predictions. To overcome these challenges, we propose Rain-PostEFNet, a multi-scale deep learning framework for post-processing NWP precipitation forecasts that captures meteorological features across scales and mitigates severe class imbalance in extreme precipitation forecasts. It integrates EfficientNet for feature extraction, ECA-Net for adaptive channel attention, and FPN-Swin-Unet for multi-scale feature refinement, along with a multi-task learning strategy that jointly optimizes classification and regression tasks using weighted focal loss and Mean Squared Error (MSE) loss to improve predictive skill for rare heavy rain events. We evaluate our approach on the PostRainBench datasets (China, Germany, Korea) using Critical Success Index (CSI), Heidke Skill Score (HSS), and Accuracy (ACC) as evaluation metrics. Experimental results demonstrate that Rain-PostEFNet significantly outperforms state-of-the-art baselines, achieving relative CSI improvements of 63.94
Multi-task recommendation systems typically rely on shared representations to improve data efficiency by jointly modeling multiple user feedback signals. However, under extreme label imbalance, such sharing can exacerbate negative transfer, where improving one task degrades another. Existing multi-task architectures, such as MMoE and Progressive Layered Extraction (PLE), attempt to mitigate task interference through expert-based routing, but they still rely on fully learned soft gating over shared and task-specific experts. Under severe imbalance, the gating networks are often dominated by high-frequency tasks, leading to unstable routing and insufficient representation learning for sparse tasks. To address this limitation, we propose H-PLE, a hierarchical multi-task recommendation framework built upon progressive layered extraction. H-PLE incorporates three complementary design elements: a hierarchical extraction structure that progressively separates shared and task-specific representations across stacked layers, a raw-input-aware gating mechanism (AGC) that injects the projected raw input as a residual candidate to stabilize routing under sparse supervision, and a lightweight expert-level interaction module that introduces FM-style second-order feature interactions to enhance inductive bias. The evaluation includes multi-seed runs, independent random splits, modern baselines, component ablations, calibration metrics, PR-AUC, GAUC, and threshold-optimized F1. Experiments on two real-world Tenrec video recommendation scenarios (QK and QB) with highly imbalanced engagement tasks (Like and Share) show that H-PLE-family variants consistently improve sparse-task ranking and minority-class retrieval over MMoE and PLE, while remaining competitive with or stronger than Cross-Stitch, PCGrad, and loss-rebalancing baselines (focal, class-balanced BCE), all of which underperform vanilla MMoE in our most imbalanced setting and therefore highlight the difficulty of the QB-Share scenario. The results also reveal an important limitation: the full H-PLE model is not uniformly best on every QB metric; AGC is the most robust component under the most extreme sparsity, while FM is beneficial when its interaction rank is carefully controlled.
Fine-grained visual classification (FGVC) aims to distinguish highly similar subcategories within the same coarse-grained category and is important in biodiversity monitoring, intelligent transportation, and industrial inspection. However, its performance is often degraded by subtle inter-class differences, large intra-class variations, scale changes, occlusion, and cluttered backgrounds. These factors weaken discriminative responses and easily induce shortcut learning from background context, thereby limiting generalization. To address this problem, a dual-view multi-scale feature enhancement and fusion network is proposed for accurate and efficient FGVC. Specifically, a lightweight multi-scale feature enhancement module is designed to strengthen discriminative cues in features through multi-branch modeling and channel attention. A dynamic feature fusion module is then designed to align and adaptively integrate enhanced multi-level features in both spatial and channel dimensions, producing more informative representations. In addition, a dual-branch self-distillation framework is constructed, in which the raw-image branch and the enhanced-view branch share the same architecture. Combined with open-vocabulary text-guided foreground cropping and an augmented-view transfer module, the soft predictions of the enhanced branch are used to guide the raw branch toward foreground-oriented decision boundaries. During inference, only a single branch is retained, so improved performance is achieved without additional inference overhead. Extensive experiments on multiple public FGVC benchmarks demonstrate the effectiveness, robustness, and generality of the proposed method.
Large language models (LLMs) give graph neural networks (GNNs) access to semantic knowledge, instruction following, and natural-language reasoning. The same integration creates new failure channels, including unsupported graph edits, prompt injection, privacy exposure, and inherited social bias. We review trustworthy LLM–GNN systems using a taxonomy that separates five trust dimensions from the operational role of the LLM. The trust dimensions are reliability, robustness, privacy, fairness, and reasoning and explainability. The operational roles are encoder, predictor, aligner, editor, and verifier or evaluator. This separation allows a method to have a primary trust objective while retaining relevant secondary effects. We use it to connect integrated methods with classical trustworthy-GNN baselines, heterophilic and hypergraph learning, few-shot settings, instruction tuning, and graph-reasoning benchmarks. The comparison focuses on threats, safeguards, evaluation settings, and reported evidence, while identifying results that cannot be compared directly. The review shows where language information helps graph learning, where it creates additional risk, and which safeguards are needed across the system lifecycle.
This paper examines multivariate realized-volatility forecasting in globally connected equity markets. Heterogeneous AutoRegressive (HAR) models and graph signal processing (GSP)-based extensions such as the graph signal processing HAR model (GSPHAR) embed spillover structure via the magnetic Laplacian and graph Fourier transform. However, they still summarize each market’s recent history mainly through multi-scale lag averages, which can miss nonlinear within-window path information. To address this limitation, we develop a signature-enhanced graph spectral HAR framework, Sig-GSPHAR, which incorporates path-signature features extracted from each market’s 22-day volatility window and fuses them with HAR lags in the directed graph spectral domain. An empirical study on 29 stock markets and five realized-volatility proxies shows that, for one-step-ahead forecasting ( h = 1 ), Sig-GSPHAR outperforms econometric and deep learning benchmarks and delivers consistent gains over GSPHAR, with paired one-sided Wilcoxon signed-rank tests (Holm-adjusted) indicating significance across multiple proxies. For long-horizon forecasting ( h = 22 ), direct comparisons against GSPHAR indicate that the advantage of Sig-GSPHAR becomes more pronounced. We further show that the static spillover graph is structurally stable over the sample, equip the point forecasts with distribution-free conformal prediction intervals, and provide a detailed diagnostic of when signature features help.
Micro-expression recognition (MER) is a challenging task due to the transient nature of facial movements and their susceptibility to interference from rigid head motion and high-frequency sensor noise. Current approaches, which often rely on raw optical flow or purely data-driven architectures, struggle to decouple subtle muscle deformations from this global noise. To address these issues, we propose a Spectral-Spatial-Strain Network (S3Net), a unified physics-aware framework. First, we introduce an early fusion strategy utilizing Dual TV-L1 optical flow and optical strain. By calculating the strain magnitude, we provide the network with a physical prior that helps emphasize non-rigid muscle movements and reduce the influence of head rotations. Second, we design a Spectral Gating Block (SGB) integrated into the ResNet backbone, which leverages the frequency domain to filter sensor noise. Third, a Strain Excitation Block (SEB) is used to recalibrate high-level features in the strain-fused feature stream. Extensive experiments under the official Micro-Expression Grand Challenge 2019 (MEGC2019) 442-sample composite protocol (SAMM, CASME II, and SMIC) demonstrate the effectiveness of our method. S3Net achieves an Unweighted F1-score (UF1) of 85.38
Precise traffic flow forecasting is crucial for the development of intelligent transportation systems, enabling proactive traffic management, congestion alleviation, and efficient resource allocation. Traditional predictive models frequently fail to accurately capture the intricate nonlinear dynamics and multiscale temporal correlations present in traffic data. Multiscale Temporal- Spatial Diffusion Model informed by Wavelet Transform (MTSDM/WT) is proposed to overcome these constraints. The MTSDM/WT architecture begins with wavelet decomposition to decompose traffic time series into high-frequency components (capturing short-term fluctuations such as sudden congestion) and low-frequency components (reflecting long-term trends like rush-hour patterns). Specifically, the High-Frequency Refinement Module (HFRM) leverages multiscale convolution and attention mechanisms to precisely model localized temporal variations in the high-frequency domain, enhancing the capture of transient dynamics. Meanwhile, the low-frequency components undergo a forward diffusion process that incrementally adds noise, followed by a reverse denoising diffusion process to learn and reconstruct their temporal distribution, enabling accurate modeling of long-term trends. Finally, the refined high-frequency components and reconstructed low-frequency components are fused via inverse wavelet transform to generate comprehensive and robust traffic flow predictions. Extensive experiments were conducted on real-world traffic flow datasets from key urban intersections in Linyi City, Shandong Province, China. The MTSDM/WT framework exhibits notable improvements across critical evaluation metrics-Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), Root Mean Squared Error (RMSE), and Continuous Ranked Probability Score (CRPS)-with a 3
Electric vehicle drive system (EVDS) operates under complex multi-physical field coupling load conditions, posing significant challenges to modeling its performance degradation and assessing its remaining useful life (RUL). Accurate RUL prediction is essential for implementing predictive maintenance and enhancing driving safety. This study proposes an innovative physics-informed machine learning framework for RUL prediction and uncertainty quantification in EVDS. Initially, multidimensional time-domain and frequency-domain features are constructed based on actual operational load data, while multiscale cumulative damage features are developed by integrating the failure mechanisms of core components, followed by dimensionality reduction using the sparse autoencoder models. Subsequently, a RUL prediction model is developed by integrating Bayesian optimization with the convolutional neural network and bidirectional long short-term memory architecture, enhanced by an attention mechanism. To capture realistic degradation patterns of EVDS, accelerated durability bench tests are conducted, with the root mean square of vibration signals used to characterize degradation trajectories. By accounting for differences in degradation rates between user data and the accelerated test spectrum, nonlinear degradation trajectories for various users are generated. Finally, model training and parameter optimization are performed to predict RUL. The results demonstrate that RUL prediction errors remain within 5
In recent years, outlier-detection approaches that integrate density estimation with clustering have received increasing attention, among which methods based on density-peak clustering (DPC) are the most prominent. Nevertheless, conventional DPC neglects local structural information when estimating density and still requires manual parameter tuning. Moreover, existing DPC-based outlier detectors lack an effective relative-distance mechanism for reliably discriminating normal instances from outliers. To address these limitations, this paper proposes an outlier detection algorithm based on affinity density peak clustering (ADPOD). First, an affinity density measure is defined based on both the k-nearest neighbor intersection size between samples and their corresponding distances. Second, cluster centers are automatically determined using the product of affinity density and decision distance. Third, a cluster relative distance is employed as the distance metric for data objects. Finally, a novel outlier factor is devised to quantify the outlier degree. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed algorithm significantly outperforms state-of-the-art competitors.
Diffusion has recently attracted attention in time series imputation for its potential to model the uncertainty that deterministic methods often fail to capture. However, diffusion-based approaches often exhibit suboptimal imputation performance, as they overlook the distribution shift introduced by padding and rely on noise-driven optimization that ignores direct supervision signals from observed values. This causes the models to train on biased samples and struggle to approximate the observations, ultimately leading to inconsistent imputations. To address these limitations, we propose the Implicit Trajectory-Constrained Diffusion network (ITCD), which employs a two-stage diffusion architecture to mitigate the distribution shift caused by padding, thereby enabling better adaptation to realistic data. It also implicitly guides the diffusion sampling along a coherent denoising trajectory through intermediate and terminal constraints that consider both distribution and numerical accuracy, thereby achieving more consistent imputations. Extensive experiments illustrate that ITCD outperforms baselines across various scenarios with an average improvement of 13
Accurately modeling learners’ evolving knowledge states is a long-standing challenge in intelligent education systems, especially under fine-grained temporal dynamics and inherently noisy learning interactions. Most existing knowledge tracing methods rely on deterministic graph-based or sequential modeling, which often struggle to disentangle genuine knowledge mastery from interaction noise such as guessing and careless errors. To address these challenges, this paper proposes DiffKT, a diffusion-based framework for fine-grained knowledge tracing that models learners’ knowledge states as probabilistic distributions rather than fixed point estimates. DiffKT integrates a dual-graph representation to capture student–question interactions and question–skill associations, together with a state-space sequence model that efficiently encodes long-range learning dependencies with linear complexity. Building upon these representations, a conditional diffusion model with an adaptive noise scheduling strategy is introduced to explicitly distinguish different types of interaction noise, enabling robust denoising and more accurate estimation of latent knowledge states. Extensive experiments on three real-world educational datasets demonstrate that DiffKT consistently outperforms advanced knowledge tracking methods in terms of prediction accuracy and stability, highlighting its effectiveness in modeling noisy and fine-grained learning behaviors.
Data-driven fault diagnosis methods have demonstrated remarkable success by leveraging neural networks to automatically learn discriminative features from raw data. However, existing approaches often exhibit limited feature extraction capabilities, leading to suboptimal performance. To address this limitation, this paper proposes a novel fault diagnosis framework based on recurrent attentional reinforcement learning, which is designed to adaptively search optimal temporal features and identify fault types. The framework operates in three stages: firstly, a one-dimensional convolutional neural network (1D-CNN) extracts preliminary local temporal features. Secondly, a recurrent attentional module (RAM) iteratively localizes and samples the most informative temporal fragments. Finally, the attended fragments are fused to predict the health state of the machinery. A key advantage of the framework is its ability to search for informative fragments and explicitly model long-range dependencies across the selected temporal fragments, effectively capturing historical fault patterns. A reinforcement learning paradigm is introduced to optimize the entire feature extraction process in an end-to-end manner. The proposed method is evaluated on two mechanical systems, a rolling-bearing system and a hydraulic system, achieving average accuracies of 76.4
Power transformers are critical components in electrical grids, and their operational reliability directly affects grid stability and power quality. Dissolved Gas Analysis (DGA) is a widely adopted technique for transformer fault diagnosis by analyzing decomposition gases dissolved in insulating oil. Existing data-driven DGA methods typically treat fault categories independently, neglecting the inherent hierarchical relationships among different fault types and severity levels, which limits model generalization capability. To address this issue, we propose MeFD, a multi-grained power transformer fault diagnosis framework based on enhanced dissolved gas features. Multi-grained learning refers to learning representations at multiple semantic granularity levels simultaneously, enabling the model to exploit hierarchical relationships among labels. Specifically, MeFD organizes transformer faults into a two-level hierarchy consisting of coarse-grained fault categories (normal, overheating, and discharging) and fine-grained severity levels. The diagnosis task is accordingly decomposed into coarse-grained fault classification and fine-grained severity regression. Furthermore, two discriminative feature enhancement strategies are introduced: (1) relative concentration ratios to characterize inter-gas relationships for fault classification, and (2) gas concentration deviations to quantify abnormality magnitude for severity regression. Experimental results on real-world datasets demonstrate that MeFD consistently improves the performance of multiple baseline machine learning algorithms and achieves superior accuracy in both fault recognition and severity regression.
Variational autoencoders (VAE) construct latent space by optimizing the prior distribution and posterior distribution of the model. Existing methods exhibit limited interpretability during the construction of the latent space, which hinders their capacity to effectively capture the disentangled representation of concepts. To construct an interpretable latent space, we propose the Multi-Decoder Concept Embedding Variational Autoencoder (MD-VAE), which enhances latent space interpretability by learning distinct latent variables through multiple decoders. Firstly, the MD-VAE model learns prior concept by training on generated data that represent this concept, thereby embedding the prior into the latent space. Subsequently, we propose a variational inference framework utilizing multiple decoders. In this framework, encoders map multiple latent variables into the latent space, and each corresponding set of latent variables is reconstructed by its dedicated decoder. On the basis of this, a theoretical derivation of variational lower bound of multiple decodes is combined with variation method to obtain the optimal model parameter estimates. Finally, experiments on MNIST, FashionMNIST, COIL20, and USPS datasets show that MD-VAE can improve the prediction performance of VAE while discovering differences between different concepts.
The manufacturing of complex products, prevalent in sectors such as electronics and automotive assembly, involves heterogeneous units which are connected in series and operate with interdependent scheduling decisions. This shift toward highly integrated production frameworks has rendered traditional single flowshop scheduling models inadequate. These models fail to meet the demands for global optimization across interconnected stages, making the coordinated scheduling of such multi-workshop systems a critical yet under-addressed challenge. To address this gap, we study the cascaded flowshop joint scheduling problem (CFJSP), which integrates a distributed permutation flowshop with a hybrid flowshop. Our proposed Adaptive Population-Based Iterated Greedy (APIG) algorithm begins with a collaborative initialization mechanism that blends diverse solution generation strategies. During the construction phase, an experience-driven skipping mechanism learns to evaluate operators adaptively. It intelligently prioritizes high-performance operations with a probabilistic set, effectively directing computational budget towards regions with higher payoff potential. To exploit neighborhood complementarity, the local search phase employs a hybrid strategy that alternates between insertion and swap operations. The efficacy of APIG is computationally confirmed by a 46.40
Object detection has been widely adopted in applications such as autonomous driving, security surveillance, industrial inspection, and UAV-based monitoring. However, the computational, memory, and power costs of modern detectors remain major bottlenecks for real-time deployment. This survey focuses on how complexity reductions induced by model pruning can be reliably translated into end-to-end latency gains in practical deployments. We provide a systematic review of deployment-oriented pruning for object detection and establish a unified comparative framework across three layers—algorithm, system, and evaluation. At the algorithm level, we summarize pruning granularity, importance criteria, pruning workflows, and accuracy recovery strategies. At the system level, we analyze static-graph export, operator coverage and fusion, tensor-shape constraints, and their alignment with hardware parallelism granularity. At the evaluation level, we consolidate reproducible reporting elements, including timing scope, inference configurations, and percentile latency metrics. By examining the end-to-end deployment pipeline, we highlight that reductions in complexity proxies such as FLOPs and parameter count do not necessarily lead to proportional decreases in inference time. Instead, system-level performance is often dominated by the optimization capability of inference engines, the regularity of tensor shapes after pruning, and runtime overhead introduced by dynamic shapes. Finally, we discuss open limitations in existing methods, including budget allocation mechanisms, hardware/software support for sparse acceleration, and worst-case latency constraints for dynamic strategies, and we outline future directions toward toolchain-aware pruning, reproducible deployment evaluation, and hardware-aware closed-loop optimization.
Predictive Maintenance (PdM) remains an important component in maintaining the long-term reliability and performance of industrial systems, especially in Internet of Things (IoT) environments. Existing PdM methods often fail when dealing with high-dimensional, noisy, and heterogeneous sensor data due to poor feature selection and the inability to model multi-scale temporal dependencies. Thus, the Predictive Maintenance Fault Network (PdM-FaultNet) is designed to enhance fault prediction accuracy. PdM-FaultNet is the combination of the Enhanced Wombat Optimization Algorithm (EWOA) for the feature selection and Dual Quantum-inspired Denoising Autoencoder Transformer (DQDAT) for the predictive modeling. The EWOA enhances the search diversity, convergence speed, and identifies informative feature subsets. The DQDAT employs a dual-path decomposition strategy to capture long-term and short-term temporal patterns. It integrates Quantum-inspired Long Short-Term Memory (LSTM) into a denoising autoencoder to derive nonlinear representation and a transformer to derive long-range dependencies. It attained a precision of 95.12
Extreme multi-label classification (XMLC) aims to assign the most relevant labels from an extremely large label space to each document. Existing graph-based XMLC methods model document-label interactions, but their graph construction often relies on random sampling, which may introduce noisy or weakly relevant neighbors. To address this issue, we propose SIAGXC (Code is publicly available at https://github.com/bjtu-yzh/siagxc .), a prediction-augmented graph framework for XMLC that leverages auxiliary relational signals derived from an upstream XMLC model. Specifically, upstream predictions are transformed into an auxiliary document-label graph, which is jointly encoded with the original document-label graph through a dual-branch graph encoder and an adaptive fusion module. This design enriches document and label representations and improves candidate label ranking. During inference, the prediction-derived auxiliary graph is further used to support graph-based refinement for unseen documents. Experiments on three public XMLC datasets show that SIAGXC consistently outperforms strong baselines, including graph-based and neural XMLC methods. These results demonstrate that prediction-derived auxiliary relations provide an effective way to enhance graph-based XMLC.
Multimodal recommendation models leverage item content and user–item interaction data to predict user preferences. However, observed interactions may reflect not only users’ intrinsic interest in multimodal content but also conformity-driven behavior influenced by popularity and social feedback, leading to the interest–conformity confusion problem. To address this issue, this paper proposes Dual-Channel Graph Embedding for Multimodal Recommendation with a Causal Perspective (DCGCE), a representation learning framework for multimodal recommendation inspired by a causal perspective on user behavior. The model introduces two complementary embedding channels to capture intrinsic interest and conformity-related preference signals, respectively. To mitigate semantic noise, multimodal deep features and semantic entities are extracted from item content and incorporated into graph-based representations. The framework further models collaborative signals, semantic-level preferences, and disentangled multimodal representations, which are jointly integrated for final prediction under a unified optimization objective. Extensive experiments on three real-world datasets demonstrate that DCGCE consistently outperforms state-of-the-art baselines across multiple evaluation metrics, while providing improved interpretability in user preference modeling.
Many cyber-threat knowledge databases possess complementary information. For example, the popular Common Vulnerabilities and Exposures (CVE) repository catalogs known software weaknesses. Meanwhile the MITRE ATT CK framework documents the tactics and techniques that adversaries use to exploit these vulnerabilities. Despite a natural correspondence, linking vulnerability descriptions with attacker techniques remains largely manual. This is due to differences in abstraction levels and linguistic representation. Existing automated approaches have made progress by relying on multi-stage pipelines or surface textual similarity. However, they often fail to capture the underlying attack events. Moreover, these approaches limit generalization for newly disclosed vulnerabilities. This paper introduces SemanticLink, an event-centric semantic representation framework that enables automated alignment between CVE vulnerability descriptions and ATT CK techniques through a single integrated learning pipeline. The approach uses semantic role labeling to extract structured attack events, capturing actions, targets, and exploitation context. These event representations are combined with a Siamese transformer architecture and cross-sentence attention to enable semantically grounded similarity learning. A contrastive learning strategy further improves discrimination among closely related attack techniques. We evaluate our approach on a dataset of 6,602 curated CVE-ATT CK mappings. SemanticLink achieves a mean accuracy of 93.43