
Severe noise in industrial environments often leads to distribution shifts, posing a critical challenge for deep learning-based bearing fault diagnosis. Models trained on clean data typically degrade under such conditions due to their inability to effectively disentangle noise from fault-related features. To address this issue, we propose a Physics-Inspired Heterogeneous Mixture-of-Experts (PI-HMoE) network for robust fault diagnosis. Specifically, a multi-scale shape embedding module is designed to align convolutional kernels with fault characteristics, capturing complementary temporal patterns at different receptive-field scales to facilitate the separation of fault-relevant and noise-dominated representations. Furthermore, a global linear temporal mixer captures long-range periodic dependencies with linear computational complexity. In addition, a heterogeneous expert module integrates diverse operators through cluster-based gating to enhance model robustness. Extensive experiments on two standard bearing benchmark datasets demonstrate the superior noise robustness of our approach. Under severe zero-decibel noise conditions, the proposed model achieves diagnostic accuracies of 94.40% and 85.00%, outperforming state-of-the-art baselines by over 15% and 31%, respectively.
Lipschitz regularization controls how strongly a denoising autoencoder propagates an input perturbation to its output, and has been used for photoplethysmography (PPG) denoising. We revisit the gradient-norm penalty that applies this constraint. The stability principle that motivates such regularization makes the constraint intrinsically an upper bound on a Lipschitz-related sensitivity of the reconstruction map, which implies a one-sided penalty. The two-sided squared penalty used in the previous formulation additionally penalizes already-contractive maps and therefore drives the measured sensitivity upward, away from the stability objective. We show, by analysis and by a controlled ablation on a PPG semi-simulation, that the two-sided penalty has an empirical instability mode that the one-sided penalty removes: under matched settings the one-sided penalty lowers reconstruction error (MSE) by about 23%, with no divergent runs. We further find that, in our experiments, this operating point depends on the penalty form rather than on how the Lipschitz target is parameterized. We recommend an anchored constant target with a one-sided penalty as a simple and stable operating point for this task.
Most deep learning–based sound source localization methods rely on specific microphone array geometries, which require retraining for different configurations and lead to high cost. IPDnet is our previous work which addresses this limitation through pair-wise processing with a mean pooling scheme. However, processing narrow-band independently leads to high computational complexity. In this work, we propose IPDnet2A, an efficient model which improves localization performance while reducing computational complexity. IPDnet2A adopts oSpatialNet as the backbone to extract spatial features, and a frequency pooling mechanism is used to compress the frequency dimension, thus reducing computational cost. Additionally, an interaction module is designed to process the mean pair-wise representation, which further improves interaction between microphone pairs. Experiments on multiple datasets demonstrate that IPDnet2A achieves state-of-the-art localization performance while significantly reducing computational cost compared to IPDnet.
Pillow-based ballistocardiography (BCG) enables unobtrusive cardiac monitoring, but J-peak detection is commonly learned as dense sequence labeling although the desired output is a sparse event set. This letter asks a narrower question: how do dense and query-set outputs differ when the data, encoder, validation protocol, and event evaluator are controlled? We formulate one-dimensional point-set prediction with 64 learned queries and Hungarian assignment, and compare it with a shared-backbone dense Transformer and a U-Net–BiLSTM. Evaluation uses five-subject leave-one-subject-out testing, three independently seeded runs, validation-only postprocessing selection, and strict one-to-one peak association. The dense Transformer attains the highest subject-wise pooled F1 ($0.786\pm 0.105$) and precision ($0.817\pm 0.089$), whereas U-Net–BiLSTM obtains $0.780\pm 0.094$ F1. Set+DN reaches $0.764\pm 0.112$ F1 but the lowest beat-count error ($0.862\pm 0.584$ beats/epoch), compared with $2.790\pm 1.356$ for the dense Transformer. DN changes set-model F1 by only $+0.004$. The results identify distinct event-accuracy and count-fidelity operating points; they do not establish universal superiority of either output formulation.
Uncertain boundary labels can make defect-segmentation probabilities overconfident near thin cracks and surface transitions. We study Boundary-Local Soft Supervision, a training-only, boundary-gated label-smoothing loss for binary defect masks. The loss adds local-mean soft-target supervision only inside a morphological boundary band, while the main binary cross-entropy plus Dice loss keeps hard-label supervision for all pixels. Adaptive Local Calibration is retained only as an optional inference correction. Across four datasets and 20 random seeds, Boundary-Local Soft Supervision at 80 epochs lowers expected calibration error relative to matched Adaptive Local Calibration at 80 epochs on all datasets, significantly on RedBrick, DeepCrack, and NEU-Seg, without evidence of Dice or Boundary-Dice degradation. Region diagnostics show boundary-band calibration-error reductions of $28\%$–$49\%$ and high-confidence boundary-error reductions of $31\%$–$58\%$. Against focal loss, global label smoothing, full-image spatially varying label smoothing, boundary binary cross-entropy, and temperature scaling, the loss gives competitive calibration without a held-out calibration split. Gains attenuate on Magnetic Tile. The optional correction adds $0.06\%$ parameters and $3.8\%$ additional floating-point operations; Boundary-Local Soft Supervision is inference-free.
Current deep watermarking frameworks typically consist of an encoder, a noise layer, and a decoder (E-N-D), in which jointly optimize the encoder and decoder over a large training set. However, this learned global embedding strategy compromises across diverse images, leaving the embedding potential of individual images under-exploited and introducing an inherent amortization gap. To address this issue, a novel paradigm termed Instance-Aware Encoder Adaptation (IAEA) is proposed in this letter. Built upon the global model trained in the first stage, IAEA freezes the decoder and fine-tunes only the encoder for each cover image and the to-be-embedded watermark in the second stage. This transforms the global optimization into an instance-level refinement under practical decoding constraints, effectively narrowing the amortization gap. The proposed IAEA can exploit the embedding potential of individual images and can be integrated with existing E-N-D methods for performance enhancement, and its effectiveness is experimentally verified.
When underwater acoustic waves propagate to the water-air interface, they induce minute surface displacements and periodic vibrations. By remotely sensing the associated phase variations, millimeter-wave radar enables cross-medium acquisition of underwater acoustic source information. However, the radar echoes are susceptible to phase wrapping, background noise, and strong interference, making the frequency characteristics of the source signal difficult to extract reliably. To address these issues, this letter proposes a water-surface micro-vibration detection method based on the phase-frequency characteristics of radar echoes. First, temporal continuity and spatial-gradient information of the echo phase are jointly exploited to construct a time-gradient two-dimensional Kalman recursion, thereby achieving phase unwrapping and distortion suppression. Second, rotational-gradient processing is employed to enhance the time-frequency representation, while instantaneous-frequency fragment linking is introduced to restore interrupted trajectories and improve source-frequency estimation accuracy. Finally, a two-stage constant false alarm rate detection method is developed to sequentially suppress strong interference and background noise, enabling reliable detection and decoding of underwater acoustic source information. Parameter studies, ablation experiments, and comparative evaluations on measured data validate the effectiveness of the proposed method. The results demonstrate that it improves frequency-estimation accuracy while effectively reducing the bit error rate and false-alarm probability.
This letter proposes an exclusive pulse clique (EPC)-based method for PRI-pattern recovery in multi emitter interleaved scenarios using only TOA information. By exploiting a prior lower bound on PRI, the method constructs exclusive pulse cliques, in which any two pulses cannot originate from the same emitter. Based on this structure, candidate PRI states are first identified, after which local PRI state transition relations are generated and fused to recover the complete PRI-pattern. Simulation results show that the proposed method remains effective in a three-emitter scenario with overlapping PRI ranges and a total of 22 PRI states, even when the missing pulse ratio reaches 40% or the interferential pulse ratio reaches 20%.
Multi-scale graph neural networks have shown strong performance in multivariate time-series forecasting by integrating representations at different temporal resolutions. However, existing fusion mechanisms typically assign the same scale weights to all nodes, implicitly assuming spatially homogeneous temporal preferences. This assumption is restrictive for real-world graph-structured systems, where different nodes may exhibit substantially different dynamics and noise characteristics. To address this limitation, we propose Node-Wise Adaptive Scale Fusion (NAF), a lightweight plug-in module that learns node-specific fusion weights through two complementary views. The adaptive propagation (AP) branch balances node-local evidence with graph-shared information, while the self-attention aggregation (SAAG) branch evaluates the reliability of each scale according to its consistency with a robust full-window trend representation. We further show that globally shared fusion is optimal only under homogeneous or proportionally aligned node reliability profiles, whereas node-wise adaptation yields a strict advantage when such profiles become heterogeneous; the proposed similarity gate also preserves the ordering of branch signal-to-noise ratios under mild conditions. Experiments on four multivariate time-series benchmarks demonstrate that NAF-GNN achieves the best performance in most forecasting settings, with particularly clear gains on heterogeneous traffic data and under additive input corruption. Moreover, the fusion overhead of NAF is independent of the number of graph nodes.
Audio-Visual Speech Recognition (AVSR) has shown robustness in noisy environments, but dynamically varying modality reliability between audio and video causes severe distribution shifts and performance degradation. Existing AVSR methods have mainly relied on data-level restoration or representation-level fusion, but cannot continuously adapt model behavior under evolving test-time conditions. Although Test-Time Adaptation (TTA) provides a natural solution, conventional TTA methods are designed for unimodal settings, leading to noisy modality gradients that dominate the adaptation process. We propose Reliability-Aware Gradient Surgery (RAGS), a gradient-level TTA framework specialized for multimodal environments that estimates modality-wise gradient reliability and employs multiple update strategies. Experimental results demonstrate that RAGS effectively suppresses noisy gradients and consistently improves robustness under severe multimodal corruptions.
This letter addresses target detection for intra-pulse polarization-coding radar, where conventional fixed-polarization detectors cannot fully exploit the polarization variation within a pulse. A polarization-agile generalized likelihood ratio test (PA-GLRT) is developed to jointly process the observations from different polarization segments and adapt to correlated disturbance using training data. The proposed method does not require prior knowledge of the target polarimetric scattering matrix, and practical false-alarm control is achieved through threshold calibration. Numerical results show that PA-GLRT outperforms fixed-polarization and known-template benchmarks, benefits from increasing the number of segments, and remains robust to target scattering variations.
Artificial neural networks (ANN) have significantly advanced speech bandwidth extension (BWE) but suffer from high computational complexity, limiting deployment on power constrained edge devices. While spiking neural networks (SNNs) offer an energy-efficient alternative, they often struggle to capture the long-range temporal dependencies required for high fidelity reconstruction. To address this, we propose SpikeBWE+, a novel framework that synergizes SNNs efficiency with the global modeling capabilities of the Mamba architecture. Central to our approach is the mamba spiking neuron (MSN), which integrates a dual-path Selective State Space Model into spiking architecture to process local and global contexts simultaneously. We further introduce a residual Spiking Connection (RSC) that injects global features directly into the spiking unit to effectively modulate threshold dynamics and optimize sparsity. Experimental evaluations on the TIMIT and VCTK datasets demonstrate that SpikeBWE+ achieves superior reconstruction quality (LSD of 1.14) comparable to complex ANN baselines while significantly reducing energy consumption (energy cost of 1.24 mJ), establishing a new benchmark for BWE. Demo is available at here.
Multimodal sentiment analysis is vulnerable to unstable local cues and temporal asynchrony among textual, visual, and acoustic signals. Existing fusion methods extensively model cross-modal interaction, but lightweight mechanisms that regularize intermediate representations before temporal alignment remain limited. This letter proposes TV-CMF, a time-varying mediation inspired fusion framework for perturbation-stable multimodal sentiment analysis. TV-CMF establishes a constrained routing pathway in which modality-specific temporal states are first reconstructed within a prototype regularized space and subsequently aligned through learnable continuous offsets. A temporal predictability regularizer further constrains the evolution of the routed representations. Here, “mediation-inspired” denotes an enforced representation-routing bottleneck rather than identifiable causal mediation. Experiments on CMU-MOSI, CMU-MOSEI, and perturbations controlled demonstrate visual/acoustic competitive clean-set performance, higher absolute perturbed-set accuracy, and consistent component-wise gains.
When deployed on unreliable hardware, deep neural networks (DNNs) encounter weight perturbations that affect model outputs, increasing uncertainty and posing significant risks to the application of deep learning in safety-critical domains. Current uncertainty estimation approaches generally ignore the influence of weight perturbations. Lacking explicit theoretical modeling and quantification of this effect, these methods remain inadequate for noisy DNNs. To address these limitations, we develop a hardware-noise-aware mean-variance propagation framework for DNNs under weight perturbations. The deployed weights are modeled as trained weights corrupted by hardware noises, and closed-form recursive updates are derived to prop agate the output mean and variance based on assumed density filtering. Moreover, an uncertainty-aware training is proposed to reduce the model output uncertainty by integrating the derived propagated variance into the loss function. Extensive experiments demonstrate that our method exhibits superior performance in uncertainty estimation compared with the state-of-the-art methods, particularly in high-noise environments, while also improving model robustness.
Fixed-position antenna (FPA) arrays provide limited spatial degrees of freedom for integrated radar and jamming (IRAJ), especially in multi-target scenarios. In this letter, we study a movable-antenna (MA)-enabled multi-target IRAJ system, where the transmit waveform and antenna positions are jointly optimized for deceptive jamming under sensing signal-to-interference-plus-noise ratio constraints. The resulting problem is highly nonconvex due to the coupling between waveform synthesis and array reconfiguration. To solve it, a block successive convex approximation-based alternating optimization algorithm is developed. Numerical results show that the proposed MA-enabled design outperforms the conventional FPA benchmark while satisfying the sensing constraints.
Perceptual image hashing aims to generate compact representations that are robust to content-preserving distortions. This letter presents a deep learning-based framework that learns invariant hash codes from a feature disentanglement perspective. The algorithm separates invariant features from distortion related factors by minimizing their statistical dependency, and a learnable mutual information estimator is trained to quantify the degree of disentanglement. Auxiliary image reconstruction tasks are introduced to facilitate training, enabling the disentangled components to accurately capture invariant and distortion features. In addition, a training loss is designed to identify hard examples and constrain their distribution in the hash space. Comparative experiments on a large benchmark dataset show that the proposed method achieves state-of-the-art performance.
Stereo networks trained on synthetic data often degrade in real-world scenes. Within an IGEV-style pipeline, domain-sensitive features may disturb correspondence cues, average matching-dimension pooling may attenuate reliable responses, and local motion encoding does not explicitly exploit rectified stereo geometry. We propose Evidence Preservation and Propagation Stereo (EPP-Stereo), which coordinates feature refinement, volume construction, and recurrent aggregation. Frequency-decoupled refinement regulates low- and high-frequency cues; a stereo-oriented Peak-Preserving Volume Pyramid applies channel-wise temperature-scaled softmax weighting to adjacent disparity responses in both the geometry-encoding and all-pairs correlation pyramids; and Epipolar Large-Kernel Aggregation injects local, epipolar, and orthogonal contexts before the original ConvGRU. Component and directional-kernel ablations further support the complementary roles of the proposed modules. Although the resulting design requires additional computation, EPP-Stereo, trained only on SceneFlow, reduces zero-shot Bad 2.0 on Middlebury from 7.1% to 6.0% at half resolution and from 6.2% to 5.1% at quarter resolution, and lowers Bad 1.0 on ETH3D from 3.6% to 3.4% relative to IGEV-Stereo.
The normalized least-mean-square (NLMS) algorithm is widely used in adaptive filtering due to its simplicity and robustness. Conventional convergence analyses of NLMS focus on one-step update behavior, and characterize step-size optimality only in a local sense. However, a locally optimal step-size does not necessarily yield the best global mean-square error (MSE) trajectory, which practically governs the overall convergence performance. This letter derives a closed-form characterization of the global mean-square error trajectory and establishes a comparison framework using the Loewner partial order. Under a commutativity condition on the relevant matrices, this letter proves that among all variable step-size LMS algorithms with step sizes $0 < \mu _{k} < 2\Vert \mathbf {x}_{k}\Vert ^{2}$, NLMS with $\mu _{k}^\star = 1\Vert \mathbf {x}_{k}\Vert ^{2}$ provides the fastest global convergence speed, which theoretically justifies the empirical numerical observations of fast convergence behavior for NLMS.
Scientific relation extraction aims to identify fine-grained semantic relations between domain-specific entities in scientific documents. However, complex scientific expressions, domain adaptation challenges, and subtle relation distinctions make this task difficult. To address these challenges, we propose a novel natural language inference-based framework for scientific relation extraction. The method reformulates relation classification as textual entailment prediction by converting each candidate relation into a natural language hypothesis and pairing it with a premise constructed from the original scientific context. To exploit the hierarchical semantics of pretrained language models, we introduce an inter-layer semantic attention fusion module that adaptively aggregates contextual representations from multiple layers instead of relying only on the final layer. Moreover, we design a semantic feasibility-constrained focal contrastive learning strategy to reduce the interference of semantically infeasible negative hypotheses and strengthen the learning of hard-to-distinguish relation instances. Experiments on the public SciER dataset show that our framework outperforms strong supervised and existing NLI-based baselines, demonstrating its effectiveness and competitive performance.