Industrial anomaly detection is essential for quality assurance in manufacturing, yet existing methods rely heavily on large-scale annotated data and often exhibit poor cross-domain generalization. These limitations hinder their applicability in real-world scenarios with scarce labeled anomalies and varying data distributions. To address these challenges, we propose CMF-AD (Cross-Modal Few-shot Anomaly Detection), a unified framework for few-shot anomaly detection and diagnosis. Specifically, we introduce a Bidirectional Cross-Modal Attention (BCMA) mechanism to strengthen feature interactions between visual and textual modalities, enabling fine-grained perception and semantic understanding of defect patterns. In addition, we design an Adaptive Feature Fusion (AFF) strategy that integrates multi-scale visual features through learned scale weighting and input-dependent spatial weighting. This design prioritizes anomaly-sensitive representations to enhance both detection accuracy and localization precision. Beyond accurate detection, we further extend the framework to support interpretable anomaly diagnosis without modifying the core detection pipeline. A Feature-to-Text bridging module converts visual features into structured textual representations, which are then fed into a large language model (LLM) to generate detailed diagnosis. This design establishes a unified framework from detection to diagnosis, alleviating the black-box limitation of conventional methods. Extensive experiments on industrial benchmark datasets demonstrate that CMF-AD achieves highly competitive performance in few-shot settings, delivers state-of-the-art fine-grained anomaly localization, especially in terms of Pixel AUPRO, and shows competitive cross-domain generalization with more pronounced advantages in pixel-level localization. Moreover, the generated diagnosis bridges the gap between low-level anomaly localization and high-level semantic interpretation, providing human-readable descriptions.
In modern industry, fault diagnosis is critical for ensuring production safety and efficiency. Intelligent fault diagnosis (IFD) methods often suffer performance degradation under varying real-world operating conditions, primarily due to distribution shifts. Domain generalization (DG) techniques have been introduced to address this issue, aiming to enhance model performance under unseen working conditions. However, existing methods often overlook the significance of multilevel feature representations in learning generalizable information. Furthermore, most methods disregard the fact that not all source domains contribute equally to generalization, where certain domains provide more transferable features than others. To address these challenges, this article proposes a self-distillation-based domain weighted generalization method for fault diagnosis. With the Vision Transformer (ViT) as the backbone, the method first incorporates a self-distillation mechanism that distills class-specific knowledge within class tokens from the final Transformer block to intermediate blocks, enabling the learning of more generalized features. Second, a dynamic domain weighting strategy based on self-distillation loss is proposed, which assigns higher weights to domains with more generalized features. Finally, these weights are integrated into domain adversarial learning and fault classification process. Extensive experiments on motor fault diagnosis are conducted, covering multiple conditions, including constant conditions, start-up and shut-down conditions, and the New European Driving Cycle (NEDC) conditions. Experimental results on 33 cross-condition tasks, including constant-to-constant, constant-to-variable, and variable-to-variable tasks, demonstrate that the proposed method achieves superior accuracy and robustness compared with several representative methods. Ablation studies further validate the effectiveness of each component in the proposed method.
Reliable health assessment of hydraulic power components is essential for autonomous construction machinery. As the core power component, the health of axial piston pumps directly impacts the safety of automated process. However, traditional single-source data-driven methods are limited by incomplete information and vulnerable to sensor malfunctions or data loss. To address this, we propose a dual-path physics-informed heterogeneous information fusion (DP-PIHIF) framework, which consists of a Virtual Health Indicator (VHI) path and a Physical Health Indicator (PHI) path, integrating high-frequency vibration and relatively low-frequency physical data. In the VHI path, multi-domain vibration features are extracted and compressed using a Variational Autoencoder to characterize dynamic fault signatures. In the PHI path, a Physics-Informed Transformer is developed to infer interpretable leakage-related coefficients from multi-sensor hydraulic measurements, with physical consistency enforced through a volumetric flow loss model. The resulting VHI and PHI representations are integrated using an adaptive heterogeneous fusion module with dynamic weight allocation, followed by a multi-task learning scheme for simultaneous health state classification and ordinal health-level regression. Experimental results show superior assessment accuracy and regression stability, outperforming single-source baselines. This indicates the promising potential of the proposed framework for remote health assessment, which may support the predictive maintenance and operational continuity in autonomous construction machinery.
Existing industrial anomaly detection methods primarily concentrate on unsupervised learning with pristine RGB images. Yet, both RGB and 3D data are crucial for anomaly detection, and the datasets are seldom completely clean in practical scenarios. To address above challenges, this paper initially delves into the RGB-3D multi-modal noisy anomaly detection, proposing a novel noise-resistant M3DM-NR framework to leveraging strong multi-modal discriminative capabilities of CLIP. M3DM-NR consists of three stages: Stage-I introduces the Suspected References Selection module to filter a few normal samples from the training dataset, using the multimodal features extracted by the Initial Feature Extraction, and a Suspected Anomaly Map Computation module to generate a suspected anomaly map to focus on abnormal regions as reference. Stage-II uses the suspected anomaly maps of the reference samples as reference, and inputs image, point cloud, and text information to achieve denoising of the training samples through intra-modal comparison and multi-scale aggregation operations. Finally, Stage-III proposes the Point Feature Alignment, Unsupervised Feature Fusion, Noise Discriminative Coreset Selection, and Decision Layer Fusion modules to learn the pattern of the training dataset, enabling anomaly detection and segmentation while filtering out noise. Extensive experiments show that M3DM-NR outperforms state-of-the-art methods in 3D-RGB multi-modal noisy anomaly detection.
Stroke-based Rendering (SBR) aims to decompose an input image into a sequence of parameterized strokes, which can be rendered into a painting that resembles the input image. Recently, Neural Painting methods that utilize deep learning and reinforcement learning models to predict the stroke sequences have been developed, but suffer from longer inference time or unstable training. To address these issues, we propose AttentionPainter, an efficient and adaptive model for single-step neural painting. First, we propose a novel scalable stroke predictor, which predicts a large number of stroke parameters within a single forward process, instead of the iterative prediction of previous Reinforcement Learning or auto-regressive methods, which makes AttentionPainter faster than previous neural painting methods. To further increase the training efficiency, we propose a Fast Stroke Stacking algorithm, which brings 13 times acceleration for training. Moreover, we propose Stroke-density Loss, which encourages the model to use small strokes for detailed information, to help improve the reconstruction quality. Finally, we design a Stroke Diffusion Model as an application of AttentionPainter, which conducts the denoising process in the stroke parameter space and facilitates stroke-based inpainting and editing applications helpful for human artists' design. Extensive experiments show that AttentionPainter outperforms the state-of-the-art neural painting methods.
To address the challenges of verifying MR-based attenuation correction (MRAC) in PET/MR due to CT positional mismatches and alignment issues, this study utilized a flatbed insert and arms-down positioning during PET/CT scans to achieve precise MR-CT matching for accurate MRAC evaluation. A validation dataset of 21 patients underwent whole-body [18F]FDG PET/CT followed by [18F]FDG PET/MR. A flatbed insert ensured consistent positioning, allowing direct comparison of four MRAC methods—four-tissue and five-tissue models with discrete and continuous μ-maps—against CT-based attenuation correction (CTAC). A deep learning-based framework, trained on a dataset of 300 patients, was used to generate synthesized-CTs from MR images, forming the basis for all MRAC methods. Quantitative analyses were conducted at the whole-body, region of interest, and lesion levels, with lesion-distance analysis evaluating the impact of bone proximity on standardized uptake value (SUV) quantification. Distinct differences were observed among MRAC methods in spine and femur regions. Joint histogram analysis showed MRAC-4 (continuous μ-map) closely aligned with CTAC. Lesion-distance analysis revealed MRAC-4 minimized bone-induced SUV interference (r = 0.01, p = 0.8643). However, tissues prone to bone segmentation interference, such as the spine and liver, exhibited greater SUV variability and lower reproducibility in MRAC-4 compared to MRAC-2 (2D bone segmentation, discrete μ-map) and MRAC-3 (3D bone segmentation, discrete μ-map). Using a flatbed insert, this study validated MRAC with high precision. Continuous μ-value MRAC method (MRAC-4) demonstrated superior accuracy and minimized bone-related SUV errors but faced challenges in reproducibility, particularly in bone-rich regions.
Neural style transfer is a prominent AI technique for creating captivating visual effects and enhancing user experiences. However, most current methods inadequately handle panoramic images, leading to a loss of original visual semantics and emotions due to insufficient structural feature consideration. To address this, a novel panorama arbitrary style transfer method named PAST-Renderer is proposed by integrating deformable convolutions and distortion constraints. The proposed method can dynamically adjust the position of the convolutional kernels according to the geometric structure of the input image, thereby better adapting to the spatial distortions and deformations in panoramic images. Deformable convolutions enable adaptive transformations on a twodimensional plane, enhancing content and style feature extraction and fusion in panoramic images. Distortion constraints adjust content and style losses, ensuring semantic consistency in salience, edge, and depth of field with the original image. Experimental results show significant improvements, with the PSNR (Peak Signal-toNoise Ratio) and SSIM (Structural Similarity Index Measure) of stylized panoramic images' semantic maps increasing by approximately 2-4 dB and 0.1-0.3, respectively. Our method PAST-Renderer performs better in both artistic and realistic style transfer, preserving semantic integrity with natural colors, realistic edge details, and rich thematic content.
The accurate and efficient operation of industrial systems heavily depends on motor fault diagnosis. To address the issue of limited model generalization in motor fault diagnosis under unknown working conditions, this paper proposes a novel domain generalization framework. The method employs Continuous Wavelet Transform (CWT) to extract features from vibration signals and utilizes a ResNet-18 network for feature encoding. It further incorporates contrastive learning and entropy regularization to learn domain-discriminative representations, enabling the model to effectively distinguish between source domains. Subsequently, a dynamic weighting mechanism based on feature centroids leverages these learned domain characteristics to improve the model’s robustness under unseen conditions. Experimental results show that the proposed method achieves an average accuracy improvement of 4.57% over baseline models across multiple target domains, demonstrating its effectiveness and generalization capability.
To address the issue where signals in fault diagnosis are heavily disturbed by external noise, making feature extraction difficult, this paper proposes a signal denoising method based on an improved diffusion model. The method uses a diffusion model to take pure signals mixed with noise as samples and employs a Bi-LSTM network as the signal denoiser for training. Weak signals subjected to external interference are collected for denoising, and the denoised signals are classified using a one-dimensional convolutional neural network. Results indicate that this approach significantly enhances detection capabilities in fault diagnosis.
Parameter-efficient fine-tuning methods adjust a small subset of parameters in large models, achieving performance comparable to or even surpassing that of models fine-tuned with the full parameter set, and significantly reducing the time and computational costs associated with the fine-tuning process. Despite the developments of parameter-efficient fine-tuning methods for large models, we observe significant performance disparities across different vision tasks. We attribute this pronounced performance variability to the insufficient robustness of current parameter-efficient fine-tuning methods. In this paper, we propose a robust reparameterization framework for parameter-efficient fine-tuning. This framework has a dynamic training structure and introduces no additional computational overhead during the inference stage. Specifically, we propose Dropout-Mixture Low-Rank Adaptation (DMLoRA), yrrwhich incorporates multiple up and down branches, to provide the model with a more robust gradient descent path. As training proceeds, DMLoRA gradually drops out branches to achieve a balance between accuracy and regularization. Additionally, we employ a 2-Stage Learning Scalar (LS) strategy to optimize the scale factor for each layer's DMLoRA module. Experimental results demonstrate that our method achieves state-of-the-art performance on the benchmark VTAB-1k and FGVC datasets for parameter-efficient fine-tuning.
Network traffic anomaly detection, as an effective analysis method for network security, can identify differentiated traffic information and provide secure operation in complex and changing network environments. To avoid information loss caused when handling traffic data while improving the detection performance of traffic feature information, this paper proposes a multi-information fusion model based on a convolutional neural network and AutoEncoder. The model uses a convolutional neural network to extract features directly from the raw traffic data, and a AutoEncoder to encode the statistical features extracted from the raw traffic data, which are used to supplement the information loss due to cropping. These two features are combined to form a new integrated feature for network traffic, which has the load information from the original traffic data and the global information of the original traffic data obtained from the statistical features, thus providing a complete representation of the information contained in the network traffic and improving the detection performance of the model. The experiments show that the classification accuracy of network traffic anomaly detection using this model outperforms that of classical machine learning methods.
AI-Generated Content (AIGC) has recently gained a surge in popularity, powered by its high efficiency and consistency in production, and its capability of being customized and diversified. The cross-modality nature of the representation learning mechanism in most AIGC technology allows for more freedom and flexibility in exploring new types of art that would be impossible in the past. Inspired by the pictogram subset of Chinese characters, we proposed PaCaNet, a CycleGAN-based pipeline for producing novel artworks that fuse two different art types, traditional Chinese painting and calligraphy. In an effort to produce stable and diversified output, we adopted three main technical innovations: 1. Using one-shot learning to increase the creativity of pre-trained models and diversify the content of the fused images. 2. Controlling the preference over generated Chinese calligraphy by freezing randomly sampled parameters in pre-trained models. 3. Using a regularization method to encourage the models to produce images similar to Chinese paintings. Furthermore, we conducted a systematic study to explore the performance of PaCaNet in diversifying fused Chinese painting and calligraphy, which showed satisfying results. In conclusion, we provide a new direction of creating arts by fusing the visual information in paintings and the stroke features in Chinese calligraphy. Our approach creates a unique aesthetic experience rooted in the origination of Chinese hieroglyph characters. It is also a unique opportunity to delve deeper into traditional artwork and, in doing so, to create a meaningful impact on preserving and revitalizing traditional heritage.
2D-based Industrial Anomaly Detection has been widely discussed, however, multimodal industrial anomaly detection based on 3D point clouds and RGB images still has many untouched fields. Existing multimodal industrial anomaly detection methods directly concatenate the multimodal features, which leads to a strong disturbance between features and harms the detection performance. In this paper, we propose Multi-3D-Memory (M3DM), a novel multimodal anomaly detection method with hybrid fusion scheme: firstly, we design an unsupervised feature fusion with patch-wise contrastive learning to encourage the interaction of different modal features; secondly, we use a decision layer fusion with multiple memory banks to avoid loss of information and additional novelty classifiers to make the final decision. We further propose a point feature alignment operation to better align the point cloud and RGB features. Extensive experiments show that our multi-modal industrial anomaly detection model outperforms the state-of-the-art (SOTA) methods on both detection and segmentation precision on MVTec-3D AD dataset. Code at github.com/nomewang/M3DM.
Face analysis tasks have a wide range of applications, but the universal facial representation has only been explored in a few works. In this paper, we explore high-performance pre-training methods to boost the face analysis tasks such as face alignment and face parsing. We propose a self-supervised pre-training framework, called Mask Contrastive Face (MCF), with mask image modeling and a contrastive strategy specially adjusted for face domain tasks. To improve the facial representation quality, we use feature map of a pre-trained visual backbone as a supervision item and use a partially pre-trained decoder for mask image modeling. To handle the face identity during the pre-training stage, we further use random masks to build contrastive learning pairs. We conduct the pre-training on the LAION-FACE-cropped dataset, a variants of LAION-FACE 20M, which contains more than 20 million face images from Internet websites. For efficiency pre-training, we explore our framework pre-training performance on a small part of LAION-FACE-cropped and verify the superiority with different pre-training settings. Our model pre-trained with the full pre-training dataset outperforms the state-of-the-art methods on multiple downstream tasks. Our model achieves 0.932 NME_diag for AFLW-19 face alignment and 93.96 F1 score for LaPa face parsing. Code is available at https://github.com/nomewang/MCF.
Servo systems are widely used in aerospace and other fields. To ensure the safety of the system operation, the fault detection of servo systems is very important. Servo system fault detection often uses comparative differential judgment method, which requires the establishment of a reliable model. On the one hand, servo system models are usually built using mathematical and physical methods by combining the transfer functions of the various components the actual servo system have into an algorithm that calculates the input–output relationship. On the other hand, machine learning algorithms can also be used to model the servo system. In this paper, we address this issue by developing a rocket data prediction model based on the LSTM algorithm using publicly available data provided by Shanghai Aerospace Control Technology Institute. After deriving the results, the results are evaluated by using R-squared metrics and some improvement outlooks are provided for the subsequent research.
Generating artistic portraits is a challenging problem in computer vision. Existing portrait stylization models that generate good quality results are based on Image-to-Image Translation and require abundant data from both source and target domains. However, without enough data, these methods would result in overfitting. In this work, we propose CtlGAN, a new few-shot artistic portraits generation model with a novel contrastive transfer learning strategy. We adapt a pretrained StyleGAN in the source domain to a target artistic domain with no more than 10 artistic faces. To reduce overfitting to the few training examples, we introduce a novel Cross-Domain Triplet loss which explicitly encourages the target instances generated from different latent codes to be distinguishable. We propose a new encoder which embeds real faces into Z+ space and proposes a dual-path training strategy to better cope with the adapted decoder and eliminate the artifacts. Extensive qualitative, quantitative comparisons and a user study show our method significantly outperforms state-of-the-arts under 10-shot and 1-shot settings and generates high quality artistic portraits. The code will be made publicly available.