The primary challenges in image-level weakly supervised semantic segmentation (WSSS) lie in addressing the under-activation issue of target pixels and mitigating the co-occurrence phenomenon in class activation maps. In recent years, Vision-Language Models (VLM) have demonstrated exceptional performance across various vision tasks, primarily attributed to their cross-modal semantic alignment capabilities achieved through contrastive learning mechanisms. Leveraging VLM's capability to capture fine-grained visual-textual correspondences, this paper proposes a novel Vision-Language Driven Prompt Learning (VLD-PL) framework that addresses two fundamental challenges in WSSS by establishing explicit semantic correspondences between textual descriptors and visual components, ultimately enabling efficient semantic segmentation. The VLD-PL framework consists of two core components Auxiliary Class Matching (ACM) and Background Class Filtering (BCF). The ACM module dynamically identifies semantically relevant auxiliary classes through feature alignment between image and textual embeddings, effectively enlarging target activation while mitigating co-occurrence interference by expanding semantic coverage. Simultaneously, the BCF constructs image-specific background prompts and adaptively refines background feature representations, achieving precise suppression of irrelevant background regions. These dual mechanisms synergistically address both target localization accuracy and background noise suppression, achieving state-of-the-art performance on both the PASCAL VOC 2012 and MS COCO 2014 benchmarks.
Magnetic Resonance Imaging (MRI) plays a crucial role in clinical diagnosis but suffers from a time-consuming data acquisition process. While under-sampling accelerates imaging, it inevitably introduces aliasing artifacts and detail loss. Recently, reference-based reconstruction methods have emerged as effective approaches to enhance accelerated reconstruction quality by leveraging high-resolution anatomical information from other modalities. However, inevitable spatial misalignment limits their potential, as conventional approaches cannot gradually align structures or iteratively refine features. Furthermore, relying solely on the image domain ignores k-space prior knowledge. To address these issues, this paper proposes a Multi-Modal Iterative Refinement Network with K-space Posterior Correction for MRI reconstruction (MMIR-Net). Specifically, we first design an image domain Iterative Refinement Network (IR-Net) incorporating a Residual Registration Module (RRM) to perform progressive feature refinement and dynamically align structures. Subsequently, we introduce a K-space Posterior Correction Module (KPCM) to use physical priors to correct frequency distribution and improve data consistency. Finally, we develop an Adaptive Fusion Module (AFM) to effectively integrate the features from both domains. Experimental results on the IXI and fastMRI datasets demonstrate that the proposed method outperforms existing methods across various sampling patterns and ratios, providing a new way to address multi-modal MRI reconstruction challenges.
Recently, Vision-Language Models (VLMs) have achieved remarkable success in vision tasks via contrastive learning. As a representative VLM, Contrastive Language-Image Pre-training (CLIP) has shown great potential for Weakly Supervised Semantic Segmentation (WSSS). However, bridging the gap between CLIP’s global representations and dense segmentation remains challenging. In this paper, we propose a Vision-Language Feature Calibration (VLFC) framework to generate complete and accurate class activation maps. VLFC consists of two complementary branches: an image feature learning branch with a Multi-Scale Gated Attention (MSGA) module to diffuse activation to integral object regions, and a text feature learning branch with a Text Knowledge Bank Construction (TKBC) strategy to provide fine-grained textual guidance. In addition, to adapt CLIP to dense segmentation tasks without losing its generalization, we introduce a Low-rank Projection-based Incremental Calibration (LP-IC) strategy, which injects learnable low-rank updates into the frozen encoders of both branches, efficiently calibrating visual and textual distributions. The calibrated features are then aligned in an image-text contrastive learning space with two designed losses that enhance foreground activation and suppress class-related background noise. Experiments on PASCAL VOC 2012 and MS COCO 2014 demonstrate superior performance over state-of-the-art methods.
Point cloud completion referring to completing 3D shapes from partial 3D point clouds is a fundamental problem for 3D point cloud analysis tasks. Benefiting from the development of deep neural networks, researches on point cloud completion have made great progress in recent years. However, the explicit local region partition like kNNs involved in existing methods makes them sensitive to the density distribution of point clouds. Moreover, it serves limited receptive fields that prevent capturing features from long-range context information. To solve the problems, we leverage the cross-attention and self-attention mechanisms to design novel neural network for point cloud completion with implicit local region partition. Two basic units Geometric Details Perception (GDP) and Self-Feature Augment (SFA) are proposed to establish the structural relationships directly among points in a simple yet effective way via attention mechanism. Then based on GDP and SFA, we construct a new framework with popular encoder-decoder architecture for point cloud completion. The proposed framework, namely PointAttN, is simple, neat and effective, which can precisely capture the structural information of 3D shapes and predict complete point clouds with detailed geometry. Experimental results demonstrate that our PointAttN outperforms state-of-the-art methods on multiple challenging benchmarks. Code is available at: https://github.com/ohhhyeahhh/PointAttN
Class activation maps generated by image classifiers are widely used as priors for image-level weakly supervised semantic segmentation. However, these activation maps mainly focus on the sparse discriminative regions, which has been a bottleneck for the segmentation task. Based on our observations, the activation maps actually capture almost the entire target regions, and some regions with lower activation values are easily to be neglected. Thus, to solve the issue, we propose an adaptive activation network with two branches to recalibrate the low-confidence regions in the activation maps. Specifically, an activation enhancement branch is designed to redistribute the activation values by leveraging attention mechanism. Since multi-scale images can provide complementary information, a scale adaptation branch is paralleled to supervise the activation enhancement branch. The mutual supervision and fusion of the two branches can promote the less-discriminative parts, and deactivate the background regions. Based on them, a simple yet effective denoising module is proposed to further improve the quality of pseudo masks, which makes use of the large scale predictions of the trained segmentation network. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 benchmarks show that our method achieves state-of-the-art performance, demonstrating the effectiveness of our algorithm. Code will be made publicly available.
With the advancement of social life, the aging of building walls has become an unavoidable phenomenon. Due to the limited efficiency of manually detecting cracks, it is especially necessary to explore intelligent detection techniques. Currently, deep learning has garnered growing attention in crack detection, leading to the development of numerous feature learning methods. Although the technology in this area has been progressing, it still faces problems such as insufficient feature extraction and instability of prediction results. To address the shortcomings in the current research, this paper proposes a new Adaptive Attention-Enhanced Yolo. The method employs a Swin Transformer-based Cross-Stage Partial Bottleneck with a three-convolution structure, introduces an adaptive sensory field module in the neck network, and processes the features through a multi-head attention structure during the prediction process. The introduction of these modules greatly improves the performance of the model, thus effectively improving the precision of crack detection.
Though deep learning-based saliency detection methods have achieved gratifying performance recently, the predicted saliency maps still suffer from the boundary challenge. From the perspective of foreground-background separation, this article attempts to extract the edge information of objects by exploiting the difference between different color channels in the RGB color space and establishes a novel multicolor contrast extraction (MCE) mechanism to improve the learning ability of exquisite boundary information of the network. To make full use of the MCE outputs and RGB colors, and well depict and capture the complementary information between them, we devise a novel Siamese densely cooperative fusion (DCF) network (SDFNet) for saliency detection, which consists of two effective components: boundary-directed feature learning (BDFL) and DCF. The BDFL provides joint learning for both MCE and RGB modalities through a Siamese network, while the DCF module is devised for complementary feature discovery, in order to effectively combine the features learned from two modalities. Experiments on five well-known benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches in terms of different evaluation metrics. We provide a detailed analysis of these results and indicate that our joint modeling of MCE and RGB colors helps to better capture the object details, especially in the object boundaries.
The airborne two-dimensional stereo (2D-S) optical array probe has been operating for more than 10 yr, accumulating a large amount of cloud particle image data. However, due to the lack of reliable and unbiased classification tools, our ability to extract meaningful morphological information related to cloud microphysical processes is limited. To solve this issue, we propose a novel classification algorithm for 2D-S cloud particle images based on a convolutional neural network (CNN), named CNN-2DS. A 2D-S cloud particle shape dataset was established by using the 2D-S cloud particle images observed from 13 aircraft detection flights in 6 regions of China (Northeast, Northwest, North, East, Central, and South China). This dataset contains 33,300 cloud particle images with 8 types of cloud particle shape (linear, sphere, dendrite, aggregate, graupel, plate, donut, and irregular). The CNN-2DS model was trained and tested based on the established 2D-S dataset. Experimental results show that the CNN-2DS model can accurately identify cloud particles with an average classification accuracy of 97%. Compared with other common classification models [e.g., Vision Transformer (ViT) and Residual Neural Network (ResNet)], the CNN-2DS model is lightweight (few parameters) and fast in calculations, and has the highest classification accuracy. In a word, the proposed CNN-2DS model is effective and reliable for the classification of cloud particles detected by the 2D-S probe.
One-stage multi-object tracking methods have achieved promising results by showing their great balance between accuracy and speed. However, the internal differences and relationships between detection and re-identification (re-ID) lead to worse performance. In this work, we propose a one-stage multi-object tracking method with attention boosting, namely AeMOT, which can effectively improve the collaboration and performance in detection and re-ID. Specifically, a discriminability enhancement module is designed to enhance the discriminative feature representations for detection and tracking, and an identity preserving module is properly designed to preserve the semantic alignment of id-embedding and improve the adaptiveness of object matching with scale variation for re-ID association. Experimental results on challenging benchmarks including MOT17 and MOT20 demonstrate that our proposed method achieves leading performance and outperforms state-of-the-art trackers.
Liquid water content (LWC) in clouds determines the precipitable water of clouds, which is a crucial factor for aircraft safety and weather modification operations. More importantly, it influences the optical depth of clouds in the visible wavelength range, thus determining their climate cooling effects. Identifying and quantifying the LWC in mixed-phase clouds via remote sensing techniques remains challenging owing to the large variability of hydrometeor sizes in the cloud. In this study, we used in-situ aircraft measured full size distributions and collocated airborne radar reflectivity (Z) to explicitly fractionate the contributions of hydrometeors (cloud liquid droplets, ice, and precipitation particles) at different size ranges from the measured total Z. A linearly decreasing contribution of non-precipitation hydrometeors with increasing total Z was discovered for a range of cloud types, including cumulus, status, and deep convection clouds. The relationship between the mass and Z for each type of hydrometeor derived from in situ measurements was then applied. This approach of apportioning the contribution of cloud liquid droplets from the measured total Z as the first step significantly reduced the scattering of the correlation between the LWC and Z, as presented in previous studies; thus, the LWC was more accurately determined. The derived liquid water path exhibited high agreement with the microwave radiometer measurements. Our method of deriving the LWC from the total Z stemming from the in situ measured size distributions may be applied in other situations to derive the cloud liquid droplets, ice, and precipitation masses for clouds with a given radar Z.
Video salient object detection (VSOD) aims at distinguishing the salient objects from the complex background and highlighting them uniformly in the spatiotemporal domain. One of the fundamental challenges in VSOD is how to make the most use of the temporal information to boost the performance. We propose a dual temporal memory network (DTMNet) which stores short- and long-term video sequence information preceding the current frame as the temporal memories to address the temporal modeling in VSOD. The proposed network consists of two temporal modules including a short-term co-inference learning (SCL) sub-module and a long-range memory learning (LML) sub-module. The SCL is designed for inferencing spatiotemporal interactions between neighboring frames of the current input video clip. The LML aims to satisfy the logical reasoning sequence in timeline and learn the long-time range information between current clip and the previous video clips. Comprehensive evaluations well demonstrate the effectiveness and robustness of our proposed architecture.
Video snapshot compressive imaging (SCI) system enables high-frame-rate imaging by projecting multiple frames into a 2D snapshot measurement during a single exposure, and the original video frames can be reconstructed by solving an optimization problem. However, existing methods usually cannot achieve a good balance between reconstruction time and reconstruction quality, which has become a major obstacle for practical application of video SCI. In order to cope with this issue, we propose a residual ensemble network to learn the explicit inverse mapping from the 2D snapshot measurement to the original video. Specifically, the proposed network aims to exploit the spatiotemporal correlations between video frames for improving reconstruction quality. The spatiotemporal correlations of video frames demonstrate multiple types, including intra-frame spatial correlation, inter-frame forward and backward temporal correlation. With the purpose of fully capturing these differentiated correlations, we design four sub-networks, namely, a pseudo-3D U-shape sub-network, two residual sub-networks, and a serial forward and backward recurrent sub-network, and further assemble these four sub-networks into an ensemble network through alternate residual links. This ensemble network can effectively fuse the predictions of each sub-network and maintain spatiotemporal consistency between video frames. We further design a compound loss function to guide the network learning, and the new video can be fast reconstructed by simply feeding its 2D snapshot measurement into the learned network. The experimental results demonstrate that our network can significantly improve the reconstruction quality while maintaining low computational cost.
Most of learning targets for multi-person pose estimation are based on the likelihood P ( Y | X ) . However, if we construct the causal assumption for keypoints, named a Structure Causal Model (SCM) for the causality, P ( Y | X ) will introduce the bias via spurious correlations in the SCM. In practice, it appears as that networks may make biased decisions in the dense area of keypoints. Therefore, we propose a novel learning method, named Causal Intervention pose Network (CIposeNet). Causal intervention is a learning method towards solving bias in the SCM of keypoints. Specifically, under the consideration of causal inference, CIposeNet is developed based on the backdoor adjustment and the learning target will change into causal intervention P ( Y | d o ( X ) ) instead of the likelihood P ( Y | X ) . The experiments conducted on multi-person datasets show that CIposeNet indeed releases bias in the networks.
In this article, we tackle the saliency detection task from an interesting perspective: we focus both on salient regions (or foreground) detection and nonsalient regions (or background) detection instead of only the foreground and propose a novel complementarity-aware attention network. It is a unified framework with two branches, namely, positive attention module (PAM) and negative attention module (NAM), for the foreground and background detection, respectively. More specifically, the PAM exploits a position self-attention mechanism to enhance the discriminant ability of feature representation, which can detect most of the salient object regions. Meanwhile, the NAM is designed to detect the background regions, aiming to pop out the missing object parts and details in the prediction map produced by the PAM. By fusing these two attention modules together, NAM can provide complementary cues to assist PAM for precise object detection. Furthermore, in order to capture more multiscale contextual information, we introduce a bidirectional structure with multisupervision to the proposed complementarity-aware attention module for performance improvement. Experiments on five benchmark datasets show that the proposed framework achieves comparable results compared with the state-of-the-art saliency detection methods.
Purpose In novelty detection, the autoencoder based image reconstruction strategy is one of the mainstream solutions. The basic idea is that once the autoencoder is trained on normal data, it has a low reconstruction error on normal data. However, when faced with complex natural images, the conventional pixel-level reconstruction becomes poor and does not show the promising results. This paper aims to provide a new method for improving the performance of novelty detection based autoencoder. Design/methodology/approach To solve the problem that conventional pixel-level reconstruction cannot effectively extract the global semantic information of the image, a novel model with the combination of attention mechanism and self-supervised learning method is proposed. First, an auxiliary task, reconstruct rotated image, is set to enable the network to learn global semantic feature information. Then, the channel attention mechanism is introduced to perform adaptive feature refinement on the intermediate feature map to optimize the correspondingly passed feature map. Findings Experimental results on three public data sets show that the proposed method has potential performance for novelty detection. Originality/value This study explores the ability of self-supervised learning methods and attention mechanism to extract features on a single class of images. In this way, the performance of novelty detection can be improved.
Recently, most existing human pose estimation methods fuse multi-stage convolutional modules to learn a shared feature representation. In this paper, we propose a expectation–maximization (EM) mapping-based network to learn specific related body parts for human pose estimation, named EMposeNet. It maps specific feature of related parts from the original fully shared feature space. From the perspective of multi-task learning, we can regard the task of human pose estimation as a homogeneous multi-task learning. Sharing features among related tasks can result in a more compact model and better generalization ability. However, sharing features for those unrelated or weakly related tasks will deteriorate the estimation performance. Our proposed method aims at performing EM algorithm to learn the related body part, where the predicted keypoint heatmap is potentially more accurate and spatially more precise. We conduct extensive experiments on two benchmark datasets, including the MSCOCO keypoint detection dataset and the MPII human pose dataset, to empirically demonstrate the validity of the proposed method. The results on such two benchmark datasets show that the proposed approach achieves a competitive performance.
This paper addresses the core issue of how to learn powerful features for saliency. We have two major observations. First, feature maps of different layers in convolutional neural networks play different roles in saliency detection. Second, different feature channels in the same layer are not of equal importance to saliency, and they often have different response to foreground or background. To address these problems, a stacked U-shape network with channel-wise attention is presented to effectively utilize these features, which mainly consists of a parallel dilated convolution (PDC) module and a multi-level attention cascaded feedback (MACF) module. More specifically, PDC aims to enlarge the receptive field without increasing the computation and effectively avoid the gridding problem. MACF is innovatively designed to adaptively select the cross-layer complementary information, and the inter-dependencies between different channel maps in the same layer can be depicted well. Finally, we adopt a multi-layer loss function to improve the commonly used binary cross entropy loss which treats all pixels equally. The extensive experiments on five saliency detection datasets demonstrate that the proposed method outperforms the state-of-the-art approaches.
Abstract Existing deep‐learning–based saliency detection methods mainly design a sophisticated architecture to integrate multi‐level convolutional features, for example, recurrent network or bi‐directional message passing model, and achieve gratifying performance. However, the direct transmission of information with large span, for example, the information from the deepest layer and the shallowest layer, may lead to antagonistic problems. To address this problem, we propose a local bi‐directional funnel network (LBDFN) to effectively integrate multi‐level features for salient object detection. In the proposed local bi‐directional funnel network, a local bi‐directional feature integration (LBDFI) module is designed to fuse the features with similar properties from adjacent convolution layers, which can solve the problem of mismatch between fusion features well. The extensive experiments on five saliency detection datasets clearly demonstrate that the proposed method outperforms the state‐of‐the‐art approaches.
Recently, optical flow guided video saliency detection methods have achieved high performance. However, the computation cost of optical flow is usually expensive, which limits the applications of these methods in time-critical scenarios. In this article, we propose an end-to-end cross complementary network (CCNet) based on fully convolutional network for video saliency detection. The CCNet consists of two effective components: single-image representation enhancement (SRE) module and spatiotemporal information learning (STIL) module. The SRE module provides robust saliency feature learning for a single image through a pyramid pooling module followed by a lightweight channel attention module. As an effective alternative operation of optical flow to extract spatiotemporal information, the STIL introduces a spatiotemporal information fusion module and a video correlation filter to learn the spatiotemporal information, the inner collaborative and interactive information between consecutive input groups. In addition to enhancing the feature representation of a single image, the combination of SRE and STIL can learn the spatiotemporal information and the correlation between consecutive images well. Extensive experimental results demonstrate the effectiveness of our method in comparison with 14 state-of-the-art approaches.
Recently, benefiting from the fast development of deep convolutional neural networks, salient object detection (SOD) has achieved gratifying performance in a variety of challenging scenarios. Among them, how to learn more discriminative features plays a key role. In this paper, we propose a novel network architecture that progressively fuses the rich multi-level contextual features from top to bottom to learn a more effective feature presentation for robust SOD. Concretely, we first design a multi-receptive field block (MRFB) to capture multi-scale contextual information. Then, we develop a feature fusion block that progressively fuses different outputs of MRFBs from top to bottom, which can effectively filter out the non-complementary parts of the high-level and low-level features. Afterwards, we leverage a refinement residual block to refine the results further. Finally, we leverage an edge-aware loss as an aid to guide the network to learn more sharpen details of the salient objects. The whole network is trained end-to-end without any pre-processing and post-processing. Exhaustive evaluations on six benchmark datasets demonstrate superiority of the proposed method against state-of-the-arts in terms of all metrics.