Current state-of-the-art deep multi-view clustering methods resort to contrastive learning to learn consensus representations with Cross-View Consistency (CVC). However, contrastive learning has inherent limitations when being applied to the multi-view clustering. On one hand, contrastive learning suffers from class collision issue, compromising the discriminability of consensus representation. On the other hand, contrastive alignment of two views of different quality could lead to representation degradation for the higher-quality view, weakening the robustness of the consensus representation. To alleviate these issues, this paper presents an Adaptive Multi-view consistency clustering method via structure-enhanced contrastive learning (AdaM), which learns multi-faceted consensus representation that balances view-consistency, discriminability and robustness, forming an optimal consensus representation. Specifically, we first design a view fusion module and a structural learning module to learn view weights and structural relationships among samples, respectively, to derive the consensus representation. Second, beyond CVC, we propose a novel clustering framework called Adaptive Multi-View Consistency (AMVC), which adaptively aligns specific view representation with consensus representation based on the learned view weights. Furthermore, compared to CVC, we theoretically demonstrate the superiority of AMVC in learning robust consensus representation. Third, AdaM leverages the structural relationships among samples to refine the conventional contrastive loss, further enhancing the discriminability of the consensus representation. Extensive experimental results on eight datasets demonstrate the superior performance of AdaM over eight advanced multi-view clustering baselines.
Image restoration, which aims to recover and enhance degraded images, is a fundamental task across diverse applications. While conventional deep learning approaches have achieved notable success for specific degradations such as noise removal or deblurring, there remains a strong demand for an all-in-one model that can: (i) address multiple degradation types, (ii) reduce the storage overhead of maintaining separate models, and (iii) provide interactive and flexible usage. To this end, we propose a novel Prompt-In-Prompt learning for all-in-one image restoration, termed PIP. Our approach is interpretable, flexible, efficient, and easy to use. Specifically, we first introduce two complementary prompts: a degradation-aware prompt that encodes high-level degradation knowledge and a basic restoration prompt that captures essential low-level information. Second, we design a prompt-to-prompt interaction module to fuse the two prompts into an universal restoration prompt. Third, we further develop a selective prompt-to-feature interaction module to modulate degradation-related features. With these designs, PIP functions as a lightweight, plug-and-play module that can be seamlessly integrated into existing restoration models for all-in-one image restoration. Extensive experimental results on seven simulated and four real-world datasets demonstrate its superior performance across diverse restoration tasks, including image denoising, deraining, dehazing, deblurring, and low-light enhancement. The source code is available at https://github.com/longzilicart/pip_universal.
Computed tomography (CT) metal artifact reduction (MAR) aims to reduce the severe streaking artifacts induced by metallic implants and other high-density objects. Effective MAR generally requires both accurate artifact localization and artifact removal. Sinogram-domain methods can exploit explicit geometric cues, such as metal traces, to identify metal-corrupted measurements, while requiring raw projection data, which is often unavailable in clinical and practical scenarios. Image-domain methods are more flexible and widely applicable, yet they usually lack comparable geometric guidance, limiting their ability to localize artifacts and leading to suboptimal results. To address this limitation, we propose GraphMAR, a geometry-aware learning framework for explicit artifact identification and spatially adaptive MAR in the image domain. The key idea is to introduce graph-based geometric modeling as an image-domain analogue of sinogram metal traces. Specifically, we first construct a geometric graph from the metal mask and derive a geometric density graph that coarsely localizes artifact-prone regions according to inter-implant geometry. We then design GraphMoE, a graph-routed mixture-of-experts module that builds a polar-coordinate artifact graph in feature space and adaptively routes different experts to different spatial regions for MAR. By aligning the learned routing maps with the geometric density graph, GraphMAR provides explicit and interpretable artifact localization while enabling region-adaptive artifact reduction. Experiments on both simulated and real-world datasets demonstrate that GraphMAR achieves superior MAR performance compared with existing methods. To the best of our knowledge, this is the first work to introduce graph-based modeling for CT MAR and to enable explicit artifact identification in the image domain, improving both restoration quality and interpretability.
While contrastive multi-view clustering has achieved remarkable success, it implicitly assumes balanced class distribution. However, real-world multi-view data primarily exhibits class imbalance distribution. Consequently, existing methods suffer performance degradation due to their inability to perceive and model such imbalance. To address this challenge, we present the first systematic study of imbalanced multi-view clustering, focusing on two fundamental problems: i. perceiving class imbalance distribution, and ii. mitigating representation degradation of minority samples. We propose PROTOCOL, a novel PaRtial Optimal TranspOrt-enhanced COntrastive Learning framework for imbalanced multi-view clustering. First, for class imbalance perception, we map multi-view features into a consensus space and reformulate the imbalanced clustering as a partial optimal transport (POT) problem, augmented with progressive mass constraints and weighted KL divergence for class distributions. Second, we develop a POT-enhanced class-rebalanced contrastive learning at both feature and class levels, incorporating logit adjustment and class-sensitive learning to enhance minority sample representations. Extensive experiments demonstrate that PROTOCOL significantly improves clustering performance on imbalanced multi-view data, filling a critical research gap in this field.
Most conventional crowd counting methods utilize a fully-supervised learning framework to establish a mapping between scene images and crowd density maps. They usually rely on a large quantity of costly and time-intensive pixel-level annotations for training supervision. One way to mitigate the intensive labeling effort and improve counting accuracy is to leverage large amounts of unlabeled images. This is attributed to the inherent self-structural information and rank consistency within a single image, offering additional qualitative relation supervision during training. Contrary to earlier methods that utilized the rank relations at the original image level, we explore such rank-consistency relation within the latent feature spaces. This approach enables the incorporation of numerous pyramid partial orders, strengthening the model representation capability. A notable advantage is that it can also increase the utilization ratio of unlabeled samples. Specifically, we propose a Deep Rank-consist Ent pyrAmid Model (DREAM), which makes full use of rank consistency across coarse-to-fine pyramid features in latent spaces for enhanced crowd counting with massive unlabeled images. In addition, we have collected a new unlabeled crowd counting dataset, FUDAN-UCC, comprising 4000 images for training purposes. Extensive experiments on four benchmark datasets, namely UCF-QNRF, ShanghaiTech PartA and PartB, and UCF-CC-50, show the effectiveness of our method compared with previous semi-supervised methods. The codes are available at https://github.com/bridgeqiqi/DREAM.
Chest X-rays play a pivotal role in diagnosing respiratory diseases such as pneumonia, tuberculosis, and COVID-19, which are prevalent and present unique diagnostic challenges due to overlapping visual features and variability in image quality. Severe class imbalance and the complexity of medical images hinder automated analysis. This study leverages deep learning techniques, including transfer learning on pre-trained models (AlexNet, ResNet, and InceptionNet), to enhance disease detection and classification. By fine-tuning these models and incorporating focal loss to address class imbalance, significant performance improvements were achieved. Grad-CAM visualizations further enhance model interpretability, providing insights into clinically relevant regions influencing predictions. The InceptionV3 model, for instance, achieved a 28 findings highlight the potential of deep learning to improve diagnostic workflows and support clinical decision-making.
Background Identification of difficult laryngoscopy is a frequent demand in cervical spondylosis clinical surgery. This work aims to develop a hybrid architecture for identifying difficult laryngoscopy based on new indexes. Methods Initially, two new indexes for identifying difficult laryngoscopy are proposed, and their efficacy for predicting difficult laryngoscopy is compared to that of two conventional indexes. Second, a hybrid adaptive architecture with convolutional layers, spatial extraction, and a vision transformer is proposed for predicting difficult laryngoscopy. The proposed adaptive hybrid architecture is then optimized by determining the optimal location for extracting spatial information. Results The test accuracy of four indexes using simple model is 0.8320. The test accuracy of optimized hybrid architecture using four indexes is 0.8482. Conclusion The newly proposed two indexes, the angle between the lower margins of the second and sixth cervical spines and the vertical direction, are validated to be effective for recognizing difficult laryngoscopy. In addition, the optimized hybrid architecture employing four indexes demonstrates improved efficacy in detecting difficult laryngoscopy. Trial registration Ethics permission for this research was obtained from the Medical Scientific Research Ethics Committee of Peking University Third Hospital (IRB00006761-2015021) on 30 March 2015. A well-informed agreement has been received from all participants. Patients were enrolled in this research at the Chinese Clinical Trial Registry ( http://www.chictr.org.cn , identifier: ChiCTR-ROC-16008598) on 6 June 2016.
Medical image segmentation has been significantly advanced with the rapid development of deep learning (DL) techniques. Existing DL-based segmentation models are typically discriminative; i.e ., they aim to learn a mapping from the input image to segmentation masks. However, these discriminative methods neglect the underlying data distribution and intrinsic class characteristics, suffering from unstable feature space. In this work, we propose to complement discriminative segmentation methods with the knowledge of underlying data distribution from generative models. To that end, we propose a novel hybrid diffusion framework for medical image segmentation, termed HiDiff, which can synergize the strengths of existing discriminative segmentation models and new generative diffusion models. HiDiff comprises two key components: discriminative segmentor and diffusion refiner. First, we utilize any conventional trained segmentation models as discriminative segmentor, which can provide a segmentation mask prior for diffusion refiner. Second, we propose a novel binary Bernoulli diffusion model (BBDM) as the diffusion refiner, which can effectively, efficiently, and interactively refine the segmentation mask by modeling the underlying data distribution. Third, we train the segmentor and BBDM in an alternate-collaborative manner to mutually boost each other. Extensive experimental results on abdomen organ, brain tumor, polyps, and retinal vessels segmentation datasets, covering four widely-used modalities, demonstrate the superior performance of HiDiff over existing medical segmentation algorithms, including the state-of-the-art transformer- and diffusion-based ones. In addition, HiDiff excels at segmenting small objects and generalizing to new datasets. Source codes are made available at https://github.com/takimailto/HiDiff.
Mild cognitive impairment (MCI) is often at high risk of progression toAlzheimer's disease (AD). Existing works to identify the progressive MCI (pMCI)typically require MCI subtype labels, pMCI vs. stable MCI (sMCI), determined bywhether or not an MCI patient will progress to AD after a long follow-up.However, prospectively acquiring MCI subtype data is time-consuming andresource-intensive; the resultant small datasets could lead to severeoverfitting and difficulty in extracting discriminative information. Inspiredby that various longitudinal biomarkers and cognitive measurements present anordinal pathway on AD progression, we propose a novel Hybrid-granularityOrdinal PrototypE learning (HOPE) method to characterize AD ordinal progressionfor MCI progression prediction. First, HOPE learns an ordinal metric space thatenables progression prediction by prototype comparison. Second, HOPE leveragesa novel hybrid-granularity ordinal loss to learn the ordinal nature of AD viaeffectively integrating instance-to-instance ordinality, instance-to-classcompactness, and class-to-class separation. Third, to make the prototypelearning more stable, HOPE employs an exponential moving average strategy tolearn the global prototypes of NC and AD dynamically. Experimental results onthe internal ADNI and the external NACC datasets demonstrate the superiority ofthe proposed HOPE over existing state-of-the-art methods as well as itsinterpretability. Source code is made available athttps://github.com/thibault-wch/HOPE-for-mild-cognitive-impairment.
The explainability of deep neural networks (DNNs) is critical for trust and reliability in AI systems. Path-based attribution methods, such as integrated gradients (IG), aim to explain predictions by accumulating gradients along a path from a baseline to the target image. However, noise accumulated during this process can significantly distort the explanation. While existing methods primarily concentrate on finding alternative paths to circumvent noise, they overlook a critical issue: intermediate-step images frequently diverge from the distribution of training data, further intensifying the impact of noise. This work presents a novel Denoising Diffusion Path (DDPath) to tackle this challenge by harnessing the power of diffusionmodels for denoising. By exploiting the inherent ability of diffusion models to progressively remove noise from an image, DDPath constructs a piece-wise linear path. Each segment of this path ensures that samples drawn from a Gaussian distribution are centered around the target image. This approach facilitates a gradual reduction of noise along the path. We further demonstrate that DDPath adheres to essential axiomatic properties for attribution methods and can be seamlessly integrated with existing methods such as IG. Extensive experimental results demonstrate that DDPath can significantly reduce noise in the attributions—resulting in clearer explanations—and achieves better quantitative results than traditional path-based methods.
Image ordinal estimation is to estimate the ordinal label of a given image. Existing methods primarily rely on ordinal regression, mapping feature representations directly to ordinal labels. However, these methods often struggle to preserve the inherent order within the learned feature representations. To this end, this paper proposes learning intrinsic Consistent Ordinal REpresentations (CORE), a novel approach that learns intrinsic ordinal relationships directly from ground-truth labels. First, it constructs an ordinal manifold using an ordinal totally ordered set (toset) distribution (OTD), capturing the inherent order of labels while regularizing feature embeddings. Second, the CORE leverages the toset distribution to convert both feature representations and labels into a unified embedding space, enabling consistent manifold alignment. Third, CORE employs an ordinal prototype-constrained convex programming formulation with dual decomposition, minimizing the Kullback–Leibler (KL) divergence between the toset distributions of labels and feature representations. Extensive experiments demonstrate that CORE, when combined with existing deep ordinal regression methods, significantly improves their performance in preserving ordinal relationships and achieves superior quantitative results across four real-world scenarios.
Signal clipping is a well-established method employed in orthogonal frequency division multiplexing (OFDM) systems to mitigate peak-to-average power ratio. The utilization of this technique is widespread in electronic devices with limited power or resource capabilities due to its high efficiency and low complexity. While clipping effectively diminishes nonlinear distortion stemming from power amplifiers (PAs), it introduces additional distortion known as clipping distortion. The optimization of system performance, considering both clipping distortions and the nonlinearity of PAs, remains an unresolved challenge due to the intricate modeling of PAs. In this article, we undertake an analysis of PA nonlinearity utilizing the Bessel–Fourier PA model and simplify its power expression through intermodulation product analysis. We mathematically derive expressions for the receiver signal-to-noise ratio and system symbol error rate (SER) for nonlinear clipped OFDM systems. Using these derivations, we explore the optimal system configuration required to achieve the lower bound of SER in practical OFDM systems, taking into account both PA nonlinearity and clipping distortion. The results and methodologies presented in this article contribute to an improved comprehension of system-level optimization in nonlinear OFDM systems employing clipping technology.
This article reviews the deep learning methods for computed tomography image denoising and deblurring separately and simultaneously. Then, we discuss promising directions in this field, such as a combination with large-scale pretrained models and large language models. Currently, deep learning is revolutionizing medical imaging in a data-driven manner. With rapidly evolving learning paradigms, related algorithms and models are making rapid progress toward clinical applications.
Prompt learning has attracted broad attention in computer vision since the large pre-trained vision-language models (VLMs) exploded. Based on the close relationship between vision and language information built by VLM, prompt learning becomes a crucial technique in many important applications such as artificial intelligence generated content (AIGC). In this survey, we provide a progressive and comprehensive review of visual prompt learning as related to AIGC. We begin by introducing VLM, the foundation of visual prompt learning. Then, we review the vision prompt learning methods and prompt-guided generative models, and discuss how to improve the efficiency of adapting AIGC models to specific downstream tasks. Finally, we provide some promising research directions concerning prompt learning.
Signal clipping is one of the most efficient signal peak-to-average power ratio (PAPR) reduction schemes for MIMO-OFDM receivers. It significantly reduces the nonlinear distortion caused by the power amplifiers (PAs) in the transmitter but introduces clipping distortion at the same time. Although the optimization of PA nonlinear distortion or clipping distortion has been well studied in previous research, the joint system optimization with both the nonlinear distortion from the transmitter and clipping distortion from the receiver is still a blank due to the complex PA modeling. This paper simplified the PA model with inter-modulation product (IMP) analysis and derived the expression of the symbol error rate (SER) in nonlinear clipped MIMO-OFDM systems. Based on the derivation, we found the optimal system setting to reach SER lower bound when PA nonlinearity and clipping distortion are simultaneously considered. The functions and methodologies in this paper can be a useful reference for system-level optimization for clipped MIMO-OFDM systems.
Image restoration, which aims to retrieve and enhance degraded images, is fundamental across a wide range of applications. While conventional deep learning approaches have notably improved the image quality across various tasks, they still suffer from (i) the high storage cost needed for various task-specific models and (ii) the lack of interactivity and flexibility, hindering their wider application. Drawing inspiration from the pronounced success of prompts in both linguistic and visual domains, we propose novel Prompt-In-Prompt learning for universal image restoration, named PIP. First, we present two novel prompts, a degradation-aware prompt to encode high-level degradation knowledge and a basic restoration prompt to provide essential low-level information. Second, we devise a novel prompt-to-prompt interaction module to fuse these two prompts into a universal restoration prompt. Third, we introduce a selective prompt-to-feature interaction module to modulate the degradation-related feature. By doing so, the resultant PIP works as a plug-and-play module to enhance existing restoration models for universal image restoration. Extensive experimental results demonstrate the superior performance of PIP on multiple restoration tasks, including image denoising, deraining, dehazing, deblurring, and low-light enhancement. Remarkably, PIP is interpretable, flexible, efficient, and easy-to-use, showing promising potential for real-world applications. The code is available at https://github.com/longzilicart/pip_universal.
Lung nodule malignancy prediction has been enhanced by advanced deep-learning techniques and effective tricks. Nevertheless, current methods are mainly trained with cross-entropy loss using one-hot categorical labels, which results in difficulty in distinguishing those nodules with closer progression labels. Interestingly, we observe that clinical text information annotated by radiologists provides us with discriminative knowledge to identify challenging samples. Drawing on the capability of the contrastive language-image pre-training (CLIP) model to learn generalized visual representations from text annotations, in this paper, we propose CLIP-Lung, a textual knowledge-guided framework for lung nodule malignancy prediction. First, CLIP-Lung introduces both class and attribute annotations into the training of the lung nodule classifier without any additional overheads in inference. Second, we design a channel-wise conditional prompt (CCP) module to establish consistent relationships between learnable context prompts and specific feature maps. Third, we align image features with both class and attribute features via contrastive learning, rectifying false positives and false negatives in latent space. Experimental results on the benchmark LIDC-IDRI dataset demonstrate the superiority of CLIP-Lung, in both classification performance and interpretability of attention maps. Source code is available at https://github.com/ymLeiFDU/CLIP-Lung .
Since stroke is the main cause of various cerebrovascular diseases, deep learning-based stroke lesion segmentation on magnetic resonance (MR) images has attracted considerable attention. However, the existing methods often neglect the domain shift among MR images collected from different sites, which has limited performance improvement. To address this problem, we intend to change style information without affecting high-level semantics via adaptively changing the low-frequency amplitude components of the Fourier transform so as to enhance model robustness to varying domains. Thus, we propose a novel FAN-Net, a U-Net–based segmentation network incorporated with a Fourier-based adaptive normalization (FAN) and a domain classifier with a gradient reversal layer. The FAN module is tailored for learning adaptive affine parameters for the amplitude components of different domains, which can dynamically normalize the style information of source images. Then, the domain classifier provides domain-agnostic knowledge to endow FAN with strong domain generalizability. The experimental results on the ATLAS dataset, which consists of MR images from 9 sites, show the superior performance of the proposed FAN-Net compared with baseline methods.
Power amplifiers (PAs) lead to nonlinear distortion in index modulation assisted OFDM (OFDM-IM) systems when the peak-to-average power ratio (PAPR) is high. But PAPR reduction in OFDM-IM is more complicated than that in traditional OFDM owing to the frequent switch of subcarrier states between active and idle. When implementing OFDM-IM on resource-limited platforms like battery-powered Internet of Things (IoT) devices, PAPR cannot always be ideally eliminated due to the complexity of the reduction algorithm and the limitation of computation resources. Hence, performance optimization under non-negligible PAPR becomes an important issue. This paper performs symbol error rate (SER) optimization in OFDM-IM systems with high PAPR. First, we derive a low complexity expression of receiver signal-to-noise ratio (SNR) in nonlinear OFDM-IM systems. Second, we find that SER in OFDM-IM is affected by the nonlinear distortion in both constellation mapping and index estimation. With this observation, we derive a closed-form expression of SER and prove that there is a lower bound. Third, we do SER optimization and identify the optimal system operating point with respect to the minimum SER. Finally, we validate our analysis through simulations.
In this paper, a pruning method based on the least absolute shrinkage and selection operator (LASSO) regression is proposed for pruning the redundant terms of the generalized memory polynomial (GMP) model. As an advanced machine learning algorithm, the LASSO regression is also an effective option for behavioral models of power amplifiers (PAs). The GMP model increases the number of model coefficients significantly with increasing memory depth and non-linear order, which can lead to a significant increase in model complexity and the risk of overfitting. The pruning method takes advantage of the fact that the LASSO regression can prevent overfitting and yield sparse solutions by introducing the L1-norm as a penalty factor into the coefficient extraction process of the GMP model, which can directly reduce the number of coefficients in the model to prune it while maintaining modelling performance. Compared to other model pruning techniques, this pruning method allows for better control over the number of coefficients in the model, determining the appropriate model complexity and accuracy based on the actual application scenario. The experimental results show that the proposed method maintains good modeling performance with low model complexity, requiring only about 60% of the number of coefficients from the original GMP model.