Diffusion bridges have recently shown promise for multi-contrast MRI translation by explicitly modeling transformations between source and target modalities. However, existing methods independently predict each timestep without effectively utilizing intermediate predictions, causing temporal inconsistency and reduced structural fidelity. To address this, we propose dual conditioning and trajectory consistency strategies. Dual conditioning jointly leverages external conditioning from the source modality and self-conditioning from intermediate predictions via a two-pass forward scheme, while trajectory consistency further improves temporal coherence across adjacent timesteps via an on-policy replay mechanism. Experiments on IXI and BraTS datasets demonstrate that our method achieves consistent improvements in structural fidelity and effectively reduces inference drift.
Real-world data are typically long-tailed, causing neural networks to over-fit head classes and underperform on rare tails. We propose Dual-Level Balanced Learning (DBL), an efficient training framework that balances gradients at both the class and instance levels. DBL combines Class-aware Balancing (CB), which corrects class-level imbalance by re-weighting gradients according to prediction bias; Instance-aware Balancing (IB), which alleviates instance-level imbalance by emphasising the learning of hard examples; and a lightweight Cross-Level Collaboration (CC) scheme that harmonises the two losses. By jointly addressing class-and instance-level imbalance, DBL delivers consistent gains across all classes and most individual samples. Extensive experiments on CIFAR10/100-LT, ImageNet-LT, Places-LT, and iNaturalist18 show that DBL sets new state-of-the-art accuracy on all five benchmarks, confirming its robustness to severe long-tailed distributions.
Single domain generalization aims to train a model on a single source domain that generalizes to unseentarget domains, which is critical in multimedia applications. Current methods typically use adversarial data augmentation to enrich the source domain distribution with novel samples. However, these methods typically rely on labeled data and require adversarial training between generators and classifiers, which may limit sample diversity and introduce spurious correlations. To tackle these problems, we propose a method that integrates Contrastive clustering regularization with an Unsupervised Diversity Augmentation (UDA), termed C-UDA. Specifically, UDA is a flexible and general framework in which two customized models iteratively optimize a novel adversarial loss to enable fully unsupervised data augmentation. Within UDA, we design a lightweight generator that diversifies each input image along three distinct visual attributes. Based on both original and augmented images, we further introduce contrastive clustering regularization to encourage the model to learn domain-invariant representations, resulting in robust decision boundaries. Extensive experiments on four challenging benchmarks demonstrate that C-UDA significantly outperforms 22 state-of-the-art methods.
Recent advances in AI-computing networks (ACN) provide a timely backdrop for assessing deep-learning (DL) research in healthcare. This systematic review synthesizes DL applications in multiple sclerosis (MS) and evaluates their readiness for ACN-enabled deployment. A Web of Science Core Collection search (2014-2024) retrieved 438 records; 264 met stringent inclusion criteria. Bibliometric and knowledge-mapping analyses reveal steady growth in publications and citations, with MRIbased UNet variants dominating lesion-segmentation and diseaseclassification tasks. Emerging themes include gait-sensor analytics, longitudinal progression modelling, quantitative susceptibility mapping, and the growing use of transfer and federated learning to overcome data scarcity and privacy barriers. These resourceaware strategies signal a shift toward distributed training and inference paradigms that align naturally with ACN architectures. Nevertheless, few studies report multi-institutional experiments or network-level performance metrics, underscoring the need for tighter integration between DL methods and ACN infrastructure. We highlight research gaps-particularly in cross-site model orchestration and low-latency, on-device inference-that AIcomputing networks are well positioned to address, enabling scalable and interoperable DL services for MS diagnosis and prognosis.
Computerised recognition of low-resolution scene text images has been a persistent challenge. To improve the recognition performance, image quality enhancement via image super-resolution technology provides an intuitive solution. Typical deep learning-based scene text image super-resolution methods assume that the image quality degradation from high-resolution images to their corresponding low-resolution counterparts can be represented by mapping well-distributed samples, which limits their reconstruction performance in a practical text recognition system. For real-world scenarios this assumption typically does not hold since image degradations arise from multiple sources during image capture and processing. In this paper, to alleviate this problem, we propose a novel self-supervised end-to-end memory network model for scene text image super- resolution. In particular, after extracting enriched and finer representations from low-resolution text images via a spatial refinement block, we introduce a memory-based network to yield an improved super-resolution model that can handle complex degradation sources. Furthermore, to boost the effectiveness of our method, we design a multi-term loss to exploit textual structure information, where, in addition to the traditional reconstruction loss, we embed a character perceptual loss and a boundary enhancement loss. Extensive experiments on different datasets demonstrate that our proposed MNTSR method effectively improves the recognition accuracy for several scene text image recognition models and achieves state-of-the-art results. The source code is made available at https://github.com/xyzhu1/MNTSR.
Single domain generalization (single-DG) is a realistic yet challenging domain generalization scenario where a model trained on a single domain generalization scenario where a model trained on a single domain generalizes well to multiple unseen domains. Unlike typical single-DG methods that are essentially supervised data augmentation and focus mainly on the novelty of images, we propose a simple adversarial augmentation method, termed Progressive Diversity Generation (PDG), to synthesize novel and diverse images in a fully unsupervised manner. Specifically, PDG minimizes the uncertainty coefficient to ensure that synthesized images are novel. By modeling conditional probabilities with an auxiliary network, we transfer the adversarial process from semantics to images, thus eliminating dependency on labels. To enhance diversity, we propose the $f$ -diversity, a collection of correlation or similarity measures, to allow our model to generate potential images from diverse perspectives. The proposed architecture combines a multi-attribute generator with a progressive generation framework to improve model performance. PDG is the unsupervised and easy-to-implement method that solves single-DG with only synthesized (source) images. Extensive experiments on multiple single-DG benchmarks show that PDG achieves remarkable results and outperforms existing supervised and unsupervised methods by a large margin in single domain generalization. Source code and data are available: https://github.com/Ruiding1/PDG .
BackgroundIn the last decade, long-tail learning has become a popular research focus in deep learning applications in medicine. However, no scientometric reports have provided a systematic overview of this scientific field. We utilized bibliometric techniques to identify and analyze the literature on long-tailed learning in deep learning applications in medicine and investigate research trends, core authors, and core journals. We expanded our understanding of the primary components and principal methodologies of long-tail learning research in the medical field.MethodsWeb of Science was utilized to collect all articles on long-tailed learning in medicine published until December 2023. The suitability of all retrieved titles and abstracts was evaluated. For bibliometric analysis, all numerical data were extracted. CiteSpace was used to create clustered and visual knowledge graphs based on keywords.ResultsA total of 579 articles met the evaluation criteria. Over the last decade, the annual number of publications and citation frequency both showed significant growth, following a power-law and exponential trend, respectively. Noteworthy contributors to this field include Husanbir Singh Pannu, Fadi Thabtah, and Talha Mahboob Alam, while leading journals such as IEEE ACCESS, COMPUTERS IN BIOLOGY AND MEDICINE, IEEE TRANSACTIONS ON MEDICAL IMAGING, and COMPUTERIZED MEDICAL IMAGING AND GRAPHICS have emerged as pivotal platforms for disseminating research in this area. The core of long-tailed learning research within the medical domain is encapsulated in six principal themes: deep learning for imbalanced data, model optimization, neural networks in image analysis, data imbalance in health records, CNN in diagnostics and risk assessment, and genetic information in disease mechanisms.ConclusionThis study summarizes recent advancements in applying long-tail learning to deep learning in medicine through bibliometric analysis and visual knowledge graphs. It explains new trends, sources, core authors, journals, and research hotspots. Although this field has shown great promise in medical deep learning research, our findings will provide pertinent and valuable insights for future research and clinical practice.
Long-tailed learning has been a widespread issue for practical application research in deep neural networks (DNNs). However, a current status review of long-tailed learning in the medical profession needs to be revised. In this paper, we propose a method for analyzing the long-tailed learning literature in the medical area based on statistics, bibliometric analytic methodologies, and scientific knowledge graphs, employing a sample of 475 valid literature from the Web of Science database. Firstly, we use statistical methods to analyze the chronological distribution of publications and citations. Then, we use bibliometric methods to analyze the distribution of core authors and journals in long-tailed learning in medicine. Finally, we construct a scientific knowledge graph based on 381 keywords in medicine long-tailed learning; based on the visual analysis of the clustered knowledge graph and the study of related literature, we summarize six research hotspots of long-tailed learning-related research in the medical field, namely: breast cancer classification, brain lesion segmentation, research on other diseases such as lung cancer or skin diseases, research based on resampling method, and research based on function reweighting methods, and other advanced long-tailed learning methods.
Image classification has witnessed a remarkable advancement in class-balanced benchmarks. However, the natural distribution of datasets in real-world scenarios are long-tailed. Long-tailed classification has become a significant challenge in critical real-world image classification applications. A deep learning network trained on a long-tailed dataset tends to classify tail classes with few samples as head classes with many samples. The severe sample imbalance leads to the overwhelming dominance of negative samples on the tail classes; then, the massive gradient descent of negative samples leads to the classifier’s performance poorly. To tackle this problem, we propose a gradient re-balanced (GREB) loss with two synergistic factors, i.e., balance factor and correction factor. First, GREB estimates the balance and correction factors by accumulating the classifier outputs and their corresponding labels during the training process. Then, GREB dynamically reweights the gradients of positive and negative samples based on the balance factor to minimize the classification bias and improve the classifier performance. Finally, GREB compensates for sample gradients based on the correction factor to minimize the occurrence of misclassifications and improve the precision rate. Experiment results show that our GREB loss achieves state-of-the-art performance on long-tailed multi-label classification datasets (MSCOCO and MultiMNIST) and long-tailed single-label classification datasets (CIFAR10-LT and CIFAR100-LT).
Single domain generalization (SDG) is a realistic yet challenging domain generalization scenario that aims to generalize a model trained on a single domain to multiple unseen domains. Typical SDG methods are essentially supervised data augmentation strategies, which tend to enhance the novelty rather than the diversity of augmented samples. Insufficient diversity may jeopardize the model generalization ability. In this paper, we propose a novel adversarial method, termed Unsupervised Diversity Probe (UDP), to synthesize novel and diverse samples in fully unsupervised settings. More specifically, to ensure that samples are novel, we study SDG from an information-theoretic perspective that minimizes the uncertainty coefficients between synthesized and source samples. Considering that the variation in a single source domain is limited, we introduce a regularization imposed on the auxiliary module that synthesizes variable samples, incorporated with uncertainty coefficients in an adversarial manner to complement the diversity. Subsequently, an available region is utilized to guarantee the samples' safety. For the network architecture, we design a simple probe module that can synthesize samples in several different aspects. UDP is an unsupervised and easy-to-implement method that solves SDG using only synthetic (source) samples, thus reducing the dependence on task models. Extensive experiments on three benchmark datasets show that UDP achieves remarkable results and outperforms existing supervised and unsupervised methods by a large margin in single domain generalization.
Scene text image super-resolution (STISR) in the wild has been shown to be beneficial to support improved vision-based text recognition from low-resolution imagery. An intuitive way to enhance STISR performance is to explore the well-structured and repetitive layout characteristics of text and exploit these as prior knowledge to guide model convergence. In this paper, we propose a novel gradient-based graph attention method to embed patch-wise text layout contexts into image feature representations for high-resolution text image reconstruction in an implicit and elegant manner. We introduce a non-local group-wise attention module to extract text features which are then enhanced by a cascaded channel attention module and a novel gradient-based graph attention module in order to obtain more effective representations by exploring correlations of regional and local patch-wise text layout properties. Extensive experiments on the benchmark TextZoom dataset convincingly demonstrate that our method supports excellent text recognition and outperforms the current state-of-the-art in STISR. The source code is available at https://github.com/xyzhu1/TSAN.
We propose ComGAN, a simple unsupervised generative model, which simultaneously generates realistic images and high semantic masks under an adversarial loss and a binary regularization. In this paper, we first investigate two kinds of trivial solutions in the compositional generation process, and demonstrate their source is vanishing gradients on the mask. Then, we solve trivial solutions from the perspective of architecture. Furthermore, we redesign two fully unsupervised modules based on ComGAN (DS-ComGAN), where the disentanglement module associates the foreground, background and mask with three independent variables, and the segmentation module learns object segmentation. Experimental results show that (i) ComGAN's network architecture effectively avoids trivial solutions without any supervised information and regularization; (ii) DS-ComGAN achieves remarkable results and outperforms existing semi-supervised and weakly supervised methods by a large margin in both the image disentanglement and unsupervised segmentation tasks. It implies that the redesign of ComGAN is a possible direction for future unsupervised work.