Diffusion models generalize well in practice. However, an optimal diffusion model fully memorizes the training data and therefore fails to generalize, raising the question of what induces generalization in a real diffusion model. We show that, despite generalizing at the sample level, diffusion models progressively overfit the denoising training objective and thereby create a generalization gap between the performance on validation and training samples. This gap is most pronounced at intermediate noise levels. Using a fully analytic error-prone toy model, we trace the factors affecting the generalization gap. We find that the optimal denoising flow field localizes sharply around training points, but the model error suppresses the exact recall of training points, yielding a smooth, generalizing flow field. Finally, we find that the generalization gap observed in training does not translate to inference, which would result in a strong similarity between generated samples and training samples. This is because the intermediate states of sampling trajectories are sufficiently far from the distribution of noisy training samples the model is trained on. Together, these findings reveal a novel picture of how diffusion models generalize: the flow field generalizes through model error, which moves sampling trajectories outside the domain of noisy training samples and thereby naturally prevents overfitting.
Large-scale visual out-of-distribution (OOD) detection has witnessed remarkable progress by leveraging vision-language models such as CLIP. However, a significant limitation of current methods is their reliance on a pre-defined set of in-distribution (ID) ground-truth label names (positives). These fixed label names can be unavailable, unreliable at scale, or become less relevant due to in-distribution shifts after deployment. Towards truly unsupervised OOD detection, we utilize widely available text corpora for positive label mining, bypassing the need for positives. In this paper, we utilize widely available text corpora for positive label mining under a general concept mining paradigm. Within this framework, we propose ClusterMine, a novel positive label mining method. ClusterMine is the first method to achieve state-of-the-art OOD detection performance without access to positive labels. It extracts positive concepts from a large text corpus by combining visual-only sample consistency (via clustering) and zero-shot image-text consistency. Our experimental study reveals that ClusterMine is scalable across a plethora of CLIP models and achieves state-of-the-art robustness to covariate in-distribution shifts. The code is available at https://github.com/HHU-MMBS/clustermine_wacv_official.
Circadian clocks regulate biological activities, providing organisms with a fitness advantage under diurnal conditions by enabling anticipation and adaptation to recurring external changes. Three proteins, KaiA, KaiB, and KaiC, constitute the circadian clock in the cyanobacterial model Synechococcus elongatus PCC 7942. Several techniques established to measure circadian output in Synechococcus yielded comparably weak signals in Synechocystis sp. PCC 6803, a strain important for biotechnological applications. We applied an approach that does not require genetic modifications to monitor the circadian rhythms in Synechococcus and Synechocystis. We placed batch cultures in shake flasks on a sensor detecting backscattered light via noninvasive online measurements. Backscattering oscillated with a period of ∼24 h around the average growth. Wavelet and Fourier transformations are applied to determine the period's significance and length. In Synechocystis, oscillations fulfilled the circadian criteria of temperature compensation and entrainment by external stimuli. Remarkably, dilution alone synchronized oscillations. Western blotting revealed that the backscatter was ∼6.5 h phase-delayed in comparison to KaiC3 phosphorylation.
Diffusion-based image generation models can enhance image quality when conditioned on ground truth labels. Here, we conduct a comprehensive experimental study on image-level conditioning for diffusion models using cluster assignments. We investigate how individual clustering determinants, such as the number of clusters and the clustering method, impact image synthesis across three different datasets. Given the optimal number of clusters with respect to image synthesis, we show that cluster-conditioning can achieve state-of-the-art performance, with an FID of 1.67 for CIFAR10 and 2.17 for CIFAR100, along with a strong increase in training sample efficiency. We further propose a novel empirical method to estimate an upper bound for the optimal number of clusters. Unlike existing approaches, we find no significant association between clustering performance and the corresponding cluster-conditional FID scores. Code is available at https://github.com/HHU—MMBS/cedm—official—wavc2025
Aims Recurrent congestive episodes are a primary cause of hospitalizations in patients with heart failure. Hitherto, outpatient management adopts a reactive approach, assessing patients clinically through frequent follow-up visits to detect congestion early. This study aims to assess the capabilities of a self-supervised contrastive learning-derived risk index to detect episodes of acute decompensated heart failure (ADHF) in patients using continuously recorded wearable time-series data. Methods and results This is the protocol for a single-arm, prospective cohort pilot study that will include 290 patients with ADHF. Acute decompensated heart failure is diagnosed by clinical signs and symptoms, as well as additional diagnostics (e.g. NT-proBNP). Patients will receive standard-of-care treatment, supplemented by continuous wearable-based monitoring of vital signs and physical activity, and are followed for 90 days. During follow-up, study visits will be conducted and presentations without clinical ADHF will be referred to as ‘regular’ and data from these episodes will be presented to a deep neural network that is trained by a self-supervised contrastive learning objective to extract features from the time-series that are typical in regular periods. The model is used to calculate a risk index measuring the dissimilarity of observed features from those of regular periods. The primary outcome of this study will be the risk index’s accuracy in detecting episodes with ADHF. As secondary outcome data integrity and the score in the validated questionnaire System Usability Scale will be evaluated. Conclusion Demonstrating reliable congestion detection through continuous monitoring with a wearable and self-supervised contrastive learning could assist in pre-emptive heart failure management in clinical care. Clinical trial registration The study was registered in the German clinical trials register (DRKS00034502).
We present a comprehensive experimental study on pretrained feature extractors for visual out-of-distribution (OOD) detection, focusing on adapting contrastive language-image pretrained (CLIP) models. Without fine-tuning on the training data, we are able to establish a positive correlation ($R^2\geq0.92$) between in-distribution classification and unsupervised OOD detection for CLIP models in $4$ benchmarks. We further propose a new simple and scalable method called \textit{pseudo-label probing} (PLP) that adapts vision-language models for OOD detection. Given a set of label names of the training set, PLP trains a linear layer using the pseudo-labels derived from the text encoder of CLIP. To test the OOD detection robustness of pretrained models, we develop a novel feature-based adversarial OOD data manipulation approach to create adversarial samples. Intriguingly, we show that (i) PLP outperforms the previous state-of-the-art \citep{ming2022mcm} on all $5$ large-scale benchmarks based on ImageNet, specifically by an average AUROC gain of 3.4\% using the largest CLIP model (ViT-G), (ii) we show that linear probing outperforms fine-tuning by large margins for CLIP architectures (i.e. CLIP ViT-H achieves a mean gain of 7.3\% AUROC on average on all ImageNet-based benchmarks), and (iii) billion-parameter CLIP models still fail at detecting adversarially manipulated OOD images. The code and adversarially created datasets will be made publicly available.
RNA-Puzzles is a collective endeavor dedicated to the advancement and improvement of RNA three-dimensional structure prediction. With agreement from structural biologists, RNA structures are predicted by modeling groups before publication of the experimental structures. We report a large-scale set of predictions by 18 groups for 23 RNA-Puzzles: 4 RNA elements, 2 Aptamers, 4 Viral elements, 5 Ribozymes and 8 Riboswitches. We describe automatic assessment protocols for comparisons between prediction and experiment. Our analyses reveal some critical steps to be overcome to achieve good accuracy in modeling RNA structures: identification of helix-forming pairs and of non-Watson-Crick modules, correct coaxial stacking between helices and avoidance of entanglements. Three of the top four modeling groups in this round also ranked among the top four in the CASP15 contest. The results of the Fifth RNA-Puzzles contest highlights advances in RNA three-dimensional structure prediction and uncovers new insights into RNA folding and structure.
We present a Deep Learning approach to predict 3D folding structures of RNAs from their nucleic acid sequence. Our approach combines an autoregressive Deep Generative Model, Monte Carlo Tree Search, and a score model to find and rank the most likely folding structures for a given RNA sequence. We show that RNA de novo structure prediction by deep learning is possible at atom resolution, despite the low number of experimentally measured structures that can be used for training. We confirm the predictive power of our approach by achieving competitive results in a retrospective evaluation of the RNA-Puzzles prediction challenges, without using structural contact information from multiple sequence alignments or additional data from chemical probing experiments. Blind predictions for recent RNA-Puzzle challenges under the name "Dfold" further support the competitive performance of our approach.
Guidance is a widely used technique for diffusion models to enhance sample quality. Technically, guidance is realised by using an auxiliary model that generalises more broadly than the primary model. Using a 2D toy example, we first show that it is highly beneficial when the auxiliary model exhibits similar but stronger generalisation errors than the primary model. Based on this insight, we introduce masked sliding window guidance (M-SWG), a novel, training-free method. M-SWG upweights long-range spatial dependencies by guiding the primary model with itself by selectively restricting its receptive field. M-SWG requires neither access to model weights from previous iterations, additional training, nor class conditioning. M-SWG achieves a superior Inception score (IS) compared to previous state-of-the-art training-free approaches, without introducing sample oversaturation. In conjunction with existing guidance methods, M-SWG reaches state-of-the-art Frechet DINOv2 distance on ImageNet using EDM2-XXL and DiT-XL. The code is available at https://github.com/HHU-MMBS/swg_bmvc2025_official.
It has been found empirically that diffusion-based generative models strongly ben- efit from weighting the score-matching objective in the training process and from redirecting trajectories in the sampling process to closer match the training dis- tribution. Here we show that a beneficial loss weight arises naturally when the training objective is derived from first principles by enforcing detailed balance between the forward and the reverse diffusion trajectories. We find that deter- ministic sampling by diffusion models induces a strong bias, favoring features of some training examples while ignoring others. To correct for the strong sampling bias, we introduce an efficient and controllable rejection sampling approach. We achieve a new state-of-the-art FID of 1.42 for CIFAR-10 in a class-conditional setting.
Deep image clustering methods are typically evaluated on small-scale balanced classification datasets while feature-based k-means has been applied on proprietary billion-scale datasets. In this work, we explore the performance of feature-based deep clustering approaches on large-scale benchmarks whilst disentangling the impact of the following data-related factors: i) class imbalance, ii) class granularity, iii) easy-to-recognize classes, and iv) the ability to capture multiple classes. Consequently, we develop multiple new benchmarks based on ImageNet21K. Our experimental analysis reveals that feature-based k-means is often unfairly evaluated on balanced datasets. However, deep clustering methods outperform k-means across most large-scale benchmarks. Interestingly, k-means underperforms on easy-to-classify benchmarks by large margins. The performance gap, however, diminishes on the highest data regimes such as ImageNet21K. Finally, we find that non-primary cluster predictions capture meaningful classes (i.e. coarser classes).
Serious clinical complications (SCC; CTCAE grade ≥ 3) occur frequently in patients treated for hematological malignancies. Early diagnosis and treatment of SCC are essential to improve outcomes. Here we report a deep learning model-derived SCC-Score to detect and predict SCC from time-series data recorded continuously by a medical wearable. In this single-arm, single-center, observational cohort study, vital signs and physical activity were recorded with a wearable for 31,234 h in 79 patients (54 Inpatient Cohort (IC)/25 Outpatient Cohort (OC)). Hours with normal physical functioning without evidence of SCC (regular hours) were presented to a deep neural network that was trained by a self-supervised contrastive learning objective to extract features from the time series that are typical in regular periods. The model was used to calculate a SCC-Score that measures the dissimilarity to regular features. Detection and prediction performance of the SCC-Score was compared to clinical documentation of SCC (AUROC ± SD). In total 124 clinically documented SCC occurred in the IC, 16 in the OC. Detection of SCC was achieved in the IC with a sensitivity of 79.7% and specificity of 87.9%, with AUROC of 0.91 ± 0.01 (OC sensitivity 77.4%, specificity 81.8%, AUROC 0.87 ± 0.02). Prediction of infectious SCC was possible up to 2 days before clinical diagnosis (AUROC 0.90 at −24 h and 0.88 at −48 h). We provide proof of principle for the detection and prediction of SCC in patients treated for hematological malignancies using wearable data and a deep learning model. As a consequence, remote patient monitoring may enable pre-emptive complication management.
We present a general methodology that learns to classify images without labels by leveraging pretrained feature extractors. Our approach involves self-distillation training of clustering heads based on the fact that nearest neighbours in the pretrained feature space are likely to share the same label. We propose a novel objective that learns associations between image features by introducing a variant of pointwise mutual information together with instance weighting. We demonstrate that the proposed objective is able to attenuate the effect of false positive pairs while efficiently exploiting the structure in the pretrained feature space. As a result, we improve the clustering accuracy over $k$-means on $17$ different pretrained models by $6.1$\% and $12.2$\% on ImageNet and CIFAR100, respectively. Finally, using self-supervised vision transformers, we achieve a clustering accuracy of $61.6$\% on ImageNet. The code is available at https://github.com/HHU-MMBS/TEMI-official-BMVC2023.
Recent progress in computer-aided technologies has had a considerable impact on helping experts with a reliable and fast diagnosis of abnormal samples. In particular, self-supervised and self-distillation techniques have advanced automated out-of-distribution (OOD) detection in the image domain. Further improvements in OOD detection have been observed by including negative samples derived from shifting transformations of natural images. In this work, we study different ways of creating negative samples for medical images and how effective they are when leveraging them in a self-supervised self-distillation framework. We investigate the impact of various types of negative examples by applying different shifting transformations on samples when they are derived from in-distribution training data, an auxiliary dataset, or a combination of both. For the case of the auxiliary dataset, we compare the OOD detection performance when auxiliary samples are extracted from an in-domain or an out-domain. Our approach uses only data belonging to healthy people during the training procedure and does not require any additional information from labels. We demonstrate the efficiency of our technique by comparing abnormality detection performance on diverse medical datasets, setting new benchmarks for pneumonia, polyp, and glaucoma detection from X-ray, colonoscopy, and ophthalmology images.
The ability to design RNA molecules with specific structures and functions could facilitate research and developments in biotechnology, biology and pharmacy. Here we present a flexible RNA design framework based on deep learning that locally optimizes sequences by gradient-guided search methods. We demonstrate its effectiveness by designing bi-stable RNA molecules by superimposing conformer target structures.
X-ray images have been widely used for medical diagnoses of cardiothoracic and pulmonary abnormalities due to their noninvasiveness. Advancement in computer-aided diagnostic technologies, such as deep supervised methods, can help radiologists with a reliable early treatment and reduce diagnosis time. Nevertheless, these methods are prone to the small number of labeled samples and are limited to a specific abnormality. In this paper, we combined a selfsupervised contrastive method with a Mahalanobis distance score to develop an abnormality detection method that uses only healthy images during the training procedure. We were able to outperform previous unsupervised methods for the task of Pneumonia detection. We show that representation learned by the self-supervised method improves the supervised tasks for Pneumonia detection.
Detecting whether examples belong to a given in-distribution or are Out-Of-Distribution (OOD) requires identifying features specific to the in-distribution. In the absence of labels, these features can be learned by self-supervised techniques under the generic assumption that the most abstract features are those which are statistically most over-represented in comparison to other distributions from the same domain. In this work, we show that self-distillation of the in-distribution training set together with contrasting against negative examples derived from shifting transformation of auxiliary data strongly improves OOD detection. We find that this improvement depends on how the negative samples are generated. In particular, we observe that by leveraging negative samples, which keep the statistics of low-level features while changing the high-level semantics, higher average detection performance is obtained. Furthermore, good negative sampling strategies can be identified from the sensitivity of the OOD detection score. The efficiency of our approach is demonstrated across a diverse range of OOD detection problems, setting new benchmarks for unsupervised OOD detection in the visual domain.
Purpose: Different imaging sequences (T1 etc.) depict different aspects of a brain tumor. As clinical MRI examinations of the brain might be terminated prematurely, not all sequences may be acquired, decreasing the performance of automated tumor segmentation. We attempt to optimize the order of sequences, to maximize information gain in case of incomplete examination. Methods: For segmentation we used the winner algorithm of the Brain Tumor Segmentation challenge 2018, trained on the BraTS 2020 dataset, with the objective to segment necrotic core, peritumoral edema, and enhancing tumor. We compared the segmentation performance for all combinations of sequences, using the Dice score (DS) as the primary metric. We compare the results with those which would be obtained by attempting to follow the consensus recommendations for brain tumor imaging [T1, FLAIR, T2, T1CE]. Results: The average segmentation accuracy varies between 0.476 for T1 only and 0.751 for the full set of sequences. T1CE has a high information content, even regarding peritumoral edema and information of T2 and FLAIR were highly redundant. The optimal order of sequences appears to be [T1, T2, T1CE, FLAIR]. Comparing segmentation accuracy after each fully acquired sequence, the first sequence (T1) is the same for both, DS for [T1, T2] (proposed) is 6.2% higher than [T1, FLAIR] (aborted recommendations), and [T1, T2, T1CE] (proposed) is 34.8% higher than [T1, FLAIR, T2] (aborted recommendations). Conclusion: For the purpose of optimal deep-learning-based segmentation purposes in potentially incomplete MRI examinations, the T1CE sequence should be acquired as early as possible.
PURPOSE Intensive treatment protocols for aggressive hematologic malignancies harbor a high risk of serious clinical complications, such as infections. Current techniques of monitoring vital signs to detect such complications are cumbersome and often fail to diagnose them early. Continuous monitoring of vital signs and physical activity by means of an upper arm medical wearable allowing 24/7 streaming of such parameters may be a promising alternative. METHODS This single-arm, single-center observational trial evaluated symptom-related patient-reported outcomes and feasibility of a wearable-based remote patient monitoring. All wearable data were reviewed retrospectively and were not available to the patient or clinical staff. A total of 79 patients (54 inpatients and 25 outpatients) participated and received standard-of-care treatment for a hematologic malignancy. In addition, the wearable was continuously worn and self-managed by the patient to record multiple parameters such as heart rate, oxygen saturation, and physical activity. RESULTS Fifty-one patients (94.4%) in the inpatient cohort and 16 (64.0%) in the outpatient cohort reported gastrointestinal symptoms (diarrhea, nausea, and emesis), pain, dyspnea, or shivering in at least one visit. With the wearable, vital signs and physical activity were recorded for a total of 1,304.8 days. Recordings accounted for 78.0% (63.0-88.5; median [interquartile range]) of the potential recording time for the inpatient cohort and 84.6% (76.3-90.2) for the outpatient cohort. Adherence to the wearable was comparable in both cohorts, but decreased moderately over time during the trial. CONCLUSION A high adherence to the wearable was observed in patients on intensive treatment protocols for a hematologic malignancy who experience high symptom burden. Remote patient monitoring of vital signs and physical activity was demonstrated to be feasible and of primarily sufficient quality.