Multiple instance learning (MIL) is a crucial paradigm addressing weakly supervised classification in histopathological images. However, existing MIL methods struggle to model tile/patch interactions, which can capture important contextual information. Moreover, MIL is limited by suboptimal bag embeddings, as traditional methods focus primarily on extracting distinct embeddings for individual instances rather than for the entire bag. These limitations restrict MIL’s ability to leverage contextual information and generate discriminative aggregated representations. To address these limitations, we propose Interactive Contrastive Multiple Instance Learning (ICMIL), a novel MIL framework that integrates graph learning (GL) and contrastive learning (CL) to enhance MIL’s contextual awareness and discriminative power. ICMIL introduces two key innovations: (i) Transformer-based Graph Attention (TransGAT) models comprehensive patch interactions by constructing fully connected graphs and generating long-range edge attention for information aggregation. This approach addresses the problem of limited patch interactions in MIL and boosts MIL’s contextual awareness. (ii) Reinforced Contrastive MIL (ReCMIL) refines the bag embedding space by selecting the most advantageous contrastive bag pairs using a policy network. ReCMIL addresses the problem of suboptimal bag embeddings, enabling MIL to generate discriminative aggregated embeddings. Experimental results demonstrate the superiority of our proposed ICMIL method over state-of-the-art approaches on four publicly available datasets, which span three anatomical sites and encompass five classification tasks (including both binary and multiclass classification). Furthermore, we extend ICMIL with features from a pretrained foundation model and achieve the best performance. Specifically, ICMIL achieves validation accuracies of 92.39% on CRC-DX, 86.86% on CRC-KR, 100.00% on BRACS (binary), 66.67% on BRACS (multiclass), and 96.04% on TCGA-Lung. These findings highlight the strong potential of integrating ICMIL with pretrained foundation models for histopathology image analysis. The code will be available at: https://github.com/JingjiaoLou/ICMIL.
Multiple instance learning (MIL) has proven effective in classifying whole slide images (WSIs), owing to its weakly supervised learning framework. However, existing MIL methods still face challenges, particularly over-fitting due to small sample sizes or limited WSIs (bags). Pseudo-bags enhance MIL's classification performance by increasing the number of training bags. However, these methods struggle with noisy labels, as positive patches often occupy small portions of tissue, and pseudo-bags are typically generated by random splitting. Additionally, they face difficulties with non-discriminative instance embeddings due to the lack of domain-specific feature extractors. To address these limitations, we propose Phenotype Clustering Reinforced Multiple Instance Learning (PCR-MIL), a novel MIL framework that integrates clusteringbased pseudo-bags to improve MIL's noise robustness and the discriminative power of instance embeddings. PCR-MIL introduces two key innovations: (i) Phenotype Clustering-based Feature Selection (PCFS) selects relevant instance embeddings for prediction. It clusters instances into phenotype-specific groups, assigns positive instances to each pseudo-bag, and then uses Grad-CAM to select the most relevant positive embeddings. This approach mitigates noisy label challenges and enhances MIL's robustness to noise; (ii) Reinforced Feature Extractor (RFE) uses reinforcement learning to train an extractor based on selected clean pseudobags instead of noisy samples. This approach improves the discriminative power of extracted instance embeddings and enhances the feature representation capabilities of MIL. Experimental results on the publicly available BRACS and CRC-DX datasets demonstrate that PCR-MIL outperforms state-of-the-art methods. The code is available at: https:// github.com/JingjiaoLou/PCR-MIL.
Semi-supervised medical image segmentation (SSMIS) uses consistency learning to regularize model training, which alleviates the burden of pixel-wise manual annotations. However, it often suffers from error supervision from low-quality pseudo labels. Vision-Language Model (VLM) has great potential to enhance pseudo labels by introducing text prompt guided multimodal supervision information. It nevertheless faces the cross-modal problem: the obtained messages tend to correspond to multiple targets. To address aforementioned problems, we propose a Dual Semantic Similarity-Supervised VLM (DuSSS) for SSMIS. Specifically, 1) a Dual Contrastive Learning (DCL) is designed to improve cross-modal semantic consistency by capturing intrinsic representations within each modality and semantic correlations across modalities. 2) To encourage the learning of multiple semantic correspondences, a Semantic Similarity-Supervision strategy (SSS) is proposed and injected into each contrastive learning process in DCL, supervising semantic similarity via the distribution-based uncertainty levels. Furthermore, a novel VLM-based SSMIS network is designed to compensate for the quality deficiencies of pseudo-labels. It utilizes the pretrained VLM to generate text prompt guided supervision information, refining the pseudo label for better consistency regularization. Experimental results demonstrate that our DuSSS achieves outstanding performance with Dice of 82.52%, 74.61% and 78.03% on three public datasets (QaTa-COV19, BM-Seg and MoNuSeg).
Low-quality pseudo labels pose a significant obstacle in semi-supervised medical image segmentation (SSMIS), impeding consistency learning on unlabeled data. Leveraging vision-language model (VLM) holds promise in ameliorating pseudo label quality by employing textual prompts to delineate segmentation regions, but it faces the challenge of cross-modal alignment uncertainty due to multiple correspondences (multiple images/texts tend to correspond to one text/image). Existing VLMs address this challenge by modeling semantics as distributions but such distributions lead to semantic degradation. To address these problems, we propose Alignment-Multiplicity Aware Vision-Language Model (AMVLM), a new VLM pre-training paradigm with two novel similarity metric strategies. (i) Cross-modal Similarity Supervision (CSS) proposes a probability distribution transformer to supervise similarity scores across fine-granularity semantics through measuring cross-modal distribution disparities, thus learning cross-modal multiple alignments. (ii) Intra-modal Contrastive Learning (ICL) takes into account the similarity metric of coarse-fine granularity information within each modality to encourage cross-modal semantic consistency. Furthermore, using the pretrained AMVLM, we propose a pioneering text-guided SSMIS network to compensate for the quality deficiencies of pseudo-labels. This network incorporates a text mask generator to produce multimodal supervision information, enhancing pseudo label quality and the model’s consistency learning. Extensive experimentation validates the efficacy of our AMVLM-driven SSMIS, showcasing superior performance across four publicly available datasets. The code will be available at: https://github.com/QingtaoPan/AMVLM.
Automatic movement analysis utilizing surveillance video is believed to be an important and convenient way for timely delirium detection in an Intensive Care Unit (ICU). However, video-based delirium movement detection (DMD) faces inherent challenges: 1) Irregular movements with large differences in the four limbs of a patient; 2) similar movements in delirium and normal situations, and large movement variations between patients experiencing delirium. To address the challenges 1, this paper proposes a Long-Short-View Aware Multi-Agent Reinforcement Learning (LS-MARL) method to identify the most representative movement snippet for DMD, considering that the global state provided by the long-view is important for guiding the agent's decision-making, but ignored in existing MARL methods. The proposed LS-MARL has two novel designs. First, a novel Teacher Auxiliary Policy (TAP) is developed for direction preperception of the representative movement snippet. Second, a new reward mechanism, Team Intrinsic Reward (TIR), is introduced to quantify the contribution of each agent. Experiments demonstrate that the proposed LS-MARL method outperforms state-of-the-art methods. Furthermore, to handle the second challenge, a new Self-Adjusting Ensemble Learning (SAEL) strategy is built to adaptively integrate an optimal classifier combination from multidomain features, which further improves the performance of classification tasks in the proposed LS-MARL method.
Background and ObjectiveRecent studies have shown that colorectal cancer (CRC) patients with microsatellite instability high (MSI-H) are more likely to benefit from immunotherapy. However, current MSI testing methods are not available for all patients due to the lack of available equipment and trained personnel, as well as the high cost of the assay. Here, we developed an improved deep learning model to predict MSI-H in CRC from whole slide images (WSIs).MethodsWe established the MSI-H prediction model based on two stages: tumor detection and MSI classification. Previous works applied fine-tuning strategy directly for tumor detection, but ignoring the challenge of vanishing gradient due to the large number of convolutional layers. We added auxiliary classifiers to intermediate layers of pre-trained models to help propagate gradients back through in an effective manner. To predict MSI status, we constructed a pair-wise learning model with a synergic network, named parameter partial sharing network (PPsNet), where partial parameters are shared among two deep convolutional neural networks (DCNNs). The proposed PPsNet contained fewer parameters and reduced the problem of intra-class variation and inter-class similarity. We validated the proposed model on a holdout test set and two external test sets.Results144 H&E-stained WSIs from 144 CRC patients (81 cases with MSI-H and 63 cases with MSI-L/MSS) were collected retrospectively from three hospitals. The experimental results indicate that deep supervision based fine-tuning almost outperforms training from scratch and utilizing fine-tuning directly. The proposed PPsNet always achieves better accuracy and area under the receiver operating characteristic curve (AUC) than other solutions with four different neural network architectures on validation. The proposed method finally achieves obvious improvements than other state-of-the-art methods on the validation dataset with an accuracy of 87.28% and AUC of 94.29%.ConclusionsThe proposed method can obviously increase model performance and our model yields better performance than other methods. Additionally, this work also demonstrates the feasibility of MSI-H prediction using digital pathology images based on deep learning in the Asian population. It is hoped that this model could serve as an auxiliary tool to identify CRC patients with MSI-H more time-saving and efficiently.
Background:According to the WHO, anemia is a highly prevalent disease, especially for patients in the emergency department. The pathophysiological mechanism by which anemia can affect facial characteristics, such as membrane pallor, has been proven to detect anemia with the help of deep learning technology. The quick prediction method for the patient in the emergency department is important to screen the anemic state and judge the necessity of blood transfusion treatment.Method:We trained a deep learning system to predict anemia using videos of 316 patients. All the videos were taken with the same portable pad in the ambient environment of the emergency department. The video extraction and face recognition methods were used to highlight the facial area for analysis. Accuracy and area under the curve were used to assess the performance of the machine learning system at the image level and the patient level.Results:Three tasks were applied for performance evaluation. The objective of Task 1 was to predict patients' anemic states [hemoglobin (Hb) <13 g/dl in men and Hb <12 g/dl in women]. The accuracy of the image level was 82.37%, the area under the curve (AUC) of the image level was 0.84, the accuracy of the patient level was 84.02%, the sensitivity of the patient level was 92.59%, and the specificity of the patient level was 69.23%. The objective of Task 2 was to predict mild anemia (Hb <9 g/dl). The accuracy of the image level was 68.37%, the AUC of the image level was 0.69, the accuracy of the patient level was 70.58%, the sensitivity was 73.52%, and the specificity was 67.64%. The aim of task 3 was to predict severe anemia (Hb <7 g/dl). The accuracy of the image level was 74.01%, the AUC of the image level was 0.82, the accuracy of the patient level was 68.42%, the sensitivity was 61.53%, and the specificity was 83.33%.Conclusion:The machine learning system could quickly and accurately predict the anemia of patients in the emergency department and aid in the treatment decision for urgent blood transfusion. It offers great clinical value and practical significance in expediting diagnosis, improving medical resource allocation, and providing appropriate treatment in the future.
Fetal brain segmentation from Magnetic Resonance (MR) images is a fundamental step in brain development study and early diagnosis. Although progress has been made, performance still needs to be improved especially for the images with motion artifacts (due to unpredictable fetal movement) and/or changes of magnetic field. In this paper, we propose a novel confidence-aware cascaded framework to accurately extract fetal brain from MR image. Different from the existing coarse-to-fine techniques, our two-stage strategy aims to segment brain region and simultaneously produce segmentation confidence for each slice in 3D MR image. Then, the image slices with high-confidence scores are leveraged to guide brain segmentation of low-confidence image slices, especially on the brain regions with blurred boundaries. Furthermore, a slice consistency loss is also proposed to enhance the relationship among the segmentations of adjacent slices. Experimental results on fetal brain MRI dataset show that our proposed model achieves superior performance, and outperforms several state-of-the-art methods.
PURPOSE:Leukemia is a lethal disease that is harmful to bone marrow and overall blood health. The classification of white blood cell images is crucial for leukemia diagnosis. The purpose of this study is to classify white blood cells by extracting discriminative information from cell segmentation and combining it with the fine-grained features. We propose a hybrid adversarial residual network with support vector machine (SVM), which utilizes the extracted features to improve the classification accuracy for human peripheral white cells. METHODS:Firstly, we segment the cell and nucleus by utilizing an adversarial residual network, which contains a segmentation network and a discriminator network. To extract features that can handle the inter-class consistency problem effectively, we introduce the adversarial residual network. Then, we utilize convolutional neural network (CNN) features and histogram of oriented gradient (HOG) features, which can extract discriminative features from images of segmented cell nuclei. To utilize the representative features fully, a discriminative network is introduced to deal with neighboring information at different scales. Finally, we combine the vectors of HOG features with those of CNN features and feed them into a linear SVM to classify white blood cells into six types. RESULTS:We used three methods to evaluate the effect of leukocyte classification based on 5000 leukocyte images acquired from a local hospital. The first approach is to use the CNN features as the input of SVM to classify leukocytes, which achieved 94.23% specificity, 95.10% sensitivity, and 94.41% accuracy. The use of the HOG features for SVM achieved 83.50% specificity, 87.50% sensitivity, and 85.00% accuracy. The use of combined CNN and HOG features achieved 94.57% specificity, 96.11% sensitivity, and 95.93% accuracy. CONCLUSIONS:We propose a novel hybrid adversarial-discriminative network for the classification of microscopic leukocyte images. It improves the accuracy of cell classification, reduces the difficulty and time pressure of doctors' work, and economizes the valuable time of doctors in daily clinical diagnosis.
MRI-based fetal brain age prediction is crucial for fetal brain development analysis and early diagnosis of congenital anomalies. The locations and directions of fetal brain are randomly variable and disturbed by adjacent organs, thus imposing great challenges to the fetal brain age prediction. To address this problem, we propose an effective framework based on a deformable convolutional neural network for fetal brain age prediction. Considering the fact of insufficient data, we introduce label distribution learning (LDL), which is able to deal with the small sample problem. We integrate the LDL information into our end-to-end network. Moreover, to fully utilize the complementary multi-view data of fetal brain MRI stacks, a multi-branch CNN is proposed to aggregate multi-view information. We evaluate our method on a fetal brain MRI dataset with 289 subjects and achieve promising age prediction performance.
Fetal brain extraction is one of the most essential steps for prenatal brain MRI reconstruction and analysis. However, due to the fetal movement within the womb, it is a challenging task to extract fetal brains from sparsely-acquired imaging stacks typically with motion artifacts. To address this problem, we propose an automatic brain extraction method for fetal magnetic resonance imaging (MRI) using multi-stage 2D U-Net with deep supervision (DS U-net). Specifically, we initially employ a coarse segmentation derived from DS U-net to define a 3D bounding box for localizing the position of the brain. The DS U-net is trained with deep supervision loss to acquire more powerful discrimination capability. Then, another DS U-net focuses on the extracted region to produce finer segmentation. The final segmentation results are obtained by performing refined segmentation. We validate the proposed method on 80 stacks of training images and 43 testing stacks. The experimental results demonstrate the precision and robustness of our method with the average Dice coefficient of 91.69%, outperforming the existing methods.
Abstract Xerostomia induced by radiotherapy is a common toxicity for head and neck carcinoma patients. In this study, the deformable image registration of planning computed tomography (CT) and weekly cone‐beam CT (CBCT) was used to override the Hounsfield unit value of CBCT, and the modified CBCT was introduced to estimate the radiation dose delivered during the course of treatment. Herein, the beams from each patient's treatment plan were applied to the modified CBCT to construct the weekly delivered dose. Then, weekly doses were summed together to obtain the accumulated dose. A total of 42 parotid glands (PGs) of 21 nasopharyngeal carcinoma patients were analyzed. Doses delivered to the parotid glands significantly increased compared with the planning doses. V20, V30, V40, Dmean, and D50 increased by 11.3%, 28.6%, 44.4%, 9.5%, and 8.4% respectively. Of the 21 patients included in the study, eight developed xerostomia and the remaining 13 did not. Both planning and delivered PG Dmean for all patients exceeded tolerance (26 Gy). Among the 21 patients, the planning dose and delivered dose of Dmean were 30.6 Gy and 33.6 Gy, respectively, for patients with xerostomia, and 26.3 Gy and 28.0 Gy, respectively, for patients without xerostomia. The D50 of the planning and delivered dose for patients was below tolerance (30 Gy). The results demonstrated that the p‐value of V20, V30, D50, and Dmean difference of the delivery dose between patients with xerostomia and patients without xerostomia was less than 0.05. However, for the planning dose, the significant dosimetric difference between the two groups only existed in D50 and Dmean. Xerostomia is closely related to V20, V30, D50, and Dmean.
近年来,大数据环境下幂级式增长的海量训练样本为癌症的诊断带来了数据资源,同时互联网的发展促进了深度学习开源框架的应用水平,推动了图像数据的精细化自动分类进入深度挖掘阶段.基于深度学习量化的核特征和派生特征可解决肿瘤细胞样本分类问题,因此肿瘤细胞病理学的研究为癌症的早期筛查和准确诊断提供条件.如何学习出更高层次的可视化特征网络模型,以及如何习得快速高效特异性强的新学习方法,需要高判别性、高稳定性及较好鲁棒性的肿瘤细胞自动分类学习算法应于临床诊断治疗中.
Objective To investigate the effects of numerous re-planning strategies on the anatomic and dosimetric outcomes of target volume and organs at risk (OARs) in patients with head and neck cancer receiving fractionated radiotherapy.Methods From 2015 to 2016,28 patients with head and neck cancer were enrolled in this study with Shandong Cancer Hospital,consisting of 19 patients with nasopharyngeal carcinoma, 4 patients with laryngocarcinoma, and 5 patients with carcinoma of the maxillary sinus.All of them received conventionally fractionated radiotherapy.Each patient had six weekly cone-beam CT (CBCT) scans, which were performed on the first day of every week, to obtain reference images.A virtual CT image was generated by registration of planning CT and each weekly CBCT image.The four re-planning strategies were used for the reconstruction of re-planned dose, while the initial planning was used as a reference.The weekly doses calculated using virtual CT were summed together to obtain the actual dose.The actual and initial planned doses were evaluated.The nonparametric Friedman test was used to evaluate the differences between multiple groups, and the differences between any two groups were analyzed by paired t test.Results The sizes of planning target volume, clinical target volume, and left/right parotid glands (PGs) changed significantly within the six weeks (P=0.041, 0.046, 0.024, and 0.017, respectively).For these four re-planning strategies, there were significant differences between the actual dose and the initial planned dose to the PGs (all P<0.05), with average values decreased by 5.02%, 11.17%, 12.08%, and 13.19%, respectively, compared with that in the reference strategy.Conclusions Re-planning during treatment course could ensure the sparing of OARs and allow for sufficient dose to the target volume.The higher the number of re-planning strategies, the more the actual dose is close to the initial planed dose;the efficiency of two re-planning strategies is the highest.