Traditional Chinese Medicine (TCM) plays an important role in global medical practices. Syndrome differentiation (SD) is a key step in the diagnosis and treatment of TCM, which involves a comprehensive analysis of patient clinical information. However, the process of SD involves a complex mapping of various symptoms and signs to their corresponding syndrome types. It requires models to have strong non-linear feature modeling capabilities, emphasizing the semantic associations and feature differences between syndrome types. Additionally, the models must be able to effectively distinguish rare syndrome types, thereby enhancing both accuracy and interpretability. To this end, a multi knowledge enhanced framework combined with Kolmogorov-Arnold, named SD-MKEK, is proposed. SD-MKEK effectively captures the complex relationships between syndrome types and symptoms through a hierarchical structure, enabling accurate SD. In the feature extraction phase, multiple knowledge enhancement module is designed to extract context-sensitive features and significantly enhance the discriminability of the features through a label-guided mechanism. In the decision-making phase, a cross-attention mechanism is combined with the Kolmogorov-Arnold classifier, and a learnable activation function is used to better capture the complex relationships in high-dimensional data. Experimental results on the multi-disease multi-syndrome TCM-SD dataset show that the performance of SD-MKEK is superior to existing state-of-the-art baselines. Experiments on the single-disease multi-syndrome COPD-SD dataset also demonstrate the effectiveness of the proposed algorithm. This study can effectively perform the task of identifying TCM syndromes and has important value in promoting the deep integration of traditional medicine with modern computing technologies.
In this paper, we introduce agentic unlearning which removes specified information from both model parameters and persistent memory in agents with closed-loop interaction. Existing unlearning methods target parameters alone, leaving two critical gaps: (i) parameter-memory backflow, where retrieval reactivates parametric remnants or memory artifacts reintroduce sensitive content, and (ii) the absence of a unified strategy that covers both parameter and memory pathways. We present Synchronized Backflow Unlearning (SBU), a framework that unlearns jointly across parameter and memory pathways. The memory pathway performs dependency closure-based unlearning that prunes isolated entities while logically invalidating shared artifacts. The parameter pathway employs stochastic reference alignment to guide model outputs toward a high-entropy prior. These pathways are integrated via a synchronized dual-update protocol, forming a closed-loop mechanism where memory unlearning and parametric suppression reinforce each other to prevent cross-pathway recontamination. Experiments on medical QA benchmarks show that SBU reduces traces of targeted private information across both pathways with limited degradation on retained data.
Recently, polyethylene glycol modification has become key to improving biopharmaceutical pharmacokinetics and clinical applicability. This study aims to build a comprehensive analytical framework that integrates current status analysis, technology flow, and value assessment, in order to provide a step-by-step and thorough characterization of the patent landscape for PEG-modified drugs. Using the Derwent patent database, this study compiled 99,540 PEG-related patents worldwide from 2014 to 2023. Descriptive statistics, social network analysis, machine learning, and deep learning methods were applied to analyze these patents. The number of related patents increased dramatically over the past decade. China filed the most patents (24303) but exhibited a narrower technological breadth, while the United States led in numbers of inventors (86208) and assignees (37045). Patents from developed regions are more likely to be cited, and patent transfer activities mainly occur between commercial institutions. PEG-modified proteins and peptides represent the most commercially active category, highlighting their current market relevance. For patent transfer prediction, the XGBoost model achieved an average accuracy of 88.15
BACKGROUND:Accurate spinal MRI segmentation is essential for computer-aided diagnosis of spinal diseases. Existing methods have limitations in global semantic modeling and boundary delineation due to complex anatomy and imaging artifacts. PURPOSE:Our work aimed to propose a novel Graph-Guided Frequency-Enhanced State Space Network (GF-SSNet) method to achieve more accurate 3D multi-modal spine MRI automatic segmentation, addressing the limitations of existing algorithms in global semantic modeling of high-dimensional voxel space, cross-modal information synergistic perception, and fine boundary identification of anatomically similar tissues, thereby providing technical support for intelligent diagnosis and precision medicine of spinal diseases. METHODS:The proposed network is based on the GF-SSNet architecture. During encoding, a dual frequency-spatial feature enhancement mechanism is employed, which adaptively fuses local frequency dynamic features and global spatial long-range dependencies through Frequency Dynamic Convolution (FDConv) and Three-Directional Mamba-based state space model (TD-Mamba). At the bottleneck, Position-Aware Attention Fusion (PAAF) and Graph Convolutional Networks (GCN) are integrated to explicitly encode topological anatomical constraints between vertebrae, enhancing the global perception capability of spinal continuity structures. During decoding, a Depth-aware Progressive Upsampling (DAPU) strategy is introduced to effectively alleviate the reconstruction loss of fine-grained spatial information. The entire framework achieves end-to-end automatic segmentation of multi-modal MR images. RESULTS:On the normal test set, GF-SSNet outperformed all baselines across all metrics. Specifically, Dice and IoU Means reached 92.04 ± 0.06% and 85.29 ± 0.10%, exceeding the best baseline results of 89.81 ± 0.61% and 81.51 ± 0.91%. HD95 and ASSD were significantly reduced to 3.06 ± 0.46 mm and 0.612 ± 0.018 mm, compared to top-tier baseline values of 4.76 ± 1.09 mm and 1.14 ± 0.01 mm, respectively. On an independent pathological test set with various spinal pathologies, GF-SSNet maintained superior performance with Dice Mean of 87.60 ± 0.10%, still outperforming all baseline methods. The 4.4 percentage point performance decline from normal cases primarily stemmed from intervertebral disc segmentation challenges in degenerative conditions, while vertebrae segmentation remained robust. Ablation studies confirmed significant contributions of all proposed components. The proposed HFD-Tversky loss outperformed conventional losses. All performance differences were statistically significant after correction for multiple comparisons. CONCLUSION:GF-SSNet demonstrates performance in spinal segmentation through adaptive fusion of frequency features and global dependencies, providing technical support for intelligent spinal disease diagnosis.
Central lumbar spinal stenosis, a prevalent degenerative spinal disorder, severely impacts the quality of life for those affected. Axial and sagittal MRI images offer diverse information on tissue structure and lesions, which is crucial for accurate diagnosis. However, MRI-based diagnostic approaches still have poor lesion localization, insufficient cross-view alignment, underutilization of multi-view MRI information, and limited generalization across patient variability. To address these problems, we proposed an Encompassing Lumbar Central Spinal Stenosis Grading Model via Multi-view MRI Image Fusion called ELSG-MF. ELSG-MF consists of three stages: the first stage utilizes the extraction of robust pseudo-labels through a contrast-driven consistency reinforcement technique to guide Med-SAM in localizing and segmenting spinal tissue components. The Sagittal-Axial Pairing (SAP) Algorithm was developed by stage2 to integrate the spatial anatomical relationship between the vertebral body and the intervertebral disc, facilitating the correlation pairing between sagittal and axial images. Stage3 subsequently innovated the multi-view Adaptive Fusion (M2AF) module, which enables adaptive dynamic fusion of anatomical features across views. M2AF enhances the extraction of contextual complementary information, and significantly improves the model's capacity to detect subtle variations in the degree of narrowness. A series of studies show that our model achieves an overall accuracy of 0.8631, AUC of 0.96, and F1-score of 0.8614. These results indicate that our model substantially outperforms mainstream approaches, attaining superior segmentation and grading accuracy, exhibiting robust generalization and clinical application potential.
Semi-Supervised Learning (SSL) has emerged as a critical enabler for data-efficient diagnostic tasks in medical imaging. However, existing SSL approaches generally neglect inherent inter-class learning difficulty disparities, introducing fairness limitations. Such models often tend to overfit to the class that is easy to classify, perform poorly when dealing with challenging samples (such as minority/complex classes), and fail to fully utilize the valuable knowledge contained in low-confidence samples. To bridge these gaps, we propose ProFair: a proactive fairness prototype-driven framework with adaptive perception for SSL. It provides fairer learning opportunities for samples with different difficulty levels. Specifically, the difficulty-aware adaptive soft label mechanism dynamically generates class-specific thresholds based on learning difficulty and optimizes label distribution through prototype similarity to model inter-class differences. The dual-path validated negative pseudo-label strategy employs dual-path noise-resistant verification with probability distribution and feature space. This strategy converts low-confidence samples into high-value training signals, thereby clearly reinforcing the decision boundary. The multi-objective global constraint loss integrates various constraint terms to jointly optimize diagnostic accuracy and model robustness. Extensive experiments on multiple public medical image datasets validate the effectiveness and superiority of ProFair. This research provides a theoretically rigorous and clinically verifiable solution for class-imbalanced medical data analysis.
Precise automated segmentation of anatomical structures is a prerequisite for computer-aided diagnosis, radiotherapy planning, and quantitative medical analysis. However, existing models, whether based on convolutional neural networks (CNN) or transformer architectures, are primarily centered on the extraction and processing of spatial features. These approaches lead to spectral feature entanglement, where low-frequency global structures, mid-frequency contours, and high-frequency textures are indiscriminately mixed, degrading segmentation accuracy, particularly at object boundaries critical for clinical delineation. To address this, we introduce the FD-SSGNet, a framework that performs frequency disentanglement with State-Space gating. Our model first employs the Fast Fourier Transform (FFT) to explicitly decompose feature maps into low-, mid-, and high-frequency components. It then leverages the Shift Bidirectional Selective Gate Mamba (SBSGM), with parallel, heterogeneously configured pathways to effectively model long-range dependencies specific to each frequency band. Finally, a dynamic fusion module adaptively reintegrates the processed multi-band features to produce a refined segmentation map. Extensive experiments on the challenging BTCV multi-organ and ACDC cardiac segmentation datasets demonstrate that FD-SSGNet achieves new state-of-the-art performance, validating the significant benefits of explicit frequency domain modeling for robust and accurate medical image analysis in clinical workflows. Our implementation is available at https://github.com/singinghz/FD-SSGNet .
Chinese herbal slices (CHS) serve as the core carrier of Traditional Chinese Medicine, and their identification represents the primary step in quality control and medication safety. Because herbal slices exhibit a variety of morphological forms, some species share similar morphological traits, making precise identification a significant challenge. Existing algorithms have a problem: they can’t capture enough global dependencies, and their local feature discriminability is not sufficient in fine-grained CHS recognition. To address these problems, we propose SFE-CA, a unified fine-grained classification framework that synergistically integrates Coordinate Attention (CA) and Sequential Feature Enhancement (SFE). Unlike conventional methods that treat architecture and optimization separately, SFE-CA incorporates a customized dual-loss optimization strategy as an intrinsic component, designed to maximize the discriminative power of the SFE-extracted features. This holistic design allows the model to address the inter-class confusion issues that arise from highly similar target categories. Feature discrimination further improves through multi-task optimization with dual losses. On a 14-category herbal slice dataset, SFE-CA achieves significant performance, with 98.91
Pancreatic cancer is characterized by insidious onset and extremely poor prognosis. Accurate segmentation of the pancreas and its tumors is crucial for achieving precision medicine. The pancreas presents as a small, low-contrast target, while pancreatic tumors appear even smaller with highly irregular shapes. Existing methods, such as Convolutional Neural Networks, are limited by their small receptive fields and thus struggle to model long-range dependencies. Meanwhile, the fixed computational path of Transformers presents challenges for their adaptation to 3D pancreas and tumor segmentation. This study proposes the Global Awareness and Local Refinement (GALR) model, which addresses the challenges of long-range dependencies and multi-scale modeling, achieving precise 3D segmentation of the pancreas and its tumor boundaries. The model consists of two core innovative modules: The Multi-scale Dual Attention Mamba (MDAM) module dynamically captures the key information of the remodeling sequence through the Selective State Space Model, and combines the dilated convolution multi-scale fusion to enhance the multi-scale object recognition ability. The module integrates the dual attention mechanism to capture the cross-regional anatomical correlation, strengthens the segmentation sensitivity of low contrast regions, and realizes global perception of precision medical image segmentation. The Lightweight Grouped Convolutional Enhancement (LGCE) module uses a parameter-efficient linkage structure to construct a feature refinement flow for small target boundaries in the decoding stage, which effectively suppresses the interference of noise such as artifacts on the target contour. This module significantly improves the localization accuracy through the Local Refinement mechanism, and achieves accurate boundary contour segmentation. Experimental results on two datasets, CT and MRI, show that GALR has the best performance in pancreas and tumor segmentation tasks. On dataset 1, our method achieves Dice coefficients of 80.27 % and 51.72 % in pancreas and tumor segmentation tasks, respectively, which are 0.46 % and 1.18 % higher than the previous best model. The performance of other indicators is also excellent. The experiment on dataset 2 also leads to the same conclusion. The experimental results indicate that GALR provides an efficient clinical solution for computeraided diagnosis of pancreatic diseases.
During the global pandemic of coronavirus disease 2019 (COVID-19), mRNA vaccines have demonstrated great potential. In 2023, mRNA won the Nobel Prize in Physiology or Medicine. Currently, the global research and development of mRNA technology is accelerating. It is urgently necessary to use patent analysis to clarify the technological competitive landscape, providing a basis for technological innovation and industrial development in this field. This study is based on the Derwent patent database and employs social network analysis and patent quality assessment methods to conduct a quantitative analysis of mRNA therapeutic patents over the past 27 years. Deep learning and machine learning methods are used to predict future core technologies and patentees. The study found that mRNA drugs are currently primarily used in the fields of infectious diseases and cancer. Delivery technology remains one of the critical challenges, while targeted drug research and vector technology will be one of the key future directions for the field. Meanwhile, New organizations have developed novel delivery technologies to break through the patent thickets established by giant companies. The global landscape of mRNA therapy is undergoing a multifaceted developmental pattern, and the monopoly of giant companies is being challenged. The patent landscape of mRNA therapies constructed based on deep learning methods in this study can not only serve as a knowledge tool for comprehensive integration and use, but also inspire the development of efficient production methods for mRNA therapies.
Medical Large Vision-Language Models (MLVLMs) show encouraging results in medical diagnostics but easily in-herit biases from pretraining data, leading to bias catastrophic inheritance, where data biases persist and distort predictions. In this work, we present the first systematic study of this issue in MLVLMs, revealing how inherited biases affect classification and free-text reasoning tasks. We propose Logit Fairness Adjustment (LFA), a training-free debiasing method that operates at the logits level to recalibrate biased predictions. LFA quantifies bias by computing logit margins between valid and invalid medical images, applying logit smoothing when the margin is small to reduce overconfidence and bias compensation when the margin is large to reinforce valid features. We introduce the Medical Multimodal Bias Benchmark to assess bias severity across binary classification, multi-class classification, and free-text reasoning. Experiments on LLaVA-Med, SkinGPT-4, and Qwen-VL-7B show that LFA effectively mitigates bias for MLVLMs.
BackgroundStanford Type A aortic dissection (TAAD) is a life-threatening condition involving the ascending aorta and requires urgent surgery. This study developed 11 machine learning regression models to predict operative duration and identify key clinical factors influencing surgical time in TAAD.Materials and methodsIn this single-center retrospective cohort study of 505 patients who underwent surgery from December 2017 to March 2023. Specifically, 11 machine learning models were construct using 47 preoperative and intraoperative features to predict operative duration. Model performance was assessed by R2, RMSE, and MAE, and SHAP analysis enhanced interpretability.ResultsThe study primarily consisted of middle-aged patients, comprising 73.4% males and 26.6% females. Furthermore, most patients underwent complex aortic procedures under time-constrained preoperative conditions. Procedures involving root replacement and total arch replacement were associated with longer surgical durations. The ExtraTrees Regressor had the highest predictive accuracy. SHAP analysis revealed five key features: Duration of extracorporeal circulation, Duration of aortic occlusion, Intraoperative blood transfusion, Treatment method for the aortic arch, and Treatment method for the aortic root.ConclusionThis study developed high-performance predictive models to identify key features affecting operative duration in TAAD surgery. Complex reconstructions prolong procedures, and longer aortic occlusion further contributes to this effect. The findings highlight the major influence of surgical strategies and intraoperative management on surgical duration. Special consideration remains warranted for specific patient subgroups.
Objective: Large Language models (LLMs) have a wide range of medical applications, especially in scenarios such as question-answering. However, existing models face the challenge of accurately assessing the quality of information when generating medical information, which may lead to the inability to effectively distinguish beneficial and harmful information, thus affecting the quality of question-answering. This study aims to improve the information quality and practicability of medical question-answering. Methods: This study proposes MedicalGLM, a fine-tuning model based on a quality evaluation mechanism. Specifically, MedicalGLM contains a reward model for assessing the quality of medical QA. It adjusts its training process by returning the assessment scores to the QA model as penalties through a quality score loss function. Results: The experimental results indicate that MedicalGLM achieved the highest scores among the evaluated models in the Rouge-1, Rouge-2, Rouge-L, and BLEU metrics, with values of 54.90, 28.02, 44.50, and 32.61, respectively. Its proficiency in generating responses for the pediatric medical quiz task is notably superior to other prevailing LLMs in the medical domain. Conclusion: MedicalGLM significantly improves the quality and practicability of the generated information of the medical question-answering model by introducing a quality evaluation mechanism, which provides an effective improvement idea for researching medical large language models. Our code and model are publicly available for further research on https://github.com/wangxinwwang/MedicalGLM.
BACKGROUND:Myocardial pathology (scar and edema) segmentation plays a crucial role in the diagnosis, treatment, and prognosis of myocardial infarction (MI). However, the current mainstream models for myocardial pathology segmentation have the following limitations when faced with cardiac magnetic resonance(CMR) images with multiple objects and large changes in object scale: the remote modeling ability of convolutional neural networks is insufficient, and the computational complexity of transformers is high, which makes myocardial pathology segmentation challenging. PURPOSE:This study aims to develop a novel model to address the image characteristics and algorithmic challenges faced in the myocardial pathology segmentation task and improve the accuracy and efficiency of myocardial pathology segmentation. METHODS:We developed a novel visual state space (VSS)-based deep neural network, MPS-Mamba. In order to accurately and adequately extract CMR image features, the encoder employs a dual-branch structure to extract global and local features of the image. Among them, the VSS branch overcomes the limitations of the current mainstream models for myocardial pathology segmentation by modeling remote relationships through linear computability, while the convolutional-based branch provides complementary local information. Given the unique properties of the dual branches, we design a modular dual-branch fusion module for fusing dual branches to enhance the feature representation of the dual encoder. To improve the ability to model objects of different scales in cardiac magnetic resonance (CMR) images, a multi-scale feature fusion (MSF) module is designed to achieve effective integration and fine expression of multi-scale information. To further incorporate anatomical knowledge to optimize segmentation results, a decoder with three decoding branches is designed to output segmentation results of scar, edema, and myocardium, respectively. In addition, multiple sets of constraint functions are used to not only improve the segmentation accuracy of myocardial pathology but also effectively model the spatial position relationship between myocardium, scar, and edema. RESULTS:The proposed method was comprehensively evaluated on the MyoPS 2020 dataset, and the results showed that MPS-Mamba achieved an average Dice score of 0.717 ± $\pm$ 0.169 in myocardial scar segmentation, which is superior to the current mainstream methods. In addition, MPS-Mamba also performed well in the edema segmentation task, with an average Dice score of 0.735 ± $\pm$ 0.073. The experimental results further demonstrate the effectiveness of MPS-Mamba in segmenting myocardial pathologies in multi-sequence CMR images, verifying its advantages in myocardial pathology segmentation tasks. CONCLUSIONS:Given the effectiveness and superiority of MPS-Mamba, this method is expected to become a potential myocardial pathology segmentation tool that can effectively assist clinical diagnosis.
As the demand for high-resolution medical images increases, super-resolution (SR) technology becomes particularly important. In recent years, SR technology based on deep learning has achieved remarkable achievements, and its application in medical images is also growing. Since brain MRI is prone to artifacts during long-term scanning, SR technology has become an effective means to improve image clarity. However, traditional SR methods are usually computationally complex and time-consuming, making them unsuitable for real-time applications. To solve this problem, this paper proposes a lightweight SR model with BSRN as the backbone network and combined with structural re-parameterization to achieve lightweight and efficient SR. The model uses a multi-branch structure during training and integrates the multiple branches into a 3×3 convolution during inference, effectively retaining important feature information. At the same time, the computational complexity and storage requirements are significantly reduced. Through experimental verification on the IXI dataset, this method shows excellent super-resolution reconstruction effects, especially when processing noisy and blurred images, and can effectively improve image clarity and details. Research results show that this method improves model performance and has good application potential, providing new ideas for future medical image processing technology development.
Objective: Large Language models (LLMs) have a wide range of medical applications, especially in scenarios such as question-answering. However, existing models face the challenge of accurately assessing the quality of information when generating medical information, which may lead to the inability to effectively distinguish beneficial and harmful information, thus affecting the quality of question-answering. This study aims to improve the information quality and practicability of medical question-answering. Methods: This study proposes MedicalGLM, a fine-tuning model based on a quality evaluation mechanism. Specifically, MedicalGLM contains a reward model for assessing the quality of medical QA. It adjusts its training process by returning the assessment scores to the QA model as penalties through a quality score loss function. Results: The experimental results indicate that MedicalGLM achieved the highest scores among the evaluated models in the Rouge-1, Rouge-2, Rouge-L, and BLEU metrics, with values of 54.90, 28.02, 44.50, and 32.61, respectively. Its proficiency in generating responses for the pediatric medical quiz task is notably superior to other prevailing LLMs in the medical domain. Conclusion: MedicalGLM significantly improves the quality and practicability of the generated information of the medical question-answering model by introducing a quality evaluation mechanism, which provides an effective improvement idea for researching medical large language models. Our code and model are publicly available for further research on https://github.com/wangxinwwang/MedicalGLM.
When analyzing screening chest X-ray (CXR) images, radiologists can naturally consider information about the relationships between pathologies, namely pathological co-occurrence, and interdependence. Such topologically meaningful pathological relationships provide complementary a priori knowledge and can improve the radiologist’s classification accuracy. Unfortunately, most existing deep learning systems focus only on regression from input to binary labels, lack the ability to jointly analyze and integrate global and local information from these multiple labels, and the ability to exploit this valuable prior knowledge of topology. We suppose that the algorithm will distinguish co-occurring pathologies easily if it is guided to learn potentially valuable information between multiple pathologies. We propose a novel multi-label global–local analysis method named CXR × MLAGCPL for disease recognition in CXR images. We first designed two methods for constructing graphs: building adaptive local graph on a single image and creating global co-occurrence graph in a data-driven manner. Then two modules, local awareness module (LAM) and global co-occurrence priori learning module (GCPL), were designed to jointly learn local class correlation and global inter-class dependency patterns. We have evaluated the effectiveness of the proposed CXR × MLAGCPL in two large-scale CXR datasets (ChestX-Ray14 and CheXpert). Our model achieved state-of-the-art performance: the mean AUC scores for the 14 pathologies were 0.828 and 0.813, respectively. The proposed LAM and GCPL can jointly effectively improve the recognition performance. Our model is accurate in terms of classification accuracy and generalization due to competing approaches.
In light of the situation and the characteristics of Omicron, the country has continuously optimized the rules for the prevention and control of COVID-19. The global epidemic is still spreading, and new cases of infection continue to emerge in China. To facilitate the infected person to estimate the course of virus infection, a prediction model for predicting negative conversion time is proposed in this article. The clinical features of Omicron-infected patients in Shandong Province in the first half of 2022 are retrospectively studied. These features are grouped by disease diagnosis result, clinical sign, traditional Chinese medicine symptoms, and drug use. These features are input to the eXtreme Gradient Boosting (XGBoost) model, and the output is the predicted number of negative conversion days. At the same time, XGBoost is used as the underlying algorithm of the conformal prediction (CP) framework, which can realize the probability interval estimation with a controllable error rate. The results show that the proposed model has a mean absolute error of 3.54 days and has the shortest interval prediction result. This shows that the method in this paper can carry more decision-making information and help people better understand the disease and self-estimate the course of the disease to a certain extent.
Precise segmentation for skin cancer lesions at different stages is conducive to early detection and further treatment. Considering the huge cost of obtaining pixel-perfect annotations for this task, segmentation using less expensive image-level labels has become a research direction. Most image-level label weakly supervised segmentation uses class activation mapping (CAM) methods. A common consequence of this method is incomplete foreground segmentation, insufficient segmentation, or false negatives. At the same time, when performing weakly supervised segmentation of skin cancer lesions, ulcers, redness, and swelling may appear near the segmented areas of individual disease categories. This co-occurrence problem affects the model’s accuracy in segmenting class-related tissue boundaries to a certain extent. The above two issues are determined by the loosely constrained nature of image-level labels that penalize the entire image space. Therefore, providing pixel-level constraints for weak supervision of image-level labels is the key to improving performance. To solve the above problems, this paper proposes a joint unsupervised constraint-assisted weakly supervised segmentation model(UCA-WSS). The weakly supervised part of the model adopts a dual-branch adversarial erasure mechanism to generate higher-quality CAM. The unsupervised part uses contrastive learning and clustering algorithms to generate foreground labels and fine boundary labels to assist segmentation and solve common co-occurrence problems in weakly supervised skin cancer lesion segmentation through unsupervised constraints. The model proposed in the article is evaluated comparatively with other related models on some public dermatology data sets. Experimental results show that our model performs better on the skin cancer segmentation task than other weakly supervised segmentation models, showing the potential of combining unsupervised constraint methods on weakly supervised segmentation.
Objective. Bladder cancer is a common malignant urinary carcinoma, with muscle-invasive and non-muscle-invasive as its two major subtypes. This paper aims to achieve automated bladder cancer invasiveness localization and classification based on MRI.Approach. Different from previous efforts that segment bladder wall and tumor, we propose a novel end-to-end multi-scale multi-task spatial feature encoder network (MM-SFENet) for locating and classifying bladder cancer, according to the classification criteria of the spatial relationship between the tumor and bladder wall. First, we built a backbone with residual blocks to distinguish bladder wall and tumor; then, a spatial feature encoder is designed to encode the multi-level features of the backbone to learn the criteria.Main Results. We substitute Smooth-L1 Loss with IoU Loss for multi-task learning, to improve the accuracy of the classification task. By learning two datasets collected from bladder cancer patients at the hospital, the mAP, IoU, Acc, Sen and Spec are used as the evaluation metrics. The experimental result could reach 93.34%, 83.16%, 85.65%, 81.51%, 89.23% on test set1 and 80.21%, 75.43%, 79.52%, 71.87%, 77.86% on test set2.Significance. The experimental result demonstrates the effectiveness of the proposed MM-SFENet on the localization and classification of bladder cancer. It may provide an effective supplementary diagnosis method for bladder cancer staging.