RATIONALE AND OBJECTIVE:Stroke-associated pneumonia (SAP) often appears as a complication following intracerebral hemorrhage (ICH), leading to poor prognosis and increased mortality rates. Previous studies have typically developed prediction models based on clinical data alone, without considering that ICH patients often undergo CT scans immediately upon admission. As a result, these models are subjective and lack real-time applicability, with low accuracy that does not meet clinical needs. Therefore, there is an urgent need for a quick and reliable model to timely predict SAP. METHODS:In this retrospective study, we developed an image-based model (DeepSAP) using brain CT scans from 244 ICH patients to classify the presence and severity of SAP. First, DeepSAP employs MRI-template-based image registration technology to eliminate structural differences between samples, achieving statistical quantification and spatial standardization of cerebral hemorrhage. Subsequently, the processed images and filtered clinical data were simultaneously input into a deep-learning neural network for training and analysis. The model was tested on a test set to evaluate diagnostic performance, including accuracy, specificity, and sensitivity. RESULTS:Brain CT scans from 244 ICH patients (mean age, 60.24; 66 female) were divided into a training set (n = 170) and a test set (n = 74). The cohort included 143 SAP patients, accounting for 58.6% of the total, with 66 cases classified as moderate or above, representing 27% of the total. Experimental results showed an AUC of 0.93, an accuracy of 0.84, a sensitivity of 0.79, and a precision of 0.95 for classifying the presence of SAP. In comparison, the model relying solely on clinical data showed an AUC of only 0.76, while the radiomics method had an AUC of 0.74. Additionally, DeepSAP achieved an optimal AUC of 0.84 for the SAP grading task. CONCLUSION:DeepSAP's accuracy in predicting SAP stems from its spatial normalization and statistical quantification of the ICH region. DeepSAP is expected to be an effective tool for predicting and grading SAP in clinic.
Landmark detection using X-ray imaging is vital for disease screening, treatment, and prognosis. It provides a framework for subsequent tasks, including segmentation, classification, and target detection. Existing landmark detection methods exhibit limits with the increasing number of X-ray examiners. Not only they serve specific data types, but also their ability to learn the complete semantic information from this dataset is limited, potentially affecting their widespread adoption. We address these challenges by proposing a universal landmark detection model for multidomain X-ray imaging, UniverDetect. By progressively acquiring local semantic information, global semantic insights, and anatomical structure knowledge at various levels, the model ensured that the detected landmarks were closely aligned with the ground-truth labels. UniverDetect consists of three key components. The landmark detection module (LDM) utilizes a U-net network equipped with our innovative pyramid depthwise separable (PDS) convolution module for initial landmark detection. The landmark refinement module (LRM) integrates a fine-tuning module comprising continuously extended convolution blocks. Finally, the landmark correction module (LCM) incorporates a graph convolutional network (GCN) to rectify offset errors during partial landmark detection. An inherent feature of this model is its domain- generalization capability, which enables continuous learning across diverse domains. This model can concurrently learn from eight domains, covering 118 landmarks within a diverse dataset of 5969 images. Extensive experiments on multiple datasets demonstrate that this method consistently outperforms state-of-the-art approaches. This study presents a versatile and effective solution to reduce doctors' workload while providing precise quantitative analysis.
To compare and analyze the diagnostic value of different enhancement stages in distinguishing low and high nuclear grade clear cell renal cell carcinoma (ccRCC) based on enhanced computed tomography (CT) images by building machine learning classifiers. A total of 51 patients (Dateset1, including 41 low-grade and 10 high-grade) and 27 patients (Independent Dateset2, including 16 low-grade and 11 high-grade) with pathologically proven ccRCC were enrolled in this retrospective study. Radiomic features were extracted from the corticomedullary phase (CMP), nephrographic phase (NP), and excretory phase (EP) CT images, and selected using the recursive feature elimination cross-validation (RFECV) algorithm, the group differences were assessed using T-test and Mann–Whitney U test for continuous variables. The support vector machine (SVM), random forest (RF), XGBoost (XGB), VGG11, ResNet18, and GoogLeNet classifiers are established to distinguish low-grade and high-grade ccRCC. The classifiers based on CT images of NP (Dateset1, RF: AUC = 0.82 ± 0.05, ResNet18: AUC = 0.81 ± 0.02; Dateset2, XGB: AUC = 0.95 ± 0.02, ResNet18: AUC = 0.87 ± 0.07) obtained the best performance and robustness in distinguishing low-grade and high-grade ccRCC, while the EP-based classifier performance in poorer results. The CT images of enhanced phase NP had the best performance in diagnosing low and high nuclear grade ccRCC. Firstorder_Kurtosis and firstorder_90Percentile feature play a vital role in the classification task.
Background: Currently, the diagnosis of gastrointestinal diseases relies heavily on electronic endoscopy, which is not suitable for all individuals and insufficient for dynamic patient monitoring. Therefore, there is an urgent need for a convenient and noninvasive examination method to enhance the diagnosis and monitoring of gastrointestinal diseases. Methods: Since the tongue is a significant reflection of the digestive system and closely associated with gastrointestinal diseases, our study aims to develop a diagnostic framework based on tongue features. We have collected a dataset of 2,167 tongue coating images from 949 consecutive patients, including 905 images from healthy individuals and 1,262 images from patients with various gastrointestinal conditions. Notably, each patient sample may represent distinct gastrointestinal diseases. This study introduces a novel approach to information fusion detection, integrating hand-crafted and auto-encoded features, alongside disease detection employing Squeeze and Excitation (SE) and Slot attention mechanisms. These two branches respectively analyze image feature and spatial position information when utilizing tongue images for diagnosing gastrointestinal diseases. The amalgamation of predictions from both aspects yields the final detection outcome. Results: For both healthy and diseased cohorts, the optimal classification metrics were achieved: AUC = 0.886, ACC = 0.849, and TPR = 0.965, indicating satisfactory performance. Notably, our framework demonstrated promising results in detecting five common gastrointestinal diseases: Helicobacter pylori infection, bile reflux, reflux esophagitis, gastric erosion, and duodenal erosion. Ablation experiments underscored the efficacy of mixed features in multi-color space, showcasing an increase in AUC and ACC by 13.48 % and 12.05 % respectively. Furthermore, the attention mechanism, by focusing on the middle region of tongue images, contributed to a 6.12 % increase in AUC and a 6.35 % increase in ACC. Conclusions: In this study, we successfully develop and validate a diagnosis framework for gastrointestinal diseases based on tongue image features. By utilizing tongue images, we achieve intelligent diagnosis of gastrointestinal diseases, establishing tongue diagnosis as a convenient, cost-effective, and noninvasive method for diagnosing and screening gastrointestinal diseases.
Computer-Aided Sperm Analysis (CASA) has received great attention in the diagnosis of male reproductive health for its objectivity and accuracy. But its accuracy may be greatly affected by algorithms used in microscopic video preprocessing, especially algorithms for microscopic videos filtering. Owing to the Peak Signal to Noise Ratio (PSNR) and Structural Similarity (SSIM) indicators are not suitable for microscopic images, objective indicators for evaluating filtering of microscopic videos are scarce at present, leading to filtering algorithms are chosen arbitrarily. In addition, it is critical to establish quality control indicators for CASA from the perspective of image processing, but they are rarely mentioned for the reason that algorithms adopted in image processing aren't public. To reduce influences of filtering on diagnostic results, it is imperative to establish an objective quality control indicator for evaluating filtering of microscopic video from the perspective of image preprocessing. The principle of formatting sperm tracking trajectory and the rules of grading trajectory continuity were introduced firstly. Then, impacts of filtering algorithms on the continuity of sperm trajectory were investigated and the relation between the continuity of sperm trajectory and the fluctuation of sperm segmentation thresholds was further explored. Finally, an objective indicator for appraising the quality of filtering for microscopic video was established based on SADTM (Sum of Absolute Difference between Threshold and the Mean).The proposed indicator can objectively evaluate the quality of microscopic video filtering and provide a guideline for algorithm selection, laying foundation for quality control for CASA.
BACKGROUND:Multicenter non-small cell lung cancer (NSCLC) patient data is information-rich. However, its direct integration becomes exceptionally challenging due to constraints involving different healthcare organizations and regulations. Traditional centralized machine learning methods require centralizing these sensitive medical data for training, posing risks of patient privacy leakage and data security issues. In this context, federated learning (FL) has attracted much attention as a distributed machine learning framework. It effectively addresses this contradiction by preserving data locally, conducting local model training, and aggregating model parameters. This approach enables the utilization of multicenter data with maximum benefit while ensuring privacy safeguards. Based on pre-radiotherapy planning target volume images of NSCLC patients, a multicenter treatment response prediction model is designed by FL for predicting the probability of remission of NSCLC patients. This approach ensures medical data privacy, high prediction accuracy and computing efficiency, offering valuable insights for clinical decision-making.METHODS:We retrospectively collected CT images from 245 NSCLC patients undergoing chemotherapy and radiotherapy (CRT) in four Chinese hospitals. In a simulation environment, we compared the performance of the centralized deep learning (DL) model with that of the FL model using data from two sites. Additionally, due to the unavailability of data from one hospital, we established a real-world FL model using data from three sites. Assessments were conducted using measures such as accuracy, receiver operating characteristic curve, and confusion matrices.RESULTS:The model's prediction performance obtained using FL methods outperforms that of traditional centralized learning methods. In the comparative experiment, the DL model achieves an AUC of 0.718/0.695, while the FL model demonstrates an AUC of 0.725/0.689, with real-world FL model achieving an AUC of 0.698/0.672.CONCLUSIONS:We demonstrate that the performance of a FL predictive model, developed by combining convolutional neural networks (CNNs) with data from multiple medical centers, is comparable to that of a traditional DL model obtained through centralized training. It can efficiently predict CRT treatment response in NSCLC patients while preserving privacy.
The fundamental premise of numerous clinical applications is automatic landmark detection of the pelvis and quantitative analysis of deformity. As a result, automatic pelvic landmark identification and quantitative analysis can give doctors with more intuitive and effective information to guide the creation of treatment programs and prognoses, which is extremely important. Previous research either employed less noisy infant pelvis datasets or the tested response time and detection accuracy that were insufficient for clinical usage. To address these issues, we propose a two-stage method for automatic detection of pelvic landmarks. The first stage uses heat map regression, the second stage uses coordinate regression, and both stages incorporate attention mechanisms. The framework comprehensively studies our data. The detection performance of the framework is evaluated: there is a 3.724mm MRE accuracy with an accuracy of 74.176, 78.117, 80.588, and 84.706% in the ranges of 4, 4.5, 5, and 6mm. As a result, the technology is expected to become an outstanding pelvic landmark detection tool, allowing doctors with minimum experience to quickly and accurately identify landmark.
Introduction:Stroke-associated pneumonia (SAP) is a common complication of stroke that can increase the mortality rate of patients and the burden on their families. In contrast to prior clinical scoring models that rely on baseline data, we propose constructing models based on brain CT scans due to their accessibility and clinical universality.Methods:Our study aims to explore the mechanism behind the distribution and lesion areas of intracerebral hemorrhage (ICH) in relation to pneumonia, we utilized an MRI atlas that could present brain structures and a registration method in our program to extract features that may represent this relationship. We developed three machine learning models to predict the occurrence of SAP using these features. Ten-fold cross-validation was applied to evaluate the performance of models. Additionally, we constructed a probability map through statistical analysis that could display which brain regions are more frequently impacted by hematoma in patients with SAP based on four types of pneumonia.Results:Our study included a cohort of 244 patients, and we extracted 35 features that captured the invasion of ICH to different brain regions for model development. We evaluated the performance of three machine learning models, namely, logistic regression, support vector machine, and random forest, in predicting SAP, and the AUCs for these models ranged from 0.77 to 0.82. The probability map revealed that the distribution of ICH varied between the left and right brain hemispheres in patients with moderate and severe SAP, and we identified several brain structures, including the left-choroid-plexus, right-choroid-plexus, right-hippocampus, and left-hippocampus, that were more closely related to SAP based on feature selection. Additionally, we observed that some statistical indicators of ICH volume, such as mean and maximum values, were proportional to the severity of SAP.Discussion:Our findings suggest that our method is effective in classifying the development of pneumonia based on brain CT scans. Furthermore, we identified distinct characteristics, such as volume and distribution, of ICH in four different types of SAP.
Ultrasound (US) is widely used in the clinical diagnosis and treatment of musculoskeletal diseases. However, the low efficiency and non-uniformity of artificial recognition hinder the application and popularization of US for this purpose. Herein, we developed an automatic muscle boundary segmentation tool for US image recognition and tested its accuracy and clinical applicability. Our dataset was constructed from a total of 465 US images of the flexor digitorum superficialis (FDS) from 19 participants (10 men and 9 women, age 27.4 ± 6.3 years). We used the U-net model for US image segmentation. The U-net output often includes several disconnected regions. Anatomically, the target muscle usually only has one connected region. Based on this principle, we designed an algorithm written in C++ to eliminate redundantly connected regions of outputs. The muscle boundary images generated by the tool were compared with those obtained by professionals and junior physicians to analyze their accuracy and clinical applicability. The dataset was divided into five groups for experimentation, and the average Dice coefficient, recall, and accuracy, as well as the intersection over union (IoU) of the prediction set in each group were all about 90%. Furthermore, we propose a new standard to judge the segmentation results. Under this standard, 99% of the total 150 predicted images by U-net are excellent, which is very close to the segmentation result obtained by professional doctors. In this study, we developed an automatic muscle segmentation tool for US-guided muscle injections. The accuracy of the recognition of the muscle boundary was similar to that of manual labeling by a specialist sonographer, providing a reliable auxiliary tool for clinicians to shorten the US learning cycle, reduce the clinical workload, and improve injection safety.
Pelvic landmark detection is a significant pre-task to measure the clinical measurement in pelvic abnormality analysis. Accurate pelvic landmark detection could provide reliable clinical parameter measurement results, which are helpful for doctors to diagnose and treat pelvic diseases. However, the multi-scale characteristics, temporal diversity, and pathological abnormalities of different pelvic X-rays bring enormous challenges to the landmark detection task. In order to retain strong robustness in irregular pelvic X-rays, we propose a novel, flexible two-stage framework. In the initial stage, a single neural network is employed to estimate the locations of every landmark simultaneously, enabling the identification of potential landmark regions. Then, the receptive field of candidate region proposals is expanded by 4 times through the receptive field amplification module. In the second stage, the landmark detection module fuses semantically rich features at different scales through a multi-scale semantic fusion module. So that the framework can fully learn the strongly relevant semantic information around the landmark at high resolution. We collected a data set of 430 pelvic X-rays, including a large number of irregular pelvic X-rays, to evaluate our framework. The experimental results demonstrate that our framework achieves a state-of-the-art detection mean radial error of 3.724 ± 4.247-mm. The experimental results show that the proposed method can help doctors quickly and accurately find the characteristic points of the pelvis and could be applied to clinical diagnosis.
Chronic rhinosinusitis (CRS) is characterized by poor prognosis and propensity for recurrence even after surgery. Identification of those CRS patients with high risk of relapse preoperatively will contribute to personalized treatment recommendations. In this paper, we proposed a multi-task deep learning network for sinus segmentation and CRS recurrence prediction simultaneously to develop and validate a deep learning radiomics-based nomogram for preoperatively predicting recurrence in CRS patients who needed surgical treatment. 265 paranasal sinuses computed tomography (CT) images of CRS from two independent medical centers were analyzed to build and test models. The sinus segmentation model achieved good segmentation results. Furthermore, the nomogram combining a deep learning signature and clinical factors also showed excellent recurrence prediction ability for CRS. Our study not only facilitates a technique for sinus segmentation but also provides a noninvasive method for preoperatively predicting recurrence in patients with CRS.
Background Stereotactic radiosurgery (SRS) treatment planning requires accurate delineation of brain metastases, a task that can be tedious and time-consuming. Although studies have explored the use of convolutional neural networks (CNNs) in magnetic resonance imaging (MRI) for automatic brain metastases delineation, none of these studies have performed clinical evaluation, raising concerns about clinical applicability. This study aimed to develop an artificial intelligence (AI) tool for the automatic delineation of single brain metastasis that could be integrated into clinical practice. Methods Data from 426 patients with postcontrast T1-weighted MRIs who underwent SRS between March 2007 and August 2019 were retrospectively collected and divided into training, validation, and testing cohorts of 299, 42, and 85 patients, respectively. Two Gamma Knife (GK) surgeons contoured the brain metastases as the ground truth. A novel 2.5D CNN network was developed for single brain metastasis delineation. The mean Dice similarity coefficient (DSC) and average surface distance (ASD) were used to assess the performance of this method. Results The mean DSC and ASD values were 88.34%±5.00% and 0.35±0.21 mm, respectively, for the contours generated with the AI tool based on the testing set. The DSC measure of the AI tool’s performance was dependent on metastatic shape, reinforcement shape, and the existence of peritumoral edema (all P values <0.05). The clinical experts’ subjective assessments showed that 415 out of 572 slices (72.6%) in the testing cohort were acceptable for clinical usage without revision. The average time spent editing an AI-generated contour compared with time spent with manual contouring was 74 vs. 196 seconds, respectively (P<0.01). Conclusions The contours delineated with the AI tool for single brain metastasis were in close agreement with the ground truth. The developed AI tool can effectively reduce contouring time and aid in GK treatment planning of single brain metastasis in clinical practice.
Measurement of pelvic landmarks is a standard analytical tool for diagnosis and treatment prognosis. The purpose of this study is to develop and verify an automatic detection system for detecting pelvic measurement landmarks on pelycograms. Pelycograms for 430 subjects were available. All images have 17 landmarks. In this paper, a deep neural network system with Active Shape Model (ASM) prior knowledge is proposed that can be used to automatically locate landmarks on pelycograms. The outline of the pelvis is obtained by ASM, and the landmarks are roughly detected. Then the landmarks are accurately detected by a depth neural network based on this prior knowledge. The system is evaluated by experiments, and the average radial error of the system is 4.159 mm and the standard deviation is 5.015 mm. 81.353
This article discusses the connotation of medical ethics, the necessity of medical ethics education in clinical teaching, and its application in clinical teaching. Combining some cases of internal medicine, the application of medical ethics in clinical teaching practice of internal medicine is discussed from three aspects: the principle of medical ethics respect, the principle of informed consent, and the treatment of doctor-patient relationship. In view of the current status of Internet medical information quality, the necessity of carrying out evaluation work is proposed; the evaluation methods and evaluation standards of online information resources are summarized in detail, and some suggestions are put forward for the problems existing in the evaluation of Internet medical information resources in my country.
Detecting significant signaling pathways in disease progression highlights the dysfunctions and pathogenic mechanisms of complex disease development. Since tensor decomposition has been proven effective for multi-dimensional data representation and reconstruction, differences between original and tensor-processed data are expected to extract crucial information and differential indication. This paper provides a tensor-based gene set enrichment analysis, called tensorGSEA, based on a data reconstruction method to identify relevant significant pathways during disease development. As a proof-of-concept study, we identify the differential pathways of diabetes in rats. Specifically, we first arrange gene expression profiles of each documented pathway as tensors with three dimensions: genes, samples, and periods. Then we compress tensors into core tensors with lower ranks. The pathways with lower reconstruction rates are obtained after reconstructing gene expression profiles in another state via these cores. Thus, differences underlying pathways are extracted by cross-state data reconstruction between controls and diseases. The experiments reveal several critical pathways with diabetes-specific functions which otherwise cannot be identified by alternative methods. Our proposed tensorGSEA is efficient in evaluating pathways by achieving their empirical statistical significance, respectively. The classification experiments demonstrate that the selected pathways can be implemented as biomarkers to identify the diabetic state. The code of tensorGSEA is available at https://github.com/zhxr37/tensorGSEA .
Purpose:The objectives of our study were to assess the association of radiological imaging and gene expression with patient outcomes in non-small cell lung cancer (NSCLC) and construct a nomogram by combining selected radiomic, genomic, and clinical risk factors to improve the performance of the risk model.Methods:A total of 116 cases of NSCLC with CT images, gene expression, and clinical factors were studied, wherein 87 patients were used as the training cohort, and 29 patients were used as an independent testing cohort. Handcrafted radiomic features and deep-learning genomic features were extracted and selected from CT images and gene expression analysis, respectively. Two risk scores were calculated through Cox regression models for each patient based on radiomic features and genomic features to predict overall survival (OS). Finally, a fusion survival model was constructed by incorporating these two risk scores and clinical factors.Results:The fusion model that combined CT images, gene expression data, and clinical factors effectively stratified patients into low- and high-risk groups. The C-indexes for OS prediction were 0.85 and 0.736 in the training and testing cohorts, respectively, which was better than that based on unimodal data.Conclusions:Combining radiomics and genomics can effectively improve OS prediction for NSCLC patients.
Stroke-associated Pneumonia (SAP) is a common complication of intracerebral hemorrhage (ICH), which increases the burden of hospitalization and the difficulty of nursing, and significantly increases the mortality of patients. To help doctors get risk scores of SAP when ICH patients are just admitted, predicting SAP rapidly and effectively is important.A total of 244 subjects were involved in this research aiming to use radiomics method to extract features from the CT of ICH patients, and use recursive methods to select features and grid search to establish three machine learning models. Two binary classification implementations were carried out to predict the development of SAP, which were to classify with-or-without SAP, and SAP above moderate level.In the two classification tasks, selected different radiomics features for each model, and the AUCs and accuracies of the three constructed models on the test set were all above 0.7.Our experimental results showed that based on brain CT of ICH patients and without collecting the baseline data of patients, radiomics features can be used to quickly and effectively predict the development of SAP, and shape (3D) feature and gray-level feature account for a lot of weight.
Research question: Which characteristics of patients with a thin endometrium (endometrial thickness [EMT]<= 75 mm on human chorionic gonadotrophin [HCG] trigger day) suggest the possibility of an EMT >75 mm in the subsequent frozen cycle? Design: Data were collected from the university-affiliated Centre for Reproductive Medicine between January 2013 and September 2019. Multivariable logistic regression was used to generate the final prediction model and construct the nomogram. Model performances were quantified by discrimination and calibration. Results: The predictive variables that entered the final model were: hysteroscopic adhesiolysis history, polycystic ovary syndrome status, application of clomiphene in the ovarian stimulation process, the ovarian stimulation protocol and the endometrial preparation protocol. The receiver operating characteristic (ROC) curve for the final model and validation cohort was 0.760 (95% confidence interval [CI] 0.722-0.797) and 0.713 (95% CI 0.664-0.759), respectively. Discrimination performed well in both the modelling and validation cohorts. Conclusions: In women with a thin endometrium (EMT <= 75 mm on HCG trigger day), the absence of a hysteroscopic adhesiolysis history, the presence of polycystic ovary syndrome, the application of clomiphene in the ovarian stimulation process, the application of a gonadotrophin-releasing hormone agonist short protocol, mild stimulation protocol, natural cycle protocol, and natural cycle for endometrial preparation are prognostic for an increased possibility of an EMT >75 mm in the subsequent frozen cycle.
目前多数实时语义分割网络不仅同时处理边界和纹理等细节信息而且还忽略了语义边界区域特征,从而导致物体边界分割质量下降.针对该问题,提出一种边界感知的实时语义分割网络,主要从三个方面提高边界语义分割质量.提出了边界感知学习机制利用位置信息降低边界特征和轮廓附近细节的耦合度使边界感知和位置关系相互促进.设计轻量级区域自适应模块增强卷积网络对复杂语义边界区域的建模能力.根据采样区域像素贡献值不同设计了高效的空洞空间金字塔池化模块以增强重要的细节和语义特征.实验方面,与基准相比,在Cityscapes验证集上精度提升了约5.8个百分点,在Cityscapes测试集上以47.2 FPS的推理速度使精度达到了74.9%.在CamVid数据集上与BiSeNetV2算法相比mIoU提升了约3.96个百分点.
Background: Colposcopy is widely used to detect cervical cancer, but developing countries lack the experienced colposcopists necessary for accurate diagnosis. Artificial intelligence (AI) is being widely used in computer-aided diagnosis (CAD) systems. In this study, we developed and validated a CAD model based on deep learning to classify cervical lesions on colposcopy images. Methods: Patient data, including clinical information, colposcopy images, and pathological results, were collected from Qilu Hospital. The study included 15,276 images from 7,530 patients. We performed two tasks in this study: normal cervix (NC) vs. low grade squamous intraepithelial lesion or worse (LSIL+) and high-grade squamous intraepithelial lesion (HSIL)- vs. HSIL+. The residual neural network (ResNet) probability was calculated for each patient to reflect the probability of lesions through a ResNet model. Next, a combination model was constructed by incorporating the ResNet probability and clinical features. We divided the dataset into a training set, validation set, and testing set at a ratio of 7:1:2. Finally, we randomly selected 300 patients from the testing set and compared the results with the diagnosis of a senior colposcopist and a junior colposcopist. Results: The model that combines ResNet and clinical features performs better than ResNet alone. In the classification of NC and LSIL+, the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) were 0.953, 0.886, 0.932, 0.846, 0.838, and 0.936, respectively. In the classification of HSIL- and HSIL+, the AUC, accuracy, sensitivity, specificity, PPV, and NPV were 0.900, 0.807, 0.823, 0.800, 0.618, and 0.920, respectively. In the two classification tasks, the diagnostic performance of the model was determined to be comparable to that of the senior colposcopist and exhibited a stronger diagnostic performance than the junior colposcopist. Conclusions: The CAD system for cervical lesion diagnosis based on deep learning performs well in the classification of cervical lesions and can provide an objective diagnostic basis for colposcopists.