Congenital heart disease (CHD) is a major cause of neonatal death, and prenatal detection depends on standard planes. However, the poor image quality of echocardiography limits the application of AI in fetal cardiac ultrasound. To solve this, we developed the AortaQA echocardiography analysis system using duct-dependant CHD as a case. AortaQA uses key anatomical structures and traditional image quality metrics for quality control of input ultrasound images and then conducts intelligent disease screening. It was trained and validated with 7,310 images and 55 video clips from four tertiary medical centers. Experimental results show it effectively enhances CHD screening, achieving high AUROC scores (0.942, 0.865, 0.885, 0.908) across four centers. It attains a 0.896 mAP in detecting key anatomical structures, outperforming general-purpose models. The ablation study demonstrates that the Anatomically-Guided Dual-Quality Control (AGD-QC) method, by integrating clinical knowledge with image quality metrics, outperforms traditional CNN approaches. The quality threshold evaluation experiment showed that when the quality evaluation cutoff value was 8, the screening performance of the model was significantly different, which was consistent with the clinical conclusion. Human-computer interaction experiments showed that the system reduced the echocardiography scanning time of doctors with different professional titles (1.83 s, 4.09 s and 9.61 s per case) and improved the disease screening performance of doctors (0.6%, 1.85%, 10.7% and 10% in four centers). The AortaQA system employs the AGD-QC method, demonstrating superior generalization capabilities compared to conventional CNN models. This breakthrough provides an AI-driven solution for precise and comprehensive screening of DUCT-dependent CHD, with potential for expansion into other CHD domains.
International Classification of Diseases (ICD) coding is a crucial multi-label medical text classification task. Itmaps clinical text to codes for clinical management, billing, and statistics. While pre-trained language models (PLMs) have advanced automated ICD coding, they often lack deep medical understanding. This limitaion hinders semantic modeling and knowledge integration. We introduce MedLink, a knowledge-enhanced pre-training and fine-tuning framework for automatic ICD coding. MedLink integrates medical knowledge with language models to improve clinical text conceptual recognition and code assignment precision. Pre-training involves constructing a medical entity lexicon and performing robust entity linking on large-scale biomedical texts to create knowledge-enriched corpora. A dynamic dual-task mechanism then jointly optimizes medical entity type prediction and enhanced masked language modeling, guiding the model towards aligned semantic and knowledge representations. For fine-tuning, chunk-based semantic aggregation and a gated hierarchical classifier address ICD coding’s long-text and hierarchical nature, enhancing performance on complex documents. Experiments on MIMIC-III and MIMIC-IV datasets show MedLink significantly surpasses state-of-the-art methods on key metrics (F1-Micro, P@5). Ablation studies confirm the medical entity prediction task’s critical role. This research offers a scalable, interpretable knowledge fusion approach for automatic ICD coding, enhancing its accuracy and applicability.
Background While Auditory Brainstem Response (ABR) provides a non-invasive window into auditory brainstem function, prior studies of ASD have primarily focused on localized waveform features (e.g., waves I, III, and V), potentially overlooking subtle but informative patterns in the full-band signal. This study introduces a deep learning framework to comprehensively characterize time-frequency signatures of auditory brainstem activity in ASD, with the goal of identifying neurophysiologically meaningful ABR features associated with ASD. Methods We analyzed a clinical dataset of 1,209 ABR recordings (ASD: 961; Typically Developing Controls: 248). A dual-branch Time-Frequency Fusion and Transformer-Based Network (TF-TBN) was developed. The temporal branch utilizes a Transformer-enhanced 1D-CNN to analyze raw ABR waveforms, while the frequency branch employs a Vision Transformer to analyze spectrograms generated via Continuous Wavelet Transform. A fusion module integrates these features for final classification. Model interpretability was analyzed to identify critical ABR features. Results The TF-TBN model achieved a classification accuracy of 96.62%, significantly outperforming conventional deep learning baselines. Interpretability analysis revealed that the model’s decision was heavily influenced by prolonged absolute latencies of waves III and V, and interpeak latencies of I-III and I-V, which were confirmed as statistically significant in the ASD cohort. This suggests that the model successfully learned biologically plausible biomarkers of auditory pathway dysfunction. Conclusions This study provides the first comprehensive characterization of full-band ABR abnormalities in ASD using a deep learning framework. The TF-TBN model identifies prolonged wave III- and wave V-related timing features as prominent contributors to ASD-TD discrimination, with wave III-related delay emerging as an important component of the observed ABR abnormality. By linking AI-driven feature discovery to interpretable neurophysiological biomarkers, our work advances the analytical framework for ABR and contributes to understanding the neural basis of auditory processing deficits in ASD.
Congenital heart disease (CHD) is the most common major congenital anomaly and a leading cause of perinatal mortality and long-term neurodevelopmental disability globally, among which severe ductal dependent CHD (D-CHD) requires urgent clinical intervention due to the high risk of catastrophic cardiovascular collapse if undiagnosed. Despite the advancement of fetal echocardiography-based analysis models, there remains a lack of comprehensive and accurate evaluation approaches for D-CHD, especially for the precise differentiation of critical subtypes including coarctation of the aorta (CoA), interrupted aortic arch (IAA), and transposition of the great arteries (TGA). To address this gap, we developed the D-CHD Precision Screening System (TLEUDS) using 9,142 fetal aortic arch echocardiographic long-axis view images and 58 ultrasound clips from three medical institutions. Aiming to solve the challenges of uneven quality consistency and long-tail distribution in real-world ultrasound imaging data, TLEUDS adopts a cascaded Dual-Transfer learning framework for quality and knowledge enhancement, which implements a closed-loop screening process from initial disease detection (via the TEN-S module, a transfer version of EfficientNetV2-Screening) to detailed subtype-specific malformation screening (via the TY-PS module, a transfer version of YOLOv11-Precise Screening). Experimental results show that TLEUDS achieved a sensitivity of 0.988, 0.941, 0.966, and 0.900 for D-CHD overall, CoA, IAA, and TGA respectively, with corresponding specificities of 0.988, 0.987, 1.000, and 0.956, demonstrating state-of-the-art (SOTA) performance compared with general models. TLEUDS is expected to serve as a potential automated screening tool for tiered healthcare and computer-aided diagnosis of fetal D-CHD. Furthermore, the proposed quality and knowledge-enhanced cascaded Dual-Transfer learning framework fully leverages the "image quality distribution" and "structural prior knowledge" of fetal echocardiography, holding great potential for extension to other ultrasound image analysis domains.
OBJECTIVE:The growing number of studies directly comparing artificial intelligence (AI) to physicians in diagnostic tasks often focuses on performance outcomes, overlooking fundamental methodological rigor. This scoping review aims to critically appraise the methodological quality of this body of literature, identifying key challenges and proposing a framework to enhance the fairness, standardization, and clinical relevance of future comparisons. MATERIALS AND METHODS:We conducted a systematic search of PubMed, Scopus, and Web of Science for studies published between January 1, 2020, and October 31, 2025, following the PRISMA-ScR guidelines. From 8,851 screened records, 120 studies met the inclusion criteria for direct AI-physician comparison. Data on study characteristics, dataset quality, task design, physician configuration, and reporting transparency were extracted and synthesized narratively. RESULTS:Our analysis of 120 studies revealed a field characterized by significant methodological heterogeneity. Key issues include a predominant focus on retrospective studies (75.8%), frequent information asymmetry between AI and physicians (20.8%), limited clinical relevance in task design despite superficial fidelity, and insufficient physician sample sizes (60.8% had ≤ 10 readers). Furthermore, we found a widespread neglect of time constraints (absent in 50.8% of studies) and a critical lack of transparency regarding code and data availability. CONCLUSION:Current research on AI-physician diagnostic comparisons is often hampered by methodological weaknesses that undermine the validity and generalizability of its findings. To ensure the generation of reliable and clinically meaningful evidence, future studies must prioritize prospective designs, ensure fairness in experimental conditions, and adhere to higher standards of transparency. We propose the AI vs. Physician Study Checklist (AIPSC) as a practical tool to guide the design and reporting of more robust and systematic evaluations, ultimately fostering the responsible integration of AI into clinical practice.
Fetal coarctation of the aorta (CoA) is a prevalent congenital heart disease with a high false positive rate (49%-94%), which easily causes confusion in early diagnosis and leads to overmedicalization. Clinically, distinguishing Pseudocoarctation of the Aorta (PCoA) from CoA in fetal echocardiography remains a critical challenge. Existing studies are limited by insufficient perception of global cardiac structure, weak anti-noise capability, and poor adaptability to sample imbalance, and have failed to include follow-up-confirmed PCoA cases. This study aims to develop an intelligent deep diagnostic framework aligned with clinicians' cognitive logic in order to accurately differentiate false positives in congenital heart disease screening and achieve precise CoA identification. We established a multi-view fetal echocardiography dataset of 2737 cases (Normal, CoA, PCoA). Then we constructed a novel model termed Progressive Transfer Learning tailored Multi-view Vision Transformer (PTLMV-ViT), which adopts parallel Vision Transformer branches as the backbone to extract high-dimensional semantic features with global receptive fields from three key fetal cardiac ultrasound views via the multi-head self-attention mechanism. Cross-view features are integrated through a concatenate-and-project fusion strategy, mimicking expert synthesis. To address data imbalance, a progressive transfer learning strategy was designed, transferring knowledge from a large binary classification task (normal vs. suspicion for CoA) to initialize the fine-grained CoA vs. PCoA classifier. Experiments showed PTLMV-ViT achieved an area under the receiver operating characteristic curve (AUC) of 0.997, sensitivity of 0.959, and specificity of 0.996 in distinguishing PCoA from CoA cases, which significantly outperformed traditional single-view models. This framework offers a powerful tool to assist clinicians in reducing CoA misdiagnosis.
Endoglin (ENG) is a single-pass transmembrane protein highly expressed in vascular endothelial cells (ECs), where it plays fundamental roles in EC functions. ENG is implicated in several cardiovascular disorders including hereditary haemorrhagic telangiectasia, pulmonary arterial hypertension (PAH) and preeclampsia. However, molecular mechanisms underlying ENG function are not fully understood. Initially identified as a co-receptor for TGF-β signalling, ENG's extracellular domain was later found to only bind BMP9 and BMP10 with high affinity. The relationship between these two observations is unclear. Here, we provide evidence for two primary functions of co-receptor ENG. First, ENG efficiently displaces prodomains from BMP9 and BMP10, enabling effective capturing of both ligands from the circulation. Second, ENG binds to and recruits TGFBRII into the BMP9 signalling complex, thereby explaining ENG's involvement in both TGF-β and BMP9 pathways. We identify BMP9 target genes NOG and ADAMTSL2 as preferentially dependent on ENG and show that their transcript levels have strong positive correlation with ENG in human lung tissues; the expression levels of all three genes are significantly reduced in PAH. Our findings address an important gap in our understanding on ENG biology and provide crucial insight for therapeutic targeting these pathways in vascular diseases.
Pediatric liver tumors are one of the most common solid tumors in pediatrics, with differentiation of benign or malignant status and pathological classification critical for clinical treatment. While pathological examination is the gold standard, the invasive biopsy has notable limitations: the highly vascular pediatric liver and fragile tumor tissue raise complication risks such as bleeding; additionally, young children with poor compliance require anesthesia for biopsy, increasing medical costs or psychological trauma. Although many efforts have been made to utilize AI in clinical settings, most researchers have overlooked its importance in pediatric liver tumors. To establish a non-invasive examination procedure, we developed a multi-stage deep learning (DL) framework for automated pediatric liver tumor diagnosis using multi-phase contrast-enhanced CT. Two retrospective and prospective cohorts were enrolled. We established a novel PKCP-MixUp data augmentation method to address data scarcity and class imbalance. We also trained a tumor detection model to extract ROIs, and then set a two-stage diagnosis pipeline with three backbones with ROI-masked images. Our tumor detection model has achieved high performance (mAP=0.871), and the first stage classification model between benign and malignant tumors reached an excellent performance (AUC=0.989). Final diagnosis models also exhibited robustness, including benign subtype classification (AUC=0.915) and malignant subtype classification (AUC=0.979). We also conducted multi-level comparative analyses, such as ablation studies on data and training pipelines, as well as Shapley-Value and CAM interpretability analyses. This framework fills the pediatric-specific DL diagnostic gap, provides actionable insights for CT phase selection and model design, and paves the way for precise, accessible pediatric liver tumor diagnosis.
Background With the rapid development of information technology and the digitization of medical devices, various diseases require the use of medical imaging equipment for diagnosis. At present, various medical imaging diagnostic equipment such as CT and nuclear magnetic resonance can provide two-dimensional planar images of diseases. Doctors urgently need to accurately determine the spatial location, size, geometry, and spatial relationship with the surrounding tissue. Therefore, it is very important to use computer technology to segment 3D MRI images, determine the location of lesions, and then perform 3D reconstruction. Method At present, automatic recognition and marking of brain images are displayed in two dimensions. Therefore, it is necessary to use 3D visualization technology for reconstruction. In addition, it can be combined with virtual and real, and some additional information is superimposed on the brain image for integrated display. In addition, a combination of virtual and real needs to be superimposed, and some additional information is superimposed on the brain image for integrated display. The research focus of this paper includes two main parts: disease segmentation and 3D reconstruction visualization. Firstly, the disease segmentation method based on 3D MRI brain image files was designed, and then the feature extraction and 3D reconstruction functions were designed. Thereby forming a complete process of disease region segmentation and three-dimensional reconstruction. Results This study is based on a three-dimensional MRI brain image segmentation algorithm. The algorithm is advanced in technology, high in accuracy, and can effectively identify the location of the disease. Then, this study used the Unity tool to implement a three-dimensional reconstruction and visual display program for brain image disease segmentation. Therefore, the doctor can quickly and intuitively grasp the spatial information inside the brain and the information of the lesion area. Conclusion This study uses advanced disease segmentation algorithm and the latest 3D reconstruction visualization technology to initially realize the brain disease region segmentation and 3D visualization display function based on augmented reality. It has improved the diagnostic efficiency of brain diseases and has certain practical application value. And it provides a good solution for a large number of information and data display problems.
Early detection and treatment can slow the progression of Alzheimer's Disease (AD), one of the most common neurodegenerative diseases. Recent studies have demonstrated the value of multimodal fusion in early AD detection. However, most approaches to this have failed to consider data modality domains, their relationships, and variations in their relative importance. To address these challenges, we propose a Hierarchical Attention-Based Multimodal Fusion framework (HAMF) that utilizes imaging, genetic and clinical data for early AD detection. In the HAMF model, attention mechanisms are utilized to learn the appropriate weights for each modality and to understand the interaction between modalities through hierarchical attention. HAMF performs better than state-of-the-art methods, achieving an accuracy of 87.2% and an AUC of 0.913, which are superior to unimodal models. By comparing the results of different unimodal and multimodal models, we find that multimodal fusion can improve model performance more than unimodal models and clinical data is the most important modality. Our ablation experiment confirmed the effectiveness of HAMF. Finally, we used SHapley Additive exPlanations (SHAP) to improve the model's interpretability. We provide the model as a guide for future research in the field, and as a framework for generating actional advice and decision support system for clinical practitioners.
Duct-dependent congenital heart diseases (CHDs) are a serious form of CHD with a low detection rate, especially in underdeveloped countries and areas. Although existing studies have developed models for fetal heart structure identification, there is a lack of comprehensive evaluation of the long axis of the aorta. In this study, a total of 6698 images and 48 videos are collected to develop and test a two-stage deep transfer learning model named DDCHD-DenseNet for screening critical duct-dependent CHDs. The model achieves a sensitivity of 0.973, 0.843, 0.769, and 0.759, and a specificity of 0.985, 0.967, 0.956, and 0.759, respectively, on the four multicenter test sets. It is expected to be employed as a potential automatic screening tool for hierarchical care and computer-aided diagnosis. Our two-stage strategy effectively improves the robustness of the model and can be extended to screen for other fetal heart development defects.
The double dividend of the carbon tax policy has been a controversial topic. To comprehensively evaluate the benefits and risks brought by the carbon tax policy and contribute to China’s emission reduction goals, this paper establishes a carbon tax policy cycle simulation model based on China’s economic and energy data from 2010 to 2020 to explore the winner-curse phenomenons of the policy. To alleviate the winner’s curse of the carbon tax policy, this paper introduces a consumer behavior model to explore the optimization degree of loss aversion effect on the carbon tax policy. The research results show that the carbon tax policy has three kinds of winner’s curse phenomenons, namely, the improvement of environmental quality and the reduction of market capital, the decline of national carbon intensity and the increase of carbon intensity of three major industries, and the reuse of the tax revenue and the increase of economic loss. The loss aversion of consumers can alleviate the negative effect of the carbon tax policy and strengthen the positive effect. In addition, during the implementation of the carbon tax policy, the loss aversion effect can also reduce the polluted population by about 2%. Finally, based on the research results, the paper puts forward some feasible policy suggestions.
The outbreak of COVID-19 provides a rare opportunity for the implementation of the carbon tax. To determine which stage is the most appropriate for introducing the policy, a simulation model based on China’s panel data is established to analyze the impact of the carbon tax on government revenue and residents’ income from five scenarios. A new GM-SD modeling method is proposed to ensure the accuracy of the model. The results show that the impact of the carbon tax on the government and the public is significantly different at different stages, and even the implementation of the carbon tax in the early stage of COVID-19 will reduce the government’s tax revenue. The score analysis of government tax revenue, residents’ surplus disposable income, residents’ emotional value, and government administrative power finds that the middle period of COVID-19 is the best time to implement the policy. In addition, a more detailed analysis of five aspects, including total population, energy consumption, and national income, shows that the best time to implement the carbon tax policy is when the damage degree of COVID-19 is moderate. The analysis results can provide a reference and basis for China to introduce the carbon tax in the event of similar events as COVID-19, and have reference significance for other countries that have not implemented a carbon tax.
A global survey indicates that genetic syndromes affect approximately 8% of the population, but most genetic diagnoses can only be performed after babies are born. Abnormal facial characteristics have been identified in various genetic diseases; however, current facial identification technologies cannot be applied to prenatal diagnosis. We developed Pgds-ResNet, a fully automated prenatal screening algorithm based on deep neural networks, to detect high-risk fetuses affected by a variety of genetic diseases. In screening for Trisomy 21, Trisomy 18, Trisomy 13, and rare genetic diseases, Pgds-ResNet achieved sensitivities of 0.83, 0.92, 0.75, and 0.96, and specificities of 0.94, 0.93, 0.95, and 0.92, respectively. As shown in heatmaps, the abnormalities detected by Pgds-ResNet are consistent with clinical reports. In a comparative experiment, the performance of Pgds-ResNet is comparable to that of experienced sonographers. This fetal genetic screening technology offers an opportunity for early risk assessment and presents a non-invasive, affordable, and complementary method to identify high-risk fetuses affected by genetic diseases. Additionally, it has the capability to screen for certain rare genetic conditions, thereby enhancing the clinic's detection rate.
Early identification and intervention of abnormal brain development individual subjects are of great significance, especially during the earliest and most active stage of brain development in children aged under 3. Neuroimage-based brain’s biological age has been associated with health, ability, and remaining life. However, the existing brain age prediction models based on neuroimage are predominantly adult-oriented. Here, we collected 658 T1-weighted MRI scans from 0 to 3 years old healthy controls and developed an accurate brain age prediction model for young children using deep learning techniques with high accuracy in capturing age-related changes. The performance of the deep learning-based model is comparable to that of the SVR-based model, showcasing remarkable precision and yielding a noteworthy correlation of 91% between the predicted brain age and the chronological age. Our results demonstrate the accuracy of convolutional neural network (CNN) brain-predicted age using raw T1-weighted MRI data with minimum preprocessing necessary. We also applied our model to children with low birth weight, premature delivery history, autism, and ADHD, and discovered that the brain age was delayed in children with extremely low birth weight (less than 1000 g) while ADHD may cause accelerated aging of the brain. Our child-specific brain age prediction model can be a valuable quantitative tool to detect abnormal brain development and can be helpful in the early identification and intervention of age-related brain disorders.
A global survey has revealed that genetic syndromes affect approximately 8% of the population, but most genetic diagnoses are typically made after birth. Facial deformities are commonly associated with chromosomal disorders. Prenatal diagnosis through ultrasound imaging is vital for identifying abnormal fetal facial features. However, this approach faces challenges such as inconsistent diagnostic criteria and limited coverage. To address this gap, we have developed FGDS, a three-stage model that utilizes fetal ultrasound images to detect genetic disorders. Our model was trained on a dataset of 2554 images. Specifically, FGDS employs object detection technology to extract key regions and integrates disease information from each region through ensemble learning. Experimental results demonstrate that FGDS accurately recognizes the anatomical structure of the fetal face, achieving an average precision of 0.988 across all classes. In the internal test set, FGDS achieves a sensitivity of 0.753 and a specificity of 0.889. Moreover, in the external test set, FGDS outperforms mainstream deep learning models with a sensitivity of 0.768 and a specificity of 0.837. This study highlights the potential of our proposed three-stage ensemble learning model for screening fetal genetic disorders. It showcases the model's ability to enhance detection rates in clinical practice and alleviate the burden on medical professionals.
With the advancement of medicine, more and more researchers have turned their attention to the study of fetal genetic diseases in recent years. However, it is still a challenge to detect genetic diseases in the fetus, especially in an area lacking access to healthcare. The existing research primarily focuses on using teenagers’ or adults’ face information to screen for genetic diseases, but there are no relevant directions on disease detection using fetal facial information. To fill the vacancy, we designed a two-stage ensemble learning model based on sonography, Fgds-EL, to identify genetic diseases with 932 images. Concretely speaking, we use aggregated information of facial regions to detect anomalies, such as the jaw, frontal bone, and nasal bone areas. Our experiments show that our model yields a sensitivity of 0.92 and a specificity of 0.97 in the test set, on par with the senior sonographer, and outperforming other popular deep learning algorithms. Moreover, our model has the potential to be an effective noninvasive screening tool for the early screening of genetic diseases in the fetus.
The carbon tax is a policy tool that internalizes external costs through a tax mechanism, which helps to reduce the consumption of fossil energy and lower carbon dioxide emissions. China, as the largest carbon emitter, introducing a carbon tax can further enhance the effectiveness of emission reduction. However, the introduction of a carbon tax may exacerbate contradictions in other aspects of the social system. To this end, the paper establishes a dynamic model of the carbon tax system by combining grey system theory and the IPAT model and then explores the coupling effect of the carbon tax on the economy, energy, and environment under the premise of China's resource endowment. It is found that carbon tax will not only distort consumer behavior but also aggravate the degree of capital market distortion. In the time-series simulation, it is found that the emission reduction efficiency of the carbon tax will show an oscillation decline. The carbon tax undermines the carbon peak target by dampening demand for energy consumption. In addition, we also find that the change of energy structure is the root of driving the failure of the "Jevons Paradox" and the realization of the "environmental Kuznets curve," and the panel data of energy and economy are only the manifestation of these two phenomena. China needs to adjust its energy structure to achieve its carbon peaking target. These results are helpful for policymakers to rationally view the carbon peaking target and formulate reasonable emission reduction policies.
The accurate, quantitative, and objective prediction of the brain age for premature infants will contribute to the exploration of brain maturity and catch-up growth. Traditional approaches rely heavily on a pediatrician's clinical experience, which makes the whole process time-consuming and labor-intensive. To solve this problem, we propose a deep learning-based brain age prediction model for preterm infants via neonatal MRI for this purpose, and it is called as BAPNET for short. First of all, we collected a specific dataset including MR images of 281 preterm infants. Then, a pretraining model (DeepBrainNet) is applied as the main backbone, and transfer learning is utilized to enhance the baseline model by making knowledge transfer from the ImageNet dataset. The proposal can be viewed as a specific prediction model by absorbing knowledge enhancement from peripheral visual features. On a test set of 70 preterm infants held out from the original dataset, 2D-BPANET achieved results with an mean square error (MAE) of 1.15 and the 95% - 95% content tolerance interval for a difference (prediction and ground truth) of [-3.82, 3.39], whereas 3D-BPANET achieved better results with an MAE of 1.8 and a difference of [0.51, 3.09]. Meanwhile, we leverage heatmaps to verify the consistency between hindbrain regions and cortical fold regions outputted by our model and the latest studies of brain development in preterm infants. In conclusion, BPANET demonstrates that deep learning can estimate brain maturity in preterm infants and provides a reference standard for preterm infant brain development, which could be applied as a promising tool.