OBJECTIVES:The application of large language models (LLMs) to systematic review tasks is rapidly expanding, yet the transparency and methodological rigor of these evaluations remain unclear. We aimed to assess reporting transparency, methodological quality, and how authors frame claims and caveats in studies applying LLMs to systematic review tasks. STUDY DESIGN AND SETTING:We conducted a cross-sectional meta-epidemiological study by searching PubMed, Embase, Web of Science Core Collection, IEEE Xplore, and five other databases from inception to December 1, 2025, for peer-reviewed articles and preprints. We included empirical studies evaluating generative transformer-based LLMs (eg, Gemini) for core systematic review tasks (eg, screening, data extraction) against a reference standard. We assessed reporting transparency using an adapted Chatbot Assessment Reporting Tool and methodological quality using an adapted Quality Assessment of Diagnostic Accuracy Studies 2 tool. We also analyzed the frequency and strength of claims and caveats mentioned by the authors. The study is registered with the Open Science Framework (https://osf.io/8edhb). RESULTS:We identified and included 229 studies comprising 440 empirical tasks. Reporting transparency was moderate, with a mean item score of 0.52 (standard deviation 0.30) on a 0-1 scale, where higher values indicate more complete reporting. We observed substantial gaps in reproducibility-essential domains, including protocol information (mean score 0.12) and model details (0.30). Although 60.6% of assessments were rated as having a low risk, key safeguards against overfitting and data leakage were rarely reported; for example, locking the test set before prompt optimization, a basic protection against information leakage, was not reported in 99.8% of tasks. We identified 837 claims and 693 caveats. Authors framed claims weakly more often than strongly (66.8% vs 33.2%). Performance superiority over a comparator was the most common claim (64.8% of tasks). Readiness for practical use was claimed in 47.5% of tasks, almost always in qualified terms (93.3%). CONCLUSION:Studies applying LLMs to systematic review tasks are reported with moderate transparency but often omit reproducibility-critical details necessary to assess leakage and overfitting. Although authors frequently make claims about performance and practice readiness, these are typically expressed cautiously. Improved reporting standards and clearer safeguards are urgently needed before routine use of LLMs in evidence synthesis can be recommended. PLAIN LANGUAGE SUMMARY:LLMs, such as ChatGPT, are increasingly used to help carry out parts of systematic reviews, which summarize evidence to inform healthcare decisions. We examined 229 studies that tested LLMs on tasks such as screening articles, extracting data, and assessing study quality, covering 440 evaluations in total. On average, these studies reported their methods with moderate clarity, but often omitted information needed to repeat the work or judge whether the results were trustworthy. Notably, almost none described safeguards to ensure that the test data had not already influenced how the model was set up, a key step for avoiding overly optimistic results. Authors frequently described LLMs as performing well and nearly ready for practical use, though usually in cautious terms. Clearer reporting standards and stronger safeguards are needed before LLMs can be routinely relied upon in evidence synthesis.
Background: Porto-sinusoidal vascular disorder (PSVD) is often misdiagnosed as liver cirrhosis due to overlapping clinical presentations and imaging features. This study conducted a blinded, independent imaging review to identify and compare the distinct imaging features between the two diseases, and to develop and validate a predictive model for differentiating PSVD from cirrhosis. Methods: Patients with histologically and clinically confirmed PSVD or cirrhosis and available contrast-enhanced computed tomography (CT) or magnetic resonance imaging (MRI) scans were retrospectively enrolled in the study. Imaging features were independently and systematically analyzed by two abdominal radiologists, who were blinded to the case grouping. Inter-reader discrepancies were resolved by consensus. The key features for analysis included liver surface nodularity (LSN), regenerative nodules (RNs), signs of portal hypertension (PH), and reticular delayed enhancement of the hepatic parenchyma. CT, MRI, and combined predictive models were developed to identify the top performing model, which was then selected and validated on an independent test cohort. Model performance was evaluated based on the area under the curve (AUC), sensitivity, and specificity. Results: In total, 106 patients with PSVD and 104 patients with cirrhosis were included for imaging evaluation and model development. The data of an additional 36 patients with PSVD and 51 patients with cirrhosis were collected for independent model testing. PSVD exhibited the same pronounced PH imaging features as cirrhosis, including grade 1-3 splenomegaly (77/106, 72.6% vs. 67/104, 64.4%; P>0.05), collateral vessels (103/106, 97.2% vs. 94/104, 90.4%; P>0.05), and ascites (35/106, 33.0% vs. 32/104, 30.8%; P>0.05). In addition, PSVD showed more increased small branches of intrahepatic blood vessels than cirrhosis (79/106, 74.5% vs. 41/104, 39.4%; P<0.001). Conversely, PSVD exhibited fewer cirrhosis-specific imaging features, such as reticular delayed enhancement of the hepatic parenchyma (10/48, 20.8% vs. 50/53, 94.3%; P<0.001), RNs (3/48, 6.3% vs. 29/53, 54.7%; P<0.001), and LSN (27/106, 25.5% vs. 75/104, 72.1%; P<0.001). In the validation set, the MRI model [AUC: 0.970, 95% confidence interval (CI): 0.912-1.0], which incorporated four imaging features (reticular delayed enhancement, RNs, LSN and increased small intrahepatic vascular branches), showed superior discriminatory performance compared to the CT model (AUC: 0.825, 95% CI: 0.646-1.0), with a sensitivity of 0.818 and a specificity of 0.889. While the combined model (AUC: 0.97, 95% CI: 0.912-1.0) did not improve upon the performance of the MRI model alone, with sensitivity of 0.821 and specificity of 0.889. Thus, we recommend the MRI model as the preferred modality for diagnosing PSVD. The MRI model also achieved optimal performance in the independent test set, with an AUC of 0.988 (95% CI: 0.967-1), sensitivity of 0.972, and specificity of 0.826. Conclusions: Patients presenting with severe PH imaging features but lacking typical cirrhosis imaging features should be thoroughly evaluated for PSVD. MRI is the preferred imaging modality. Our MRI-based predictive model reliably differentiates PSVD from cirrhosis, offering a non-invasive method for enhancing the suspicion of PSVD.
Despite the centrality of syndrome differentiation in guiding personalized traditional Chinese medicine (TCM) interventions for coronary heart disease (CHD), existing studies of TCM syndrome distribution are constrained by fragmented methodologies and limited spatiotemporal resolution. In this study, we employed an artificial intelligence (AI)-assisted literature mining framework combined with statistical and pattern-mining analyses to characterize the spatiotemporal patterns and evolutionary dynamics of CHD syndrome elements across China. We mined data from 651 peer-reviewed publications reporting 234,656 CHD patients, published between January 2000 and December 2023. A customized constrained Sequential PAttern Discovery using Equivalence classes algorithm was applied to quantify syndrome transition trajectories, while hierarchical clustering integrated with similarity-matrix heatmaps was utilized to detect syndrome co-occurrence patterns. Analyses were stratified by four disease stages, six geographic regions and three periods (2000-2009, 2010-2019, post-2020). Meta-analysis and clustering results revealed that blood stasis, qi deficiency and phlegm emerged as conserved core syndrome elements across all stages, dynamically interacting with qi-yin dual deficiency and phlegm-turbidity as the disease progressed. Geospatial analysis identified a Northern China predominance of qi stagnation-blood stasis, whereas humid subtropical regions exhibited accelerated phlegm-stasis synergy. Temporal analysis uncovered a paradigm shift toward blood stasis-phlegm synergy, coinciding with modern dietary transitions and sedentary lifestyle trends. By leveraging AI-driven methodologies, this study reveals the spatiotemporal heterogeneity of CHD syndrome patterns and underscores the necessity for geographically personalized and temporally adaptive TCM therapeutic strategies for CHD management.
Evidence-based medicine (EBM), formalized in the 1990s, has redefined clinical practice by advocating the integration of research evidence, clinical expertise, and patient values. This paradigm has introduced methodological rigor through randomized controlled trials (RCTs) to establish causality, systematic reviews to synthesize findings, and the GRADE approach to evaluate evidence based on risk of bias, inconsistency, indirectness, imprecision, and publication bias. These advancements have shaped clinical guidelines, reduced practice variability, and influenced medical education toward evidence-based inquiry. Despite its contributions, EBM faces challenges in the evolving landscape of modern medicine. The lengthy process of evidence generation, often requiring years for trials and guideline updates, limits responsiveness to emerging health needs, as observed during the COVID-19 pandemic. The external validity of RCT results is constrained by strict inclusion criteria, posing difficulties in applying findings to diverse patient populations with comorbidities. Additionally, the siloed nature of evidence complicates comprehensive care for multifactorial conditions, while the annual influx of over one million medical publications overwhelms traditional synthesis methods. Artificial intelligence (AI) presents a promising avenue to address these issues, leveraging capabilities in processing heterogeneous data. Natural language processing may enhance literature analysis, machine learning could identify patterns in complex datasets, and causal inference might improve the reliability of observational data insights. These technologies hold potential to accelerate evidence development and tailor it to individual needs. This paper proposes digital intelligent evidence-based medicine (i-EBM) as a conceptual evolution of EBM, designed for the AI era. i-EBM envisions a three-layered framework. The data foundation layer aims to integrate structured evidence from RCTs, domain knowledge such as biomedical ontologies and traditional Chinese medicine principles, and multi- modal patient data, including electronic health records, genomics, and wearable device outputs. Knowledge graphs are proposed to link these elements into a unified, computable knowledge network. The intelligent processing layer seeks to apply AI for evidence retrieval, data extraction, quality assessment, and synthesis, potentially using large language models to assist these processes. The knowledge service layer intends to provide dynamic guidelines and individualized predictions, supported by ongoing human-machine collaboration to ensure clinical relevance and ethical considerations. i-EBM has the potential to mitigate EBM's limitations by facilitating real-time evidence updates, reducing knowledge fragmentation through integrated data, and offering personalized decision support. For instance, it may support precision medicine by connecting diverse data sources, with applications possibly extending to fields like oncology or traditional Chinese medicine. Future research could explore autonomous AI systems, optimized clinical workflows, and governance frameworks to address data privacy, bias, and global standardization. In conclusion, i-EBM offers a theoretical framework to extend EBM principles, harnessing AI's potential alongside human expertise to advance medical research and practice. Meanwhile, for issues such as the quantitative study of the complex intervention characteristics and syndrome differentiation patterns of traditional Chinese medicine, i-EBM can provide methodological support in data integration, pattern recognition, and causal inference, offering potential tools and insights for uncovering the intrinsic regularities of TCM evidence and optimizing its evaluative framework.
Abstract: Traditional Chinese medicine (TCM) clinical practice guidelines are important for standardizing practice and supporting evidence-informed decision-making, but their applicability and real-world use remain limited. As clinical needs vary across settings, conventional guideline development approaches are insufficient. This article explores the development of TCM clinical practice guidelines in varied scenarios, focusing on rapid and living guidelines in public health emergencies and clinical application guidelines in drug use scenarios, particularly for Chinese patent medicines (CPMs). It summarizes the current situation, highlights key methodological challenges and proposes scenario-specific strategies, including rapid evidence synthesis, dynamic updating, conflict-of-interest management, and clearer, more implementable recommendations. A scenario-oriented approach may improve the rigor, responsiveness, and applicability of TCM guidelines, thereby promoting the translation of evidence into practice and the high-quality development of TCM healthcare.
Quercetin is a flavonoid compound that has demonstrated substantial potential in the treatment of diabetic kidney disease (DKD). However, there is still a lack of systematic research on the exact mechanism of action of quercetin. This review discusses the druggability, molecular targets, and signaling pathways of quercetin in DKD treatment. We retrieved the latest research on the pharmacological effects and mechanisms related to quercetin from PubMed and Scopus as of June 2025 (2012–2025). Evidence suggests that quercetin has the potential to eliminate senescent cells in DKD. Network pharmacology was used to predict the targets and pathways of quercetin in targeting cellular senescence to treat DKD. Using on existing research, it was further confirmed that quercetin can effectively act on hub target and pathway. The mechanism of quercetin therapy in DKD was summarized from three dimensions: inflammation, oxidative stress, and cell death. This review highlights the potential of quercetin for treating DKD by providing a biological basis for its mechanism of action and its use as a senolytic drug for this disease.
Pediatric influenza is a major public health concern. While Chinese patent medicines play a significant role in its treatment, information regarding drug recommendations in existing clinical guidelines is often fragmented, and the logic of syndrome differentiation-based medication is complex, limiting the efficiency of clinical decision-making. The objective of this study is to construct a structured Knowledge Graph for the treatment of pediatric influenza with Chinese patent medicines by integrating multiple authoritative guidelines and consensuses, and to develop an intelligent Question-Answering System based on this graph to provide precise clinical decision support. Using three guidelines and consensuses—including the Clinical Practice Guidelines for the Treatment of Pediatric Influenza with Chinese Patent Medicines (2024)—as core data sources, we manually and systematically extracted and standardized 11 types of entities, including influenza diagnosis, symptoms, Traditional Chinese Medicine (TCM) syndromes, therapeutic drugs, usage methods, and evidence levels. A domain ontology was constructed to define multidimensional semantic relationships between entities, and the Neo4j graph database was utilized for knowledge storage and visualization. The constructed Knowledge Graph consists of 433 nodes and 604 relationships, integrating 22 recommended Chinese patent medicines, 21 symptoms, and 6 types of TCM syndromes. The graph visualizes the dynamic decision-making path of "symptom-syndrome-drug" and embeds GRADE evidence levels and recommendation strengths. It supports intelligent symptom-based queries and provides comprehensive decision support, including drug recommendations, specific usage and dosage, combination medication suggestions, and risk warnings for contraindications. This study successfully constructed a structured Knowledge Graph for the treatment of pediatric influenza with Chinese patent medicines, effectively addressing the fragmentation of knowledge in traditional guidelines. This graph assists physicians, particularly those in primary care and Western medicine practitioners, in rapidly understanding the logic of syndrome differentiation and making standardized medication decisions, thereby reducing the barrier to diagnosis and treatment of pediatric influenza. This study provides a feasible paradigm for the digital transformation of TCM guideline knowledge and the development of clinical intelligent auxiliary tools.
Despite years of methodological progress, how far AI has come in liver fibrosis staging has never been systematically evaluated under the heterogeneous, multi-center conditions that define clinical practice. To address this gap, we introduce LiFS, a large-scale dataset and benchmark derived from the MICCAI 2025 CARE-Liver challenge, comprising 610 patients across multiple centers and scanners with multi-sequence MRI. To the best of our knowledge, LiFS is the first benchmark providing complete gadoxetic acid-enhanced sequences with histopathology-confirmed annotations from diverse real-world scanners. Through systematic evaluation of 9 independently developed methods selected from 96 registered teams against in-cohort radiologist reference results, our findings address how far current AI has progressed toward clinical-level liver fibrosis staging from three complementary perspectives. First, against radiologists, the best AI methods were broadly comparable to the senior radiologist and significantly exceeded the junior radiologist in selected settings, while median AI performance generally approached junior-radiologist levels. Second, from a data perspective, cross-center heterogeneity, label imbalance, and contrast-enhanced sequence variability emerge as the dominant challenges for AI methods. Third, from a technical perspective, methodological design choices, including spatial registration, input dimensionality, multi-modal fusion strategy, and backbone architecture, appear to modulate cross-center robustness, although no single choice alone closes the gap. Overall, LiFS provides a rigorous real-world benchmark for positioning the current state of AI in liver fibrosis staging and for enabling future research on the key challenges that limit clinically reliable deployment.
BACKGROUND The Liver Imaging Reporting and Data System (LI-RADS) is widely used for the diagnosis of hepatocellular carcinoma, but feature scoring by radiologists is subjective and time-consuming. An urgent need exists for an objective and efficient radiologist-supervised automated LI-RADS categorization system. AIM To develop an evidence-based radiologist-supervised automated LI-RADS grade 3 (LR-3), 4 (LR-4) and 5 (LR-5) categorization system (Evi-LIRADS) through quantitative feature characterization, following LI-RADS v2018. METHODS This retrospective multicenter study (April 2012-November 2022) included untreated patients with suspected hepatocellular carcinoma undergoing gadoxetic acid-enhanced magnetic resonance imaging. Lesions from center 1 were partitioned into a development set (275 lesions used for five-fold cross-validation) and an internal testing set (62 lesions). Lesions from centers 2 (85 lesions) and 3 (104 lesions) constituted two external testing sets. Evi-LIRADS was designed by emulating the decision-making process of radiologists through a series of image processing algorithms, to recognize nonrim arterial phase hyper-enhancement, nonperipheral washout, and enhancing capsule, which provided detailed assessments of feature locations and patterns, improving the transparency of feature classification. Based on the three major image features and the automatically segmented lesion size, LI-RADS categories were assigned using LI-RADS v2018 algorithm. Feature classification was evaluated using area under the receiver operating characteristic curve. LI-RADS categorization was assessed by accuracy. RESULTS The internal dataset included 337 patients from center 1, while external datasets comprised 76 patients from center 2 and 97 patients from center 3. For feature classification, areas under the receiver operating characteristic curves were 0.975, 0.898, and 0.940 for arterial phase hyper-enhancement; 0.803, 0.824, and 0.850 for washout; 0.759, 0.800, and 0.784 for capsule across three datasets. Three-class LI-RADS categorization among LR-3, LR-4 and LR-5 achieved accuracies of 80.6%, 74.1%, and 77.9%, respectively, surpassing comparison methods (58.6%-69.6%). LI-RADS categorization between LR-3 and combined LR-4/LR-5 achieved 95.2%, 88.2%, and 90.4% accuracies for the three datasets, respectively. The visualization provided detailed feature locations and patterns. Evi-LIRADS saved an average of 21.1 seconds per patient (58.8% of the time) compared with radiologists, excluding radiologists' quality control time. CONCLUSION Following LI-RADS guidelines and radiologists' decision-making process, Evi-LIRADS was developed through quantitative feature characterization, demonstrating good accuracy, robust generalization, improved efficiency, enhanced clinical relevance, and improved transparency.
Numerous mathematical models use various predictors and modeling approaches to forecast pandemics; however, given the diversity of these models, it remains unclear which type offers the most advantage. This study aimed to systematically evaluate the accuracy of models that incorporate natural, social, and pollution-related environmental factors to predict COVID-19 transmission, and to explore the potential role of predictive models in managing future large-scale epidemics. The results revealed that, although only six studies used mathematical models with both natural and social environmental predictors, these models outperformed those with only natural or pollution predictors. Currently, most models focus primarily on temperature and humidity, with limited emphasis on other environmental factors, such as wind speed and pollution. In future pandemic responses, utilizing more multidimensional and diverse predictors can improve the accuracy and clinical applicability of the model and reduce the detrimental effects of disease transmission and spread.
BackgroundChildren are the main group affected by the influenza virus, posing challenges to their health. The high risk of viral variability, drug resistance, and drug development leads to a scarcity of therapeutic drugs. Baikening (BKN) granules are a marketed traditional Chinese medicine used to treat children’s lung heat, asthma, whooping cough, etc. Therefore, exploring the potential mechanisms of BKN in treating pediatric influenza is of great significance for discovering new drugs.MethodsThrough the database, we obtained differentially expressed genes (DEGs) between pediatric influenza and healthy samples, identified the components of BKN, and collected the targets. Target networks were built with the purpose of screening both targets and key components. Pathway and function enrichment were conducted on the relevant targets of BKN for treating pediatric influenza. BKN-related hub genes for influenza were discovered through DEGs, weighted gene co-expression network analysis (WGCNA), BKN-cluster WGCNA, and machine learning model. The accuracy of prediction efficiency and the value of BKN-related hub gene were validated through analysis of external datasets and receiver operating characteristics. Ultimately, simulations using molecular docking and molecular dynamics were used to forecast how active components will bind to hub genes.ResultA total of 20 candidate active compounds, 58 potential targets, and 3,819 DEGs were identified. The target network screened the top 10 key components and 6 core targets (PPARG, MMP2, GSK3B, PARP1, CCNA2, and IGF1). Potential target enrichment analysis indicated that BKN may be involved in AMPK signaling pathway, PI3K Akt signaling pathway, etc., to combat pediatric influenza. Subsequently, two hub genes (OTOF, IFI27) were obtained through WGCNA, BKN-cluster WGCNA, and machine learning models as potential biomarkers for BKN-related pediatric influenza. Two hub genes were found to have primary diagnostic value based on ROC curve analysis. Molecular docking confirmed the binding between BKN and hub gene. Molecular dynamics further revealed the stable binding between Peimisine and hub genes.ConclusionBKN may alleviate pediatric influenza via key components targeting core targets (PPARG, MMP2, GSK3B, PARP1, CCNA2, and IGF1) and hub genes (OTOF, IFI27), with the involvement of feature genes-related pathways. These results have potential consequences for future research and clinical practice.
ABSTRACT Objective To explore patients' perceptions and attitudes towards patient guidelines (PGs) and to identify specific factors related to PG content, design, presentation, and management that may influence patients' use or adoption of PGs. Methods An exploratory sequential mixed‐methods design was employed. Initial semi‐structured interviews were conducted with a diverse group of individuals, including people with diabetes or oncology, and clinicians. These interviews were analysed through directed content analysis. Findings from the qualitative study were used to develop a questionnaire. The questionnaire was circulated to patients with diabetes and cancer and asked them to report their awareness, attitudes, and the PG‐related factors influencing their use and adoption of PGs. Results In total, 25 participants were interviewed qualitatively, and 400 participated in the quantitative survey. Analysis of interviews yielded three themes: perception of PGs, attitude towards PGs, and key PG attributes influencing patients' use or adoption of PG. Qualitative findings indicated limited awareness of PGs, supported quantitatively by only 26.5% of patients being aware of PGs. Attitudes varied, with 73.0% expressing an overall positive attitude towards PG, but only 17.3% preferred PGs for evidence‐based answers and 32.3% favoured them for decision support, citing concerns that general recommendations may not meet individual needs. Participants suggested tailoring recommendations based on subgroups considering age, comorbidity, and weight, explaining why treatments work or don't work in different populations. Eight PG attributes influencing their use or adoption were found: accessibility, identifiability, attractiveness, credibility, usability, timeliness, relevance and simplicity. Lack of credibility was the most frequently mentioned hindrance, with 34.8% identifying unverifiable information as a barrier. Further qualitative and quantitative analyses revealed that medical staff were trusted sources for conveying PGs to patients. Conclusions and Practice Implications This study underscores the necessity for PGs to acknowledge individual differences and provide recommendations that are more tailored considering age, comorbidities, weight, and other factors influencing decision‐making, ensuring that they address patients' specific needs and support informed decision‐making. Additionally, there was a significant need to improve the dissemination of PGs, using medical staff as key channels to improve patients' use or adoption of PGs. Patient or Public Contribution In this study, patients were actively involved in several stages. During the development of the interview guide, feedback from two patients, alongside one patient guidelines (PGs) developer and three clinicians, was incorporated to ensure the guide's relevance and comprehensiveness. Patients' insights were integral to refining the interview questions, ensuring they were appropriate and effective. Additionally, the survey questionnaire was pre‐tested among 20 patients using the Think‐Aloud method, which led to significant revisions for better comprehension and response quality. These steps highlight the essential role of patients in shaping the data collection instruments and enhancing the overall quality and relevance of the study.
Experimental evidence suggests that alkaloids have anti-influenza and anti-inflammatory effects. However, the risk of translating existing evidence into clinical practice is relatively high. We conducted a systematic review and meta-analysis of animal studies to evaluate the therapeutic effects of alkaloids in treating influenza, providing valuable references for future studies. Seven electronic databases were searched until October 2024 for relevant studies. The Review Manager 5.2 software was utilized to perform the meta-analysis. Our study was registered within the International Prospective Register of Systematic Reviews (PROSPERO) as number CRD42024607535. Alkaloids are significantly correlated with viral titers, pulmonary inflammation scores, survival rates, lung indices, and body weight. However, alkaloid therapy is not effective in reducing the levels of tumor necrosis factor-α (TNF-α) and interleukin-6 (IL-6). In addition, the therapeutic effects of alkaloids may be related to the inhibition of the Toll-like receptor 4 or 7/Nuclear factor (NF)-κB signaling pathway, NACHT, LRR, and PYD domains-containing protein 3 (NLRP3) inflammasome pathway, and the Antiviral innate immune response receptor RIG-I (RIG-I) pathway. Alkaloids are potential candidates for the prevention and treatment of influenza. However, extensive preclinical studies and clinical studies are needed to confirm the anti-influenza and anti-inflammatory properties of alkaloids.
Introduction: Retinal vein occlusion (RVO) is the second most prevalent retinal vascular disease. In China, Fufang Xueshuantong Capsule (FXC) is a common compound Chinese medicine used to treat RVO. This review aims to illustrate the therapeutic efficacy and pharmacological mechanisms of FXC for RVO. Methods: Eight databases were searched until 17 September 2024 to collect literature meeting the criteria. Two investigators independently assessed the quality of the review and used Stata 16.0 to meta-analyse each metric. Then, network pharmacology was performed to explore the underlying mechanisms of FXC in the treatment of RVO. Results: A meta-analysis included 18 randomized control trials (RCTs) with 1574 patients. Our results indicated that compared to conventional/traditional Chinese medicine injection (TCM) therapies, FXC alone or in combination can enhance clinical effectiveness, raise best-corrected visual acuity (BCVA), reduce central macular thickness (CMT), improve hemorheology, ameliorate hemodynamic indices, enhance fundus morphology, and decrease TCM scores with a lower adverse reaction incident. Network pharmacology revealed core components included quercetin, (3-sitosterol, formononetin, luteolin, tanshinone IIA, etc., and the core targets such as TGF(31, PKC, JNK, ERK1/2, VEGF, MCP-1, VACM-1, IL-6, TNF-alpha, and Akt. Several of these targets and pathways are related to RVO, and the pharmacological mechanism is mainly related to the AGE-RAGE signaling pathway in diabetic complications. Conclusion: FXC is clinically effective and safe for treating RVO. It synergistically treats RVO through a multicomponent, multi-target, and multi-pathway network. The AGE-RAGE signaling pathway in diabetic complications may be one of the key pathways of FXC to ameliorate inflammation and thrombosis in RVO. Further basic experiments and high-quality RCTs are still needed to enrich the evidence chain for FXC in treating RVO.
Liver fibrosis represents a significant global health burden, necessitating accurate staging for effective clinical management. This report introduces the LiQA (Liver Fibrosis Quantification and Analysis) dataset, established as part of the CARE 2024 challenge. Comprising 440 patients with multi-phase, multi-center MRI scans, the dataset is curated to benchmark algorithms for Liver Segmentation (LiSeg) and Liver Fibrosis Staging (LiFS) under complex real-world conditions, including domain shifts, missing modalities, and spatial misalignment. We further describe the challenge's top-performing methodology, which integrates a semi-supervised learning framework with external data for robust segmentation, and utilizes a multi-view consensus approach with Class Activation Map (CAM)-based regularization for staging. Evaluation of this baseline demonstrates that leveraging multi-source data and anatomical constraints significantly enhances model robustness in clinical settings.
Abstract Background The integration of traditional Chinese medicine (TCM) into emergency health systems in China serves as a model for global policy development and refining the inclusion of traditional medicine in health emergencies. Methods This study investigated 13 public health emergency policies related to TCM released by the Chinese central government from 2003–2023. A PMC(Policy Modeling Consistency) index model was developed combining ROSTCM text mining analysis software. The contents of these policy documents were quantitatively assessed using 10 first- and 40 s-level indicators. Results The content analysis results showed that current policies focus on emergency treatment, and that the State Administration of Traditional Chinese Medicine is the issuing authority of the main policies, most of which are issued in the form of a notice. The scoring results for the 13 policies showed that two, five, three, and three policies were rated as excellent, good, qualified, and unqualified, respectively. This indicates that the policy quality related to TCM use in emergency response was normally distributed and generally qualified, although room for further improvement exists; policies should follow the principles of science, reasonableness, and operability, and should be updated in a timely manner with continuous development of the governance period while focusing on the policy content, safeguards, and role measures. Conclusion Effective integration of traditional medicine into health emergency policies backed by state institutions is vital. This includes enforcing relevant laws and regulations, establishing multidisciplinary medical teams, and developing integrated medicine strategies that support clinical research and maximize the unique benefits of traditional medicine.
Objective:Whether large language models (LLMs) can effectively facilitate CM knowledge acquisition remains uncertain. This study aims to assess the adherence of LLMs to Clinical Practice Guidelines (CPGs) in CM. Methods:This cross-sectional study randomly selected ten CPGs in CM and constructed 150 questions across three categories: medication based on differential diagnosis (MDD), specific prescription consultation (SPC), and CM theory analysis (CTA). Eight LLMs (GPT-4o, Claude-3.5 Sonnet, Moonshot-v1, ChatGLM-4, DeepSeek-v3, DeepSeek-r1, Claude-4 sonnet, and Claude-4 sonnet thinking) were evaluated using both English and Chinese queries. The main evaluation metrics included accuracy, readability, and use of safety disclaimers. Results:Overall, DeepSeek-v3 and DeepSeek-r1 demonstrated superior performance in both English (median 5.00, interquartile range (IQR) 4.00-5.00 vs. median 5.00, IQR 3.70-5.00) and Chinese (both median 5.00, IQR 4.30-5.00), significantly outperforming all other models. All models achieved significantly higher accuracy in Chinese versus English responses (all p < 0.05). Significant variations in accuracy were observed across the categories of questions, with MDD and SPC questions presenting more challenges than CTA questions. English responses had lower readability (mean flesch reading ease score 32.7) compared to Chinese responses. Moonshot-v1 provided the highest rate of safety disclaimers (98.7% English, 100% Chinese). Conclusion:LLMs showed varying degrees of potential for acquiring CM knowledge. The performance of DeepSeek-v3 and DeepSeek-r1 is satisfactory. Optimizing LLMs to become effective tools for disseminating CM information is an important direction for future development.
[This corrects the article DOI: 10.1016/j.jot.2025.05.004.].
Background:Post-COVID-19 condition presents complex symptomatology involving multifaceted interactions, which has resulted in a current lack of comprehensive understanding of its disease trajectory. This knowledge gap significantly compromises the efficiency of symptom management and adversely affects patients' quality of life. Objective:This study aims to comprehensively characterize the temporal evolution of post-COVID-19 condition by identifying core symptom clusters and clinical phenotypes, thereby enhancing understanding of the disease trajectory. Methods:The PubMed, Web of Science, and Embase databases were searched from December 1, 2019, to March 1, 2024. Observational studies related to the prevalence of symptoms in post-COVID-19 condition had been included. We conducted a meta-analysis to synthesize symptom prevalence across different follow-up intervals following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines and used a network to explore interrelationships and co-occurrence patterns among symptoms, enabling the identification of core symptoms and changes over time. Clustering analysis was used to classify included studies into distinct clinical subtypes. Results:This study analyzed 155 sets of macrolevel data from 108 clinical studies, encompassing 63,771 patients. Fatigue was the most prevalent symptom across all 4 follow-up points (52%, 48%, 46%, and 54%). Dyspnea peaked at the third and sixth follow-ups (36% and 31%) and then declined steadily (28% and 22%). Subgroup analysis revealed that Africa reported the fewest symptoms overall, yet showed high early incidences of fatigue (68%, 95% CI 50%-85%) and dyspnea (56%, 95% CI 15%-98%). The Americas placed greater emphasis on symptom evolution within the first postinfection year, with notably higher prevalence of anxiety (60%, 95% CI 54%-66%) and depression (36%, 95% CI 16%-55%). Asia and Europe documented the most comprehensive symptom profiles, with Asia reporting lower early dyspnea rates (29%, 95% CI 18%-40%) and Europe exhibiting more complex multisystem involvement during long-term follow-up. Network analysis showed that core post-COVID-19 symptoms evolved from early respiratory-neurological manifestations to chronic multisystem symptoms dominated by dizziness. Clustering analysis further indicated a progressive convergence of 2 initially distinct post-COVID-19 subtypes, with the acute inflammatory type becoming less prominent and gradually transitioning into a more chronic, persistent pattern. Conclusions:This study provides a comprehensive characterization of the dynamic evolution of post-COVID-19 condition symptoms and clinical subtypes, highlighting their multisystem involvement. The results reveal a progressive decline in respiratory symptoms over time, while neurological manifestations emerge as the most persistent and systemically impactful core symptoms. Our findings emphasize the need for region-specific surveillance and early warning systems informed by symptom progression patterns. By continuously monitoring the trajectories of symptom clusters, this approach offers valuable insights for identifying early warning signals and targeted intervention points in the management of postinfectious sequelae arising from future large-scale epidemics.