
Background:Biomedical informatics increasingly reuses clinical artificial intelligence (AI) benchmarks, yet successor releases are often treated as independent external corpora. If the same participants reappear under new corpus names, features, or partition files, evaluation can become internal while reported as external. Objectives:The objective of this study is to define successor-release participant leakage as a benchmark-integrity threat, operationalize a preevaluation audit and leak-free protocol, and demonstrate its impact using the Distress Analysis Interview Corpus, Wizard-of-Oz condition (DAIC-WOZ), and the Extended DAIC (E-DAIC) depression-screening benchmarks. Methods:We audited release lineage and participant identity using persistent identifiers, content hashing, Patient Health Questionnaire-8 (PHQ-8) label reconciliation, and fold-transition cross-tabulation. We compared leaky and participant-disjoint E-DAIC-to-DAIC-WOZ protocols on the same heldout participants using paired bootstrap inference and derived conservative leak-free reference baselines across standard pipelines. Results:All 189 DAIC-WOZ participants reappeared in E-DAIC with byte-identical recordings and identical PHQ-8 totals. Across official partitions, 104 of 189 changed fold. Training on E-DAIC and evaluating on the DAIC-WOZ test set placed 47/47 test participants in the model-development pool; the reverse direction was clean. The leaky acoustic protocol yielded area under the receiver operating characteristic curve (AUROC) 0.797 versus 0.569 after deduplication, with paired Delta AUROC + 0.227 (95% confidence interval + 0.106 to + 0.372; two-sided p < 0.001). At matched training size, Delta AUROC was + 0.249. A generic visual classifier reached AUROC 0.887 on a blended external set but 0.562 on the unseen AI-only portion. Leak-free reference baselines were approximately AUROC of 0.60. Conclusion:Successor-release participant leakage is a preventable evaluation failure in clinical AI benchmark reuse. Before pooling, external validation, or model comparison, biomedical informatics studies should document release lineage, verify participant identity at the signal or record level, reconcile labels, deduplicate across releases, and enforce participant-disjoint evaluation.
Background:Biomedical informatics plays a central role in supporting learning health systems (LHSs) that aim to continuously improve care by transforming data into knowledge and knowledge into action. However, health care data reflect existing disparities in access, delivery, and outcomes, raising concerns about whether the LHS cycle can improve care equitably without intentional design. Objective:This study aims to describe the recurring theme of health equity found in recent biomedical informatics literature, using a convenience sample of articles contributed to the 2025 American Medical Informatics Association (AMIA) Year in Review (YIR). Methods:In developing the 2025 AMIA YIR presentation, we found that 240 articles submitted by AMIA experts could be organized into multiple themes to allow summarization and synthesis for efficient presentation of 140 of them. Many themes could themselves be organized around the concept of the LHS, which further streamlined the presentation. We were struck by how frequently biomedical informatics literature addressed health equity in relation to LHSs. We therefore applied our own expertise to distill relevant articles from the larger set to explore this recurrent issue. Results:Here, we present 56 papers selected from 9 of the 19 themes that illustrate how biomedical informatics can support equity at each stage of the LHS. Mature LHS implementations provided large-scale examples of how learning cycles can be designed for equity. Real-World Evidence and Inclusive Data themes showed that equitable learning depends on whether data are complete enough to represent the populations the system intends to serve. Natural language processing and artificial intelligence (AI) modeling showed how data are transformed into knowledge and translated into clinical action through decision support. Trust and justice in data use, health equity as infrastructure, and equitable access emerged as interdependent conditions that influence whether the learning cycle operates equitably and effectively. Conclusion:Equity emerged as a condition influencing every stage of the LHS cycle. The value of informatics interventions cannot be judged by technical performance alone, but by whether they build learning cycles that are inclusive, accountable, and equitable for the populations they serve.
Background:Multi-informant observational data obtained from parents and educators provide rich but context-dependent information. However, these data are rarely transformed into structured representations that can be consistently shared and interpreted across contexts without relying on diagnostic classification. Objective:This study aimed to propose and implement a deterministic framework for transforming multi-informant observational data into structured, interpretable descriptor sets, and to examine its feasibility under controlled conditions. The framework targets an intermediate representation layer for structuring observational data prior to interpretation or decision-making. Methods:Parent- and educator-reported questionnaire data were used as standardized inputs to a predefined deterministic framework. Child-centered signals were derived from parent reports, while educator reports were incorporated as educator-derived contextual signals. The framework applies a transparent rule-based mapping process to generate context-preserving sets of predefined structured descriptors without reliance on machine learning or statistical modeling. Structural behavior was examined using controlled simulated input conditions. Results:The framework consistently generated a limited and prioritized set of descriptors for each input profile. Outputs were structurally stable and reproducible across all predefined input configurations, demonstrating consistent transformation of multi-informant data into structured representations. Conclusion:This study demonstrates the feasibility of a deterministic framework for organizing multi-informant observational data into structured, non-diagnostic descriptors. By introducing a reproducible intermediate organizational layer, the framework provides a transparent approach to cross-context information structuring that may be applicable to other multi-context observational settings involving multi-informant data integration.
Background:The rapid digitalization of health care services has introduced new safety risks related to patient confidentiality, the reliability of digital tools, and their integration into clinical workflows. While digital health adoption is emphasized in national and international policy, guidance on embedding digital clinical safety (DCS) principles in large organizations is limited. The Health New Zealand - Te Whatu Ora (Health NZ) Data and Digital team has developed the first DCS Framework for New Zealand, but mechanisms to support its organizational adoption have not yet been established. This article describes the methodological design of a DCS training program and identifies preimplementation considerations for integrating DCS principles into organizational practice. Methods:A qualitative exploratory study was conducted within Health NZ using purposive sampling and semistructured interviews with seven stakeholders involved in digital health, clinical informatics, and organizational learning. A structured options assessment was undertaken prior to the interviews to compare potential organizational approaches for implementing DCS capability. Stakeholder feedback on these options, together with interview findings, informed an integrated assessment of program feasibility and design. Results:Findings from the stakeholder interviews indicated that there is limited shared understanding of DCS across clinical and digital teams, strong support for DCS training that is accessible and relevant to everyday practice, and variable compliance with current mandatory training offered at Health NZ. The options assessment identified a comprehensive, modular training program as the most suitable option for embedding DCS principles within Health NZ policies and practices. Conclusion:This study presents a structured methodology informed by the DCS Framework for designing DCS training within a national health system. Findings highlight that training can serve as a key mechanism for organizational change, promoting shared understanding, engagement, and consistent adoption of safe digital practices. This approach provides transferable insights into embedding DCS practices within a complex health system to support safe, resilient, and equitable digital health care delivery.
Objectives:A publicly shared digital phenotyping dataset was recently used to train models predicting day-before-panic (DBP) events, reporting a cross-validated receiver operating characteristic area under the curve (ROC-AUC) of around 0.90. Because the stratified five-fold cross-validation (CV) allowed the same participant's days in both training and test folds, it was unclear whether this reflects between-person generalization or information leakage. We reanalyzed the public dataset to quantify this distinction. Methods:We reanalyzed the public dataset (3,969 person-days; 254 DBP events, 6.4%, across 19 of 43 participants) with the same XGBoost classifier and 71 features, comparing stratified five-fold cross-validation with leave-one-subject-out (LOSO) cross-validation. A retrospective diagnostic decomposition was applied within-subject z-scoring to dynamic features before LOSO. Results:We reproduced the original ROC-AUC at 0.895 ± 0.011. Under LOSO, pooled ROC-AUC dropped to 0.489 (bootstrap 95% CI: 0.325-0.643, including 0.50). The SHapley Additive exPlanations (SHAP) ranking under stratified CV was dominated by static traits (CTQ subscales, BRIAN, SPAQ). The diagnostic decomposition yielded LOSO ROC-AUC 0.516 (precision-recall area under the curve [PR-AUC] 0.114), suggesting a weak residual within-person signal, though this should not be interpreted as prospective performance. Conclusion:Robust between-person generalization was not confirmed under subject-held-out evaluation. The feature ranking was trait-dominated, though the data cannot definitively distinguish identity discrimination from other between-person variance sources. Digital phenotyping studies should report subject-level cross-validation alongside standard metrics, per the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) and Prediction model Risk of Bias Assessment Tool (PROBAST) guidelines.
Background:The invasive mosquito Aedes aegypti is a major vector of arboviruses, such as Dengue, Zika, Chikungunya and Yellow Fever. Objective:As a first step toward a transmission model for vector-borne diseases, a weather-dependent population dynamic simulation for Ae. aegypti was developed, suitable for high weather short-term variability in the Austrian region. Methods:We developed an agent-based model, incorporating temperature- and precipitation-dependent development, mortality, movement, feeding behavior, and egg-laying. Species-specific parameters were derived from published experimental studies. Simulations were run for observed weather data from 2024 and for EURO-CORDEX climate projections (ÖKS15) for 2050 and 2080 under RCP4.5 and RCP8.5. Daily immigration of one adult female was assumed to mimic human-mediated introduction along major transport routes. Results:Under all scenarios, population development began in late May and ceased by late September. Observed 2024 conditions produced the highest population sizes (area under the curve [AUC] = 361,343 individuals), whereas all future projections resulted in substantially lower abundances, despite higher mean temperatures. Conclusion:This discrepancy was driven by stronger fluctuations in temperature and precipitation, as well as the absence of urban heat-island effects in climate projections. Under no scenario did mosquitoes survive winter, indicating that long-term establishment is highly unlikely. The model provides a foundation for future extension toward spatially explicit virus transmission simulations, and for assessing the public health implications of climate change in alpine regions.
Background:Secondary use of health and social care registry data enables evaluation of service effectiveness. The focus is shifting from organization-based metrics to person-centered, system-level approaches that understand outcomes as the combined result of interacting actors. In this context, digital service use offers an important lens for understanding accessibility, participation, and equity within public services. Objectives:The aim of this study is to develop an interpretable analysis approach that uses a pre-existing person-centered, registry-based data model to identify individual variables associated with clients' digital service use. Methods:This registry-based study in a Finnish wellbeing services county applied a person-centered data model linking variables across health care, social services, and client experience. Digital service use was analyzed across 17 reason-for-encounter groups defined by the International Classification of Primary Care (ICPC) and one system-generated supplementary category arising from the dataset's internal structure. Subgroup analysis covered single and paired variables. Effect size was quantified using lift values from Market Basket Analysis, and generalizability was assessed aggregating uncorrected directional test outcomes across ICPC groups. Results:We developed a subgroup-based dual-loop analysis method to identify variable-specific associations across 18 ICPC categories. The method showed that digital service use was most common among younger individuals, women, and urban residents. Low service need and mental-health or family-center-related variables were associated with a higher likelihood of choosing a digital channel, whereas chronic illnesses, intensive care needs, and continuity of care were associated with a lower likelihood of digital channel selection. Conclusion:Registry data and a person-centered model reveal how individual variables influence digital service choices. The method enabled interpretable insights into single variables and their combinations. The findings support more equitable digital services and emphasize the need to tailor solutions to meet diverse client needs.
Background:Hospital discharge represents a critical moment in the inpatient care process, and its documentation, known as the discharge summary, plays a fundamental role. Incomplete discharge summaries may lead to hospital readmissions, adverse events, poor understanding of treatment and disease, patient and family dissatisfaction, lack of follow-up, and discontinuity of care. In 2017, Hospital Municipal de Agudos "Dr. Leónidas Lucero" (Bahía Blanca, Buenos Aires Province, Argentina) implemented an electronic health record system. Although substantial experience has been gained in its use, the quality of discharge summaries has never been systematically evaluated. Objective:To assess the quality of discharge summaries produced between July 2023 and June 2024 by the Internal Medicine Inpatient Service using the Physician Documentation Quality Instrument-9 Spanish Version (PDQI-9 SV). Methods:An observational, descriptive, cross-sectional study was conducted. The PDQI-9 SV consists of nine items (up-to-date, accurate, thorough, useful, organized, comprehensible, concise, synthesized, and internally consistent), each scored on a 5-point Likert scale. Prior to data collection, an instrument adaptation process and reviewer training were carried out. The assessment was done by two trained physician reviewers. Results:A total of 135 discharge summaries were analyzed. The mean overall quality score was 34.77 (standard deviation [SD] = 4.17). The highest-scoring domains were accuracy (mean = 4.32; SD = 0.43), completeness (mean = 4.02; SD = 0.85), and internal consistency (mean = 4.01; SD = 0.65). The lowest-scoring domain was conciseness (mean = 3.28; SD = 0.90). No statistically significant differences were observed according to physician seniority (p = 0.757) or day of the week (p = 0.809). Interrater agreement was moderate (intraclass correlation coefficient = 0.63; 95% confidence interval: 0.45-0.75). Conclusion:Overall discharge summary quality was good, with strengths in accuracy, completeness, and internal consistency. The absence of differences according to physician seniority or day of completion suggests homogeneous documentation practices. Lower scores in conciseness highlight a clear opportunity for targeted educational interventions.
Background:The multilevel semantic structure of traditional Chinese medicine (TCM) formulas makes their efficacy difficult to represent computationally. Objective:This study aimed to develop an interpretable, statistically rigorous model for quantitatively predicting the dominant efficacies of classical TCM herbal formulas. Methods:A knowledge graph encompassing five semantic entities-disease, syndrome, symptom, efficacy, and herb-was constructed to standardize and infer multilevel efficacy relationships. Based on this structure, the Hypergeometric Efficacy Prediction Model (HEPM) was established, using hypergeometric enrichment analysis to assess whether specific efficacies are significantly aggregated within a formula. A curated dataset of 174 classical formulas from authoritative TCM sources was used for model validation. Results:HEPM effectively reproduced characteristic efficacy patterns of classical prescriptions, achieving an average F1 score of 0.63 across 174 formulas. The knowledge graph structure resolved semantic inconsistency and incompleteness in traditional efficacy descriptions, enhancing the integrity and computability of efficacy information. Conclusion:HEPM provides a statistically grounded and interpretable framework for modeling efficacy formation in TCM herbal formulas. The method offers a replicable approach for efficacy prediction and supports the development of knowledge-driven intelligent TCM analysis and clinical decision support applications.
Feature extraction from free text medical reports is a frequently required clinical, operational, or research procedure. Large language models (LLMs) hold a promise for automating feature extraction, which can also enable category assignment tasks.To compare the groundedness of extracted features by five LLMs from magnetic resonance imaging (MRI) brain scan reports using a clinician-engineered versus an LLM-generated prompt.Five OpenAI LLMs were evaluated for their ability to extract nine binary features from synthetic MRI brain reports. Two types of prompts, a clinician-engineered and an LLM-generated, were used. Metrics including recall, precision, accuracy, and F1 score were calculated to assess model performance.For all extracted features by all studied models from both tested prompts, the overall average recall was 0.956, the average precision was 0.9347, the average accuracy was 0.982, and the average F1 score was 0.9431. Using GPT-3.5-turbo, the LLM-generated prompt had better numerical performance than the clinician-engineered prompt. For the other four GPT-4 models examined, overall recall, precision, and accuracy were higher regardless of the prompt source.This study highlights the potential of LLMs to generate prompts and accurately extract features, with newer models like GPT-4 performing consistently well. The efficacy of feature extraction by LLMs depends on the engineered prompt and model used. Our experimentation demonstrates the potential of LLMs to engineer prompts and extract features from MRI brain scan reports.
Chronic kidney disease, CKD in short, is a kind of long-term kidney illness in which rapid deterioration of kidney function is observed over a period of time. Unlike other organs, this damage in kidney function cannot be recovered and reversed as well. Moreover, in its early stages, asymptomatic renal disease is highly prevalent, making early identification with conventional clinical approaches difficult. Thus, early and accurate detection of risk factors is a very challenging step in CKD diagnosis. This research work showed earlier and effective identification of risk factors using notable feature selection techniques for the enhancement of patient care. It also aimed at the improvement of predictive diagnosis of CKD employing different supervised and ensemble machine learning classifiers. A CKD-focused dataset consisting of 1,032 patient records and 14 features was used for this research purpose. This research emphasized on identifying the risk factors of CKD using feature importance (for tree-based model) with sequential feature selector and ReliefF algorithm as feature selection process. Based on the ranking for both feature selection techniques, the top 10 features were identified. Then utilizing those features, the classifiers such as random forest, support vector machine, Naïve Bayes, decision tree, logistic regression, gradient boosting, K-nearest neighbors, and ensemble classifier voting technique were trained using stratified 5-fold and grid-based search cross-validation techniques. After that, their performances were assessed using evaluation measures, i.e., accuracy, F1 score, precision, recall, training loss, test loss, bias, and AUC, to classify the individual having presence or absence of CKD. The feature selection algorithms selected the significant data-driven top 10 features. Based on the ranking for both feature selection procedures, hemoglobin is determined to be the significant risk factor among these features. For both feature selection techniques, all the classifiers showed their best performance, having 86 to 98% of accuracy, AUC value of over 0.96 to 1.00, and bias value of 0.003 to 0.103. All the classifiers showed a very good trade-off between false positives and false negatives, with precision, recall, and F1 score ranging from 92 to 98%, 90 to 99%, and 93 to 98%, respectively, using feature importance with SFS. In both cases of the feature selection techniques, gradient boosting outperformed all other algorithms in terms of accuracy, precision, AUC, recall, F1 score, specificity, and bias. To conclude, in the suggested methodology the feature selection algorithms effectively identified the prominent features based on their importance, and the pipeline demonstrated a good performance in diagnosing individuals at risk of CKD development. Some of the classifiers showed their effectiveness in CKD prediction using the selected features by achieving higher accuracy, F1 score, precision, recall, AUC, specificity, and lower bias to ensure the diagnostic performance. Therefore, it can be inferred that this proposed methodology, combining the power of these eight machine learning models with two efficient feature selection approaches, demonstrated that people at risk of this nephrological condition can be detected earlier, more accurately identifying increased risk factors than with conventional methods. This holds a great promise toward enhancing healthcare judgment and eventually ensuring treatment for patients.
Standardized clinical terminology is essential for semantic interoperability. Typically, a hospital's terminology expert manually maps local terminology with international standards such as SNOMED CT. The manual mapping process is demanding, labor-intensive, and time-consuming, and its effectiveness relies on the expertise of the professional handling it.We developed a method to map clinical terms to SNOMED CT concept descriptions using an information retrieval (IR) approach with rich synonyms. We also provide a free mapping support service to help terminology experts alleviate the challenges of manual mapping without the need for additional manipulation.We created indexes using edge n-grams and synonyms. We adopted Elasticsearch for indexing and query processing, incorporating data from the SPECIALIST Lexicon to enrich the synonym database. Eight different indexes were initially created, but only four were retained based on performance. We tested indexes individually and in combination, using a dataset of 1,753 one-to-one mapped instances from the National Library of Medicine ICD-9-CM Procedure codes to the SNOMED CT Map. We compared our approach with MetaMap for evaluation.We found that using rich synonyms and edge n-gram indexing significantly improved the accuracy of mapping clinical terms to SNOMED CT. The indexes incorporating synonyms and edge n-grams performed better than those using either technique alone. Combining these methods captured more relevant terms and synonyms, resulting in more precise mappings. Our method outperformed the baseline provided by MetaMap, demonstrating enhanced capability in handling complex medical terminology and improving the overall mapping quality.Our study introduced an IR method with rich synonyms for mapping clinical terms to SNOMED CT, analyzing 40 unmapped terms, and identifying key issues. The approach shows promise in improving terminology mapping, and future work will explore advanced methods to enhance accuracy further, aiming to reduce manual mapping efforts and improve result evaluation.
Multicentre evaluations are key to provide generalisable results on medication safety interventions. Yet, for computerised physician order entry (CPOE) assessment, evaluations are mostly singe-centred and poorly comparable.We performed a multicentre simulation-based lab study at three independent study sites implementing varying CPOE systems with the aim to compare the effects of CPOE implementation on workflow changes, time requirement, and quality of medication documentation.At each study site, medication documentation processes with and without CPOE usage were analysed. Based on patient case scenarios, a simulation-based lab study was performed where the time required to document medication according to the analysed processes was measured. Additionally, the quality of medication documentation with and without CPOE usage was evaluated. Results were compared between the three study sites.At two hospitals, CPOE implementation led to a streamlining of the medication documentation process with less documentation systems and professional groups involved and an elimination of double documentation. However, at one hospital multiple documentation of medication orders persisted even after CPOE implementation. In total, the time required to document medication according to the case scenarios was faster without than with a CPOE system (median time without CPOE = 13:56 minutes [range = 07:29-32:17], median time with CPOE = 16:51 minutes [09:18-33:41], p = 0.047, +20.8%). One hospital was taking considerably longer for medication documentation than the two others, both with and without CPOE usage. Medication documentation quality rose significantly and to a median of 100.0% at each study site with CPOE usage.This study showed that a simulation-based lab study methodology is suitable for comparing CPOE effects across study sites. Furthermore, it provides evidence on the changes in medication documentation workflow, time requirement, and quality that occur with a CPOE implementation.
Cancer registries collect extensive data on cancer patients, including diagnoses, treatments, and disease progression. These data offer valuable insights into cancer care, but it is challenging to analyze due to its complexity. Machine learning techniques, particularly clustering, enable the exploration of treatment data to uncover previously unknown patterns and relationships. This work aimed to develop a method for clustering breast cancer patients in cancer registries based on their treatment courses, to demonstrate the usefulness of clustering for gaining insights, improving data quality, and identifying clinically relevant patterns. We developed a similarity measure adapted from the Levenshtein distance to compare treatment courses, incorporating cancer diagnosis, surgeries, radiotherapies, and systemic therapies. The method was evaluated on 17,822 breast cancer cases diagnosed in 2019 from the cancer registry of North Rhine-Westphalia. Evaluation involved two stages: first, domain experts reviewed the clustering results to assess clinical relevance and interpretability. Second, an intercluster survival analysis was performed to identify clinically relevant differences between treatment patterns. Expert evaluations confirmed that clustering produced clinically plausible groups while also uncovering unexpected treatment patterns and potential data inconsistencies. The survival analysis showed differences in survival between clusters in both prognostically favorable and unfavorable subgroups. These results demonstrate that treatment-course clustering can identify patient groups with differing survival outcomes. However, registry data incompleteness and unmeasured confounders may influence these findings. Clustering treatment courses in cancer registries can reveal data quality issues, distinguish groups with different prognostic profiles, and support exploratory analyses of treatment patterns. While these findings are not intended to guide clinical decision making or evaluate treatment effectiveness, they can help generate hypotheses, identify unexpected care pathways, and support quality monitoring within cancer registries. Future work should focus on improving treatment data completeness, incorporating additional clinical variables, and refining clustering methods for broader applicability.
BACKGROUND:Existing drug recommendation systems lack integration with up-to-date clinical guidelines (the latest diabetes association standards of care and clinical guidelines that align with local government healthcare regulations) and lack high-precision drug interaction processing, explainability, and dynamic dosage adjustment. As a result, the recommendations generated by these systems are often inaccurate and do not align with local standards, greatly limiting their practicality. OBJECTIVE:To develop a personalized drug recommendation and dosage optimization system named Diabetes Drug Recommendation System (DDRs), integrating FHIR-standardized EHR data and up-to-date clinical guidelines for accurate and practical recommendations. METHODS:We analyzed patients' EHR and ICD-10 codes and integrated them with a drug interaction database to reduce adverse reactions. ADA guidelines and Taiwan's NHI chronic disease guidelines served as data sources. Bio-GPT and RAG were used to build the clinical guideline database and ensure recommendations align with the latest standards, with references provided for interpretability. Finally, optimal dosage was dynamically calculated by integrating patient disease progression trends from the EHR. RESULT:DDRs achieved superior drug recommendation accuracy (PRAUC = 0.7951, Jaccard = 0.5632, F1-score = 0.7158), with a low DDI rate (4.73%) and dosage error (±6.21%). Faithfulness of recommendations reached 0.850. Field validation with three physicians showed that the system reduced literature review time by 30-40% and delivered clinically actionable recommendations. CONCLUSION:DDRs is the first system to integrate EHR data, LLMs, RAG, ADA guidelines, and Taiwan NHI policies for diabetes treatment. The system demonstrates high accuracy, safety, and interpretability, offering practical decision support in routine clinical settings.
Assessing treatment response in patients with myeloproliferative neoplasms is difficult because data components exist in unstructured bone marrow pathology (hematopathology) reports, which require specialized, manual annotation, and interpretation. Although natural language processing (NLP) has been successfully implemented for the extraction of features from solid tumor reports, little is known about its application to hematopathology.An open-source NLP framework called Leo was implemented to parse document segments and extract concept phrases utilized for assessing responses in myeloproliferative neoplasms. A reference standard was generated through the manual review of hematopathology notes.Compared with a reference standard (n = 300 reports), our NLP method extracted features such as aspirate myeloblasts (F1 = 98%) and biopsy reticulin fibrosis (F1 = 93%) with high accuracy. However, other values, such as myeloblasts from the biopsy (F1 = 6%) and via flow cytometry (F1 = 8%), were affected by sparsity representative of reporting conventions. The four features with the highest clinical importance were extracted with F1 scores exceeding 90%. Whereas manual annotation of 300 reports required 30 hours of staff effort, automated NLP required 3.5 hours of runtime for 34,301 reports.To the best of our knowledge, this is among the first studies to demonstrate the application of NLP to hematopathology for clinical feature extraction. The approach may inform efforts at other institutions, and the code is available at https://github.com/wcmc-research-informatics/BmrExtractor.
Syndrome is a unique and crucial concept in traditional Chinese medicine (TCM). However, much of the syndrome knowledge lacks systematic organization and correlation, and current information technologies are unsuitable for TCM ancient texts.We aimed to develop a knowledge graph that presents this knowledge in a more orderly, structured, and semantically oriented manner, providing a foundation for computer-aided diagnosis and treatment.We developed a construction framework of TCM syndrome knowledge from ancient books, using a pretrained model and rules (TCMSF). We conducted fine-tuning training on Enhanced Representation through Knowledge Integration (ERNIE), Bidirectional Encoder Representation from Transformers pretrained language models, and chatGLM3-6b large language models for named entity recognition (NER) tasks. Furthermore, we employed the progressive entity relationship extraction method based on the dual pattern feature combination to extract and standardize entities and relationships between entities in these books.We selected Yin deficiency syndrome as a case study and constructed a model layer suitable for the expression of knowledge in these books. Compared with multiple NER methods, the combination of ERNIE and Conditional Random Fields performs the best. By utilizing this combination, we completed the entity extraction of Yin deficiency syndrome, achieving an average F1 value of 0.77. The relationship extraction method we proposed reduces the number of incorrectly connected relationships compared with fully connected pattern layers. We successfully constructed a knowledge graph of ancient books on Yin deficiency syndrome, including over 120,000 entities and over 1.18 million relationships.We developed TCMSF in line with the knowledge characteristics of ancient TCM books and improved the accuracy of knowledge graph construction.