Germline BRCA2 loss-of function variants, which can be identified through clinical genetic testing, predispose to several cancers1–5. However, variants of uncertain significance limit the clinical utility of test results. Thus, there is a need for functional characterization and clinical classification of all BRCA2 variants to facilitate the clinical management of individuals with these variants. Here we analysed all possible single-nucleotide variants from exons 15 to 26 that encode the BRCA2 DNA-binding domain hotspot for pathogenic missense variants. To enable this, we used saturation genome editing CRISPR–Cas9-based knock-in endogenous targeting of human haploid HAP1 cells6. The assay was calibrated relative to nonsense and silent variants and was validated using pathogenic and benign standards from ClinVar and results from a homology-directed repair functional assay7. Variants (6,959 out of 6,960 evaluated) were assigned to seven categories of pathogenicity based on a VarCall Bayesian model8. Single-nucleotide variants that encode loss-of-function missense variants were associated with increased risks of breast cancer and ovarian cancer. The functional assay results were integrated into models from ClinGen, the American College of Medical Genetics and Genomics, and the Association for Molecular Pathology9 for clinical classification of BRCA2 variants. Using this approach, 91% were classified as pathogenic or likely pathogenic or as benign or likely benign. These classified variants can be used to improve clinical management of individuals with a BRCA2 variant. Results from a comprehensive evaluation of the function of BRCA2 variants, particularly variants of uncertain significance, provide a useful resource to improve the clinical management of individuals who carry such genetic variants.
OBJECTIVE:Despite exponential growth in artificial intelligence (AI) research for laboratory medicine and pathology, a significant gap exists between model development and clinical AI implementation. This article introduces a structured framework, the Clinical AI Readiness Evaluator (CARE), to bridge this gap and support the responsible adoption of AI in clinical laboratory settings. METHODS:Building upon the Machine Learning Technology Readiness Levels framework, we developed CARE specifically for the clinical laboratory environment by incorporating health care-specific requirements, regulatory considerations, and workflow integration needs. This framework was iteratively refined through practical application across diverse AI use cases within laboratory medicine and pathology. RESULTS:The CARE framework provides a systematic approach to AI development and implementation through 8 component workstreams: clinical use case, data, data pipeline, code, clinical user experience, clinical technology infrastructure, clinical orchestration, and regulatory compliance. Unlike generic AI frameworks, CARE distinguishes itself by emphasizing both health care and laboratory workflow integration, regulatory requirements, ethical considerations, and comprehensive validation for clinical contexts. The framework accommodates both internally developed models and commercial AI solutions, providing clear guidance through technology readiness levels and structured review processes. CONCLUSIONS:The CARE framework addresses the unique challenges of implementing AI in laboratory medicine and pathology by providing a comprehensive roadmap from initial concepts through clinical deployment and maintenance. This article, the first in a series of 4, establishes the foundational AI lifecycle framework, while subsequent articles will explore data documentation, ethical AI considerations, and governance structures. By adopting this structured approach, laboratories can responsibly harness AI's potential to enhance diagnostic accuracy and operational efficiencies and, ultimately, improve patient care.
Context.— Computational pathology combines clinical pathology with computational analysis, aiming to enhance diagnostic capabilities and improve clinical productivity. However, communication barriers between pathologists and developers often hinder the full realization of this potential. Objective.— To propose a standardized framework that improves mutual understanding of clinical objectives and computational methodologies. The goal is to enhance the development and application of computer-aided diagnostic (CAD) tools. Design.— This article suggests pivotal roles for pathologists and computer scientists in the CAD development process. It calls for increased understanding of computational terminologies, processes, and limitations among pathologists. Similarly, it argues that computer scientists should better comprehend the true use cases of the developed algorithms to avoid clinically meaningless metrics. Results.— CAD tools improve pathology practice significantly. Some tools have even received US Food and Drug Administration approval. However, improved understanding of machine learning models among pathologists is essential to prevent misuse and misinterpretation. There is also a need for a more accurate representation of the algorithms’ performance compared to that of pathologists. Conclusions.— A comprehensive understanding of computational and clinical paradigms is crucial for overcoming the translational gap in computational pathology. This mutual comprehension will improve patient care through more accurate and efficient disease diagnosis.
Context.— Artificial intelligence is a transforming technology for anatomic pathology. Involvement within the workforce will foster support for algorithm development and implementation. Objective.— To develop a supportive ecosystem that enables pathologists with variable expertise in artificial intelligence to create algorithms in a development environment with seamless transition to a production environment. Design.— Platform requirements included (1) approachable and intuitive user interface, (2) diverse algorithmic modeling options, (3) support capability for internal and external collaborations, (4) a seamless mechanism for discovery to clinical deployment transition, (5) the ability to meet minimum institutional requirements for information technology (IT) review, and (6) the ability to be scalable over time. The ecosystem required platform education, data science guidance, project management structure, and ongoing leadership. Results.— The development team considered internal development and vended solutions. Because of the extended timeline and resource requirements for internal development, a decision was made to use a vended solution. Vendor proposals were solicited and reviewed by pathologists, IT, and security groups. A vendor was selected and pipelines for development and production were established. Proposals for development were solicited from the pathology department. Eighty-four investigators were selected for the initial cohort, receiving training and access to dedicated subject matter experts. A total of 30 of 31 projects progressed through the model development process of annotating, training, and validation. Based on these projects, 15 abstracts were submitted to national meetings. Conclusions.— Democratizing artificial intelligence by creating an ecosystem to support pathologists with varying levels of expertise can break down entry barriers, reduce overall cost of algorithm development, improve algorithm quality, and enhance the speed of adoption.
BackgroundDeveloping medical software requires navigating complex regulatory, ethical, and operational challenges. A comprehensive framework that supports both technical maturity and clinical safety is essential for effective artificial intelligence and machine learning system deployment. This paper introduces the Clinical Artificial Intelligence Readiness Evaluator Lifecycle and the Clinical Artificial Intelligence Readiness Evaluator Agent-a framework and AI-driven tool designed to streamline technology readiness level assessments in medical software development.MethodsWe developed the framework using an iterative process grounded in collaborative stakeholder analysis. Key institutional stakeholders-including clinical informatics experts, data engineers, ethicists, and operational leaders-were engaged to identify and prioritize the regulatory, ethical, and technical requirements unique to clinical AI/ML development. This approach, combined with a thorough review of existing methodologies, informed the creation of a lifecycle model that guides technology maturation from initial concept to full deployment. The AI-driven tool was implemented using a retrieval-augmented generation strategy and evaluated through a synthetic use case (the Diabetes Outcome Predictor). Evaluation metrics included the proportion of correctly addressed assessment questions and the overall time required for automated review, with human adjudication validating the tool's performance.ResultsThe findings indicate that the proposed framework effectively captures the complexities of clinical AI development. In the synthetic use case, the AI-driven tool identified that 32.8% of the assessment questions remained unanswered, while human adjudication confirmed discrepancies in 19.4% of these instances. These outcomes suggest that, when fully refined, the automated assessment process can reduce the need for extensive multi-stakeholder involvement, accelerate project timelines, and enhance resource efficiency.ConclusionsThe Clinical Artificial Intelligence Readiness Evaluator Lifecycle and Agent offer a robust and methodologically sound approach for evaluating the maturity of medical AI systems. By integrating stakeholder-driven insights with an AI-based assessment process, this framework lays the groundwork for more streamlined, secure, and effective clinical AI development. Future work will focus on optimizing retrieval strategies and expanding validation across diverse clinical applications.
Clinical genetic testing identifies variants causal for hereditary cancer, information that is used for risk assessment and clinical management. Unfortunately, some variants identified are of uncertain clinical significance (VUS), complicating patient management. Case-control data is one evidence type used to classify VUS. As an initiative of the Evidence-based Network for the Interpretation of Germline Mutant Alleles (ENIGMA) Analytical Working Group we analyze germline sequencing data of BRCA1 and BRCA2 from 96,691 female breast cancer cases and 302,116 controls from three studies: the BRIDGES study of the Breast Cancer Association Consortium, the Cancer Risk Estimates Related to Susceptibility consortium, and the UK Biobank. We observe 11,207 BRCA1 and BRCA2 variants, with 6909 being coding, covering 23.4% of BRCA1 and BRCA2 VUS in ClinVar and 19.2% of ClinVar curated (likely) benign or pathogenic variants. Case-control likelihood ratio (ccLR) evidence is highly consistent with ClinVar assertions for (likely) benign or pathogenic variants; exhibiting 99.1% sensitivity and 95.3% specificity for BRCA1 and 93.3% sensitivity and 86.6% specificity for BRCA2. This approach provides case-control evidence for 787 unclassified variants; these include 579 with strong or moderate benign evidence and 10 with strong pathogenic evidence for which ccLR evidence is sufficient to alter clinical classification.
BACKGROUND:Pathogenic variants (PVs) in ATM, BRCA1, BRCA2, CHEK2, and PALB2 are associated with increased breast cancer risk. It is unknown, however, whether this risk differs by PV type or location in carriers ascertained from the general population. PATIENTS AND METHODS:To evaluate breast cancer risks associated with PV type and location in ATM, BRCA1, BRCA2, CHEK2, and PALB2, we carried out age-adjusted case-control association analysis in 32 247 women with and 32 544 age-matched women without breast cancer from the CARRIERS Consortium. PVs were grouped by type and location within genes and assessed for risks of breast cancer [odds ratios (OR), 95% confidence intervals (CI), and P values] using logistic regression. RESULTS:Compared with women carrying BRCA2 exon 11 protein truncating variants (PTVs) in the CARRIERS population-based study, women with BRCA2 ex1-10 PTVs (OR = 13.5, 95% CI 6.0-38.7, P < 0.001) and ex13-27 PTVs (OR = 9.0, 95% CI 4.9-18.5, P < 0.001) had higher breast cancer risks, lower rates of estrogen receptor (ER)-negative breast cancer (ex13-27 OR = 0.5, 95% CI 0.2-0.9, P = 0.035; ex1-10 OR = 0.5, 95% CI 0.1-1.0, P = 0.065), and earlier age at breast cancer diagnosis (ex13-27 5.5 years, P < 0.001; ex1-10 2.4 years, P = 0.169). These associations with ER-negative breast cancer and age were replicated in a high-risk clinical cohort from Ambry Genetics and the population-based UK Biobank cohort. No differences in risk by gene region were observed for PTVs in other predisposition genes. CONCLUSIONS:Population-based and clinical high-risk cohorts establish that PTVs in exon 11 of BRCA2 are associated with reduced breast cancer risk, later age at diagnosis, and greater risk of ER-negative disease. These differential risks may improve individualized risk prediction and clinical management for women carrying BRCA2 PTVs.
Context.— Generative artificial intelligence (AI) technologies are rapidly transforming numerous fields, including pathology, and hold significant potential to revolutionize educational approaches. Objective.— To explore the application of generative AI, particularly large language models and multimodal tools, for enhancing pathology education. We describe their potential to create personalized learning experiences, streamline content development, expand access to educational resources, and support both learners and educators throughout the training and practice continuum. Data Sources.— We draw on insights from existing literature on AI in education and the collective expertise of the coauthors within this rapidly evolving field. Case studies highlight practical applications of large language models, demonstrating both the potential benefits and unique challenges associated with implementing these technologies in pathology education. Conclusions.— Generative AI presents a powerful tool kit for enriching pathology education, offering opportunities for greater engagement, accessibility, and personalization. Careful consideration of ethical implications, potential risks, and appropriate mitigation strategies is essential for the responsible and effective integration of these technologies. Future success lies in fostering collaborative development between AI experts and medical educators, prioritizing ongoing human oversight and transparency to ensure that generative AI augments, rather than supplants, the vital role of educators in pathology training and practice.
PURPOSE:Germline pathogenic variants (PV) in ATM increase the risk of pancreatic ductal adenocarcinoma (PDAC), but the underlying tumor biology of PDAC associated with germline PV in ATM has not been adequately explored. EXPERIMENTAL DESIGN:Whole-genome, whole-exome, and RNA sequencing were performed on PDAC tumors from 25 germline ATM PV carriers diagnosed at Mayo Clinic between 2007 and 2017. Somatic and copy-number alterations, mutational signatures, transcriptomic subtypes, and the immune landscape were evaluated. RESULTS:High-quality whole-exome and whole-genome sequencing were obtained from 21 and 15 tumors, respectively. Biallelic inactivation of ATM was observed in 87%, KRAS PV in 90%, CDKN2A homozygous loss in 60%, and TP53 alterations in <10% of these tumors. A predominant clock-like mutational signature was present in all samples. Whole-transcriptome analysis identified that the aberrantly differentiated endocrine exocrine subtype accounted for 18% of PDAC and was consistently associated with >5-year overall survival. In addition, a 28-gene expression-based signature associated with overall survival was identified and further validated in The Cancer Genome Atlas cohort. Immune landscape analysis through CODEX identified enriched CD4 T-helper cell/tumor interactions and reduced B7H3-high cell/tumor interactions in ATM PV carriers compared with noncarriers. CONCLUSIONS:The observed absence of TP53 PV and enrichment for CDKN2A alterations in ATM tumors, along with differences in the mutational signatures, transcriptomic subtypes and immune landscape, improve our understanding of the mechanistic pathways involved in PDAC development in germline ATM PV carriers and help identify potential targeted therapeutic strategies.
Importance: Mammographic density (MD) and pathogenic variants (PVs) in breast cancer susceptibility genes are major determinants of breast cancer risk, but their association and joint effects on breast cancer risk are unclear. Objective: To investigate the association between the presence or absence of PVs in breast cancer susceptibility genes and MD measures, and their joint effects on breast cancer risk in an observational study; and to evaluate causality using Mendelian randomisation (MR) analyses. Design: Case-control analyses using data from the Breast Cancer Association Consortium (1991-2016). Sequencing and genotyping took place between 2009 and 2021. Setting Multicenter Participants A total of 6,809 cases and 18,189 controls were included, from 15 studies, comprising women aged 19 to 92 years with mammograms taken at least one year before diagnosis. Exposure MD measures, including dense area (DA), non-dense area (NDA), percentage density (PD) and absolute difference in PD between left and right breasts (ADPD), and PVs in ATM, BARD1, BRCA1, BRCA2, CHEK2, PALB2, RAD51C and RAD51D. Main outcomes and measures Breast cancer risk overall, by oestrogen receptor expression-defined subtypes, and among BRCA1 and BRCA2 PV carriers. Results No association was found between the overall burden of PVs and any MD measure. There was some evidence for a negative interaction between the burden of PVs in the eight genes and PD (OR=0.79,95%CIint=0.62,1.00, PLRT=0.047). This appears to be largely driven by a positive interaction with NDA. MR analyses indicated attenuated effects for BRCA1 (for PD, OR per standard deviation =1.02(95%CI:0.78,1.34) but not BRCA2 PV carriers (1.54,95%CI=1.08,2.24)). Conclusions and Relevance There was no evidence of association between PVs in breast cancer susceptibility genes and MD measures, but some suggestion that the association between MD and breast cancer risk may be weaker in PV carriers. Replication of these findings in further large datasets is required. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by Cancer Research UK grant: PPRPGM-Nov20\100002 and by core funding from the NIHR Cambridge Biomedical Research Centre (NIHR203312) [*]. *The views expressed are those of the author(s) and not necessarily those of the NIHR or the Department of Health and Social Care. Additional funding for BCAC is provided by the Confluence project which is funded with intramural funds from the National Cancer Institute Intramural Research Program, National Institutes of Health, the European Union's Horizon 2020 Research and Innovation Programme (grant numbers 634935 and 633784 for BRIDGES and B-CAST respectively), and the PERSPECTIVE I&I project, funded by the Government of Canada through Genome Canada and the Canadian Institutes of Health Research, the Ministere de l'Economie et de l'Innovation du Quebec through Genome Quebec, the Quebec Breast Cancer Foundation. The EU Horizon 2020 Research and Innovation Programme funding source had no role in study design, data collection, data analysis, data interpretation or writing of the report. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The organisations providing ethical approval for the contributing studies are summarised by Dorling et al, NEJM, https://www.nejm.org/doi/full/10.1056/NEJMoa1913948 I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Requests for individual-level data used in these analyses should be made via the BCAC data access coordinating committee (bcac@medschl.cam.ac.uk). Summary-level data from BCAC and CIMBA used in the Mendelian Randomisation analysis is publicly available at https://www.ccge.medschl.cam.ac.uk/breast-cancer-association-consortium-bcac), https://www.ccge.medschl.cam.ac.uk/consortium-investigators-modifiers-brca12-cimba and the GWAS Catalog (https://www.ebi.ac.uk/gwas/)
Recent advances in the field of immuno-oncology have brought transformative changes in the management of cancer patients. The immune profile of tumours has been found to have key value in predicting disease prognosis and treatment response in various cancers. Multiplex immunohistochemistry and immunofluorescence have emerged as potent tools for the simultaneous detection of multiple protein biomarkers in a single tissue section, thereby expanding opportunities for molecular and immune profiling while preserving tissue samples. By establishing the phenotype of individual tumour cells when distributed within a mixed cell population, the identification of clinically relevant biomarkers with high-throughput multiplex immunophenotyping of tumour samples has great potential to guide appropriate treatment choices. Moreover, the emergence of novel multi-marker imaging approaches can now provide unprecedented insights into the tumour microenvironment, including the potential interplay between various cell types. However, there are significant challenges to widespread integration of these technologies in daily research and clinical practice. This review addresses the challenges and potential solutions within a structured framework of action from a regulatory and clinical trial perspective. New developments within the field of immunophenotyping using multiplexed tissue imaging platforms and associated digital pathology are also described, with a specific focus on translational implications across different subtypes of cancer. © 2024 The Authors. The Journal of Pathology published by John Wiley & Sons Ltd on behalf of The Pathological Society of Great Britain and Ireland.
Background The loss-of-function (LOF) classification of most missense variants in tumor suppressor breast cancer genes BRCA1, BRCA2, PALB2, and RAD51C remains unclassified and confounds clinical actionability. Classifying these variants is challenging due to their rarity, leading clinicians to rely on in silico predictive methods. Protein stability changes are associated with function, making stability predictors valuable. Stability predictions upon missense variant perturbations require high-resolution protein structures. However, the availability of these high-resolution structures is lacking. This study explores using generative AI to predict high-resolution protein structures, which can then be analyzed with in silico protein stability prediction methods to assess LOF activity in ordered regions of the protein. This study also determines the appropriate in silico protein stability and dedicated in silico missense prediction methods in dbNSFP v4.7 database to predict LOF activity in ordered regions of these four genes. Functional classifications from homology recombination DNA repair (HDR) assays and variant classifications from the ClinVar database provide a reliable dataset for evaluating the performance of these in silico prediction methods. Results Complex AlphaFold2 structures of the BRCA1-C terminal (BRCT) domain and the DNA-binding (DB) domain of BRCA2, analyzed using protein stability tool FoldX predicts LOF activity from missense variants significantly better than experimentally-derived structures in ordered regions. The BRCT domain achieved an Area Under the Curve (AUC)= 0.861 (95 % CI:0.858–0.863) and AUC= 0.842 (95 % CI:0.840–0.845), while the DB domain achieved an AUC= 0.836 (95 % CI:0.8322–0.841), compared to AUC= 0.847 (95 % CI:0.844–0.850) and AUC= 0.835 (95 % CI:0.832–0.837) from the BRCT domain, and AUC= 0.830 (95 % CI:0.821–0.8320) from the DB domain from experimentally-derived structures. Protein stability does not predict LOF activity from missense variants better than dedicated in silico missense predictors. Overall, we find that AlphaMissense ranks highly, with an average AUC= 0.890 (95 % CI 0.886–0.895) from ordered regions across these four cancer genes, compared to all other in silico missense predictors present in the dbNSFP database. Conclusions The study reveals that generative AI protein predicted structures can outperform experimentally-derived structures in evaluating LOF activity from predicted protein stability in ordered regions of genes BRCA1, BRCA2, PALB2 and RAD51C. The study also highlights the predictive performance of AlphaMissense as the premier in silico missense prediction method to predict LOF activity from missense variants in these four tumor suppressor breast cancer genes. The code for this study can be downloaded for free on GitHub (https://github.com/rohandavidg/CarePred)
Unsupervised embeddings are fundamental to numerous machine learning applications, yet their evaluation remains a challenging task. Traditional assessment methods often rely on extrinsic variables, such as performance in downstream tasks, which can introduce confounding factors and mask the true quality of embeddings. This paper introduces the Intrinsic Distance Preservation Evaluation (IDPE) method, a novel approach for assessing embedding quality based on the preservation of Mahalanobis distances between data points in the original and embedded spaces. We demonstrate the limitations of extrinsic evaluation methods through a simple example, highlighting how they can lead to misleading conclusions about embedding quality. IDPE addresses these issues by providing a task-independent measure of how well embeddings preserve the intrinsic structure of the original data. Our method leverages efficient similarity search techniques to make it applicable to large-scale datasets. We compare IDPE with established intrinsic metrics like trustworthiness and continuity, as well as extrinsic metrics such as Average Rank and Mean Reciprocal Rank. Our results show that IDPE offers a more comprehensive and reliable assessment of embedding quality across various scenarios. We evaluate PCA and t-SNE embeddings using IDPE, revealing insights into their performance that are not captured by traditional metrics. This work contributes to the field by providing a robust, efficient, and interpretable method for embedding evaluation. IDPE's focus on intrinsic properties offers a valuable tool for researchers and practitioners seeking to develop and assess high-quality embeddings for diverse machine learning applications.
The functional classification of a missense variant in cancer predisposition genes is often challenging due to how rare the variant is observed in the population. When available, clinicians utilize a combination of family history, in vitro functional assays and in silico methods to infer protein function. In silico methods, such as missense predictors (predict changes in protein function) and protein stability predictors (predict changes in free energy) have been used to help classify a missense variant in accordance with the American College of Medical Genetics and Genomics (ACMG) guideline. To measure protein stability, many in silico algorithms predict stability based on the change of free energy and most accurate protein stability predictors require a wild-type protein template. In this study, we examine the use of generative AI to predict high-resolution protein structures as templates analyzed with protein stability methods to evaluate loss of function (LOF) activity in cancer predisposition genes BRCA1, BRCA2, PALB2 and RAD51C upon the presence of missense variant. Utilizing multiplexed assay of variant effect measurements and variant classifications from ClinVar, we find that prediction of Gibbs free energy (ΔΔG) from AlphaFold2 (AF2) structures analyzed with FoldX predicts LOF better than experimental-derived wild type structures in the BRCT domain of BRCA1 and the DNA binding domain (DBD) of BRCA2 , but not in PALB2 and RAD51C . We also find that AF2 structures in the BRCT domain of BRCA1 and DBD-Dss1 domain of BRCA2 analyzed with FoldX measure homologous DNA recombination (HDR) activity significantly better than Rosetta and DDGun3D. Our study also revealed that there are other factors that contribute to predicting loss of function activity other than protein stability, with AlphaMissense ranking the best overall predictor of LOF activity in these tumor suppressor breast cancer genes. Author Summary The stability of a protein, often expressed in terms of Gibbs free energy (ΔΔG), is a critical factor in predicting loss of function (LOF) activity when a missense variant is present. The effect is higher in haploinsufficient genes like the tumor suppressor genes BRCA1, BRCA2, PALB2 and RAD51C . Protein stability predictors that utilizes a wild-type structure to make its predictions is often limited by the availability of experimentally-derived protein structures. Here, in our study we show that generative AI, like AlphaFold2 (AF2) can predict structures similar to experimentally-derived structures with high similarity. Furthermore, protein stability tools such as FoldX, Rosetta, and DDGun3D can be used in conjunction to measure changes in stability. From our study, we find that complex AF2 structures representing the BRCT domain of BRCA1 and DBD domain of BRCA2 analyzed by FoldX predicts function significantly better than the experimentally-derived structures. However, predicted |ΔΔG| does not predict function better than purpose-built in silico missense predictors for protein function. Overall, we find the AlphaMissense is the best predictor to predict function in these tumor suppressor breast cancer genes. ### Competing Interest Statement The authors have declared no competing interest.
With the increasing utilization of exome and genome sequencing in clinical and research genetics, accurate and automated extraction of human phenotype ontology (HPO) terms from clinical texts has become imperative. Traditional methods for HPO term extraction, such as PhenoTagger, often face limitations in coverage and precision. In this study, we propose a novel approach that leverages large language models (LLMs) to generate synthetic sentences with clinical context, which were semantically encoded into vector embeddings. These embeddings are linked to HPO terms, creating a robust knowledgebase that facilitates precise information retrieval. Our method circumvents the known issue of LLM hallucinations by storing and querying these embeddings within a true database, ensuring accurate context matching without the need for a predictive model. We evaluated the performance of three different embedding models, all of which demonstrated substantial improvements over PhenoTagger. Top recall (sensitivity), precision (positive-predictive value, PPV), and F1 are 0.64, 0.64, and 0.64, respectively, which were 31%, 10%, and 21% better than PhenoTagger. Furthermore, optimal performance was achieved when we combined the best performing embedding model with PhenoTagger (a.k.a. Fused model), resulting in recall (sensitivity), precision (PPV), and F1 values of 0.7, 0.7, and 0.7, respectively, which are 10%, 10%, and 10% better than the best embedding models. Our findings underscore the potential of this integrated approach to enhance the precision and reliability of HPO term extraction, offering a scalable and effective solution for biomedical data annotation.
Germline mutations in BRCA1 and BRCA2 (gBRCA1/2) are required for a PARP inhibitor therapy in patients with HER2-negative (HER2−) advanced breast cancer (aBC). However, little is known about the prognostic impact of gBRCA1/2 mutations in aBC patients treated with chemotherapy. This study aimed to investigate the frequencies and prognosis of germline and somatic BRCA1/2 mutations in HER2- aBC patients receiving the first chemotherapy in the advanced setting. Patients receiving their first chemotherapy for HER2- aBC were retrospectively selected from the prospective PRAEGNANT registry (NCT02338167). Genotyping of 26 cancer predisposition genes was performed with germline DNA of 471 patients and somatic tumor DNA of 94 patients. Mutation frequencies, progression-free and overall survival (PFS, OS) according to germline mutation status were assessed. gBRCA1/2 mutations were present in 23 patients (4.9%), and 33 patients (7.0%) had mutations in other cancer risk genes. Patients with a gBRCA1/2 mutation had a better OS compared to non-mutation carriers (HR: 0.38; 95%CI: 0.17–0.86). PFS comparison was not statistically significant. Mutations in other risk genes did not affect prognosis. Two somatic BRCA2 mutations were found in 94 patients without gBRCA1/2 mutations. Most frequently somatic mutated genes were TP53 (44.7%), CDH1 (10.6%) and PTEN (6.4%). In conclusion, aBC patients with gBRCA1/2 mutations had a more favorable prognosis under chemotherapy compared to non-mutation carriers. The mutation frequency of ~5% with gBRCA1/2 mutations together with improved outcome indicates that germline genotyping of all metastatic patients for whom a PARP inhibitor therapy is indicated should be considered.
AbstractContextBreast cancer is one of the most common cancers in women. With early diagnosis, some breast cancers are highly curable. However, the concordance rate of breast cancer diagnosis from histology slides by pathologists is unacceptably low. Classifying normal versus tumor breast tissues from microscopy images of breast histology is an ideal case to use for deep learning and could help to more reproducibly diagnose breast cancer. Since data preprocessing and hyperparameter configurations have impacts on breast cancer classification accuracies of deep learning models, training a deep learning classifier with appropriate data preprocessing approaches and optimized hyperparameter configurations could improve breast cancer classification accuracy.Methods and MaterialUsing 12 combinations of deep learning model architectures (i.e., including 5 non-specialized and 7 digital pathology-specialized model architectures), image data preprocessing, and hyperparameter configurations, the validation accuracy of tumor versus normal classification were calculated using theBreAstCancerHistology (BACH) dataset.ResultsThe DenseNet201, a non-specialized model architecture, with transfer learning approach achieved 98.61% validation accuracy compared to only 64.00% for the digital pathology-specialized model architecture.ConclusionsThe combination of image data preprocessing approaches and hyperparameter configurations have a profound impact on the performance of deep neural networks for image classification. To identify a well-performing deep neural network to classify tumor versus normal breast histology, researchers should not only focus on developing new models specifically for digital pathology, since hyperparameter tuning for existing deep neural networks in the computer vision field could also achieve a high (often better) prediction accuracy.
Aggregate effect of P/LP/D carrier status across BRCA2, ATM, NBN, and PALB2 genes on PCa risk in African ancestry men.