AI-driven chatbots have been utilized in healthcare to automate administrative tasks, improve patient education, and expand access to medical information; however, their role in genetic counseling remains underexplored. To investigate the adoption, perceptions, and potential utility of AI-based chatbots in genetic counseling practice, 217 genetic counselors and genetic counseling students from across North America were surveyed regarding chatbot usage, confidence in their application, and perceived benefits and limitations. While most participants (166/217; 76.5%) reported using general AI chatbots outside of clinical settings, far fewer (18/204; 8.8%) reported using or recommending clinical genetics chatbots in clinical practice. For those that used clinical genetics chatbots, the primary purpose was for communication with at-risk family members (11/18; 61.1%) and patient education (10/18; 55.6%). Confidence in chatbot technology varied, with highest confidence in gathering family history information (81/199; 40.7%) and lowest confidence in their ability to disclose variants of uncertain significance or positive genetic testing results (5/199; 2.5%). The greatest perceived benefits included reducing repetitive tasks (165/195, 84.6%) and allowing for time for other tasks (141/195; 72.3%), while major concerns revolved around patient comprehension (167/195; 85.6%) and having accurate, up-to-date information (145/195; 74.4%). Despite some concern about AI replacing human counselors, most participants reported they felt there was potential for chatbots to enhance workflow efficiency (128/195; 65.6%) if properly integrated and regulated. Limited AI training was identified as a barrier to adoption (16/195; 8.2% received training), highlighting a need for structured education on AI applications in genetic counseling. These findings suggest that AI chatbots hold promise as supplementary tools, but significant challenges must be addressed before widespread implementation in genetic counseling practice.
Artificial intelligence (AI) is rapidly transforming biomedicine, including genomic medicine. This comment explores three myths and controversies involving AI in genomic medicine, including: how clinicians will use AI; whether AI’s impressive on-paper performance will translate to the real world; how patient preferences may or may not impact AI’s transformation of the field. Based on this discussion, the comment concludes with three recommendations for an AI-based future of genomic medicine.
PURPOSE:Because clinical genetics artificial intelligence (AI) applications and capabilities are on a rapid rise, this study sought to assess workforce readiness. METHODS:To assess the workforce's use, knowledge, and attitudes about medical AI applications, we conducted a survey of 215 US-based genetics clinicians and trainees. RESULTS:More than half (51.2%) of participants reported little to no knowledge of AI in clinical genetics; 64.3% reported no formal training. Formal training correlated with greater self-reported knowledge of AI in clinical genetics: 69.3% of respondents with formal training (vs 37.5% without) reported intermediate to extensive knowledge of AI. Most participants reported insufficient knowledge of clinical AI (83.4%), desired more education (97.6%), and would take available training (89.3%). The majority (51.6%) of clinician participants initially reported not using AI applications in the clinic. However, after a tutorial describing clinical AI applications, 75.8% reported some use. When asked about specific applications, the majority of clinician participants used facial diagnostic applications (54.9%) and AI-generated genomic testing results (62.1%); other applications, such as chatbots, large language models, pedigree or medical summary generators, and risk assessment, were less commonly used (11.1%-12.5%). CONCLUSION:Further education is both desired and needed to optimally use AI applications in clinical genetics.
Introduction Artificial intelligence (AI) is increasingly prevalent. Patients and clinicians may use AI-based tools in many different languages.Objective To investigate AI translation tools for descriptions of genetic conditions and how AI identification of genetic conditions is affected by translations.Materials and Methods We used Neural machine translation (NMT) and large language-model (LLM) translation to translate descriptions of 40 genetic conditions into 191 and 93 languages, respectively. Excluding translations retaining English medical terms verbatim, we respectively focused on 139 and 70 languages. After assessing translations, we assessed the ability of 3 proprietary and 3 open-weight general LLMs to identify conditions in the translations. We analyzed how accuracy was affected by the conditions' prevalence in the literature, and attributes of the languages (the script, language family, and prevalence of the language in training sources). We also investigated adaptive translation for select languages.Results We found significant differences in condition identification based on the translation method, condition, language, and prediction model. The accuracy of some models was more affected than others by factors like the conditions' literature prevalence, language script, family, and language prevalence. Adaptive translation for select languages did not improve translations or diagnostic accuracy with the 3 tested LLMs. However, further analysis with 1 language showed that this approach was more effective with smaller LLMs.Conclusions AI-based translation has variable performance, which can affect the ability of AI models to recognize genetic conditions. These findings should inform safe medical AI use to support consistent performance in different languages.
FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinical diagnosis. Using annotations from 124 clinicians across 10 disorders (107 HPO terms) combined with non-syndromic controls, we generated 3D facial meshes (478 landmarks) from 2D images and trained a hierarchical PointNet-based pipeline with cascading classification and feature elimination. The best models, incorporating 3D meshes, facial outline, and demographic metadata, achieved AUROCs between 0.55 and 0.89, with higher performance at parent nodes than leaf terms. External validation showed variable generalizability across disorders. Results demonstrate that hierarchical modeling of 3D facial geometry enables interpretable, ontology-linked phenotype classification, though performance on rare leaf terms remains limited. Improved data diversity and feature selection strategies are needed to enhance robustness and clinical utility.
Purpose of review Artificial intelligence (AI) is being rapidly but unevenly integrated into many aspects of medicine as well as society more broadly. This review focuses on recent developments in AI through the lens of how it is affecting the field of pediatric genomics. Recent findings This article describes and explains general definitions of AI and its subtypes, including related to the latest AI applications. The article then outlines ways in which AI is being used and tested in multiple aspects of pediatric genomics, including: identifying individuals who may have genetic conditions and assisting in the diagnostic process; AI implementation in the laboratory-based genetic testing process; how recent AI developments may support patient management as well as discovery and research into novel ways to treat genetic conditions. Overall, AI is being applied to pediatric genomics in many different ways. Summary The field of pediatrics will undergo dramatic change because of AI. This piece traces different ways in which AI may affect one particular field, pediatric genomics. Due to rapid advances, it is likely that future AI discoveries will lead to additional ways in which AI affects this field.
Deep learning (DL) is increasingly used to analyze medical imaging, but is less refined for rare conditions, which require novel pre-processing and analytical approaches. To assess DL in the context of rare diseases, this study focused on alkaptonuria (AKU), a rare disorder that affects the spine and involves other sequelae; treatments include the medication nitisinone. Since assessing x-rays to determine disease severity can be a slow, manual process requiring considerable expertise, this study aimed to determine whether these DL methods could accurately identify overall spine severity at specific regions of the spine and whether patients were receiving nitisinone. DL performance was evaluated versus clinical experts using cervical and lumbar spine radiographs. DL models predicted global severity scores (30-point scale) within 1.72 ± 1.96 points of expert clinician scores for cervical and 2.51 ± 1.96 points for lumbar radiographs. For region-specific metrics, the degrees of narrowing, calcium, and vacuum disc phenomena at each intervertebral space (IVS) were assessed. The model's narrowing scores were within 0.191-0.557 points from clinician scores (6-point scale), calcium was predicted with 78%-90% accuracy (present, absent, or disc fusion), and vacuum disc phenomenon predictions were less consistent (41%-90%). Intriguingly, DL models predicted nitisinone treatment status with 68%-77% accuracy, while expert clinicians appeared unable to discern nitisinone status (51% accuracy) (p = 2.0 × 10-9). This highlights the potential for DL to augment certain types of clinical assessments in rare disease, as well as identifying occult features like treatment status.
Artificial intelligence (AI) is rapidly transforming numerous aspects of daily life, including clinical practice and biomedical research. In light of this rapid transformation, and in the context of medical genetics, we assembled a group of leaders in the field to respond to the question about how AI is affecting, and especially how AI will affect, medical genetics. The authors who contributed to this collection of essays intentionally represent different areas of expertise, career stages, and geographies, and include diverse types of clinicians, computer scientists, and researchers. The individual pieces cover a wide range of areas related to medical genetics; we expect that these pieces may provide helpful windows into the ways in which AI is being actively studied, used, and considered in medical genetics.
The facial gestalt (overall facial morphology) is a characteristic clinical feature in many genetic disorders that is often essential for suspecting and establishing a specific diagnosis. Therefore, publishing images of individuals affected by pathogenic variants in disease-associated genes has been an important part of scientific communication. Furthermore, medical imaging data is also crucial for teaching and training deep-learning models such as GestaltMatcher. However, medical data is often sparsely available, and sharing patient images involves risks related to privacy and re-identification. Therefore, we explored whether generative neural networks can be used to synthesize accurate portraits for rare disorders. We modified a StyleGAN architecture and trained it to produce artificial condition-specific portraits for multiple disorders. In addition, we present a technique that generates a sharp and detailed average patient portrait for a given disorder. We trained our GestaltGAN on the 20 most frequent disorders from the GestaltMatcher database. We used REAL-ESRGAN to increase the resolution of portraits from the training data with low-quality and colorized black-and-white images. To augment the model’s understanding of human facial features, an unaffected class was introduced to the training data. We tested the validity of our generated portraits with 63 human experts. Our findings demonstrate the model’s proficiency in generating photorealistic portraits that capture the characteristic features of a disorder while preserving patient privacy. Overall, the output from our approach holds promise for various applications, including visualizations for publications and educational materials and augmenting training data for deep learning.
Artificial intelligence (AI) tools are increasingly employed in clinical genetics to assist in diagnosing genetic conditions by assessing photographs of patients. For medical uses of AI, explainable AI (XAI) methods offer a promising approach by providing interpretable outputs, such as saliency maps and region relevance visualizations. XAI has been discussed as important for regulatory purposes and to enable clinicians to better understand how AI tools work in practice. However, the real-world effects of XAI on clinician performance, confidence, and trust remain underexplored. This study involved a web-based user experiment with 31 medical geneticists to assess the impact of AI-only diagnostic assistance compared to XAI-supported diagnostics. Participants were randomly assigned to either group and completed diagnostic tasks with 18 facial images of individuals with known genetic syndromes and unaffected individuals, before and after experiencing the AI outputs. The results show that both AI-only and XAI approaches improved diagnostic accuracy and clinician confidence. The effects varied according to the accuracy of AI predictions and the clarity of syndromic features (sample difficulty). While AI support was viewed positively, users approached XAI with skepticism. Interestingly, we found a positive correlation between diagnostic improvement and XAI intervention. Although XAI support did not significantly enhance overall performance relative to AI alone, it prompted users to critically evaluate images with false predictions and influenced their confidence levels. These findings highlight the complexities of trust, perceived usefulness, and interpretability in AI-assisted diagnostics, with important implications for developing and implementing clinical decision-support tools in facial phenotyping for rare genetic diseases.
Artificial intelligence (AI) has been growing more powerful and accessible, and will increasingly impact many areas, including virtually all aspects of medicine and biomedical research. This review focuses on previous, current, and especially emerging applications of AI in clinical genetics. Topics covered include a brief explanation of different general categories of AI, including machine learning, deep learning, and generative AI. After introductory explanations and examples, the review discusses AI in clinical genetics in three main categories: clinical diagnostics; management and therapeutics; clinical support. The review concludes with short, medium, and long-term predictions about the ways that AI may affect the field of clinical genetics. Overall, while the precise speed at which AI will continue to change clinical genetics is unclear, as are the overall ramifications for patients, families, clinicians, researchers, and others, it is likely that AI will result in dramatic evolution in clinical genetics. It will be important for all those involved in clinical genetics to prepare accordingly in order to minimize the risks and maximize benefits related to the use of AI in the field.
Unlike some health conditions that have been extensively delineated throughout the lifespan, many genetic conditions are largely described in pediatric populations, with a focus on early manifestations like congenital anomalies and developmental delay. An apparent gap exists in understanding clinical features and optimal management as patients age. Generative artificial intelligence is transforming biomedical disciplines including through the introduction of large language models (LLMs). Motivated by these advances, we explored how LLMs handle age with respect to 282 genetic conditions selected based on prevalence. We divided these conditions into five categories: Disorders limited to childhood; Disorders limited to adulthood; Disorders with changes in presentation across ages; Disorders with changes in management across ages; Disorders with no changes across ages. We evaluated Llama-2-70b-chat (70b) and GPT-3.5 (GPT) capabilities at generating accurate medical vignettes for these conditions based on Correctness, Completeness, and Conciseness as graded by 3 clinicians. Using accurately generated vignettes as in-context prompts, we further generated and evaluated patient-geneticist dialogues and assessed LLM performance in answering specific questions regarding age-based management plans for a subset of conditions. Results revealed impressive performances of 70b with in-context prompting and GPT in generating vignettes. We overall did not observe age-based biases, though our experiments identified statistically significant differences in some areas related to LLM output. Despite impressive capabilities, LLMs still have limitations in clinical applications.
Most genetic conditions are described in pediatric populations, leaving a gap in understanding their clinical progression and management in adulthood. Motivated by other applications of large language models (LLMs), we evaluated whether Llama-2-70b-chat (70b) and GPT-3.5 (GPT) could generate plausible medical vignettes, patient-geneticist dialogues and management plans for a hypothetical child and adult patients across 282 genetic conditions (selected by prevalence and categorized based on age-related characteristics). Results showed that LLMs provided appropriate age-based responses in both child and adult outputs based on Correctness and Completeness scores graded by clinicians. Sub-analysis of metabolic conditions including those typically presents neonatally with crisis also showed age-appropriate LLM responses. However 70b and GPT obtained low Correctness and Completeness scores at producing plausible management plans (55-66% for 70b and a wider range, 50-90%, for GPT). This suggests that LLMs still have some limitations in clinical applications.
Stargardt disease, also called ABCA4-related retinopathy (ABCA4R), is the most common form of juvenile-onset macular dystrophy and yet lacks an FDA approved treatment. Substantial progress has been made through landmark studies like that of the Progression of Atrophy Secondary to Stargardt Disease (ProgStar), but tasks like image segmentation and phenotyping still pose major challenges in terms of monitoring disease progression and categorizing patient subgroups. Furthermore, these methods are subjective and laborious. Recent advancements in machine learning (ML) and deep learning show considerable promise in automating these processes. This scoping review explores ML applications in ABCA4R, with a focus on segmentation and phenotyping. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) methodology, 15 articles were selected from 264, with 12 focused on the task of segmenting atrophic lesions, retinal flecks, retinal layer boundaries, or en-face imaging. Three studies addressed phenotyping based on electroretinography (ERG), visual acuity, and microperimetry. Several effective approaches were implemented in these studies, including ensemble modeling, self-attention mechanisms, soft-label approaches, and dynamic frameworks that consider extent of tissue damage. Excellent model performance includes segmentation DICE performances of 0.99 and ERG phenotyping accuracies 90
Background Rare bone diseases (RBDs) are an important group of conditions characterized by abnormalities in bone and cartilage. Their large number, individual rarity, and heterogeneity make accurate and timely diagnosis challenging. Establishing correlations between genotype and phenotype (mainly via imaging) is critical for diagnosing RBDs. Image recognition artificial intelligence (AI) has the potential to significantly improve the diagnostic process by assisting healthcare providers to identify and differentiate imaging patterns associated with various RBDs. This survey study sought to assess the interest of various healthcare providers worldwide in utilizing an AI-based assistant tool for the differential diagnosis of RBDs. Method Survey data were collected from March to September 2024. The survey was performed online and the link was disseminated via direct email, newsletters, and flyers at scientific talks and conferences. Results We received 103 completed surveys, representing respondents from 27 different countries covering most global regions, but mostly from Europe, the United States, and Canada. The majority of the participants are physicians (n = 92, 89%) and primarily work at academic medical centers (n = 84, 81%). While each participant could select multiple specialties, the most frequent clinician types were medical geneticists, pediatricians, and endocrinologists, accounting for 71 (69%) of the respondents. Ninety-four (91%) of the respondents find imaging to be very or extremely important, and the majority (n = 84, 81%) consider X-rays to be the most important imaging modality. Although around half of the participants (n = 45) have concerns about AI-related errors and consider the explainability of AI algorithms to be very (42/103) or extremely (9/103) important, 81% of the respondents report that they are somewhat (n = 39) or extremely (n = 45) likely to consider integrating image recognition AI into their current diagnostic workflow. Conclusions Most survey participants are open to integrating image recognition AI into their RBD diagnostic workflow. However, concerns about AI-related errors, privacy, and model interpretability highlight the importance of transparent collaboration between developers and healthcare professionals throughout the development process to ensure that such technologies are clinically trustworthy and practically adoptable.
Background: The rate of diagnosis of mast cell activation syndrome (MCAS) has increased since the disorder's original description as a mastocytosis-like phenotype. While a set of consortium MCAS criteria is well described and widely accepted, this increase occurs in the setting of a broader set of proposed alternative MCAS criteria. Objective: Effective diagnostic criteria must minimize the range of unrelated diagnoses that can be erroneously classified as the condition of interest. We sought to determine if the symptoms associated with alternative MCAS criteria result in less concise or consistent diagnostic alternatives, reducing diagnostic specificity. Methods: We used multiple large language models, including ChatGPT, Claude, and Gemini, to bootstrap the probabilities of diagnoses that are compatible with consortium or alternative MCAS criteria. We utilized diversity and network analyses to quantify diagnostic precision and specificity compared to control diagnostic criteria including systemic lupus erythematosus, Kawasaki disease, and migraines. Results: Compared to consortium MCAS criteria, alternative MCAS criteria are associated with more variable (Shannon diversity 5.8 vs 4.6, respectively; P = .004) and less precise (mean Bray-Curtis similarity 0.07 vs 0.19, respectively; P = .004) diagnoses. The diagnosis networks derived from consortium and alternative MCAS criteria had lower between- network similarity compared to the similarity between diagnosis networks derived from 2 distinct systemic lupus erythematosus criteria (cosine similarity 0.55 vs 0.86, respectively; P = .0022). Conclusion: Alternative MCAS criteria are associated with a distinct set of diagnoses compared to consortium MCAS criteria and have lower diagnostic consistency. This lack of specificity is pronounced in relation to multiple control criteria, raising the concern that alternative criteria could disproportionately contribute to MCAS overdiagnosis, to the exclusion of more appropriate diagnoses. (J Allergy Clin Immunol 2025;155:213-8.)
Large language models (LLMs) are generating interest in medical settings. For example, LLMs can respond coherently to medical queries by providing plausible differential diagnoses based on clinical notes. However, there are many questions to explore, such as evaluating differences between open- and closed-source LLMs as well as LLM performance on queries from both medical and non-medical users. In this study, we assessed multiple LLMs, including Llama-2-chat, Vicuna, Medllama2, Bard/Gemini, Claude, ChatGPT3.5, and ChatGPT-4, as well as non-LLM approaches (Google search and Phenomizer) regarding their ability to identify genetic conditions from textbook-like clinician questions and their corresponding layperson translations related to 63 genetic conditions. For open-source LLMs, larger models were more accurate than smaller LLMs: 7b, 13b, and larger than 33b parameter models obtained accuracy ranges from 21%-49%, 41%-51%, and 54%-68%, respectively. Closed-source LLMs outperformed open-source LLMs, with ChatGPT-4 performing best (89%-90%). Three of 11 LLMs and Google search had significant performance gaps between clinician and layperson prompts. We also evaluated how in-context prompting and keyword removal affected open-source LLM performance. Models were provided with 2 types of in-context prompts: list-type prompts, which improved LLM performance, and definition-type prompts, which did not. We further analyzed removal of rare terms from descriptions, which decreased accuracy for 5 of 7 evaluated LLMs. Finally, we observed much lower performance with real individuals' descriptions; LLMs answered these questions with a maximum 21% accuracy.