BackgroundThe human voice contains rich acoustic information indicative of laryngeal pathology, yet current screening relies on resource-intensive in-person laryngoscopy. While artificial intelligence has shown promise for voice analysis, progress has been limited by small, inconsistent datasets and challenges to clinical translation. The Bridge2AI-Voice initiative addresses these barriers by providing a large-scale, ethically sourced dataset with standardized, privacy-preserving derived features.ObjectiveTo determine whether the derived-feature release of Bridge2AI-Voice v3.0.0 can support a high-sensitivity screening model for laryngeal lesions and to evaluate its translational readiness using telemedicine implementation frameworks.MethodsWe analyzed data from 205 adult participants (136 controls, 52 benign vocal fold lesions, 13 precancerous lesions, 4 laryngeal cancer) drawn from the Bridge2AI-Voice v3.0.0 derived-feature release. An L2-regularized logistic regression model was fit to 131 OpenSMILE static acoustic features with age and sex at birth, evaluated under participant-level stratified 10-fold nested cross-validation. Inner-fold cross-validation was used for operating-point threshold selection. Pre-specified validity tests against age confounding included a DeLong comparison against an age-only baseline and an age-stratified label permutation test. Alternative feature modalities (SPARC articulatory features, Mel spectrogram derivatives, and multimodal combinations) and alternative classifier families were evaluated as robustness checks.ResultsThe OpenSMILE-based model achieved cross-validated AUC 0.812 (95% CI 0.744–0.876), with operating-point sensitivity 0.870 (95% CI 0.767–0.939) and specificity 0.566 (95% CI 0.479–0.651). Model discrimination significantly exceeded an age-only baseline (DeLong p = 0.0008) and survived age-stratified label permutation (observed AUC 0.812 vs. null mean 0.553, p = 0.0099). Subgroup analysis showed approximately consistent sensitivity across benign (0.865) and precancerous (0.846) lesion subgroups. Alternative feature modalities did not provide incremental discriminative information beyond OpenSMILE, and alternative classifier families produced AUCs within bootstrap confidence intervals of the primary model.ConclusionsDerived acoustic features from the Bridge2AI-Voice v3.0.0 release combined with basic demographic information support cross-validated discrimination of vocal fold lesions consistent with the upper range of published voice-based laryngeal pathology classifiers. The result is presented as a candidate signal warranting confirmatory investigation in a larger, prospectively recruited cohort.
Recent advances in large language models (LLMs) have made significant progress across multiple biomedical tasks, including biomedical question answering, lay-language summarization of the biomedical literature, and clinical note summarization. These models have demonstrated strong capabilities in processing and synthesizing complex biomedical information and in generating fluent, human-like responses. Despite these advancements, hallucinations or confabulations remain key challenges when using LLMs in biomedical and other high-stakes domains. Inaccuracies may be particularly harmful in high-risk situations, such as medical question answering, making clinical decisions, or appraising biomedical research. Studies on the evaluation of the LLMs' abilities to ground generated statements in verifiable sources have shown that models perform significantly
Objectives Clinical document metadata, such as document type, structure, author role, medical specialty, and encounter setting, is essential for accurate interpretation of information captured in clinical documents. However, vast documentation heterogeneity and drift over time challenge harmonization of document metadata. Automated extraction methods have emerged to coalesce metadata from disparate practices into target schema. This scoping review aims to catalog research on clinical document metadata extraction, identify methodological trends and applications, and highlight gaps warranting further investigation. Methods We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines to identify articles from Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science and external sources that perform clinical document metadata extraction, either primarily as a methodology study, secondarily as a feature for a downstream application, or for analysis. We initially identified and screened 342 articles published between 2011 and 2025, then comprehensively reviewed 77 we deemed relevant to our study. Results Among the 77 articles included in our full text review, 49 were methodological, 22 used document metadata as features in a downstream application, and 6 analyzed document metadata composition. We observe myriad purposes for methodological study and application types. Available labelled public data remains sparse except for structural section datasets. Methods for extracting document metadata have progressed from largely rule-based and traditional machine learning with ample feature engineering to transformer-based architectures with minimal feature engineering. Discussion and conclusion Clinical document metadata extraction research has accelerated over recent years. The emergence of large language models has enabled broader exploration of generalizability across tasks and datasets, allowing the possibility of advanced clinical text processing systems. We anticipate that research will continue to expand into richer document metadata representations and integrate further into clinical applications and workflows.
Competencies help define the skills and knowledge needed by learners. Often broad, educators integrate competencies to provide a framework for curricula or professional standards. For data science, the rate of change in the field, role variations, and specificity in key applications can be challenging. Our objective was to adapt general data science competencies for different learner roles in an emerging area: the clinical utility of Voice, Language, and Speech-based Artificial Intelligence/Machine Learning (AI/ML). Using a persona-inductive approach, we adapted competencies to support learners from varying professional and educational backgrounds and implemented these adaptations in a multi-institutional summer school. Results from these pilot efforts demonstrated feasibility, highlighted the importance of cross-role collaboration, and provided lessons for scaling to broader audiences. Our frameworks show that competency adaptation is necessary and practical in rapidly evolving AI domains.
Generative artificial intelligence (AI) has had a profound impact on biomedicine and health, both in professional work and in education. Based on large language models (LLMs), generative AI has been found to perform as well as humans in simulated situations taking medical board exams, answering clinical questions, solving clinical cases, applying clinical reasoning, and summarizing information. Generative AI is also being used widely in education, performing well in academic courses and their assessments. This review summarizes the successes of LLMs and highlights some of their challenges in the context of education, most notably aspects that may undermine the acquisition of knowledge and skills for professional work. It then provides recommendations for best practices to overcome the shortcomings of LLM use in education. Although there are challenges for the use of generative AI in education, all students and faculty, in biomedicine and health and beyond, must have understanding of it and be competent in its use.
Objective: As AI becomes increasingly central to healthcare, there is a pressing need for bioinformatics and biomedical training systems that are personalized and adaptable. Materials and Methods: The NIH Bridge2AI Training, Recruitment, and Mentoring (TRM) Working Group developed a cross-disciplinary curriculum grounded in collaborative innovation, ethical data stewardship, and professional development within an adapted Learning Health System (LHS) framework. Results: The curriculum integrates foundational AI modules, real-world projects, and a structured mentee-mentor network spanning Bridge2AI Grand Challenges and the Bridge Center. Guided by six learner personas, the program tailors educational pathways to individual needs while supporting scalability. Discussion: Iterative refinement driven by continuous feedback ensures that content remains responsive to learner progress and emerging trends. Conclusion: With over 30 scholars and 100 mentors engaged across North America, the TRM model demonstrates how adaptive, persona-informed training can build interdisciplinary competencies and foster an integrative, ethically grounded AI education in biomedical contexts.
Benign and malignant vocal fold lesions can alter voice quality and lead to significant morbidity or, in the case of malignancy, mortality. Early, noninvasive identification of these lesions using voice as a biomarker may improve diagnostic access and outcomes. In this study, we analyzed data from the initial release of the Bridge2AI-Voice dataset to evaluate which acoustic features best distinguish laryngeal cancer and benign vocal fold lesions from other vocal pathologies and healthy voice function. Seven diagnostic cohorts were grouped into two analyses: the first included participants with laryngeal cancer, benign lesions, or no voice disorder; the second included those with laryngeal cancer or benign lesions without other voice disorders, as well as individuals with spasmodic dysphonia or vocal fold paralysis. Acoustic features including fundamental frequency, jitter, shimmer, and harmonic-to-noise ratio (HNR) were extracted from standardized speech recordings and compared using nonparametric statistical methods. Among the overall sample, significant differences were identified in HNR and fundamental frequency between benign lesions and both healthy controls and laryngeal cancer. In cisgender men, these distinctions were also observed, particularly in HNR and its variability. No statistically significant differences were observed among cisgender women, likely due to the limited sample size. These findings suggest that HNR, particularly its variability, may hold promise as a voice-based marker for early detection and monitoring of vocal fold lesions. Further research with larger, more diverse populations is needed to refine these features and validate their clinical utility.
OBJECTIVE:Unlocking clinical information embedded in clinical notes has been hindered to a significant degree by domain-specific and context-sensitive language. Identification of note sections and structural document elements has been shown to improve information extraction and dependent downstream clinical natural language processing (NLP) tasks and applications. This study investigates the viability of a dynamic example selection prompting method to section classification using lightweight, open-source large language models (LLMs) as a practical solution for real-world healthcare clinical NLP systems. MATERIALS AND METHODS:We develop a dynamic few-shot prompting approach to classifying sections where section samples are first embedded using a transformer-based model and deposited in a vector store. During inference, the embedded samples with the most similar contextual embeddings to a given input section text are retrieved from the vector store and inserted into the LLM prompt. We evaluate this technique on two datasets comprising two section schemas, including varying levels of context. We compare the performance to baseline zero-shot and randomly selected few-shot scenarios. RESULTS:The dynamic few-shot prompting experiments yielded the highest F1 scores in each of the classification tasks and datasets for all seven of the LLMs included in the evaluation, averaging a macro F1 increase of 39.3% and 21.1% in our primary section classification task over the zero-shot and static few-shot baselines, respectively. DISCUSSION AND CONCLUSION:Our results showcase substantial performance improvements imparted by dynamically selecting examples for few-shot LLM prompting, and further improvement by including section context, demonstrating compelling potential for clinical applications.
Objective: Information retrieval (IR, also known as search) systems are ubiquitous in modern times. How does the emergence of generative artificial intelligence (AI), based on large language models (LLMs), fit into the IR process? Process: This perspective explores the use of generative AI in the context of the motivations, considerations, and outcomes of the IR process with a focus on the academic use of such systems. Conclusions: There are many information needs, from simple to complex, that motivate use of IR. Users of such systems, particularly academics, have concerns for authoritativeness, timeliness, and contextualization of search. While LLMs may provide functionality that aids the IR process, the continued need for search systems, and research into their improvement, remains essential.
Generative artificial intelligence (AI) systems have performed well at many biomedical tasks, but few studies have assessed their performance directly compared to students in higher-education courses. We compared student knowledge-assessment scores with prompting of 6 large-language model (LLM) systems as they would be used by typical students in a large online introductory course in biomedical and health informatics that is taken by graduate, continuing education, and medical students. The state-of-the-art LLM systems were prompted to answer multiple-choice questions (MCQs) and final exam questions. We compared the scores for 139 students (30 graduate students, 85 continuing education students, and 24 medical students) to the LLM systems. All of the LLMs scored between the 50th and 75th percentiles of students for MCQ and final exam questions. The performance of LLMs raises questions about student assessment in higher education, especially in courses that are knowledge-based and online.
Clinical information retrieval (IR) plays a vital role in modern healthcare by facilitating efficient access and analysis of medical literature for clinicians and researchers. This scoping review aims to offer a comprehensive overview of the current state of clinical IR research and identify gaps and potential opportunities for future studies in this field. The main objective was to assess and analyze the existing literature on clinical IR, focusing on the methods, techniques, and tools employed for effective retrieval and analysis of medical information. Adhering to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, we conducted an extensive search across databases such as Ovid Embase, Ovid Medline, Scopus, ACM Digital Library, IEEE Xplore, and Web of Science, covering publications from January 1, 2010, to January 4, 2023. The rigorous screening process led to the inclusion of 184 papers in our review. Our findings provide a detailed analysis of the clinical IR research landscape, covering aspects like publication trends, data sources, methodologies, evaluation metrics, and applications. The review identifies key research gaps in clinical IR methods such as indexing, ranking, and query expansion, offering insights and opportunities for future studies in clinical IR, thus serving as a guiding framework for upcoming research efforts in this rapidly evolving field. The study also underscores an imperative for innovative research on advanced clinical IR systems capable of fast semantic vector search and adoption of neural IR techniques for effective retrieval of information from unstructured electronic health records (EHRs).
The First Search Futures Workshop, in conjunction with the Fourty-sixth European Conference on Information Retrieval (ECIR) 2024, looked into the future of search to ask questions such as: • How can we harness the power of generative AI to enhance, improve and re-imagine Information Retrieval (IR)? • What are the principles and fundamental rights that the field of Information Retrieval should strive to uphold? • How can we build trustworthy IR systems in light of Large Language Models and their ability to generate content at super human speeds? • What new applications and affordances does generative AI offer and enable, and can we go back to the future, and do what we only dreamed of previously? The workshop started with seventeen lightning talks from a diverse set speakers. Instead of conventional paper presentations, the lightning talks provided a rapid and concise overview of ideas, allowing speakers to share critical points or novel concepts quickly. This format was designed to encourage discussion and introduce a wide range of topics within a short period, thereby maximising the exchange of ideas and ensuring that participants could gain insights into various future search areas without the deep dive typically required in longer presentations. This report, co-authored by the workshop's organisers and its participants, summarises the talks and discussions. This report aims to provide the broader IR community with the insights and ideas discussed and debated during the workshop - and to provide a platform for future discussion. Date: 24 March 2024. Website: https://searchfutures.github.io/.
With the advancement of large language models (LLMs), the biomedical domain has seen significant progress and improvement in multiple tasks such as biomedical question answering, lay language summarization of the biomedical literature, clinical note summarization, etc. However, hallucinations or confabulations remain one of the key challenges when using LLMs in the biomedical and other domains. Inaccuracies may be particularly harmful in high-risk situations, such as making clinical decisions or appraising biomedical research. Studies on the evaluation of the LLMs' abilities to ground generated statements in verifiable sources have shown that models perform significantly worse on lay-user generated questions, and often fail to reference relevant sources. This can be problematic when those seeking information want evidence from studies to back up the claims from LLMs[3]. Unsupported statements are a major barrier to using LLMs in any applications that may affect health. Methods for grounding generated statements in reliable sources along with practical evaluation approaches are needed to overcome this barrier. Towards this, in our pilot task organized at TREC 2024, we introduced the task of reference attribution as a means to mitigate the generation of false statements by LLMs answering biomedical questions.
The value and methods of online learning have changed tremendously over the last 25 years. The goal of this paper is to review a quarter-century of experience with online learning by the author in the field of biomedical and health informatics, describing the learners served and the lessons learned. The author details the history of the decision to pursue online education in informatics, describing the approaches taken as educational technology evolved over time. A large number of learners have been served, and the online learning approach has been well-received, with many lessons learned to optimize the educational experience. Online education in biomedical and health informatics has provided a scalable and exemplary approach to learning in this field.
Data science, machine learning and artificial intelligence applications impact clinicians, informaticians, science journalists, and researchers. Most biomedical data science training focuses on learning a programming language in addition to higher mathematics and advanced statistics. This approach is appropriate for graduate students but greatly reduces the number of individuals in healthcare who can be involved in data science. To serve these four stakeholder audiences, we describe several curricular strategies focusing on solving real problems of interest to these audiences. Relevant competencies for these audiences include using intuitive programming tools that facilitate data exploration with minimal programming background, creating data models, evaluating results of data analyses, and assessing data science research reports, among others. Offering the curricula described here more broadly could broaden the stakeholder groups knowledgeable about and engaged in data science.
Received: 22 September 2022 Accepted after revision: 02 December 2022 Accepted Manuscript online:19 December 2022
OBJECTIVES:Electronic health record (EHR) data may facilitate the identification of rare diseases in patients, such as aromatic l-amino acid decarboxylase deficiency (AADCd), an autosomal recessive disease caused by pathogenic variants in the dopa decarboxylase gene. Deficiency of the AADC enzyme results in combined severe reductions in monoamine neurotransmitters: dopamine, serotonin, epinephrine, and norepinephrine. This leads to widespread neurological complications affecting motor, behavioral, and autonomic function. The goal of this study was to use EHR data to identify previously undiagnosed patients who may have AADCd without available training cases for the disease. MATERIALS AND METHODS:A multiple symptom and related disease annotated dataset was created and used to train individual concept classifiers on annotated sentence data. A multistep algorithm was then used to combine concept predictions into a single patient rank value. RESULTS:Using an 8000-patient dataset that the algorithms had not seen before ranking, the top and bottom 200 ranked patients were manually reviewed for clinical indications of performing an AADCd diagnostic screening test. The top-ranked patients were 22.5% positively assessed for diagnostic screening, with 0% for the bottom-ranked patients. This result is statistically significant at P < .0001. CONCLUSION:This work validates the approach that large-scale rare-disease screening can be accomplished by combining predictions for relevant individual symptoms and related conditions which are much more common and for which training data is easier to create.