Clinical trials are essential for generating evidence in drug development, yet they continue to face persistent challenges including high costs, lengthy timelines, and low success rates. Recent advances in artificial intelligence (AI) and its transformative impact across medicine have highlighted its potential to address these limitations. Accordingly, AI is being increasingly integrated throughout the clinical trial lifecycle, enabling innovations in study design, patient recruitment, trial monitoring, data analysis, outcome prediction, and operational decisionmaking. In this review, we provide a comprehensive overview of current and emerging AI applications in clinical trials, emphasizing their potential to improve efficiency, enhance trial quality, and accelerate evidence generation. We further examine key practical considerations for implementation, including data quality, model interpretability, regulatory requirements, ethical concerns, and barriers to real-world adoption. Finally, we discuss future directions for AIenabled clinical trials, highlighting opportunities for improved scalability, generalizability, and clinical impact. By synthesizing recent advances and ongoing challenges, this review aims to guide researchers and practitioners navigating the rapidly evolving landscape of AI in clinical trials and to provide actionable insights for AI researchers, clinicians, patient advocates, trial investigators, and drug developers seeking to integrate AI into clinical research practice.
The rapid expansion of biomedical publications creates challenges for organizing knowledge and detecting emerging trends, underscoring the need for scalable and interpretable methods. Common clustering and topic modeling approaches such as K-means or LDA remain sensitive to initialization and prone to local optima, limiting reproducibility and evaluation. We propose a reformulation of a convex-optimization-based clustering algorithm that produces stable, fine-grained topics by selecting exemplars from the data and guaranteeing a global optimum. Applied to ~12,000 PubMed articles on aging and longevity, our method uncovers topics validated by medical experts. It yields interpretable topics spanning from molecular mechanisms to dietary supplements, physical activity, and gut microbiota. The method performs favorably, and most importantly, its reproducibility and interpretability distinguish it from common clustering approaches, including K-means, LDA, and BERTopic. This work provides a basis for developing scalable, web-accessible tools for knowledge discovery.
Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4 and DeepSeek-R1, represent a transformative class of artificial intelligence tools capable of revolutionizing various aspects of healthcare by generating human-like responses across diverse contexts and adapting to novel tasks following human instructions. Their potential application spans a broad range of medical tasks, such as clinical documentation, matching patients to clinical trials and answering medical questions. Here in this Tutorial, we discuss an actionable set of best practices to help healthcare professionals utilize LLMs more effectively and efficiently. The overall workflow follows sequential phases from formulating the task, choosing the most appropriate LLMs, engineering the prompts, fine-tuning the requests and through to model deployment. We discuss a set of critical considerations in identifying medical tasks that align with the core capabilities of LLMs and selecting models based on the required task, data, performance and model interface. We then review the strategies, such as prompt engineering and fine-tuning, to adapt standard LLMs to specialized medical tasks. We then cover deployment considerations, including regulatory compliance, ethical guidelines and continuous monitoring for fairness and bias. By providing a structured step-by-step methodology, this entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Trustworthiness and transparency are essential for the clinical adoption of artificial intelligence (AI) in healthcare and biomedical research. Recent deep research systems aim to accelerate evidence-grounded scientific discovery by integrating AI agents with multi-hop information retrieval, reasoning, and synthesis. However, most existing systems lack explicit and inspectable criteria for evidence appraisal, creating a risk of compounding errors and making it difficult for researchers and clinicians to assess the reliability of their outputs. In parallel, current benchmarking approaches rarely evaluate performance on complex, real-world medical questions. Here, we introduce DeepER-Med, a Deep Evidence-based Research framework for Medicine with an agentic AI system. DeepER-Med frames deep medical research as an explicit and inspectable workflow of evidence-based generation (EBG), consisting of three modules: research planning, agentic collaboration, and evidence synthesis. To support realistic evaluation, we also present DeepER-MedQA, an evidence-grounded dataset comprising 100 expert-level research questions derived from authentic medical research scenarios and curated by a multidisciplinary panel of 11 biomedical experts. Expert manual evaluation demonstrates that DeepER-Med consistently outperforms widely used production-grade platforms across multiple criteria, including the generation of novel scientific insights. Beyond manual assessment, we further evaluate DeepER-Med across distinct stages of EBG using quantitative metrics, including semantic similarity and information entropy, capturing both system performance and the relevance of retrieved evidence. We further demonstrate the practical utility of DeepER-Med through eight real-world clinical cases. Human clinician assessment indicates that DeepER-Med's conclusions align with clinical recommendations in seven cases, highlighting its potential for medical research and decision support.
MOTIVATION:Gene set analysis (GSA) is a foundational approach for interpreting genomic data of diseases by linking genes to biological processes. However, conventional GSA methods overlook clinical context of the analyses, often generating long lists of enriched pathways with redundant, nonspecific, or irrelevant results. Interpreting these requires extensive, ad-hoc manual effort, reducing both reliability and reproducibility. RESULTS:We introduce cGSA, a novel AI-driven framework that enhances GSA by incorporating context-aware pathway prioritization. cGSA integrates gene cluster detection, enrichment analysis, and large language models to identify pathways that are not only statistically significant but also biologically meaningful. Benchmarking on 102 curated gene sets across 19 diseases and ten disease-related biological mechanisms shows that cGSA outperforms baseline methods by over 30%, with expert validation confirming its increased precision and interpretability. Two independent case studies in melanoma and breast cancer further demonstrate its potential to uncover context-specific insights and support targeted hypothesis. AVAILABILITY AND IMPLEMENTATION:The demo website is publicly available at https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/cGSA/, while the data and code can be accessed at https://github.com/ncbi-nlp/cGSA.
Abstract Background Identifying appropriate clinical trials for oncology patients is challenging due to complex eligibility criteria, heterogeneous clinical documentation, and evolving trial landscapes, particularly in rare cancers such as cholangiocarcinoma. Manual trial screening is time consuming and may overlook relevant opportunities. This study presents an AI-assisted trial matching system customized to support efficient and clinically meaningful trial identification for cholangiocarcinoma patients. Methods We developed an AI-assisted trial matching system that supports real-world cholangiocarcinoma patient information in different formats, including both structured EHR-based data and unstructured narrative clinical summaries, and prioritizes the most relevant candidate trials for clinician review. For a given patient, the system analyzes clinical context and matches it against hundreds of cholangiocarcinoma-related trials, including biomarker-selected trials that allow cholangiocarcinoma. Subsequently, the system produces a ranked list of candidate trials with structured explanations (including eligible reasons, ineligible reasons, and missing information). System performance was manually assessed by clinicians across 19 patient cases. Results We evaluated the system on 9 structured and 10 unstructured cholangiocarcinoma patient cases. Eligibility agreement with clinician review was 73.4% (structured) and 77.8% (unstructured). Among clinician-confirmed eligible trials, 56.3% (structured) and 73.8% (unstructured) were judged clinically appropriate to recommend. The system consistently prioritized clinically meaningful trials and was effective at identifying biomarker-specific studies embedded within broader solid tumor protocols, which clinicians noted are challenging to identify through manual screening alone. Clinicians also found the explicit explanation of missing or uncertain eligibility information valuable for guiding follow-up eligibility review. From a clinician perspective, the Al-generated explanations were efficient to review, typically requiring 1-2 minutes per trial. Conclusion This study demonstrates the real-world usability and effectiveness of using AI-assisted trial matching for cholangiocarcinoma trial screening. Future directions include further refinement to improve performance, and additional testing in a real-life advocacy setting.
Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there is growing interest in using AI to support patient recruitment, existing systems largely perform eligibility assessment alone and have rarely been evaluated in real-world oncology workflows. Here we present TrialGPT 2.0, an AI-assisted clinical trial recommendation system designed for real-world deployment. Rather than asking only whether a patient may qualify, the system also assesses which trials warrant further consideration given the patient's current clinical needs and local workflow priorities, and provides structured, inspectable explanations for expert review. Importantly, we evaluated TrialGPT 2.0 retrospectively and prospectively across multiple oncology-focused settings, spanning government, academic cancer-center, patient-advocacy, and NIH referral workflows. In retrospective multicenter cohorts comprising 288 cases, TrialGPT 2.0 retrieved at least one clinician-recommended trial in its top 10 recommendations for approximately 91
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish genuine reasoning from pattern matching and that remain discriminative as model capabilities improve. Existing biomedical question answering (QA) benchmarks are limited in this respect: multiple-choice formats allow models to succeed by answer elimination rather than inference, and widely circulated exam-style datasets are subject to performance saturation and training data contamination. Multi-hop reasoning. i.e., the ability to integrate information across multiple sources to derive an answer, is central to clinically meaningful tasks such as diagnosis support, literature-based discovery, and hypothesis generation, yet remains underrepresented in current biomedical QA benchmarks. We present MedHopQA, a disease-centered multi-hop reasoning benchmark of 1,000 expert-curated question-answer pairs introduced as a shared task at BioCreative IX. Each question requires synthesis of information across two distinct Wikipedia articles, and answers are provided in open-ended free-text format rather than as multiple-choice selections. Gold annotations are augmented with ontology-grounded synonym sets (MONDO, NCBI Gene, NCBI Taxonomy) to support both lexical and concept-level evaluation. The dataset was constructed through a multi-stage human-AI pipeline combining structured human annotation, triage, iterative verification, and LLM-as-a-judge validation. To reduce leaderboard gaming and contamination risk, the 1,000 scored questions are embedded within a publicly downloadable set of 10,000 questions, with answers withheld, on a CodaBench leaderboard. Evaluation of four frontier LLMs under a zero-shot setting (GPT-5.1, Gemini 2.5 Pro, Claude Sonnet 4.5, and GPT-4o) reveals performance variation across answer types, with overall accuracy ranging from 66.3% to 83.4%. Performance is strongest on chemical and anatomical questions and most variable on disease and gene/protein categories, where fine-grained semantic discrimination is required. MedHopQA provides both a benchmark and a reusable framework for constructing future biomedical QA datasets that prioritize compositional reasoning, saturation resistance, and contamination mitigation as design constraints.
LitSense 2.0 (https://www.ncbi.nlm.nih.gov/research/litsense2/) is an advanced biomedical search system enhanced with dense vector semantic retrieval, designed for accessing literature on sentence and paragraph levels. It provides unified access to 38 million PubMed abstracts and 6.6 million full-length articles in the PubMed Central (PMC) Open Access subset, encompassing 1.4 billion sentences and ∼300 million paragraphs, and is updated weekly. Compared to PubMed and PMC, the primary platforms for biomedical information search, LitSense offers cross-platform functionality by searching seamlessly across both PubMed and PMC and returning relevant results at a more granular level. Building on the success of the original LitSense launched in 2018, LitSense 2.0 introduces two major enhancements. The first is the addition of paragraph-level search: users can now choose to search either against sentences or against paragraphs. The second is improved retrieval accuracy via a state-of-the-art biomedical text encoder, ensuring more reliable identification of relevant results across the entire biomedical literature.
To address the rapid growth of scientific publications and data in biomedical research, knowledge graphs (KGs) have become a critical tool for integrating large volumes of heterogeneous data to enable efficient information retrieval and automated knowledge discovery. However, transforming unstructured scientific literature into KGs remains a significant challenge, with previous methods unable to achieve human-level accuracy. Here we used an information extraction pipeline that won first place in the LitCoin Natural Language Processing Challenge (2022) to construct a large-scale KG named iKraph using all PubMed abstracts. The extracted information matches human expert annotations and significantly exceeds the content of manually curated public databases. To enhance the KG’s comprehensiveness, we integrated relation data from 40 public databases and relation information inferred from high-throughput genomics data. This KG facilitates rigorous performance evaluation of automated knowledge discovery, which was infeasible in previous studies. We designed an interpretable, probabilistic-based inference method to identify indirect causal relations and applied it to real-time COVID-19 drug repurposing from March 2020 to May 2023. Our method identified around 1,200 candidate drugs in the first 4 months, with one-third of those discovered in the first 2 months later supported by clinical trials or PubMed publications. These outcomes are very challenging to attain through alternative approaches that lack a thorough understanding of the existing literature. A cloud-based platform ( https://biokde.insilicom.com ) was developed for academic users to access this rich structured data and associated tools. This study presents iKraph, a large-scale biomedical knowledge graph built using an award-winning natural language processing pipeline with expert-level accuracy. Using probabilistic semantic reasoning, iKraph enables automated knowledge discovery with excellent performance.
Biological relation networks contain rich information for understanding the biological mechanisms behind the relationship of entities such as genes, proteins, diseases, and chemicals. The vast growth of biomedical literature poses significant challenges updating the network knowledge. The recent Biomedical Relation Extraction Dataset (BioRED) provides valuable manual annotations, facilitating the develop-ment of machine-learning and pre-trained language model approaches for automatically identifying novel document-level (inter-sentence context) relationships. Nonetheless, its annotations lack directionality (subject/object) for the entity roles, essential for studying complex biological networks. Herein we annotate the entity roles of the relationships in the BioRED corpus and subsequently propose a novel multi-task language model with soft-prompt learning to jointly identify the relationship, novel findings, and entity roles. Our results in-clude an enriched BioRED corpus with 10,864 directionality annotations. Moreover, our proposed method outperforms existing large language models such as the state-of-the-art GPT-4 and Llama-3 on two benchmarking tasks. Our source code and dataset are available at https://github.com/ncbi-nlp/BioREDirect.
Supplementary materials accompanying scientific articles are critical components of biomedical research, offering detailed datasets, experimental protocols, and extended analyses that complement the main text. These materials play an important role in enhancing transparency, reproducibility, and scientific impact by providing in depth analyses and the details necessary for reproducing experiments. However, the lack of consistent and standard formats has limited the access to supplementary materials in scientific investigations. In response, we propose a novel system aimed to enhance FAIR access to Supplementary MAterials for Research Transparency (FAIR-SMART). Specifically, we first aggregate supplementary files in a single location, standardize them into structured and machine-readable format, and make them accessible via web APIs. Next, we employ advanced large language models to automatically categorize the tabular data, which represents over 90% of the textual content in supplementary materials, enabling precise and efficient data retrieval. By bridging the gap between diverse file types and automated workflows, this work not only advances biomedical research but also highlights the transformative potential of accessible supplementary materials in shaping the behaviors and decision-making processes of the scientific community. FAIR-SMART is freely available for supplementary materials data retrieval via its APIs: https://www.ncbi.nlm.nih.gov/research/bionlp/APIs/FAIR-SMART/.
Gene-set analysis seeks to identify the biological mechanisms underlying groups of genes with shared functions. Large language models (LLMs) have recently shown promise in generating functional descriptions for input gene sets but may produce factually incorrect statements, commonly referred to as hallucinations in LLMs. Here we present GeneAgent, an LLM-based AI agent for gene-set analysis that reduces hallucinations by autonomously interacting with biological databases to verify its own output. Evaluation of 1,106 gene sets collected from different sources demonstrates that GeneAgent is consistently more accurate than GPT-4 by a significant margin. We further applied GeneAgent to seven novel gene sets derived from mouse B2905 melanoma cell lines. Expert review confirmed that GeneAgent produces more relevant and comprehensive functional descriptions than GPT-4, providing valuable insights into gene functions and expediting knowledge discovery.
Gene set knowledge discovery is essential for advancing human functional genomics. Recent studies have shown promising performance by harnessing the power of Large Language Models (LLMs) on this task. Nonetheless, their results are subject to several limitations common in LLMs such as hallucinations. In response, we present GeneAgent, a first-of-its-kind language agent featuring self-verification capability. It autonomously interacts with various biological databases and leverages relevant domain knowledge to improve accuracy and reduce hallucination occurrences. Benchmarking on 1,106 gene sets from different sources, GeneAgent consistently outperforms standard GPT-4 by a significant margin. Moreover, a detailed manual review confirms the effectiveness of the self-verification module in minimizing hallucinations and generating more reliable analytical narratives. To demonstrate its practical utility, we apply GeneAgent to seven novel gene sets derived from mouse B2905 melanoma cell lines, with expert evaluations showing that GeneAgent offers novel insights into gene functions and subsequently expedites knowledge discovery.
To enable electronic screening of eligible patients for clinical trials, free-text clinical trial eligibility criteria should be translated to a computable format. Natural language processing (NLP) techniques have the potential to automate this process. In this study, we explored a supervised multi-input multi-output (MIMO) sequence labelling model to parse eligibility criteria into combinations of fact and condition tuples. Our experiments on a small manually annotated training dataset showed that that the performance of the MIMO framework with a BERT-based encoder using all the input sequences achieved an overall lenient-level AUROC of 0.61. Although the per-formance is suboptimal, representing eligibility criteria into logical and semantically clear tuples can potentially make subsequent translation of these tuples into database queries more reliable.
PubTator 3.0 (https://www.ncbi.nlm.nih.gov/research/pubtator3/) is a biomedical literature resource using state-of-the-art AI techniques to offer semantic and relation searches for key concepts like proteins, genetic variants, diseases, and chemicals. It currently provides over one billion entity and relation annotations across approximately 36 million PubMed abstracts and 6 million full-text articles from the PMC open access subset, updated weekly. PubTator 3.0's online interface and API utilize these precomputed entity relations and synonyms to provide advanced search capabilities and enable large-scale analyses, streamlining many complex information needs. We showcase the retrieval quality of PubTator 3.0 using a series of entity pair queries, demonstrating that PubTator 3.0 retrieves a greater number of articles than either PubMed or Google Scholar, with higher precision in the top 20 results. We further show that integrating ChatGPT (GPT-4) with PubTator APIs dramatically improves the factuality and verifiability of its responses. In summary, PubTator 3.0 offers a comprehensive set of features and tools that allow researchers to navigate the ever-expanding wealth of biomedical literature, expediting research and unlocking valuable insights for scientific discovery.
SUMMARY:Over 55% of author names in PubMed are ambiguous: the same name is shared by different individual researchers. This poses significant challenges on precise literature retrieval for author name queries, a common behavior in biomedical literature search. In response, we present a comprehensive dataset of disambiguated authors. Specifically, we complement the automatic PubMed Computed Authors algorithm with the latest ORCID data for improved accuracy. As a result, the enhanced algorithm achieves high performance in author name disambiguation, and subsequently our dataset contains more than 21 million disambiguated authors for over 35 million PubMed articles and is incrementally updated on a weekly basis. More importantly, we make the dataset publicly available for the community such that it can be utilized in a wide variety of potential applications beyond assisting PubMed's author name queries. Finally, we propose a set of guidelines for best practices of authors pertaining to use of their names. AVAILABILITY AND IMPLEMENTATION:The PubMed Computed Authors dataset is publicly available for bulk download at: https://ftp.ncbi.nlm.nih.gov/pub/lu/ComputedAuthors/. Additionally, it is available for query through web API at: https://www.ncbi.nlm.nih.gov/research/bionlp/APIs/authors/.
PubTator 3.0 (this https URL) is a biomedical literature resource using state-of-the-art AI techniques to offer semantic and relation searches for key concepts like proteins, genetic variants, diseases, and chemicals. It currently provides over one billion entity and relation annotations across approximately 36 million PubMed abstracts and 6 million full-text articles from the PMC open access subset, updated weekly. PubTator 3.0's online interface and API utilize these precomputed entity relations and synonyms to provide advanced search capabilities and enable large-scale analyses, streamlining many complex information needs. We showcase the retrieval quality of PubTator 3.0 using a series of entity pair queries, demonstrating that PubTator 3.0 retrieves a greater number of articles than either PubMed or Google Scholar, with higher precision in the top 20 results. We further show that integrating ChatGPT (GPT-4) with PubTator APIs dramatically improves the factuality and verifiability of its responses. In summary, PubTator 3.0 offers a comprehensive set of features and tools that allow researchers to navigate the ever-expanding wealth of biomedical literature, expediting research and unlocking valuable insights for scientific discovery.