The Duke–NUS Medical School (Duke–NUS) is a graduate medical school in Singapore. The school was set up in April 2005 as the Duke–NUS Graduate Medical School, Singapore's second medical school, after the Yong Loo Lin School of Medicine, and before the Lee Kong Chian School of Medicine. It is a collaboration between Duke University from the United States and the National University of Singapore from Singapore. Duke-NUS follows the American model of post-baccalaureate medical education, in which students begin their medical studies after earning a bachelor’s degree. Students are awarded degrees from both Duke University School of Medicine and the National University of Singapore.
The field of regenerative medicine for Parkinson’s disease (PD) has reached a pivotal moment. After decades of preclinical research, recent first-in-human clinical trials demonstrated that cell replacement therapy using stem cell-derived dopaminergic neurons is not only feasible and safe but also shows promising signs of efficacy. Here we analyze three landmark 2025 studies, including the phase I/II trial of allogeneic induced pluripotent stem cell-derived dopaminergic progenitors, that mark a significant leap forward for PD therapy. We discuss principles underpinning the therapy, the historical context of fetal tissue transplants, findings from recent trials, and critical challenges. The convergence of robust cell manufacturing, precise stereotactic surgery, and advanced neuroimaging provides compelling evidence that stem cell-based therapies are potentially a viable treatment paradigm for PD.
The rapid growth of medical knowledge and the increasing complexity of clinical practice pose challenges. In this context, large language models (LLMs) demonstrate value; however, inherent limitations remain. Retrieval-augmented generation (RAG) shows potential to enhance their clinical applicability. This study reviews RAG applications in medicine. We find that research primarily relies on publicly available data, with limited use of private data. For retrieval, approaches commonly rely on English-centric embedding models, while LLMs are mostly generic, with limited use of medical-specific LLMs. For evaluation, automated metrics evaluate generation quality and task performance, whereas human evaluation focuses on accuracy, completeness, relevance, and fluency, with insufficient attention to bias and safety. RAG applications are concentrated on question answering, report generation, text summarization, and information extraction. Overall, medical RAG remains at an early stage, requiring advances in clinical validation, cross-linguistic adaptation, and support for low-resource settings to enable trustworthy and responsible global use.
Developments in large language models (LLMs) in the past 2 years have shifted the focus from text, image, and audio generation to LLMs capable of multistep reasoning (thinking). The development of LLMs is particularly important for medicine and health care, but the translation of these models has been limited by the black-box nature of previous LLMs. New reasoning-driven LLMs incorporate chain-of-thought prompting and reveal intermediate reasoning steps, offering transparency and traceability, potentially improving the clinical adoption and utility of LLMs. In this Viewpoint, we examine four emerging reasoning-driven LLMs, namely OpenAI's o1 and o3-mini, Google's Gemini 2.0 Flash Thinking, and DeepSeek R1. We compare their methodological approaches, benchmark their performance on medical question-answering tasks, and assess their potential for clinical integration. We highlight both opportunities and challenges associated with deploying reasoning-driven LLMs. Key future considerations include real-world validation, rigorous benchmarking with ethical safeguards, and advancements in improving the efficiency and sustainability of reasoning-driven LLMs. Addressing these challenges will enable the fine-tuning of these LLMs for specific medical applications, enhancing their potential clinical decision support, patient education, medical training, and evidence synthesis.
Generalist multimodal large language models (MLLMs) have achieved impressive performance across a wide range of vision-language tasks. However, their performance on medical tasks—particularly in zero-shot settings where generalization is critical—remains suboptimal. A key research gap is the limited understanding of why medical MLLMs underperform in medical image interpretation. **In this work**, we present a pioneering systematic investigation into the visual grounding capabilities of state-of-the-art medical MLLMs. To disentangle *visual grounding* from *semantic grounding*, we design VGMED, a novel evaluation dataset developed with expert clinical guidance, explicitly assessing the visual grounding capability of medical MLLMs. We introduce new quantitative metrics and conduct detailed qualitative analyses. Our study across **eight** state-of-the-art (SOTA) medical MLLMs validates that they often fail to ground their predictions in clinically relevant image regions. We note that this finding is specific to medical image analysis; in contrast, prior work has shown that MLLMs are capable of grounding their predictions in the correct image regions when applied to natural scene images. Motivated by these findings, we propose VGRefine, a simple yet effective inference-time method that refines attention distribution to improve visual grounding in medical settings. Our approach achieves SOTA performance across 6 diverse Med-VQA benchmarks (over 110K VQA samples from 8 imaging modalities) without requiring additional training or external expert models. Overall, our work, for the first time, systematically validates inadequate visual grounding as one of the key contributing factors for medical MLLMs' under-performance. Additional experiments are included in the Supp. Project Page: https://guimeng-leo-liu.github.io/Medical-MLLMs-Fail/
Despite continuous advances in medical technology, the global distribution of health care resources remains uneven. The development of large language models (LLMs) has transformed the landscape of medicine and holds promise for improving health care quality and expanding access to medical information globally. However, existing LLMs are primarily trained on high-resource languages, limiting their applicability in global medical scenarios. To address this gap, we constructed GlobMed, a large multilingual medical dataset, containing over 500,000 entries spanning 12 languages, including four low-resource languages. Building on this, we established GlobMed-Bench, which systematically assesses 56 state-of-the-art proprietary and open-weight LLMs across multiple multilingual medical tasks, revealing significant performance disparities across languages, particularly for low-resource languages. Additionally, we introduced GlobMed-LLMs, a suite of multilingual medical LLMs trained on GlobMed, with parameters ranging from 1.7B to 8B. GlobMed-LLMs achieved an average performance improvement of over 40