Large language models (LLMs) are increasingly central to clinician workflows, spanning clinical decision support, medical education, and patient communication. However, current evaluation methods for medical LLMs rely heavily on static, templated benchmarks that fail to capture the complexity and dynamics of real-world clinical practice, creating a dissonance between benchmark performance and clinical utility. To address these limitations, we present MedArena, an interactive evaluation platform that enables clinicians to directly test and compare leading LLMs using their own medical queries. Given a clinician-provided query, MedArena presents responses from two randomly selected models and asks the user to select the preferred response. Out of 1571 preferences collected across 12 LLMs up to November 1, 2025, Gemini 2.0 Flash Thinking, Gemini 2.5 Pro, and GPT-4o were the top three models by Bradley-Terry rating. Only one-third of clinician-submitted questions resembled factual recall tasks (e.g., MedQA), whereas the majority addressed topics such as treatment selection, clinical documentation, or patient communication, with 20
Modern clinical practice increasingly depends on reasoning over heterogeneous, evolving, and incomplete patient data. Although recent advances in multimodal foundation models have improved performance on various clinical tasks, most existing models remain static, opaque, and poorly aligned with real-world clinical workflows. We present Cerebra, an interactive multi-agent AI team that coordinates specialized agents for EHR, clinical notes, and medical imaging analysis. These outputs are synthesized into a clinician-facing dashboard that combines visual analytics with a conversational interface, enabling clinicians to interrogate predictions and contextualize risk at the point of care. Cerebra supports privacy-preserving deployment by operating on structured representations and remains robust when modalities are incomplete. We evaluated Cerebra using a massive multi-institutional dataset spanning 3 million patients from four independent healthcare systems. Cerebra consistently outperformed both state-of-the-art single-modality models and large multimodal language model baselines. In dementia risk prediction, it achieved AUROCs up to 0.80, compared with 0.74 for the strongest single-modality model and 0.68 for language model baselines. For dementia diagnosis, it achieved an AUROC of 0.86, and for survival prediction, a C-index of 0.81. In a reader study with experienced physicians, Cerebra significantly improved expert performance, increasing accuracy by 17.5 percentage points in prospective dementia risk estimation. These results demonstrate Cerebra's potential for interpretable, robust decision support in clinical care.
Infectious disease threats to individual and public health are numerous, varied and frequently unexpected. Artificial intelligence (AI) and related technologies, which are already supporting human decision making in economics, medicine and social science, have the potential to transform the scope and power of infectious disease epidemiology. Here we consider the application to infectious disease modelling of AI systems that combine machine learning, computational statistics, information retrieval and data science. We first outline how recent advances in AI can accelerate breakthroughs in answering key epidemiological questions and we discuss specific AI methods that can be applied to routinely collected infectious disease surveillance data. Second, we elaborate on the social context of AI for infectious disease epidemiology, including issues such as explainability, safety, accountability and ethics. Finally, we summarize some limitations of AI applications in this field and provide recommendations for how infectious disease epidemiology can harness most effectively current and future developments in AI.
With all the advances in both the science of aging and artificial intelligence (AI), we are in a propitious position to accurately and precisely determine who is at high risk of developing Alzheimer’s disease years before signs of even mild cognitive deficit. It takes at least 20 years for aggregates of misfolded β-amyloid and tau proteins to accumulate in the brain along with neuroinflammation that they incite. This provides a long window of opportunity to get ahead of the pathobiological process, both for prediction and prevention.
Single-cell RNA-seq (scRNAseq) struggles to capture the cellular heterogeneity of transcripts within individual cells due to the prevalence of highly abundant and ubiquitous transcripts, which can obscure the detection of biologically distinct transcripts expressed up to several orders of magnitude lower levels. To address this challenge, here we introduce single-cell CRISPRclean (scCLEAN), a molecular method that globally recomposes scRNAseq libraries, providing a benefit that cannot be recapitulated with deeper sequencing. scCLEAN utilizes the programmability of CRISPR/Cas9 to target and remove less than 1% of the transcriptome while redistributing approximately half of reads, shifting the focus toward less abundant transcripts. We experimentally apply scCLEAN to both heterogeneous immune cells and homogenous vascular smooth muscle cells to demonstrate its ability to uncover biological signatures in different biological contexts. We further emphasize scCLEAN's versatility by applying it to a third-generation sequencing method, single-cell MAS-Seq, to increase transcript-level detection and discovery. Here we show the possible utility of scCLEAN across a wide array of human tissues and cell types, indicating which contexts this technology proves beneficial and those in which its application is not advisable.
LLMs are bound to transform healthcare with advanced decision support and flexible chat assistants. However, LLMs are prone to generate inaccurate medical content. To ground LLMs in high-quality medical knowledge, LLMs have been equipped with external knowledge via RAG, where unstructured medical knowledge is split into small text chunks that can be selectively retrieved and integrated into the LLMs context. Yet, existing RAG pipelines rely on raw, unstructured medical text, which can be noisy, uncurated and difficult for LLMs to effectively leverage. Systematic approaches to organize medical knowledge to best surface it to LLMs are generally lacking. To address these challenges, we introduce MIRIAD, a large-scale, curated corpus of 5,821,948 medical QA pairs, each rephrased from and grounded in a passage from peer-reviewed medical literature using a semi-automated pipeline combining LLM generation, filtering, grounding, and human annotation. Unlike prior medical corpora, which rely on unstructured text, MIRIAD encapsulates web-scale medical knowledge in an operationalized query-response format, which enables more targeted retrieval. Experiments on challenging medical QA benchmarks show that augmenting LLMs with MIRIAD improves accuracy up to 6.7 the same source corpus and with the same amount of retrieved text. Moreover, MIRIAD improved the ability of LLMs to detect medical hallucinations by 22.5 to 37 map of MIRIAD spanning 56 medical disciplines, enabling clinical users to visually explore, search, and refine medical knowledge. MIRIAD promises to unlock a wealth of down-stream applications, including medical information retrievers, enhanced RAG applications, and knowledge-grounded chat interfaces, which ultimately enables more reliable LLM applications in healthcare.
In 2021, a year before ChatGPT took the world by storm amid the excitement about generative artificial intelligence (AI), AlphaFold 2 cracked the 50-year-old protein-folding problem, predicting three-dimensional (3D) structures for more than 200 million proteins from their amino acid sequences. This accomplishment was a precursor to an unprecedented burgeoning of large language models (LLMs) in the life sciences. That was just the beginning. In recent months, we have moved into a hyperaccelerated phase of new foundation models, pretrained on massive datasets, with the ability to perform a wide range of tasks that are helping us understand the structure, biology, evolution, and design of proteins, RNA, DNA, and ligands, as well as their biomolecular interactions. Unlike multimodal LLMs such as GPT-4, Gemini, and Claude, which process text, audio, and images, these large language of life models (LLLMs) are multiomic. That is to say, they are not only multimodal but pertain to different layers of molecular biology. For example, Evo, a foundation model trained on 2.7 million diverse phage and prokaryotic genomes (equivalent to about 300 billion DNA nucleotides), predicts the impact of variants in DNA, RNA, or proteins on structure and function, as well as how essential genes are to cell function, and can generate new DNA sequences.
There is a path to living longer and healthier that doesn’t require reversing the aging process
Decentralized yet coordinated networks of specialized artificial intelligence agents, multi-agent systems for healthcare (MASH), that excel in performing tasks in an assistive or autonomous manner within specific clinical and operational domains are likely to become the next paradigm in medical artificial intelligence.
Rapid advancements in artificial intelligence (AI), particularly large language models (LLMs) and multimodal AI, are transforming medicine through enhancements in diagnostics, patient interaction, and medical forecasting. LLMs enable conversational interfaces, simplify medical reports, and assist clinicians with decision making. Multimodal AI integrates diverse data like images and genetic data for superior performance in pathology and medical screening. AI-driven tools promise proactive, personalized healthcare through continuous monitoring and multiscale forecasting. However, challenges like bias, privacy, regulatory hurdles, and integration into healthcare systems must be addressed for widespread clinical adoption.
Type 2 diabetes (T2D) is a multifaceted disease associated with several factors, including diet, genetics, exercise, sleep and gut microbiome. Current diagnostic and monitoring methods based on episodic assays like glycated hemoglobin (HbA1c) fail to capture its full complexity. Here, in a prospective cohort of 1,137 participants in the United States, we analyzed multimodal data from 347 deeply phenotyped individuals (174 normoglycemic, 79 prediabetic and 94 T2D). We found significant differences in the distribution of glucose spike metrics among different diabetes states, with longer expected time for spike resolution and higher values of nocturnal hypoglycemia in T2D. We identified significant correlations between mean glucose level and gut microbiome diversity, and between expected time for spike resolution and resting heart rate. Our multimodal glycemic risk profiles, validated in 1,955 normoglycemic and 114 prediabetic individuals from an independent cohort, improved risk stratification by highlighting substantial variability among individuals with the same value of HbA1c. Such a multimodal approach provides a detailed phenotype that can potentially improve T2D prevention, diagnosis and treatment, and is more informative than HbA1c.
Clinicians frequently use conditional reasoning for treatment decisions by envisioning potential outcomes for patients. This is counterfactual thinking, exploring "what if" scenarios. Developments in generative artificial intelligence (AI) enable us to simulate this patient-level reasoning at the data level, opening new opportunities for science and health care. We term this approach counterfactual AI.
AI in Precision OncologyAhead of Print Free AccessThe State of Artificial Intelligence in Precision Oncology: An Interview with Eric TopolEric J. Topol and Douglas FloraEric J. TopolScripps Research Translational Institute, La Jolla, California, USA.Search for more papers by this author and Douglas FloraSt. Elizabeth Healthcare, Edgewood, Kentucky, USA.Editor-in-Chief, AI in Precision Oncology.Search for more papers by this authorPublished Online:25 Jan 2024https://doi.org/10.1089/aipo.2024.29004.intAboutSectionsPDF/EPUB Permissions & CitationsDownload CitationsTrack CitationsAdd to favorites Back To Publication ShareShare onFacebookXLinked InRedditEmail Eric J. Topol, MDDouglas Flora, MDIntroductionEric Topol, MD, is a world-renowned cardiologist, best-selling author of several books on personalized medicine, and the founder and director of the Scripps Research Translational Institute in La Jolla, California. But he has also been the tip of the spear for the past 10 years or more as an advocate for using digital technologies, artificial intelligence (AI), and health care. Those arguments were introduced in his 2019 book Deep Medicine: How Artificial Intelligence can make Healthcare Human Again.1In December 2023, AI in Precision Oncology Chief Editor Doug Flora, MD, sat down with Topol in the opening keynote session of the journal's inaugural virtual summit, “The State of AI in Precision Oncology,” which broadcast on December 13, 2023, and is available to view on demand.This is a lightly edited and abbreviated transcript of that conversation.Eric, how did you get into AI and digital health?Eric Topol: It started in college. I was at the University of Virginia and really into genetics. I even wrote a thesis about prospects for gene therapy in humans—that was about 40 years ago! I got back to genomics when that became possible. In the 1990s, we started accruing huge data sets, and then digital biology became a possibility with sensors and smartphones connected to the internet, which was another dimension of big data. All of a sudden, there were all these data, we're all dressed up with nowhere to go without the proper analytics.AI is the third leg of the stool. It's a progression of needing ways to analyze immense data sets over time. As you know, cancer is a genomic disease, however, it's not just a genomics disease—we tend to oversimplify things. That's why having as many layers of data, whether digital or environmental or immunological or other important layers, that orthogonal data perspective is vital.My friend Ryan Langdale called this a “Cambrian moment,” where we have so much data because things have been digitalized and digitized. In November 2022, we had the release of the first models of generative AI and the acceleration/democratization of uptake. When did you start to see that happening? You certainly alluded to it in your book Deep Medicine.Topol: When I was writing Deep Medicine, I was talking to AI experts around the world. They told me there was no model yet such as our current models like ChatGPT (GPT-4), Gemini, and imminently GPT-5, but ultimately it was going to be available. My book came out in 2019, but it took almost 5 years to see the light with ChatGPT because the precursors to ChatGPT, even though this was incubating since 2017 with a classic preprint,2 which was never published as a regular article. It took years to get to this point of having massive graphics processing units, well over 25,000 for GPT-4.It was a natural progression in the background because it wasn't until November 30, 2022, when ChatGPT came out and 100 million people got onto it, and said, “Whoa, this is really something!” This is still just the beginning of where we're headed and it doesn't stop here. The ability now is to take multimodal data, whether it's slides and images, audio from visits with patients or bedside rounds, anything in text. We don't even have medically supervised training, specialized training, or fine tuning yet. We are dealing with base models and they are doing extremely well for medical questions and even medical diagnoses right now.What are the differences between the original models and even Siri, Cortana, Alexa, Google Digital Assistant, infiltrating our lives? You authored a nice article in Science last September3 about this transition to multimodal, writing that computers or machines don't have eyes—but they do! What is the revolution brought about by multimodal models such as Gemini and Bard expanding in the new GPT-4, which are more multimodal than the originals?Topol: The first phase of this kind of AI era was deep learning, deep neural networks, inputs that were largely annotated. Ground truth experts saying it's this or it's that. You can't do that at scale because there are not enough experts to do this labeling. And it costs a fortune to do hundreds of thousands if not millions of images. So we had to get past supervised learning to get to self-supervised, unsupervised, to let the data basically move ahead through artificial neurons, or artificial neural networks. It's training, but not through experts in ground truths now.Next was: how do we go from one task like an image to multiple tasks? That was the basis of the transformer model I mentioned beginning in 2017, because instead of going back and forth with each word in a sentence or a paragraph—a recurrent neural network, a type of deep neural network—it could do the whole thing. It had the context and soon enough, it wasn't just words, it was videos and images and of course speech. So that was the big change—to improve on old school deep neural networks around 2015. And by 2020 it had really been validated to go to the next big jump, which are these multimodal, self-supervised, or unsupervised. That's taken us with enormous computing requirements to where we are right now.Recently in Harvard Business Review, Ron Adner and James Weinstein wrote about the Napster model.4 We've got a log jam now and we haven't proved the return on investment. We know it is going to be there, but health care administrators are struggling with this. As they discussed the Napster model solved the question that the music companies had, which is how do we monetize this? How do we make this digital? It opened a brand-new ecosystem. That's where health care is going. If you're a health care administrator (like me), where would you suggest systems start? Because right now they're stuck between Siri and Skynet!Topol: Well, you're alluding to the fact there hasn't been much implementation of AI and health care to date. There are over 650 algorithms that are cleared or approved by the FDA. Most of those are deep learning, unimodal one-task algorithms; no transformer models or multimodal has been cleared or approved by FDA. There's no transparency—as a medical community, administrators or whoever's making decisions, we can't even review the data. Most of these haven't ever been published, and if they are published, they're proprietary and nontransparent. So, we have a problem.The other issue, of course, is things that may not need FDA clearance, which is a good thing. So, if I was an administrator right now, I want to undo the damage done with electronic health records (EHRs), where EHRs—if it didn't ruin the patient–doctor relationship and hurt doctors' and nurses' lives, it sure didn't do any good.Most clinicians hate data clerk work because it takes them away from what they really want to do—caring for patients—and it eats up hours when they're not seeing patients. We can move now without any FDA approval to a synthetic note. And it's not just the note from the conversation with a patient in clinic (adjusted for articulating the physical examination) because otherwise that would not happen during the visit. But just with that little addition, everything else is put into notes that are far better than what you see in Epic, Cerner, and others. Once you have that note digitized, it does all the other things such as preauthorization and billing, follow-up appointments, and prescriptions. It even nudges the patients—Did you check your blood pressure? What were the results? Did you go for the test that was ordered?It also coaches physicians to be more sensitive and empathetic reading the notes saying, why did you interrupt the patient after X seconds? Patients really like this because they have the audio recording and to clear up any confusion or things they forgot; they can go from the note to the link to the audio file. So this is the future and it's now and it's going to take over in the next couple of years. Administrators who want to make their clinician group happy and patients being able to see literally face-to-face their clinicians might want to think about trying these things out if they haven't already or wide-scale adoption as some health systems are already doing.Coding and billing, rev cycle management, rote administrative tasks, fighting with insurance companies—those could all be done with generative AI just a bit of training. Let's talk about upending cancer screening. Using age is truly an anachronism in many ways. How does that change the way that we think about screening cancers using polygenic risk scores and AI-driven algorithms?Topol: This is something I feel strongly about—we've got this all wrong. We're only picking up 12–14% of all cancers that are being diagnosed through mass screening. It's wasting tens of billions if not hundreds of billions of dollars every year. It's inducing a lot of anxiety for all the false positives as in mammography, but also other screening. And it's all based on age, which is so dumb. Now when you start to think about the fact that cancers occurring in younger people, much more commonly now, people in their 20s are coming across with colon cancer and women with breast cancer in their 30s. If you just use the current criteria, we're going to miss these people. So is there a better way?I am convinced there is. There are layers of data that would define the risk of each individual. We've already seen just for pancreatic cancer using data sets from Denmark and the U.S. Veterans health data set, that you can pick up cancer risk from the notes, laboratory tests, we wouldn't [otherwise] see the trends. Then you start adding unstructured text and polygenic risk scores, which are very inexpensive to obtain and we have for most common cancers. We can define risk and with AI picking up things in images that we can't see. If we start to reboot how we do cancer screening, I think we're going to get to a point where we can narrow down the field.For example, 88% of women will never develop breast cancer. Why do those women need to have mammography every year or two? Let's define risk. Let's not miss young people who are at risk for cancer. We have cell-free tumor DNA tests. There are many ways we can do this, but we can't be complacent about how we do screening now because it isn't working. It's wasteful. And the cancers that are being picked up waiting for some symptoms or scans to be so abnormal, they're often late. And we're not changing the natural history of the cancer. We've got to get better at that too.In the United States, we're missing 94% of patients who are appropriate for screening for lung cancer. Can we screen smarter, not harder? Can we use these machine learning tools to better identify specifically who is at highest risk? The recent large Dutch Pancreatic Cancer Screening trial used a transformer AI model to examine 28,000 pancreatic cancer patients—and outperformed existing models without even using factors such as BRCA or PALB2 or you didn't even have genomics in that trial and it still had an area under the curve of 0.88. That type of data for us clinicians is better than a CT scan or an MRI or a PET. We keep banging the drum. Cancer doctors like me are really tired of finding stage 4 cancers when the tools exist to fix half of them just without invention of a new test.Topol: There was an article last year from deCODE in Iceland in the New England Journal about mutations in cancer genes that if you knew about, could extend lifespan 7 years.5 The leading one was BRCA2. So if you just knew that you were BRCA2 [positive], just that alone to prevent not just breast and ovarian cancer, but all of men's cancers too—we're not taking advantage of what's published in the literature. All this great knowledge is compartmentalized in a different orbit and it isn't being offered to patients. And we've got to get out of that mode because we're losing people by not integrating knowledge into practice.There are emerging AI tools that identify areas that might warrant closer attention from an endoscopist on colonoscopies. You've written extensively about pattern doctors being outpaced by machine learning and by training these machine learning modules to identify things in their faintest footsteps. How do we address this for your medical students?Topol: What's interesting is that the gastroenterologists have led the field of AI doing randomized trials. There were recently 33 randomized trials from around the world, many from China, but now most places around the world have done randomized trials. The pickup of polyps is substantially better when machine vision real time is being used during the colonoscopy. Interestingly, there are also studies that as the day goes on, the gastroenterologist is more likely to miss those polyps. Now, we haven't seen an article yet that shows by picking up the significantly higher rate of adenomas as polyps that it changes the natural history of cancer. But that is pretty likely. We've already seen in 80,000 women randomized with mammography, with AI or without AI, that the AI helped tremendously in accuracy of diagnosis and reducing the time of review of scans.So we're seeing some great compelling evidence for the benefit of the patterns. I would extend that to pathology slides. It's amazing that from a whole slide image, you could get the driver mutations, the structural variations in play, whether it is a metastasis or the primary source of that tumor, and even the prognosis from a slide to a reasonable level of accuracy. We're not using that. We still are in the mode of pathologists who are not in agreement about what the H&E slide shows. We can do better with patterns of slides and of all types of medical images.It is augmented intelligence. We're not replacing radiologists. I still think they're going to be there for quality control and to make sure that we are appropriately interpreting with our whole brain the nuances of these films. You were instrumental in founding the Lerner College of Medicine, and you've devoted a great deal of your entire life's work is to medical education reaching the most people you can. How are we going to bridge that gap for students? They're obviously tech savvy, but you no longer need to memorize the 15 causes of pancreatitis. They have it in their fingertips. They can find it with a quick look with the new Ray Ban glasses. How are we going to train them to be critical thinkers and to ask the right questions, even if it's just designing prompts?Topol: That's a fantastic question because that makes us rethink not only how do we educate, but how do we select future physicians? We used to pick them still today by they have to be kind of brainiacs where really high MCAT scores and GPAs, and they won't even get past the threshold without that. What about their ability to communicate and empathize and connect with other people? I'm hoping that that will change and we'll emphasize that. You're going to have so much knowledge at your fingertips that memorizing everything is not the deal. It's about reasoning and building that ability to build trust, presence with your patients and know that you have their back.I think it's going to be an education that will have to include what is AI, how does it work, what are the liabilities? The fact that it can lose performance over time, that it requires surveillance, that it has issues about bias and inequities and lots of issues about AI that medical students have to understand because they're going to be using it in their daily lives. Once we get the right people in medical school and we got to train the old docs like us too, because most of us have not gotten up to speed. We can go to ChatGPT and play around, but that's just going to be medicine in the future. Everyone in medicine needs to understand the nuances. They don't have to know how to code. In fact, you can get GPT to code for you.What you do need to know is how do you do good prompts? How do you get the output that you really are looking for? Those are the things that we have to learn about as well as when to trust it. If I have any questions with a GPT-4 response, I'll do it over again. It's like double data entry. We have got to learn when we use it that we can trust it. A lot of things are made up now so you have to be savvy enough and require that authentication. We don't want to use something that's faulty, especially involving patient care.Allie Miller draws an analogy of not using existing AI tools in health care to be akin to toasting bread with a flashlight or mowing the lawn with scissors. The tools are here and they're very simple and user friendly. Where do we get clinicians started? Where do regular doctors and clinicians dip their toes into this? Where do they find the sandbox to start to become facile with the terms that you and I are sharing?Topol: I think everyone should be going to GPT-4 or when the other ones that are comparable or better, use them. I don't do Google searches for anything important anymore. But start to get facile. You'll start to understand better prompting as you use it more. The other thing, of course, is to get administrators to start to get these ambient conversation tools so they become the norm because that's another way to become facile. I think they are going to be one of the most transformative tools that we'll see for a long time in terms of having overriding benefit to the daily practice of medicine. Beyond that, there are some good resources. The book The AI Revolution in Medicine is a good resource.6We've got to change medicine for the better. I don't know anything else that has this potential. You mentioned the retina—that should be telling because not only do you get the ability to know about every eye disease from the retina photograph, but you also get a window to hepatobiliary disease, kidney disease, Parkinson's, and Alzheimer's 5–7 years before there are any symptoms. You can get the heart-risk calcium scores from the retinal vessels. As physicians, are we going to be doing smartphone, retinal photographs, and getting AI algorithmic interpretation?We already have a smart watch that can diagnose heart rhythms, but soon we'll be seeing things such as diagnosis of urinary tract infections through an AI kit you can get at a drugstore, skin lesions and cancers that will get you a preliminary diagnosis whether to even see a dermatologist or a physician, and ear infections in children. The list is growing quickly. We're seeing parents getting an AI stethoscope to monitor their children with asthma. All sorts of things are happening in the patient space. Just because you think AI hasn't changed everything yet? Don't think of it just for clinicians, it's also for patients.Flora: As we move into the future of medicine, where do you see this going in the next 2 years?Topol: Over the next couple of years, I'm hoping we'll start to see cancer screening get upended. We won't have it finalized, but at least some of the trials are ongoing to challenge the old way of doing cancer screening. We will get a diagnosis improved, whether it's because the accuracy of scan interpretation in the next couple of years or whether it's because each doctor through their health system practice has access to a GPT support that gives them a differential diagnosis of difficult diagnoses. If you only have 7 min for a routine visit, that's not enough to hear about a patient's concerns and to think. One of the things we must work on is not let the AI make things worse, not let more patients get squeezed into any daily schedule.That's a challenge because we've got many administrators, nonphysicians making the call as to, “oh well, you're more efficient now; let's get your schedule filled up even more.” These are things we have to confront in the next couple of years because this is a very big, if not the most extraordinary, transformation of medicine that we'll see in our lifetime. But we must plan ahead. There will be tools to summarize every aspect of a patient's data before you even start to look at their chart or see the patient and they will be ready in the next couple of years. It's very hard for us to keep up with the medical literature. Two years from now, don't worry, it'll keep up for you. It'll give you the daily skinny if you want on everything in your field because the corpus of medical literature for you is something that is right in the sweet spot of generative AI.Flora: Recently leaders from 27 different countries convened in Bletchley Park, the site of the origin of computation, to derive a rule set. That's the first time we've really seen formative guidelines. We need some guidelines to keep us on the Siri side and not the Skynet side. How do we do this? It can't be self-regulated because there will always be bad actors. It's moving so fast that I don't know if government can keep up. Is there a sweet spot?Topol: I hope so. This is happening so fast, the adaptation to get to that proper middle ground where you're not stifling innovation, but you're not fulfilling the doomsday prophecies, we don't want to see anything like that. The big debate is this artificial general intelligence and whether we're going to approach that, whether we're already there, depending on how you define it and what can happen when we get to that level. As a species, we don't want to acknowledge that machines can be smarter than us and when that happens, it's very threatening. But also we're caught. It's happened so fast, we haven't really come to terms. We saw what happened with powerful tools like genome editing, we had a scientist in China suddenly doing germline genome editing and he wound up in prison.We have to think about what could go wrong and try to anticipate and prevent these very worrisome untoward uses. This is not going to be easy. Even the regulatory agencies, like the FDA, haven't even reviewed one multimodal AI algorithm. So we're in the early days. We have to deal with embedded biases in our culture, lack of diverse populations in the AI that when you put something out like a pulse oximeter with AI, but you haven't tested it and people of color, we got a problem. There are all sorts of holes that have to be worked on. Ultimately this will get on the right track, but we're not there yet.Flora: AI is a means, not an end, the patient should be the end. I hope that the clinicians starting to publish this peer-reviewed data take charge of this and serve as the final arbiters in patient care and make sure that these tools are well vetted to have included data sets that include Black and Brown patients or Appalachian patients south of me in northern Kentucky. We have to make sure that that's the case. Eric, as you look forward, what are you most excited about?Topol: The ultimate thing is that we restore that patient–clinician bond that we've lost. Finishing medical school in the late 70s, I remember in my early years practicing medicine in the 80s compared with what it is now, and it's really eroded. I know we can get that back. We have the makings of the gift of time so that medicine becomes much more reflective rather than reflexive and that we are able to execute our charge, which is caring for patients and they know that they're being looked after.That's what I'm most excited about. It won't happen tomorrow, but we're sowing the seeds of that. We know that efficiency and productivity can be improved, support is there, and there's a rescue on the way. That is the goal for me. I can see the accuracy part being enhanced, but that's not enough. We have to have the overarching goal of this humanity/humaneness in medicine.Flora: The Harvard Business Review article mentioned earlier said that AI is a tide that raises all boats. The authors also made the point that the tide is already coming in. My hope is that with interviews like this, journals such as AI in Precision Oncology, NEJM AI, JCO Informatics—clinicians start to publish in a rigorous peer-reviewed form to prove that these tools have real merit. And I'd like the oncologists and the practitioners who care for the patients to participate in the development of these tools. I'd like us to engage clinicians to make sure that we are a part of the solution.Topol: Absolutely! We have to have compelling evidence of overriding extraordinary benefit that outweighs the risk, which justifies the change in practice. We don't have that for the most part. And that's where your journal and others can really make a difference demanding that level of evidence because that'll drive it. And if we don't have it, it will be held up unfortunately.References1. Topol EJ. Deep Medicine: How Artificial Intelligence can make Healthcare Human Again. Basic Books: New York; 2019. Google Scholar2. bioRxiv 2017. Google Scholar3. Topol EJ. As artificial intelligence goes multimodal, medical applications multiply. Science 2023;381(6663):adk6139; doi: 10.1126/science.adk6139 Crossref, Medline, Google Scholar4. Adner R, Weinstein JN. GenAI could transform how health care works. Harvard Business Review; 2023. Available from: https://hbr.org/2023/11/genai-could-transform-how-health-care-works Google Scholar5. Jensson BO, Arnadottir GA, Katrinadottir H, et al. Actionable genotypes and their association with life span in Iceland. New Engl J Med 2023;389:1741–1752; doi: 10.1056/NEJMoa2300792 Crossref, Medline, Google Scholar6. Lee P, Goldberg C, Kohane I. The AI Revolution in Medicine: GPT-4 and Beyond. Pearson; 2023. Google ScholarFiguresReferencesRelatedDetails Volume 0Issue 0 InformationCopyright 2024, Mary Ann Liebert, Inc., publishersTo cite this article:Eric J. Topol and Douglas Flora.The State of Artificial Intelligence in Precision Oncology: An Interview with Eric Topol.AI in Precision Oncology.ahead of printhttp://doi.org/10.1089/aipo.2024.29004.intOnline Ahead of Print:January 25, 2024PDF download