Abstract Background Continuing medical education (CME) is a legal and ethical obligation for physicians in Germany. The rapid rise of large language models (LLMs) such as ChatGPT, Gemini, Claude, and Grok raises concerns about the integrity of CME assessments, as LLMs can already pass German CME tests. Objective This study aims to determine whether the choice of document format (searchable PDF, protected PDF, raster PDF, or vector PDF) and LLM influences the ability of LLMs to solve CME test questions at rates exceeding the passing threshold specified for each CME module (typically 70%). Methods In a fully crossed within-subjects repeated-measures design, 18 expired CME articles from 3 major German publishers across 6 specialties will be converted into 3 cheating-impeding PDF formats and processed alongside the original PDF files by 4 current LLMs (GPT-5, Claude Sonnet 4, Grok-4, and Gemini 3). This results in 16 model-format combinations. Each model will answer every article 3 times per file-format condition, with outcomes derived from aggregated run-level results. The primary outcome is the proportion of correctly answered questions; the secondary outcome is the pass/fail rate. Results The study has been approved by the Witten/Herdecke University Ethics Committee (S-260/2025; dated August 10, 2025) and is preregistered at the Open Science Framework. The study is supported by internal departmental resources only, and no external funding was received. Because this protocol evaluates LLMs using expired CME materials, no human participants are being recruited. Data collection is planned to begin in June 2026 and is expected to last approximately 4 weeks. At the time of manuscript submission, no data have been collected or analyzed. Results are expected to be available after the completion of data collection and statistical analysis in 2026. The analyses will quantify performance differences across document formats; these findings may inform the feasibility of nonsearchable document formats as a temporary measure to reduce LLM-enabled cheating risks in CME contexts. Conclusions By quantifying how document format constrains LLM performance, this study aims to evaluate simple technical safeguards that may reduce artificial intelligence–assisted manipulation of CME tests and inform regulators and CME providers about how to balance assessment validity, accessibility, and responsible LLM integration into postgraduate medical education.
Background End-of-life decision-making is a clinically and ethically challenging aspect of intensive care medicine. Contemporary data on the ethical attitudes and legal knowledge of intensive care unit physicians in postgraduate training in Germany are scarce. We aimed to assess residents’ medical-ethical knowledge, moral attitudes, and end-of-life decision-making, identify demographic factors associated with these attitudes, and compare our findings with previously published national and international cohorts. Methods We conducted an anonymous, voluntary online survey among physicians in postgraduate training attending seven German intensive care medicine training courses between March 2024 and March 2025. A 71-item questionnaire assessed demographic characteristics, self-reported knowledge of the German legal framework on active euthanasia and advance directives, ethical attitudes toward end-of-life decision-making, stakeholder involvement, and responses to clinical case vignettes. Associations between demographic characteristics and questionnaire responses were explored using non-parametric statistical methods and a mass-screening approach with Benjamini-Hochberg adjustment for multiple testing. Results Of 1564 eligible participants, 1298 completed the survey at least partially and 1152 completed all questionnaire items. Respondents were predominantly early-career physicians (mean age 31.1 ± 4.45 years), with 81.5% still in postgraduate specialist training. Religious affiliation, degree of religious practice, and population size of workplace municipality were the demographic factors most consistently associated with attitudes toward withholding and withdrawing life-sustaining treatment and active euthanasia. Patient wishes were identified as the most important determinant of treatment limitation, while involvement of relatives, senior physicians, and nursing staff was widely endorsed. Although most respondents considered both withholding and withdrawing ethically acceptable, they continued to distinguish between the two in clinical case vignettes, showing a preference for withholding over withdrawing treatment. Conclusions Among German early-career physicians with ICU training exposure, end-of-life attitudes were strongly influenced by religion and sociodemographic factors and reflect a shift toward greater respect for patient autonomy compared with historic national data. Persistent uncertainty regarding withholding versus withdrawing life-sustaining treatment and the legal framework of end-of-life care underscores the need for enhanced ethics and legal education during postgraduate training.
Background The increasing development and spread of artificial and assistive intelligence is opening up new areas of application not only in applied medicine but also in related fields such as continuing medical education (CME), which is part of the mandatory training program for medical doctors in Germany. This study aimed to determine whether medical laypersons can successfully conduct training courses specifically for physicians with the help of a large language model (LLM) such as ChatGPT-4. This study aims to qualitatively and quantitatively investigate the impact of using artificial intelligence (AI; specifically ChatGPT) on the acquisition of credit points in German postgraduate medical education. Objective Using this approach, we wanted to test further possible applications of AI in the postgraduate medical education setting and obtain results for practical use. Depending on the results, the potential influence of LLMs such as ChatGPT-4 on CME will be discussed, for example, as part of a SWOT (strengths, weaknesses, opportunities, threats) analysis. Methods We designed a randomized controlled trial, in which adult high school students attempt to solve CME tests across six medical specialties in three study arms in total with 18 CME training courses per study arm under different interventional conditions with varying amounts of permitted use of ChatGPT-4. Sample size calculation was performed including guess probability (20% correct answers, SD=40%; confidence level of 1–α=.95/α=.05; test power of 1–β=.95; P<.05). The study was registered at open scientific framework. Results As of October 2024, the acquisition of data and students to participate in the trial is ongoing. Upon analysis of our acquired data, we predict our findings to be ready for publication as soon as early 2025. Conclusions We aim to prove that the advances in AI, especially LLMs such as ChatGPT-4 have considerable effects on medical laypersons’ ability to successfully pass CME tests. The implications that this holds on how the concept of continuous medical education requires reevaluation are yet to be contemplated. Trial Registration OSF Registries 10.17605/OSF.IO/MZNUF; https://osf.io/mznuf International Registered Report Identifier (IRRID) PRR1-10.2196/63887
Aktuell sind die meisten Prozesse und diagnostischen Algorithmen im Rahmen der gültigen deutschen Chest Pain Unit-Zertifizierungskriterien auf die Untersuchung ischämischer Ursachen von Brustschmerzen ausgerichtet. Nichtkardiale Ursachen sind unzureichend abgebildet. Andere Differenzialdiagnosen, die ebenfalls mit einer erhöhten Sterblichkeit einhergehen, werden seltener in Betracht gezogen. Da im Bereich der initialen Triage das Leitsymptom des akuten Thoraxschmerzes primär ursachenunabhängig das entscheidende Kriterium bildet, wird ein symptomgeleitetes und patientenzentriertes Vorgehen vorgeschlagen. Fokus dieses patientenzentrierten Algorithmus ist die rasche Unterscheidung zwischen potenziell akut lebensbedrohlichen kardialen und nichtkardialen sowie nicht akut lebensbedrohlichen Ursachen in Abhängigkeit der Risikokonstellation des akuten Thoraxschmerzes. Chest Pain Units sollten lokale Standards bereithalten, welche auch die Erkennung potenziell lebensbedrohlicher nichtischämischer oder extrakardialer Ursachen umfassen. In diesem Kontext sollte auch eine Sensibilisierung der Öffentlichkeit zum Thema „Thoraxschmerz“ erfolgen, und mittelfristig sollten Laienschulungen und -zertifizierungsprogramme etabliert werden.
Chest pain poses a diagnostic challenge in the emergency department and requires a thorough clinical assessment. The traditional distinction between “atypical” and “typical” chest pain carries the risk of not addressing nonischemic clinical pictures. The newly conceived subdivision into cardiac, possibly cardiac, and (probably) noncardiac causes of the presenting symptom complex addresses a much more interdisciplinary approach to a symptom-oriented diagnostic algorithm. The diagnostic structures of the chest pain units in Germany do not currently reflect this. An adaptation should therefore be considered.
BackgroundChatGPT is a 175-billion-parameter natural language processing model that is already involved in scientific content and publications. Its influence ranges from providing quick access to information on medical topics, assisting in generating medical and scientific articles and papers, performing medical data analyses, and even interpreting complex data sets. ObjectiveThe future role of ChatGPT remains uncertain and a matter of debate already shortly after its release. This review aimed to analyze the role of ChatGPT in the medical literature during the first 3 months after its release. MethodsWe performed a concise review of literature published in PubMed from December 1, 2022, to March 31, 2023. To find all publications related to ChatGPT or considering ChatGPT, the search term was kept simple (“ChatGPT” in AllFields). All publications available as full text in German or English were included. All accessible publications were evaluated according to specifications by the author team (eg, impact factor, publication modus, article type, publication speed, and type of ChatGPT integration or content). The conclusions of the articles were used for later SWOT (strengths, weaknesses, opportunities, and threats) analysis. All data were analyzed on a descriptive basis. ResultsOf 178 studies in total, 160 met the inclusion criteria and were evaluated. The average impact factor was 4.423 (range 0-96.216), and the average publication speed was 16 (range 0-83) days. Among the articles, there were 77 editorials (48,1%), 43 essays (26.9%), 21 studies (13.1%), 6 reviews (3.8%), 6 case reports (3.8%), 6 news (3.8%), and 1 meta-analysis (0.6%). Of those, 54.4% (n=87) were published as open access, with 5% (n=8) provided on preprint servers. Over 400 quotes with information on strengths, weaknesses, opportunities, and threats were detected. By far, most (n=142, 34.8%) were related to weaknesses. ChatGPT excels in its ability to express ideas clearly and formulate general contexts comprehensibly. It performs so well that even experts in the field have difficulty identifying abstracts generated by ChatGPT. However, the time-limited scope and the need for corrections by experts were mentioned as weaknesses and threats of ChatGPT. Opportunities include assistance in formulating medical issues for nonnative English speakers, as well as the possibility of timely participation in the development of such artificial intelligence tools since it is in its early stages and can therefore still be influenced. ConclusionsArtificial intelligence tools such as ChatGPT are already part of the medical publishing landscape. Despite their apparent opportunities, policies and guidelines must be implemented to ensure benefits in education, clinical practice, and research and protect against threats such as scientific misconduct, plagiarism, and inaccuracy.
Current guidelines emphasize the diagnostic value of non-cardiac or possibly cardiac chest pain. The goal of this analysis was to determine whether German chest pain units (CPUs) adequately address conditions with “atypical” chest pain in existing diagnostic structures. A total of 11,734 patients from the German CPU registry were included. The analyses included mode of admission, critical time intervals, diagnostic steps, and differential diagnoses. Patients with unspecified chest pain were younger, more often female, were less likely to have classic cardiovascular risk factors and tended to present more often as self-referrals. Patients with acute coronary syndrome (ACS) mostly had prehospital medical contact. Overall, there was no difference between these two groups regarding the time from the onset of first symptoms to arrival at the CPU. In the CPU, the usual basic diagnostic measures were performed irrespective of ACS as the primary working diagnosis. In the non-ACS group, further ischemia-specific diagnostics were rarely performed. Extra-cardiac differential diagnoses were not specified. The establishment of broader awareness programs and opening CPUs for low-threshold evaluation of self-referring patients should be discussed. Regarding the rigid focus on the clarification of cardiac causes of chest pain, a stronger interdisciplinary approach should be promoted.
Summary Background ChatGPT (Chat Generative Pre-trained Transformer) has initiated widespread conversation across various human sciences. We here performed a concise review combined with a SWOT (strengths, weaknesses, opportunities, threats) analysis on ChatGPT potentials in natural science including medicine. Methods This is a concise review of literature published in PUBMED from 01.12.2022 to 31.03.2023. The only search term used was “ChatGPT”. Publications metrics (author, journal, and subdisciplines thereof) as well as findings of the SWOT analysis are presented. Findings Of 178 studies in total, 160 could be evaluated. The average impact factor was 4,423 (0 – 96,216), average publication speed was 16 days (0-83 days). Of all articles, there were 77 editorials, 43 essays, 21 studies, six reviews, six case reports, six news, and one meta-analyses. Strengths of ChatGPT include well-formulated expression as well as the ability to formulate general contexts flawlessly and comprehensibly, whereas the time-limited scope as well as the need for correction by experts were identified as weaknesses and threats. Opportunities include assistance in formulating medical issues for non-native speakers as well as the chance to be involved in the development of such AI in a timely manner. Interpretation Artificial intelligences such as ChatGPT will revolutionize more than just the medical publishing landscape. One of the biggest dangers in this is uncontrolled use, so we would do well to establish control and security measures at an early stage. Research in context Evidence before this study Since its release in 11/ 2022, only a few randomized controlled trials using ChatGPT have been published. To date, the majority of data stems from short notes or communication. Given the enormous interest (and also potential for misuse), we conducted a PUBMED literature search to create the most comprehensive evidence base currently available. We searched PUBMED for publications including the quote “ChatGPT” in English or German from 01.12.2022 until 31.03.2023. In order not risk any bias of evidence all related publications were screened initially. Added value of this study This is the most concise review for ChatGPT up to date. By means of a SWOT analysis, readers and researchers gain comprehensive insight to strengths, weaknesses, opportunities and threats of ChatGPT especially in the context of medical literature. Implications of all the available evidence Our review may well serve as origin for further research related to the topic in order to create more evidence, strict regulations and policies in dealing with ChatGPT.
Background Medical emergencies are complex and stressful, especially for the young and inexperienced. Cognitive aids (CA) have been shown to facilitate management of simulated medical emergencies by experienced teams. In this randomized trial we evaluated guideline adherence and treatment efficacy in simulated medical emergencies managed by residents with and without CA. Methods Physicians attending educational courses executed simulated medical emergencies. Teams were randomly assigned to manage emergencies with or without CA. Primary outcome was risk reduction of essential working steps. Secondary outcomes included prior experience in emergency medicine and CA, perceptions of usefulness, clinical relevance, acceptability, and accuracy in CA selection. Participants were grouped as “medical” (internal medicine and neurology) and “perioperative” (anesthesia and surgery) regarding their specialty. The study was designed as a prospective randomized single-blind study that was approved by the ethical committee of the University Duisburg-Essen (19-8966-BO). Trial registration: DRKS, DRKS00024781. Registered 16 March 2021—Retrospectively registered, http://www.drks.de/DRKS00024781 . Results Eighty teams participated in 240 simulated medical emergencies. Cognitive aid usage led to 9% absolute and 15% relative risk reduction. Per protocol analysis showed 17% absolute and 28% relative risk reduction. Wrong CA were used in 4%. Cognitive aids were judged as helpful by 94% of the participants. Teams performed significantly better when emergency CA were available (p < 0.05 for successful completion of critical work steps). Stress reduction using CA was more likely in “medical” than in “perioperative” subspecialties (3.7 ± 1.2 vs. 2.9 ± 1.2, p < 0.05). Conclusions In a high-fidelity simulation study, CA usage was associated with significant reduction of incorrect working steps in medical emergencies management and was characterized by high acceptance. These findings suggest that CA for medical emergencies may have the potential to improve emergency care.
Background Little is known about importance and implementation of end-of-life care (EOLC) in German intensive care units (ICU). This survey analyses preferences and differences in training between “medical” (internal medicine, neurology) and “surgical” (surgery, anaesthesiology) residents during intensive care rotation. Methods This is a point-prevalence study, in which intensive care medicine course participants of one educational course were surveyed. Physicians from multiple ICU and university as well as non-university hospitals and all care levels were asked to participate. The questionnaire was composed of a paper and an electronic part. Demographic and structural data were prompted and EOLC data (48 questions) were grouped into six categories considering importance and implementation: category 1 (important, always implemented), 2 (important, sometimes implemented), 3 (important, never implemented) and 4–6 (unimportant, implementation always, sometimes, never). The trial is registered at the “Deutsches Register für klinische Studien (DRKS)”, Study number DRKS00026619, registered on September 10th 2021, www.drks.de . Results Overall, 194/ 220 (88%) participants responded. Mean age was 29.7 years, 55% were female and 60% had scant ICU working experience. There were 64% medical and 35% surgical residents. Level of care and size of ICU differed significantly between medical and surgical (both p < 0.001). Sufficient implementation was stated for 66% of EOLC questions, room for improvement (category 2 and 3) was seen in 25, and 8% were classified as irrelevant (category 6). Areas with the most potential for improvement included prognosis and outcome and patient autonomy. There were no significant differences between medical and surgical residents. Conclusions Even though EOLC is predominantly regarded as sufficiently implemented in German ICU of all specialties, our survey unveiled still 25% room for improvement for medical as well as surgical ICU residents. This is important, as areas of improvement potential may be addressed with reasonable effort, like individualizing EOLC procedures or setting up EOLC teams. Health care providers as well as medical societies should emphasize EOLC training in their curricula.
Early heart attack awareness programs are thought to increase efficacy of chest pain units (CPU) by providing live-saving information to the community. We hypothesized that self-referral might be a feasible alternative to activation of emergency medical services (EMS) in selected chest pain patients with a specific low-risk profile. In this observational registry-based study, data from 4743 CPU patients were analyzed for differences between those with or without severe or fatal prehospital or in-unit events (out-of-hospital cardiac arrest and/or in-unit death, resuscitation or ventricular tachycardia). In order to identify a low-risk subset in which early self-referral might be recommended to reduce prehospital critical time intervals, the Global Registry of Acute Coronary Events (GRACE) score for in-hospital mortality and a specific low-risk CPU score developed from the data by multivariate regression analysis were applied and corresponding event rates were calculated. Male gender, cardiac symptoms other than chest pain, first onset of symptoms and a history of myocardial infarction, heart failure or cardioverter defibrillator implantation increased propensity for critical events. Event rates within the low-risk subsets varied from 0.5–2.8%. Those patients with preinfarction angina experienced fewer events. When educating patients and the general population about angina pectoris symptoms and early admission, activation of EMS remains recommended. Even in patients without any CPU-specific risk factor, self-referral bears the risk of severe or fatal pre- or in-unit events of 0.6%. However, admission should not be delayed, and self-referral might be feasible in patients with previous symptoms of preinfarction angina.
Background We aimed to analyze the 2020 standard of care in certified German chest pain units (CPU) with a special focus on non-ST-segment elevation acute coronary syndrome (NSTE-ACS) through a voluntary survey obtained from all certified units, using a prespecified questionnaire. Methods The assessment included the collection of information on diagnostic protocols, risk assessment, management and treatment strategies in suspected NSTE-ACS, the timing of invasive therapy in non-ST-segment elevation myocardial infarction (NSTEMI), and the choice of antiplatelet therapy. Results The response rate was 75%. Among all CPUs, 77% are currently using the European Society of Cardiology (ESC) 0/3‑h high-sensitive troponin protocol, and only 20% use the ESC 0/1‑h high-sensitive troponin protocol as a default strategy. Conventional ergometry is still the commonly performed stress test with a utilization rate of 47%. Among NSTEMI patients, coronary angiography is planned within 24 h in 96% of all CPUs, irrespective of the day of the week. Prasugrel is the P2Y12 inhibitor of choice in ST-segment elevation myocardial infarction (STEMI), but despite the impact of the ISAR-REACT 5 trial on selection of antiplatelet therapy, ticagrelor is still favored over prasugrel in NSTE-ACS. If triple therapy is used in NSTE-ACS with atrial fibrillation, it is maintained up to 4 weeks in 51% of these patients. Conclusion This survey provides evidence that Germany’s certified CPUs ensure a high level of guideline adherence and quality of care. The survey also identified areas in need of improvement such as the high utilization rate of stress electrocardiogram (ECG).
Introduction: Since 2008, specialized chest pain units (CPUs) were implemented across Germany ensuring structured diagnostics in acute chest pain. This study aims to analyze the management of pulmonary embolism (PE) patients in such certified CPUs. Methods: Data were retrieved from 13,902 patients enrolled in the German CPU registry and analyzed for the diagnosis of PE including patient characteristics, critical time intervals, diagnostic workup, treatment, and prognosis. PE patients were compared to the overall CPU patient cohort. Only patients with a complete 3-month follow-up were included. Results: Overall, 1.1% of all CPU patients were diagnosed with PE. Chest pain and dyspnea were the leading symptoms. Patients with PE were older, presented with higher heart rates, and more frequently exhibited signs of heart failure, despite a normal left ventricular function. PE patients showed significantly longer time delays between symptom onset and the first medical contact, while PE patients with chest pain presented earlier than PE patients with dyspnea only. Whereas more PE patients had to be transferred to the intensive care unit, in-CPU mortality and event rates over 3 months were low. Discussion/Conclusion: This study suggests a certain risk for underdiagnosis and consecutive potential undertreatment of PE patients in German Cardiac Society (GCS)-certified CPUs, which is thought to result from an anticipated focus on patients with acute coronary syndrome (ACS). Public awareness for PE beyond chest pain should be improved. Certified CPUs should be urged to implement strategic pathways for a better simultaneous diagnostic workup of differential diagnosis beyond ACS.
Sowohl Chest Pain Units (CPU) als auch Stroke Units (SU) haben sich als essenzielle Komponenten der klinischen Akutversorgung etabliert. Für beide Instanzen existieren Zertifizierungsverfahren. Bis Mitte 2020 sind 290 CPU und 335 SU zertifiziert. In der aktuellen Übersicht sollen die Strukturen und der aktuelle Zertifizierungsstand beleuchtet werden. Dabei soll die jüngere CPU-Zertifizierungs-Initiative dem langjährig etablierten SU-Konzept gegenübergestellt werden. Der Vergleich erstreckt sich auf die historischen Hintergründe, den Zertifizierungsprozess, die Qualitätserfassung, mögliche additive Strukturen, den aktuellen Zertifizierungsstand, die Übertragbarkeit der Konzepte auf die europäische Ebene sowie die wirtschaftlichen Faktoren. Beide Zertifizierungskonzepte weisen deutliche Parallelen auf. Für die SU gibt es eine positive Cochrane-Analyse, für die CPU bestehen zahlreiche positive Registerarbeiten. Wesentliche Unterschiede betreffen ein einheitliches CPU-System gegenüber einem abgestuften SU-Konzept. Darüber hinaus existieren für SU obligate Elemente der Qualitätsdokumentation. In ökonomischer Hinsicht gewährleisten OPS(Operationen- und Prozedurenschlüssel)-Ziffern eine bessere Abbildung des Ressourceneinsatzes in der Schlaganfallkomplextherapie, was für die CPU bislang nicht etabliert werden konnte. Das gut etablierte CPU-Konzept könnte von einer übergeordneten Qualitätskontrolle zusätzlich profitieren. Adäquates Benchmarking ist Grundvoraussetzung zur Potenzialanalyse sowie zur Schaffung einer gesonderten Vergütungsstruktur. Hier ist die Deutsche Gesellschaft für Kardiologie als zertifizierende Institution gefordert, im Rahmen der regelmäßigen Kriterienupdates einen entsprechenden Mechanismus zu etablieren.