• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    G

    Gemeinschaftskrankenhaus Havelhöhe

    EST. 1995
    297论文总数
    2,453引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Christian Grah
    Christian Grah
    Hospital Havelhoehe
    论文:69引用:0H-index:0
    M. Jecht
    M. Jecht
    Medizinische Klinik – Diabetologie, Gemeinschaftskrankenhaus Havelhöhe
    论文:46引用:0H-index:0
    Friedemann Schad
    Friedemann Schad
    Forschungsinstitut Havelhohe
    论文:36引用:0H-index:0
    Harald Matthes
    Harald Matthes
    Hospital Havelhoehe
    论文:24引用:0H-index:0
    Gerrit Grieb
    Gerrit Grieb
    Hospital Havelhoehe
    论文:17引用:0H-index:0
    Stephan Eggeling
    Stephan Eggeling
    Vivantes Klinikum Neukolln
    论文:15引用:0H-index:0
    Anja Thronicke
    Anja Thronicke
    Freie Universität Berlin, Humboldt-Universität zu Berlin
    论文:13引用:0H-index:0
    Joachim Pfannschmidt
    Joachim Pfannschmidt
    Chirurgische Abteilung der Thoraxklinik, Universitätsklinikum Heidelberg
    论文:12引用:0H-index:0
    Ralf Eberhardt
    Ralf Eberhardt
    Department of Respiratory Medicine and Intensive Care Medicine, the University of Heidelberg
    论文:10引用:0H-index:0

    论文(297)

    年份
    起
    –
    止
    排序
    1AI-based Burn Image Assessment: Reliability and Clinical Error Patterns of Multimodal Large Language Models in a Repeated-Inference Study.
    Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling

    Accurate assessment of burn depth and total body surface area (TBSA) is critical for clinical decision-making; however, it remains subjective and prone to interobserver variability. Multimodal large language models (MLLMs) are increasingly encountered in clinical contexts, but whether these systems can reliably assess burn images remains unclear. We evaluated four MLLMs (GPT-5.4 Pro, Grok 4.1, Gemini 3.1 Pro, and Claude Opus 4.6) on 50 clinical burn photographs using a repeated-inference design with five independent runs per model. Burn depth classification was assessed in numeric and text-based formats, alongside ordinal TBSA estimation. Performance varied across the models, with burn depth accuracy ranging from 34.0 ± 6.5% to 76.4 ± 6.8% and TBSA accuracy from 32.8 ± 9.4% to 68.4 ± 3.3%. Inter-run reliability (Fleiss' κ) ranged from slight (κ = 0.171) to almost perfect (κ = 0.916), demonstrating response variability not captured by single-query evaluations. Notably, no model combined high accuracy and high reliability, indicating a dissociation between performance and consistency. All models showed a tendency toward overestimation of burn depth, including assignment of fourth-degree burns despite their absence in the dataset. Error direction analysis revealed model-specific and task-dependent biases, including opposing patterns within the same model. Internal consistency between numeric and text classifications was near-perfect (99.6-100%), indicating format-invariant but systematically biased outputs. These findings demonstrate that MLLM performance is characterized by stochastic response instability invisible to single-query evaluations. Such inconsistency for identical inputs represents a fundamental limitation for workflows requiring consistent outputs across repeated evaluations.

    2026Journal of plastic, reconstructive & aesthetic surgery JPRAS(2026)
    引用
    AI阅读
    加入学术空间
    2Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models
    Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling

    Background: Burn extent guides triage, transfer and fluid resuscitation, yet its clinical estimation is imprecise and observer-dependent. Multimodal large language models (MLLMs) process clinical photographs without task-specific training, but their error has rarely been separated into systematic and random components or their performance across skin tones characterized. Methods: Three state-of-the-art MLLMs (Gemini 3.1 Pro, GPT-5.6 Sol, and Fable 5) each assessed 153 burn photographs five times under an identical prompt. The tasks were as follows: burned proportion of the imaged field, against an expert-guided pixel-wise segmentation (tolerance ± 10 percentage points, pp); burned percentage of total body surface area (TBSA), against physician consensus (±2 pp); and binary Fitzpatrick skin tone (FST; light I–III versus dark IV–VI). The first of these was the primary endpoint. Results: The primary endpoint was in the range of 32.5–70.2%, with TBSA at 69.5–77.5%. All models compressed the estimation range (slopes 0.58–0.76, intercepts +10.0 to +22.2 pp); one multiplicative constant per model brought errors differing more than twofold into a 1.4 pp range. Across repeated queries, the median within-image range was 5.0–25.0 pp; averaging the five answers reduced error by only 0.24–2.22 pp. FST accuracy was 83.8–91.9% against a majority-class baseline of 81.0%. Conclusions: Averaging repeated answers removes only the smaller, random component; the larger, systematic one persists and requires calibration against reference data before clinical use can be considered.

    2026Bioengineering(2026)
    引用
    AI阅读
    加入学术空间
    3Lungenkrebsscreening – Wie Kann Der „teachable Moment“ Wirksam Für Die Tabakentwöhnung Genutzt Werden?
    C Rustler, F Amawi, C Grah, P Lindinger
    2026Pneumologie 66 Kongress der Deutschen Gesellschaft für Pneumologie und Beatmungsmedizin e V(2026)
    引用
    AI阅读
    加入学术空间
    4Artificial Intelligence for Biomedical Diagnostics: Diagnostic Accuracy and Reliability of Multimodal Large Language Models in Electrocardiogram Interpretation
    Henrik Stelling, Armin Kraus, Gerrit Grieb, David Breidung, Ibrahim Güler

    The electrocardiogram (ECG) is a central tool in cardiovascular diagnostics, yet interpretation requires expertise and remains subject to variability. Multimodal large language models (MLLMs) have shown emerging capabilities in medical image analysis, but their performance in ECG interpretation remains insufficiently characterized. This study evaluated the diagnostic accuracy and inter-run reliability of five MLLMs across ECG interpretation tasks. Thirteen standard 12-lead ECGs were presented to five models (ChatGPT-5.3, Gemini 3.1 Pro, Claude Opus 4.6, Grok 4.1, and ERNIE 5.0) across five independent runs per case, yielding 2275 task-level assessments. Six categorical interpretation tasks (rhythm, electrical axis, PR/P-wave morphology, QRS duration, ST/T-wave morphology, and QTc interval) were compared with expert-consensus ground truth, while heart rate estimation was evaluated using mean absolute error (MAE). Overall categorical accuracy ranged from 52.3% to 64.9%. QRS duration classification achieved the highest accuracy (66.2-90.8%), whereas ST/T-wave assessment showed the lowest performance (20.0-41.5%). Heart rate MAE ranged from 14.8 to 46.7 bpm. A dissociation between diagnostic accuracy and inter-run reliability was observed across models. These findings indicate that current MLLMs do not achieve clinically reliable ECG interpretation performance and highlight the importance of assessing diagnostic accuracy and inter-run reliability when evaluating artificial intelligence systems in biomedical diagnostics.

    2026Life (Basel, Switzerland)(2026)
    引用
    AI阅读
    加入学术空间
    5Performance and Reliability of Large Language Models on the European Board of Hand Surgery Examination: a Multi-Model Evaluation Study.
    Ibrahim Güler, Lindsay Muir, Gerrit Grieb, Philipp Moog, Armin Kraus, Henrik Stelling

    INTRODUCTION:Artificial intelligence (AI) has demonstrated transformative potential in medical education and assessment, with large language models achieving competitive results across multiple high-stakes examinations. In this study, we evaluated the performance and inter-run reliability of 10 widely adopted large language models (LLMs) on the European Board of Hand Surgery written examination. METHODS:Ten LLMs were assessed on the complete 300-item European Board of Hand Surgery written examination using standardized zero-shot prompting. The models included five proprietary systems (GPT-5 Pro, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok-4 and ERNIE 4.5 Turbo) and five open-source architectures (DeepSeek V3.2, Qwen3 Max, Mistral Medium 3.1, Llama 3.3 and Falcon H1). Each LLM completed five independent runs, producing 15000 answers analysed for mean accuracy, 95% confidence intervals and inter-run reliability using Cohen's kappa (κ). RESULTS:Mean accuracy across the LLMs ranged from 72 to 85%, corresponding to total European Board of Hand Surgery scores between 131 and 211 points. Seven of the 10 LLMs reached or exceeded the illustrative pass threshold of 75%, equivalent to 150 of 300 points. Proprietary systems showed consistently higher mean accuracy than open-source systems. The highest-performing LLM (GPT-5 Pro) achieved 85% accuracy with a 95% confidence interval of 84 to 86% and a mean inter-run reliability measured by Cohen's κ of 0.739. The overall reliability across the LLMs was 0.821. CONCLUSIONS:Contemporary LLMs show robust and reproducible performance on a complex surgical certification examination, with proprietary architectures tending to outperform open-source counterparts. Although several models reached or exceeded an illustrative pass threshold, persistent gaps in subspecialty knowledge remain such as congenital anomalies and complex reconstructions. Therefore, LLMs may assist in structured learning and examination preparation but require specialist oversight and remain unsuitable for independent subspecialty decision-making. LEVEL OF EVIDENCE:Not applicable.

    2026The Journal of hand surgery, European volume(2026)
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 297 篇论文

    合作机构(100)

    柏林夏里特大学医学院合作论文 29
    Helios Klinikum Emil von Behring,Helios Kliniken合作论文 17
    德国海德堡大学合作论文 14
    Asklepios Fachkliniken München-Gauting合作论文 11
    马堡大学合作论文 11
    Lungenklinik Hemer合作论文 11
    Witten/Herdecke University合作论文 11
    DRK Kliniken Berlin合作论文 10
    University Hospital in Halle合作论文 8
    Evangelische Lungenklinik Berlin,Paul Gerhardt Diakonie合作论文 7

    机构统计