Background With the rising utilization of large language models (LLMs), such as OpenEvidence (OE), ChatGPT-4o (GPT), Google Gemini (GG), and DeepSeek (DS), their use in complex Orthopaedic cases remains unclear. Massive irreparable rotator cuff tears (MIRCTs) represent a challenging clinical scenario requiring nuanced, patient-specific decision-making. We evaluated the concordance of LLM-generated surgical recommendations with expert consensus derived from a Delphi study by the American Shoulder and Elbow Surgeons (ASES) Neer Circle and compared the various LLM models against each other. Methods Sixty-one MIRCT Delphi consensus scenarios were entered into the most current free version of each LLM in a standardized prompt format in June 2025. Recommendations were categorized as fully concordant or discordant with Delphi consensus. Further LLM testing included the addition of diabetes, smoking, and their combination. Accuracy (%) for each LLM and test with differences assessed using Cochran’s Q test and McNemar’s tests. A generalized linear mixed model identified significant predictors of AI-LLM accuracy, while the relationship between Delphi consensus strength and LLM concordance was assessed using Spearman's rank correlation coefficient. Results A total of 976 recommendations (61 scenarios × 4 platforms x 4 tests) were analyzed. OE demonstrated the greatest accuracy across all 4 tests (65.6%, 68.9%, 62.3%, and 65.6%, respectively) (p < 0.05), while DS consistently scoring the lowest. Accuracy was significantly positively predicted by age greater than 70 (OR 31.6, p < 0.001), dynamic instability (OR 18.8, p = 0.002), and pseudoparesis (OR 2.9, p = 0.025), and negatively predicted by intact or anatomically reparable subscapularis (OR 0.17, p = 0.001). There was a significant positive correlation between the strength of Delphi expert consensus and the number of LLM platforms concordant with that consensus (Spearman's ρ = 0.493, p < 0.001). Discussion & Conclusion This study is the first to systematically compare multiple LLMs against recommendations of an MIRCT Delphi consensus study. OE and GPT demonstrated the highest concordance; however, they did not approach levels of expert decision-making in many scenarios. LLMs have potential as adjunctive decision-support tools, particularly in resource-limited settings or for generalists managing complex shoulder pathology. Level of Evidence Basic Science Study, Computer Modeling using AI;
更多