
The deltoid ligament complex (DLC) is well-known as an ankle joint stabilizer; however, through the superficial DLC, it also acts at the subtalar joint and helps to maintain the medial longitudinal arch. DLC repair in the presence of an ankle fracture remains a topic of controversy in the current literature. There is a paucity within the literature surrounding clinical tests that could be performed to assess for hindfoot instability in patients with DLC injuries associated with ankle fracture. However, hindfoot instability has been observed intraoperatively under fluoroscopy via valgus stressing of the calcaneus, while maintaining a congruent medial clear space. In cases of hindfoot instability, the calcaneal tuberosity crosses the lateral fibular line. We aim to validate this technique as a method for detecting hindfoot instability radiographically with DLC injury when undertaking ankle fracture fixation. For this, 10 morphologically normal cadaveric leg specimens were used to validate the hindfoot instability test. Specimens were radiographed before and after transecting both the superficial and deep fibers of the DLC. A line was drawn along the lateral aspect of the fibula on the radiographs. The position of the calcaneal tuberosity in relation to this line was used to determine if the specimen had hindfoot instability. Pre-transecting the DLC, the calcaneal tuberosity did not cross the lateral fibular line on hindfoot valgus stressing. Post-transecting the DLC, the calcaneal tuberosity crossed the lateral fibular line in all specimens on repeat valgus stressing of the hindfoot. Post-DLC repair, the calcaneal tuberosity did not cross the lateral fibular line on hindfoot valgus stressing. We demonstrated that the hindfoot instability test may provide a useful intraoperative adjunct to determine hindfoot instability in the presence of a DLC injury and can be used to guide repair of the DLC in cases of ankle fractures.
This study aimed to provide a detailed anatomical mapping of proper palmar digital artery perforators, focusing on their number, distribution, and spatial relationship to the proximal (PIP) and distal interphalangeal (DIP) joints. The goal was to establish a reliable anatomical reference to support accurate, anatomically guided flap planning in fingers and fingertip reconstruction. A cadaveric study was conducted on 20 digits from five fresh-frozen upper limbs (three left, two right), excluding thumbs. Methylene blue dye was injected to visualize perforator arteries, and dissections were performed in an accredited anatomy facility. Images were processed using Adobe Photoshop, and data were analyzed in SPSS with descriptive and nonparametric tests (Mann-Whitney U, Kruskal-Wallis), with significance set at p < 0.05. A total of 441 phalangeal sites were assessed, of which 406 (92.1%) contained perforators. The proximal phalanx (P1) showed the highest concentration (69.2%), followed by the middle (P2; 22.9%) and distal (P3; 7.9%) phalanges. No significant side dominance (radial vs. ulnar) was observed. Perforator clustering was consistently found approximately 8 mm from the DIP joint and 22 mm from the PIP joint. This study highlights the anatomical consistency and clinical relevance of proper palmar digital artery perforators, providing a practical reference for flap planning and supporting safer reconstructive strategies.
Anatomy remains among the most resilient foundations of medicine, yet the professional who applies it most substantively to patient care has no defined professional identity, and the designation clinical anatomist is presently adopted by anyone who wishes to claim it. We argue that this absence is an opportunity rather than a deficiency. Drawing together two literatures that are rarely read side by side, professional identity formation and socialization in medical education and the sociology of professions concerned with expert jurisdiction, we first establish what a professional identity is and then distinguish it from competency: a competency is deployed, whereas an identity is embodied. That distinction resolves an apparent tension between a shared anatomical competency and a distinct subspecialty identity, a principle we term shared competency, asymmetric identity. It also separates three categories the field has understandably spoken of as one: the anatomist (a discipline), the anatomy educator (a role), and the clinical anatomist (an identity). We propose daily, practice-based collaboration directed at patient care as a working criterion for the identity, and the clinical anatomy fellowship as the apparatus through which the identity is formed and reproduced. The definition is offered as a starting point for debate and empirical testing rather than as a settled account.
The external ventricular drainage is a routine neurosurgical procedure, most often performed by cannulating the frontal horn of the lateral ventricle. Various entry points and insertion trajectory definitions have been proposed; nevertheless, it remains unclear which approach yields the highest success rate. Furthermore, in contemporary clinical practice, generally more efficient navigated ventriculostomies are available. This study aims to compare success rates between freehand and navigated procedures and to compare various trajectory definitions. A search of the Cochrane, Embase, PubMed, and Scopus databases was conducted. Inclusion criteria were: (1) studies presenting an unequivocal definition of entrypoint and insertion trajectory; (2) studies reporting a statistically quantifiable outcome. A total of 30 articles met the inclusion criteria. The pooled success rate was 69.7% and 85.1% for freehand and navigated procedures, respectively (3528 vs. 464 cases, p < 0.001). The highest success rate (97.8%, CI 0.949-1.000) was observed for a trajectory defined by entry 10-20 mm anterior to the coronal suture and 20-30 mm lateral to the midline, with catheter direction toward the ipsilateral tragus in the sagittal plane and the ipsilateral medial canthus in the coronal plane. Navigated external ventricular drainage demonstrates significantly higher success rates compared to the freehand technique. Therefore, navigated ventriculostomy should be regarded as the preferred standard of care. In settings where time or resources are constrained, the freehand approach remains a justifiable alternative. Due to missing direct comparison, no single trajectory definition for nonnavigated insertion could be recommended over others.
Visual identification of anatomical structures is a foundational skill in gross anatomy education. Whether multimodal large language models (LLMs) can reliably perform this task, and whether prompt engineering can meaningfully improve their accuracy, remains insufficiently investigated This cross-sectional comparative study evaluated four leading multimodal LLMs (GPT-5.2, Claude Sonnet 4.6, Gemini 3, and Grok 4) on their ability to identify anatomical structures across cadaveric dissection images. A total of 75 expert-validated anatomical structures spanning five body regions (abdomen, head and neck, lower limb, thorax, and upper limb) were evaluated using three prompt strategies: zero-shot (P0), few-shot (PF), and chain-of-thought (PCoT). Each prompt strategy was administered across three independent sessions conducted in April-May 2026, yielding 675 binary-scored responses per model (2700 total responses). Gemini 3 achieved the highest overall accuracy (67.3%), followed by GPT-5.2 (41.9%), Claude Sonnet 4.6 (27.3%), and Grok 4 (23.0%)-an ordering that inverts the hierarchy typically observed in text-based anatomy assessments, where GPT-4o has generally led, and Gemini has ranked lower. Gemini 3 significantly outperformed all other models (all Bonferroni-corrected p < 0.001). Prompt strategy had a statistically significant effect for Gemini 3 (PCoT > PF, p = 0.015) and Grok 4 (PCoT > P0, p = 0.023); no significant prompt effect was observed for GPT-5.2 or Claude Sonnet 4.6. Performance varied substantially by anatomical region: the abdomen consistently yielded the highest accuracy across all models, while the lower limb yielded the lowest. Inter-trial consistency was high (70.7%-89.3%) but dissociated from accuracy, as some models produced reproducibly incorrect responses. Current multimodal LLMs, accessed via consumer interfaces, are insufficient for reliable standalone identification of cadaveric anatomical structures. Gemini 3, when used with chain-of-thought prompting, may serve as a supplementary aid in select anatomical regions; however, critical educator supervision and verification against authoritative anatomical resources remain essential before any clinical or educational deployment.
This study, conducted among first-year undergraduates, investigates the feasibility of integrating professional-values-oriented education into the discussion-based human anatomy teaching module via the problem-based learning (PBL) model. Over a three-year reform period, PBL cases combining clinical scenarios with professional-values-oriented elements were designed and implemented specifically for this human anatomy module; the effectiveness of this approach was subsequently evaluated through a systematic assessment. The results indicate a notably high level of student acceptance (98.7%) and satisfaction (96.9%) with this instructional model. Furthermore, over 92% of students reported enhanced performance in three key areas: knowledge acquisition, skill development, and moral awareness. Overall course performance was evaluated as excellent. Integrating professional-values-oriented education into PBL anatomy discussion sessions demonstrated significant potential for promoting the concurrent development of professional knowledge and value-based literacy among early-stage undergraduates.
Artificial intelligence is increasingly being used in education, including the field of anatomy, where it could support learning. The aim of this study was to evaluate and compare the performance of two large language models (LLMs), AnatomyGPT 5.2 and ChatGPT 5.2, in answering anatomy-based questions. A standardized prompt was prepared, and anatomical questions, originally used in official anatomy tests, were posed separately to AnatomyGPT 5.2 and ChatGPT 5.2 in two independent attempts. A total of 550 questions were used, including 150 in Polish and 400 in English. ChatGPT 5.2 outperformed AnatomyGPT 5.2 in both trials, achieving accuracies of 84.36% and 86.36%, compared with 82.73% and 85.27%, respectively; however, the differences between the models were not statistically significant in either trial. In both trials, ChatGPT 5.2 and AnatomyGPT 5.2 performed better on the English question set than on the Polish question set, with the differences between the question sets being statistically significant (p < 0.05). Questions on innervation and vascularization were most frequently answered correctly by both models, while multiple-choice questions were least frequently correct; no statistically significant differences were found between models or trials across all categories of questions (p > 0.05). Cohen's kappa analysis indicated substantial agreement between repeated responses for both models across the two trials (p < 0.001). The intraclass correlation coefficient was 0.62 for AnatomyGPT 5.2 and 0.66 for ChatGPT 5.2, indicating moderate agreement between the original and repeated responses for both models. Overall, the studied models demonstrated generally high performance across most questions and relatively consistent responses over time. Higher accuracy was observed for the English question set than for the Polish question set; however, this difference cannot be attributed solely to language.
Assessment frameworks in anatomy education should provide valid, reliable, and equitable measures of student learning. This study evaluated the validity evidence and equity of a newly developed multi-component assessment framework within a Developmental Anatomy elective for third-year medical students. A cross-sectional study was conducted among 43 third-year medical students enrolled in a Developmental Anatomy elective at King Faisal University, Saudi Arabia. The assessment framework comprised three quizzes, an image-based assignment, group presentations, and a final written examination. Validity evidence was evaluated using Messick's unified validity framework. Reliability analyses included Kuder-Richardson Formula 20 (KR-20) and Cronbach's alpha. Student perceptions were assessed through a validated questionnaire and qualitative feedback. The student perception questionnaire demonstrated excellent content validity (S-CVI/Ave = 1.00) and reliability (Cronbach's α = 0.985). The combined quiz component showed good reliability (KR-20 = 0.81), while the assignment demonstrated acceptable reliability (Cronbach's α = 0.68). Most quiz items had acceptable discrimination (mean = 0.39) despite high item difficulty (mean = 0.93). The strongest correlation was observed between summative assessment and final examination performance (r = 0.92). Students reported positive perceptions (mean = 4.06/5), particularly regarding fairness. No significant differences were observed between male and female sections across any component (all p > 0.05). The multi-component assessment framework demonstrated satisfactory validity evidence, reliability, and equity within a Developmental Anatomy elective. Integrating formative and summative methods provided a meaningful approach to evaluating student learning. Similar frameworks may be valuable in anatomy electives and other medical education settings.
The present study aimed to evaluate the developmental morphometry of the paranasal sinuses using computed tomography (CT) in subjects aged 1-25 years and to characterize age-related changes in sinus dimensions, volume, and surface area throughout childhood, adolescence, and young adulthood. This retrospective CT-based study included 250 subjects (125 females and 125 males) without radiologic evidence of sinonasal pathology. Bilateral morphometric analyses of the frontal, sphenoid, maxillary, and ethmoid sinuses were performed using multiplanar CT images and three-dimensional reconstructions. Measurements included anteroposterior, mediolateral, and superoinferior diameters, sinus volume, and sinus surface area. Significant age-related differences were identified for all morphometric parameters of the frontal, sphenoid, maxillary, and ethmoid sinuses (all p < 0.05). The frontal sinus demonstrated the latest developmental onset, with no measurable pneumatization during infancy and first detectable development at 5 years of age. In contrast, the maxillary and ethmoid sinuses were identifiable in most subjects during infancy, whereas sphenoid sinus pneumatization was identifiable at 2 years of age. Progressive enlargement of sinus dimensions, volume, and surface area was observed throughout childhood and adolescence, with relative stabilization of most morphometric measurements between the postpubescent and young adult groups. Paranasal sinus morphology demonstrates substantial age-related developmental variation from infancy to young adulthood. These findings provide comprehensive morphometric reference data regarding paranasal sinus maturation and may contribute to improved radiologic interpretation, developmental assessment, and safer surgical planning in pediatric and young adult populations.
Donor-based dissection is a central component of anatomy education and is often considered a context for fostering ethical reflection, bioethics learning, and professionalism. This study examined first-year medical students' perceptions of these dimensions and explored factors associated with their evaluation of the educational value of dissection. A cross-sectional observational study was conducted using a structured questionnaire administered in two phases during a Human Anatomy course. Complete data for the primary outcome variables were obtained from 86 students. Associations were analyzed using Spearman's rank correlation coefficient with false discovery rate correction, and correlations with |ρ| ≥ 0.30 were considered practically relevant. A limited number of statistically significant associations were identified. Enthusiasm toward dissection was positively associated with perceived usefulness for bioethics learning (ρ = 0.34) and professionalism (ρ = 0.35). Perceived usefulness for professionalism was also associated with intention to pursue a surgical career (ρ = 0.32) and with perceived relevance for future career (ρ = 0.33, reverse-coded). A moderate inverse association was observed between perceived usefulness for ethical reflection and frequency of ethical reflection (ρ = -0.31). These findings indicate that students' perceptions of the educational and professional value of dissection are shaped by specific experiential factors-particularly emotional engagement and perceived career relevance-and are best understood as domain-specific patterns rather than reflecting a uniform process. These findings suggest that anatomy teaching should incorporate strategies that actively promote student engagement and explicitly link dissection with professional development.
Prolonged operative time is associated with increased surgical complications, resource utilization, and cost. While factors such as surgeon experience and patient habitus are well recognized, the impact of anatomical anomalies and variations remains poorly quantified. This systematic review evaluates the influence of abdominal structural anomalies and anatomical variations on operative time. Following PRISMA 2020 guidelines, a search of PubMed (2000-2025) identified 82 eligible studies. Operative-time impact was assessed using added operative time relative to reference benchmarks. The unit of analysis was the procedure-level observation. Observations were stratified by procedure category and surgical complexity and summarized using median added operative time and range. Among procedure-level observations with available benchmarks, 13 of 14 (92.9%) structural anomalies, 48 of 54 (88.9%) situs inversus cases, and 10 of 11 (90.9%) other anatomical variants demonstrated increased operative time relative to baseline. Structural anomalies were associated with larger median increases, particularly in colorectal procedures, whereas anatomical variations demonstrated a more heterogeneous effect across procedural systems. Higher-complexity procedures, including pancreatic operations, showed greater absolute increases in operative time. Preoperative identification of anatomical differences was associated with attenuation of operative delays. Structural anomalies are frequently associated with increased operative duration, whereas anatomical variations demonstrate a more variable and context-dependent effect. The magnitude of operative time impact is influenced by procedural complexity and preoperative planning. Incorporating anatomical variability into surgical planning may reduce avoidable operative delays and improve operative efficiency.
ABSTRACT Competency‐based medical education (CBME) has reshaped undergraduate medical training by emphasizing demonstrable performance, developmental progression, and entrustment decisions rather than time‐based advancement. Within this reform, anatomical sciences are increasingly conceptualized as longitudinal clinical competencies rather than discrete foundational courses. Although structural changes in curriculum and assessment have been widely described, the learning‐theory foundations supporting authentic competency development in anatomy remain insufficiently articulated. This conceptual communication examines how the three major learning theories (behaviorism, cognitivism, constructivism) inform anatomy education within the CBME context. Behaviorist principles support mastery and procedural reliability through deliberate practice and feedback. Cognitivist principles organize complex anatomical knowledge into durable schemas that facilitate clinical reasoning and efficient retrieval. Constructivist principles situate anatomical understanding within authentic clinical contexts, promoting transfer and adaptive expertise. The manuscript further examines how these theoretical traditions guide the integration of anatomy with clinical competencies and inform assessment strategies aligned with CBME, including structured observation, case‐based evaluation, simulation, and programmatic assessment. By linking learning theory to instructional design and assessment selection, this work clarifies how anatomical competence develops across the educational continuum and highlights the importance of theoretical coherence in maintaining the clinical relevance and disciplinary rigor of anatomical sciences.
Sustainable laboratory practice is generally focused on reducing occupational exposure in anatomy, with environmental impact due to embalming chemicals receiving less attention. We apply a sustainable chemistry methodology that reconsiders a screen of seven cadaver preservation formulations with complete for and two with incomplete data. A whole life-cycle assessment compared carbon burdens per kilogram and per liter of solution based on a cadaver weighing 64 kg required. The findings also suggest that mixtures of preservatives are not environmentally interchangeable. Finally, the burden produced by all assessment trees was lowest for modified Larssen solution (0.196 kg CO2e/kg) and highest for Erskine (1.577 kg CO2e/kg). Burden's story was about more than formaldehyde, though it wasn't the only key ingredient; glycerol, phenol, methanol and sodium sulfate were also important. This indicates that a more sustainable anatomy lab needs whole-formulation redesigns, not just add-ingredient reductions. The study presents a framework of chemistry-informed reformulation priorities to be implemented within anatomy teaching laboratories.
Generative artificial intelligence (GAI) is increasingly being applied in biomedical sciences and medical education, including anatomy, where text-to-image generators may facilitate rapid creation of visual materials. However, the anatomical accuracy of such generated illustrations remains uncertain. The present study evaluated the performance of selected GAI models in generating anatomically correct representations of human anatomical structures. Six text-to-image generators (DALL-E, Gemini, Freepik, Midjourney, DeepAI, and Canva) were assessed using a standardized prompt ("Generate the most anatomically accurate image of the human [name of anatomical structure]"). For each anatomical region, two images were generated. The evaluated structures included the liver, femur, scapula, aortic arch, arm muscles, sacrum, humerus, thigh muscles, celiac trunk, kidney, brainstem, and lumbar plexus. All images were analyzed using reference tables based on Terminologia Anatomica, with individual structures assessed for presence and physiological correctness by four independent reviewers. Statistical analysis included calculation of structure-level proportions with 95% confidence intervals, pairwise inter-rater agreement using Cohen's κ, and exploratory logistic regression with robust standard errors clustered by generated image to evaluate the effects of model, anatomical region, and evaluator. A total of 144 images were analyzed. DALL-E demonstrated the highest overall proportion of structure-level entries scored as present (56.7%, 95% CI: 53.8-59.5) as well as the highest proportion of physiologically correct entries among those present (66.0%, 95% CI: 62.2-69.5). When both criteria were combined, the proportion of fully correct structure-level entries remained limited across all models, with the highest value observed for DALL-E (37.4%). Considerable variability in performance was noted across anatomical regions, with particularly low accuracy observed for complex structures such as the lumbar plexus and celiac trunk. Inter-rater agreement ranged from moderate to almost perfect. Overall, current GAI models demonstrate substantial limitations in producing anatomically accurate illustrations. Although certain models outperform others, the reliability of generated images remains insufficient, and expert verification is necessary before their application in medical education or scientific contexts.
The aim of this study was to investigate the effects of an AI-Narrative-Stratified (ANS) model and traditional anatomy education on academic mastery and clinical reasoning. The study included 254 students who were cluster-randomized into two groups. The two groups were (n = 126) control (Group 1) and (n = 128) experimental (Group 2). The baseline academic scores of the two groups were comparable, and the difference between them was not significant. On this basis, the groups were assigned to control and experimental groups. The control group received traditional anatomy instruction, while the experimental group received the ANS intervention in addition to traditional instruction. The posttest OSPE scores of the experimental group were significantly higher than those of the control group, with an average increase of 9.32 points, and there was a statistically significant difference (p < 0.001). It is predicted that providing AI-driven stratified narrative education in addition to traditional anatomy education will have a positive effect on academic success and professional identity. The qualitative findings of the study revealed several key findings. Participants in the experimental group reported that the role-specific narratives facilitated a deeper understanding and retention of relevant anatomical concepts. They highlighted the simulation of clinical scenarios as helpful in making complex topography more relatable and applicable. In addition, students expressed that the "role enactment" approach increased their engagement and professional belonging, contributing to a more meaningful learning experience. These qualitative findings highlight the potential of the ANS model to complement traditional teaching methods and provide a precise, context-driven learning experience.
Artificial intelligence (AI) is increasingly implemented in medical education, adding new opportunities to improve learning in complex subjects like anatomy. This study assessed the perceived effectiveness of AI-powered tools in undergraduate anatomy education by evaluating the content validity of AI-generated materials, the student-perceived quality of generated questions, and students' perceptions of AI use. A cross-sectional, multi-phase mixed-methods study was conducted at the College of Medicine and Health Sciences, Sultan Qaboos University, Oman (September to December 2025). Three aims were addressed: (1) expert content validity review of AI-generated heart anatomy materials; (2) students' evaluation of AI-generated MCQs across clarity, difficulty, scientific accuracy, and educational utility; (3) a structured survey examined students' perceptions, usage patterns, and attitudes toward AI tools in anatomy education. Findings were interpreted using three complementary frameworks. These linked the study aims to technical quality, psychometric assessment, educational value, ethical considerations, and human-involved review process. AI-generated anatomy content showed favorable, small-panel-dependent validity (CVR = 0.933; 90% above threshold). Students rated the AI-generated MCQs positively, especially for clarity and scientific accuracy, while difficulty scored lowest. AI was widely used for concept clarification and self-testing. Students found it helpful and time-saving, but reported only moderate confidence in its accuracy, limited AI competence, and concerns about ethical use and long-term retention. When mapped to the three frameworks, the findings indicate favorable expert-rated content validity (Roveta), positive learner responses, and indirect learning gains (Kirkpatrick Levels 1-2). They also highlight the need for a human-in-the-loop workflow (QUEST-AI). AI-powered tools can support anatomy education by improving efficiency, engagement, and self-directed learning. However, they should complement rather than replace traditional teaching. Structured guidance and expert validation of generated content are essential to ensure safe and effective integration into medical curricula.
Terminologia Histologica, the international standard nomenclature of human histology and cytology, contains 1093 Latin words that appear to be used as adjectives. Among these, we identified 56 (5%) with a variety of linguistic issues, including typographical or spelling errors, less favored spelling variants, and several unfortunate word choices, including using the prepositions cis and trans, and the noun gigans, as adjectives. The most common adjectival form was that of nominal adjectives (39%). Other common forms were simple adjectives (6%) and participles (9%), chiefly from classical Latin, and prefixed adjectives (22%) and compound adjectives (16%), primarily from neo-Latin. In addition, to simplify several microscopic anatomy terms, in accordance with the updated rules of anatomical nomenclature, we recommend that adverbs (valde, non, nec, and neque) that modify adjectives in these terms be replaced by prefixes (per- and non-).
Terminology surrounding fascia has expanded considerably following renewed interdisciplinary interest in connective tissues, surgical anatomy, mechanotransduction, manual therapies, and movement science. Motivation for this publication arose from the observed disparity between formal anatomical nomenclature and widespread terminology usage. Commonly used expressions in current literature include "the fascial system" and "the superficial fascial system." Widespread adoption of such terminology has occurred with limited or no critical examination regarding its anatomical validity or its lack of consistency with internationally recognized anatomical nomenclature standards. This review critically evaluates the recurring classification of fascia as a "system" within scientific and clinical literature. Drawing upon established anatomical principles and the terminological frameworks promoted by the International Federation of Associations of Anatomists and the Federative International Programme for Anatomical Terminologies, it is argued that fascia does not satisfy the defining criteria associated with anatomical systems. Fascia demonstrates continuity across regional, structural, and functional boundaries, resisting reduction into discrete organ-based hierarchies. A narrative review methodology was employed to identify representative examples within peer-reviewed literature in which fascia was described as a "system" and analyzed in relation to accepted anatomical definitions and broader principles of biological organization. The paper further examines how repeated uncritical use of the term through publication and peer review has contributed to conceptual ambiguity in anatomy education, clinical communication, and fascia research. Consistent with the nomenclature principles and classifications recognized by the International Federation of Associations of Anatomists and its terminological frameworks, fascia should not be described as a "system".
Large language models (LLMs) are increasingly used in medical education and academic writing. However, concerns remain regarding reference hallucination, citation, and the reliability of LLM-generated content. This study aimed to evaluate the performance of ChatGPT 5.2, Gemini 3 Pro, and DeepSeek V3.2 in generating anatomy-related responses by assessing bibliographic reference accuracy, citation content consistency, and the readability of LLM-generated content. A total of 120 open-ended anatomy questions covering six anatomical categories (neuroanatomy, musculoskeletal, respiratory and circulatory, gastrointestinal, urogenital and endocrine, head and neck) were submitted to each model. Individual citation components, including author names, article titles, journal names, publication details, and PMIDs, were verified against indexed sources. Citation content consistency was evaluated using a three-point Likert scale. Readability was assessed using the Flesch Reading Ease score, Flesch-Kincaid Grade Level, Coleman-Liau, and Simple Measure of Gobbledygook indices. A total of 1800 references were analyzed. ChatGPT 5.2 demonstrated the lowest hallucination rate (23.2%), whereas Gemini 3 Pro and DeepSeek V3.2 exhibited substantially higher hallucination rates (45.8% and 47.5%, respectively). DeepSeek V3.2 achieved the highest accuracy for several individual bibliographic components, including author names, article titles, volumes, issues, pages, and journal names. PMID accuracy remained limited across all models, ranging from 25.1% to 57.6%. Citation content consistency differed significantly among the models (p < 0.001), with ChatGPT 5.2 demonstrating the highest proportion of fully supported citations (67.2%), compared with Gemini 3 Pro (42.5%) and DeepSeek V3.2 (41.0%). Citation accuracy differed significantly across most anatomical subcategories, with the greatest intermodel discrepancy observed in head and neck anatomy. Readability analyses indicated that the generated responses generally required college-level reading proficiency. Although LLMs can generate plausible anatomy-related responses, substantial limitations remain regarding reference accuracy, hallucination, and citation reliability. Human verification remains essential before incorporating LLM-generated references into academic or educational materials.