BACKGROUND:Accurate prediction of ischemic stroke or systemic embolism is central to atrial fibrillation (AF) management, yet the transportability of available risk scores to Chinese populations remains uncertain. OBJECTIVE:Externally validate and compare published stroke risk scores in Chinese patients with non-valvular AF. METHODS:We systematically retrieved and externally validated 17 published models in a retrospective cohort of 1,283 adults hospitalized with non-valvular AF at a Chinese tertiary-care center. Discrimination was evaluated using the 4-year Uno C-index and time-dependent areas under the curve (AUC). Models providing absolute risks also underwent calibration, Brier score, and reclassification analyses. Sensitivity analyses accounted for death as a competing event and separately evaluated patients not receiving OAC at baseline. RESULTS:During 4 years of follow-up, 103 primary outcomes occurred, corresponding to 2.85 events per 100 person-years. Across models, 4-year Uno C-indices ranged from 0.572 to 0.767, compared with 0.690 for CHA2DS2-VASc. GARFIELD-AF 2017, GARFIELD-AF 2021, and ATRIA had higher Uno C-indices than CHA2DS2-VASc. However, no model had significantly higher time-dependent AUC at any time point. For 1-year calibration, CHA2DS2-VASc had predicted and observed risks of 3.39% and 3.54%, respectively, with an observed-to-expected ratio of 1.04, whereas several other models underestimated risk. Brier scores and reclassification measures showed no consistent advantage Competing-risk analyses lowered cumulative incidence without changing comparative patterns. Among patients not receiving OAC at baseline, no model outperformed CHA2DS2-VASc. CONCLUSION:Candidate models did not consistently outperform CHA2DS2-VASc in this external validation; CHA2DS2-VASc remains practical, while other models may require population-specific validation and recalibration.