OBJECTIVES:Laboratory detection of myositis-specific autoantibodies (MSAs) utilizes ELISA and multianalyte line blot assays (LBA). We sought to evaluate the concordance and reliability of these two commercial assays. METHODS:Serum samples from patients with idiopathic inflammatory myopathies (IIMs) were obtained from seven countries across the Asia-Pacific region. Anti-Jo-1, anti-EJ, anti-PL-7, anti-PL-12, anti-MDA5, anti-Mi-2 and anti-TIF1-γ antibodies were centrally measured with commercial ELISA and LBA kits. The positive percentage agreement (PPA), negative percentage agreement (NPA) and Cohen's kappa were calculated by comparing the two assays. Sera with discordant results were subjected to 'gold-standard' immunoprecipitation (IP) assays. RESULTS:Serum samples obtained from 485 patients with IIMs, including 180 with DM, 44 with amyopathic DM, seven with JDM, 197 with PM or immune-mediated necrotizing myopathy and 57 with IBM, were subjected to ELISA and LBA. The PPA was the highest for anti-Jo-1 at 0.98, followed by 0.94 for anti-PL-7, 0.93 for anti-EJ, 0.93 for anti-MDA5, 0.89 for anti-TIF1-γ, 0.78 for anti-PL-12 and 0.67 for anti-Mi-2, whereas the NPA was high (ranging from 0.97 to 1 for all MSAs). Kappa values exceeded 0.80 for anti-Jo-1, anti-EJ, anti-MDA5 and anti-TIF1-γ, whereas anti-PL-7, anti-PL-12 and anti-Mi-2 exhibited low values. IP assays using sera with discordant results revealed a high rate of false positives for anti-PL-7 and anti-Mi-2 in LBA. CONCLUSION:Discrepancies in the measurement results were observed between commercially available ELISA and LBA, especially for anti-PL7 and anti-Mi-2. ELISA is more accurate than LBA.
AIM:Currently, global rheumatology training programs lack structured frameworks to address the digital literacy and governance of large language models (LLMs) use. This study aimed to assess the use and educational gap related to LLMs in rheumatology. METHODS:The Asia-Pacific League of Associations for Rheumatology (APLAR) Young Rheumatologists (AYR) conducted an international cross-sectional survey to assess the familiarity, usage patterns, perceptions, and educational unmet needs regarding LLMs among rheumatologists and trainees. RESULTS:A total of 767 participants completed the survey, with 89.6% reporting prior use of LLMs. ChatGPT was the most adopted tool (68.7%), followed by DeepSeek (37.2%) and Google Gemini (21.9%), with a subset of respondents integrating multiple LLMs platforms. While usage is widespread, only 9.9% reported established institutional policies on LLMs use. LLMs users were significantly younger and expressed higher confidence in verifying LLMs-generated information compared to nonusers. Although both groups shared concerns regarding the potential loss of traditional clinical skills (81.1%), active users significantly favored the integration of mandatory LLMs training into core rheumatology curricula. CONCLUSION:There is a significant disparity between the high prevalence of LLMs adoption in rheumatology and the absence of formal educational guidelines or institutional governance. These findings underscore an urgent need for the development of competency frameworks to ensure the safe, critical, and effective application of LLMs in clinical practice.
Large language models have generated intense interest in clinical artificial intelligence (AI), yet their evolution from passive chatbots into agentic systems capable of autonomous planning, tool use, and multistep execution represents a distinct technological paradigm. Unlike conversational AI, which makes suggestions and awaits review, agentic systems act independently and can chain clinical decisions across electronic health records, laboratory data, and patient communications. Rheumatology, with its decades-long disease trajectories, multimodal data streams, and iterative treatment adjustments, is a field well suited to such automation, but also particularly vulnerable to its failures. Errors in autonomous chains compound silently, verification demands exceed those of any previous clinical decision support tool, and the foundational tasks most suited to automation are precisely those through which trainees develop clinical expertise. This Viewpoint critically appraises the agentic AI landscape, analyses its applications and risks within rheumatology, and proposes a tiered governance framework to ensure meaningful human oversight.
This review aims to provide clinicians with a practical framework for distinguishing idiopathic inflammatory myopathies (IIMs) from their numerous non–immune-mediated mimics, including hereditary, toxic, metabolic, and endocrine myopathies, by integrating clinical, serologic, imaging, electrophysiologic, and histopathologic clues. IIMs represent a heterogeneous group of immune-mediated muscle diseases. Despite major advances in antibody discovery, imaging, and classification criteria, accurate diagnosis remains challenging because many non-immune-mediated myopathies can mimic IIMs. Such resemblance may lead to misdiagnosis, delay in genetic evaluation, prolonged exposure to offending agents in toxic myopathy, and unnecessary treatment with immunosuppressive therapy that carries significant adverse effects. The temporal course of weakness, pattern of muscle involvement, presence of extramuscular manifestations, and ancillary testing offer important diagnostic clues, but none are pathognomonic when interpreted in isolation. We introduce the mnemonic “MYOSITIS” to summarize key diagnostic red flags that should raise concern for mimicking disorders: Myopathic motor unit potentials without fibrillation potentials, Young age of symptom onset or positive family history, Onset atypical for IIMs, Seronegative or weakly positive myositis specific autoantibodies, Iatrogenic causes, Treatment refractoriness, Irregular weakness patterns, and Systemic features. Recognizing these pitfalls and adopting an integrated diagnostic approach that combines clinical pattern recognition with selective use of serologic, imaging, and genetic testing can help clinicians differentiate true IIMs from their mimics and ensure timely, accurate diagnosis and optimal patient management.
Background:,Cutaneous dermatomyositis (DM) poses diagnostic and monitoring challenges due to its heterogeneous skin manifestations. The Cutaneous Dermatomyositis Disease Area and Severity Index (CDASI) is the standard for assessing disease activity, yet its complexity hinders routine use. Large language models (LLMs), such as Claude v3.5 Sonnet, offer potential for automated scoring through image interpretation and structured reasoning, but their utility in dermatologic assessment remains underexplored.,Materials:,We retrospectively analyzed 30 published DM cases with standardized clinical images suitable for CDASI scoring. Two expert rheumatologists independently scored each case using CDASI criteria. Claude v3.5 Sonnet was provided with identical images and prompted using a chain-of-thought strategy incorporating structured CDASI definitions. Intraclass correlation coefficients (ICCs) were calculated to assess concordance between Claude and human raters.,Results:,Claude demonstrated strong agreement with expert assessors in global CDASI scoring (ICC vs Expert 1: 0.92; vs Expert 2: 0.87), comparable to inter-expert reliability (ICC: 0.87). Domain-specific agreement was moderate for erythema, scaling, and ulceration (ICCs: 0.57–0.61), but lower for poikiloderma (ICC: 0.47) and chronic damage (ICC: 0.37). Hand assessments showed excellent reliability, including perfect agreement in periungual changes detection (ICC: 1.0). The LLM scored cases 92% faster than experts, averaging 42 seconds per case.,Conclusions:,Claude v3.5 Sonnet achieved expert-level reliability in scoring cutaneous DM, particularly in visually distinct domains. Its efficiency and scalability suggest potential utility in clinical trial screening and decision support, though limitations remain in recognizing subtle or chronic changes. Integration into hybrid human-AI workflows may enhance dermatologic assessment.
Background Consensus guidelines for malignancy screening in idiopathic inflammatory myopathy (IIM) were recently published by the International Myositis Assessment and Clinical Studies (IMACS) working group. This international multicentre study audited historical cancer screening practices against IMACS recommendations and examined the distribution of observed malignancies across IMACS-defined risk categories in a real-world cohort.Methods This retrospective study included patients with IIM from eight centres across seven countries/regions. Patients were stratified using IMACS-defined cancer risk categories. Historical screening was compared with the guideline recommendations. Multivariable logistic regression identified predictors of cancer and underscreening.Results Of the 795 patients (72.3% female), 51.7%, 37.9% and 10.4% were classified as high risk, moderate risk and standard risk, respectively. Across all sites, few high-risk (6.2%), moderate-risk (2.7%) and standard-risk (11.1%) patients underwent a full panel of IMACS-recommended tests. Within 3 years of diagnosis, 86 patients (10.8%) developed malignancy. Among observed cancers, 77.4% occurred in high-risk and 22.6% in moderate-risk patients, with none detected in the standard-risk group. Anti-transcription intermediary factor 1-gamma (TIF1γ) positivity (OR 4.64, 95% CI 2.35 to 9.16, p<0.001), smoking (OR 3.00, 95% CI 1.64 to 5.50, p<0.001), night sweats (OR 12.76, 95% CI 2.72 to 59.87, p=0.001) and increasing age (OR 1.05 95% CI 1.03 to 1.07 per year, p<0.001) were independently associated with cancer, whereas anti-Mi2 antibodies and non-dermatomyositis subtypes were protective. Cancer risk increased with each additional high-risk feature (OR 1.70) and decreased with each low-risk feature (OR 0.76). Underscreened individuals had lower rates of immune-mediated necrotising myopathy (5.6% vs 14.3%, p=0.005) and anti-TIF1γ positivity (6.8% vs 31.4%, p<0.001).Conclusion Observed cancer distribution across IMACS risk categories was consistent with the framework’s stratification logic, but differential screening intensity precludes formal assessment of predictive validity. Variation between historical practice and current recommendations underscores the need for continued education and prospective longitudinal studies to assess cumulative risk, age thresholds, screening uptake and outcomes in real-world clinical settings.
BACKGROUND:The management of reproductive health in individuals with autoimmune rheumatic diseases (AIRDs) has evolved into a primary clinical priority. Wide variations in clinical practice and drug availability in the Asia-Pacific regions necessitate localized frameworks. These consensus statements, developed by the Asia Pacific League of Associations for Rheumatology (APLAR) Special Interest Group for Women's Health & Reproductive Issues in Rheumatic and Musculoskeletal Diseases, aim to bridge the gaps within the region, foster multidisciplinary collaboration, and address regional knowledge deficits. METHODS:An expert panel comprising 16 members of the 13 APLAR member nation organizations formulated 23 research questions using the PICO framework. A systematic review of English-language literature up to July 2025 was conducted across the MEDLINE, Scopus, Google Scholar, and Cochrane Library databases, supplemented by a manual search of journals from the APLAR regions. Evidence from the APLAR region was appraised using the GRADE system. A modified Delphi online voting process was employed to reach consensus, pre-defined as ≥ 75% agreement. RESULTS:The panel approved four overarching principles and 86 statements, organized into 16 broad categories. These statements cover all three phases of pregnancy in AIRDs, the safety of medications during pregnancy and lactation, contraception, assisted reproductive techniques, fertility preservation, and hormonal replacement therapy, covering both female and male aspects as relevant. CONCLUSION:In a field often lacking high-quality data, these consensus statements from APLAR provide expert opinion-based guidance to support clinical decision-making. It is envisaged that it will assist in educational and training purposes and help shape future research priorities.
Idiopathic inflammatory myopathies (IIM) follow a chronic, polycyclic course in a substantial proportion of patients. Unlike remission, which has recently been standardised, the definition of flare (or worsening) remains highly heterogeneous, creating a divide between clinical trials and real-world practice. Validated core set definitions of worsening exist for trials but were derived by consensus on paper patient profiles and are criticised for inadequate content validity; observational studies instead rely on simplified markers such as enzyme elevation or therapeutic escalation. This review summarises existing definitions, contrasting the IMACS and PRINTO statistical criteria with the variable definitions used in longitudinal cohorts. We examine reported flare rates and predictors, the discordance between global and organ-specific (skin, muscle, lung) worsening, and the lessons offered by flare definitions in other rheumatic and inflammatory diseases. We then propose a multi-domain, treatment-anchored framework intended to bridge statistical rigour and clinical utility.
Objectives:Serum creatine kinase (CK) is widely used in the diagnostic evaluation of neuromuscular disorders, including idiopathic inflammatory myopathies (IIM). As most laboratories define reference ranges using the central 95% of values from white populations, elevated CK levels are common in otherwise healthy individuals. We investigated the predictive value of baseline CK at defined cutoffs for diagnosing an underlying myopathic illness including IIM. Methods:A retrospective chart review was performed of consecutive adult patients attending a tertiary neuromuscular clinic from January 2018 to June 2022. We investigated the test characteristics of CK based on several a priori defined thresholds. Results:From a total of 488 patients, we identified 276 IIM cases, 89 with a non-inflammatory myopathy and 123 had an alternative/reassuring aetiology. Baseline CK levels were significantly higher in the IIM and non-inflammatory myopathy groups vs the alternative/reassuring group. The laboratory defined cutoffs for CK yielded a sensitivity of 77.0%, a specificity of 66.7% and a positive predictive value (PPV) of 87.3% for identifying IIM or other neuromuscular disorders, which only marginally improved when evaluating the 97.5th percentile (75.6, 67.5 and 87.3%) or the European Federation of Neurological Societies (EFNS) thresholds (68.8, 78.0 and 90.3%). Conclusion:The current thresholds defined by the EFNS guidelines compared with laboratory defined cutoff to pursue investigations for elevated CK only marginally improve specificity. A baseline CK >1000 IU/l is highly predictive of an IIM or other neuromuscular disorders. This study highlights the role of baseline CK levels in suspected IIM and neuromuscular disorders, clarifying its diagnostic value.
Despite improvement in treatment, rheumatoid arthritis (RA) management remains inconsistent. To evaluate the worldwide disparities in the use of biological and targeted molecules (advanced) RA therapies, focusing on differences across continents and socioeconomic strata, and to identify factors associated with their utilisation. Cross-sectional analysis of the international COVAD-2 cohort, including demographics, socioeconomic factors, disease characteristics, patient-reported outcomes, and treatments. Primary outcomes assessed treatment distribution by continent, secondary outcomes evaluated distribution by Human Development Index (HDI), with predictors analysed using multivariable logistic regression. At the time of analysis, COVAD2 included 10,739 participants; 2007 had RA, 1997 with geographical data included in this study (mean age 50.9 years, 88.1
PURPOSE OF REVIEW:The delineation of myositis-specific autoantibody (MSA) subgroups over the past two decades has reframed environmental research in idiopathic inflammatory myopathies (IIMs), making it plausible to examine discrete exposure-phenotype pairings within serologically defined subgroups. This review synthesises evidence on environmental, occupational, infectious, and pharmacological exposures, with attention to phenotypic specificity. RECENT FINDINGS:Ultraviolet radiation shapes MSA-specific geographic distributions at the population level and associates with dermatomyositis onset at the individual level. Disease onset follows MSA-specific seasonal patterns, replicated across geographically independent cohorts. Smoking amplifies antisynthetase syndrome risk through a gene-environment interaction with HLA-DRB1*03:01, whilst protecting against anti-TIF1-γ autoantibodies. Gestational and early-life tobacco smoke exposure is consistently implicated as risks for juvenile IIM. Occupational silica exposure associates with antisynthetase syndrome and overlap myositis, with synergistic lung disease risk when combined with smoking. Statins, immune checkpoint inhibitors, anti-TNFα agents, and interferons each induce immunopathologically distinct myositis phenotypes. Postpandemic surveillance has documented a rise in anti-MDA5 autoantibody-positive disease following SARS-CoV-2 infection. SUMMARY:Environmental exposures in IIM act preferentially in specific serological and genetic contexts, consistent with discrete exposure-phenotype pairings shaped by genetic background and age at exposure. Prospective, MSA-stratified studies with direct exposure quantification represent priority directions for the field.
OBJECTIVES:Medical nutrition therapy significantly impacts cardiovascular risk and overall health, but effects on muscle diseases remain unclear. This systematic review evaluates the safety and efficacy of dietary interventions and supplements on muscle disease outcomes. METHODS:A multidisciplinary team conducted a PRISMA-guided systematic review registered on PROSPERO. Searches were conducted across multiple databases and screened against pre-specified inclusion criteria. RESULTS:Of 107 full-text articles screened, 51 met inclusion criteria. Most identified interventions used dietary supplements rather than whole dietary approaches. In inflammatory myopathies, creatine (loading dose 20 g/day, maintenance 3 g/day) combined with exercise improved high-intensity functional performance in PM and DM over 6 months. In Duchenne muscular dystrophy, creatine (2-10 g/day for 8-16 weeks) improved maximal voluntary contraction and fatigue resistance. Carbohydrate-rich diets (65% CHO) reduced exercise-related symptoms in McArdle disease, while high-dose creatine (150 mg/kg/day) paradoxically worsened symptoms. Four trials of aceneuramic acid (6 g/day for 48 weeks) in GNE myopathy demonstrated dose-dependent strength improvements, leading to regulatory approval in Japan. High-protein supplementation showed positive trends for muscle preservation in critical illness myopathy. Quality assessment revealed 31% at low risk of bias, 49% with some concerns and 20% at high risk. CONCLUSION:Evidence for nutritional interventions in muscle diseases remains limited, especially for inflammatory myopathies. The strongest support emerged for mechanistically targeted approaches: creatine with exercise, carbohydrate-rich and ketogenic diets in McArdle disease and sialic acid in GNE myopathy. Future research requires adequately powered multicentre trials with standardized outcomes, with focus on inflammatory myopathies.
OBJECTIVES:Discordance between patient and physician perspectives on disease activity in idiopathic inflammatory myopathies (IIM) can compromise treatment outcomes. This study aimed to characterize patterns of discordance and distinct longitudinal trajectories in the MyoCite IIM cohort. METHODS:Discordance was defined as a difference between patient (PGA) and physician (PhGA) global assessments (0-10 cm scale) of ≥ ±1 cm. Prevalence was assessed at baseline (n = 244) and longitudinally at 6, 12 and 24 months (n = 128). Linear mixed-effects (n = 191) and bidirectional mediation models identified independent drivers and causal pathways. RESULTS:Clinically significant baseline discordance occurred in 19.3% of patients, predominantly manifesting as positive discordance (PGA > PhGA; 16.8%) across IIM subtypes (DM, PM, overlap myositis, anti-synthetase syndrome). Over 24 months, discordance narrowed with treatment. Baseline positive discordance strongly predicted poorer functional capacity (HAQ) at 1 year (P = 0.034). While objective muscle weakness statistically drove the gap (MMT-8: β = -0.013, 95% CI -0.022 to -0.004; standardized β = -0.155, P = 0.006), fluctuations had minimal absolute impact. A distinct subset (25%) maintained persistent discordance, exhibiting high subjective pain despite normalized muscle enzymes. Bidirectional mediation analysis revealed discordance unidirectionally drives future functional disability through unmanaged pain (32.1% mediated, P < 0.001) while reverse mediation (HAQ driving discordance via pain) was not statistically significant. CONCLUSION:Patient-physician discordance in IIM is an active, upstream driver of future functional loss, mediated significantly by pain, rather than a mere reflection of existing objective damage. It serves as a critical marker requiring early targeted pain intervention and integration of patient-reported outcomes into routine clinical monitoring.
Objectives Artificial intelligence (AI) is revolutionising medicine. The aim of this study was to detail its use, opinions, knowledge, and concerns in rheumatology and paediatric rheumatology. Methods A web-based survey open to all professionals working in the field was developed by the Emerging EULAR Network (EMEUNET) and disseminated between March and July 2025 in collaboration with other international rheumatology societies (AFLAR, ArLAR, CARRA, PAFLAR, PANLAR). The survey was divided into 4 sections: (i) participants’ characteristics, (ii) AI use and applications, (iii) opinions and knowledge, and (iv) concerns, needs, and expectations. Results Overall, 461 responses were collected from 59 countries. Respondents were mostly physicians who completed their training (316, 68.7%) and were based in Europe (170, 36.9%). Most participants (397, 86.7%) used AI for medical purposes, especially large language models (385, 83.7%) for grammar correction and brainstorming. Although there was broad optimism about its use (366, 79.6%), self-reported practical skills were predominantly basic or still in development (346, 75.1%), and knowledge was rarely defined as strong or expert-level (63, 13.7%). Concerns focused on ethics (314, 69%), lack of trust (316, 69.5%), and insufficient training (270, 59.3%). Disparities emerged across geographic regions in use, knowledge, and practical skills. Conclusions AI is widely used and positively perceived in rheumatology, despite limited knowledge and practical skills, and regional disparities. Addressing gaps in ethics, transparency, and insufficient training through targeted education and implementation strategies will be essential to ensure an equitable and effective integration into clinical and research practice.
BACKGROUND:Gender disparities in rheumatology persist despite increasing female representation. This comprehensive global survey, conducted by the Coalition for Health and Gender Equity (CHANGE) group, examined preferences for gender equity interventions among rheumatology professionals. METHODS:A cross-sectional online survey of rheumatologists and allied health professionals was conducted across 105 countries (January 2023-May 2024). Intervention preferences spanned four domains: conference-based, organizational, skills training, and work-based strategies. Gender differences were evaluated using logistic regression adjusted for years of work experience, with subgroup analyses by Human Development Index (HDI), Gender Inequality Index (GII), professional role, and caregiving responsibilities. Multiple comparisons were adjusted using Benjamini- Hochberg correction for false discovery rate. RESULTS:Among 1945 respondents (66.4% female), the most widely endorsed interventions overall were family- and child-friendly conference policies (71.4%), communication and scientific writing training (52.6%), and engagement of national rheumatology groups (46.9%). Women demonstrated significantly stronger support for increasing the visibility of female role models (11.89% vs 5.81%, OR = 3.29, p < 0.01), establishing gender-balanced committees (17.66% vs 16.01%, OR = 1.76, p < 0.01), and gender-sensitive editorial board policies (12.70% vs 10.66%, OR = 1.77, p < 0.01). Respondents from higher-GII countries prioritised scientific writing masterclasses (OR = 4.58, p < 0.05) and career planning training (OR = 3.90, p < 0.05) but showed lower endorsement of unconscious bias training (OR = 0.48, p < 0.05) and male allyship programmes (OR = 0.42, p < 0.05). Women in low/middle-HDI countries more frequently selected grant writing support (40.3% vs 26.4%, p < 0.01). CONCLUSION:Gender equity intervention preferences vary systematically across contexts. These findings provide empirical foundations for developing culturally responsive, evidence-informed intervention frameworks to advance gender equity in rheumatology.
Objectives Evaluating compliance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement can be time-consuming and subjective. This study compares STROBE assessments from large language models (LLMs), a human reviewer panel, and the original manuscript authors in observational rheumatology research. Methods Guided by the Guidelines for Reporting Reliability and Agreement Studies and Development, Evaluation, and Assessment of Large Language Model Pathway B frameworks, 17 rheumatology articles (11 cohort, 4 cross-sectional, and 2 case-control) were independently assessed. Evaluations used the 22-item STROBE checklist, completed by the authors, a 5-person human panel (ranging from junior to senior professionals), and 2 LLMs (ChatGPT-5.2 and Gemini 3 Pro). Interrater reliability was calculated using Gwet’s Agreement Coefficient (AC1) with 95% CIs. Results Overall agreement across all reviewers was 85.0% (AC1 = 0.826 [95% CI: 0.801-0.851]). Domain stratification showed almost perfect agreement for ‘Presentation & Context’ (AC1 = 0.841 [95% CI: 0.810-0.872]) and substantial agreement for ‘Methodological Rigor’ (AC1 = 0.803 [95% CI: 0.761-0.845]). Although LLMs achieved complete agreement with all human reviewers on standard formatting elements, their agreement declined on complex methodological items, with some pairwise comparisons yielding negative AC1 values. Intra-LLM and cross-version agreement across repeated independent runs was high, and estimates were stable across publication periods, providing no clear evidence of data leakage. Conclusions While LLMs show potential for basic STROBE screening, their lower agreement with human experts on complex methodological items likely reflects a reliance on surface-level information. These models appear more reliable for standardising straightforward checks than for replacing expert human judgement in evaluating observational research.