Mixed mood states are a common presentation of pediatric mood disorders. However, their representation in categorical and dimensional taxonomies are misaligned with empirical findings, delaying treatment, and worsening outcomes for youth with mood disorders. To mitigate these negative outcomes, taxonomic representation of mixed mood states (e.g., concomitant symptoms and placement within taxonomic structure) must be reassessed. Therefore, the present study aimed to build empirical definitions of mood disorder dimensions and categorical episode presentations, including mixed mood states, using clinically relevant and substantiated assessments of depression and mania. All participants were families with children, ages 5-18, seeking outpatient mental health services between July 2003 to July 2007. An exploratory factor analysis found that a three-factor solution with manic, depressive, and psychomotor retardation factors best represented the data. A latent profile analyses differentiated seven profiles differentiable as (a) Attention-Deficit/Hyperactivity Disorder and Disruptive Behavior Disorders, (b) Moderate Low Mood, (c) Major Depressive Episode, (d) Mixed Mood State, (e) Psychomotor Retardation Without Mood Symptoms, (f) Euthymia, (g) Adjustment Disorder and Other Mild Problems. Analyses indicate that depression, mania, and psychomotor retardation may be underlying mood dimensions that distinguish unique groups of youth. Findings also demonstrated that mixed-mood youth can be clearly identified using a dimensional approach despite categorical systems deprioritizing mixed presentation. This study contributes empirical findings that can be used to rebuild categorical or dimensional taxonomies of pediatric mood disorders, which would lead to improved, timely diagnosis and treatment for affected youth.
Bipolar disorder is a recurrent, heterogeneous condition that often begins in adolescence and typically requires lifelong, multimodal management. Advances in evidence-based assessment (EBA) offer structured frameworks for prediction, prescription, and progress monitoring, and pharmacological and psychosocial interventions supported by recent reviews and the Canadian Network for Mood and Anxiety Treatments (CANMAT)/International Society for Bipolar Disorders (ISBD) guidelines provide effective options across phases of illness. Despite these advances, the impact of evidence-based approaches remains blunted in practice: Diagnosis is often delayed, pharmacotherapy is inconsistently prescribed or monitored, psychosocial interventions are underused, and relapse prevention strategies are rarely sustained. Therefore, the field must embed prediction, treatment, and monitoring within community treatment settings—primary care, schools, digital platforms, and family systems—where risk can be identified early, preventive strategies can be delivered, and long-term maintenance can be supported. Framing EBA as a dynamic, community-anchored cycle offers the best chance of translating evidence into improved outcomes, bridging the gap between research efficacy and real-world effectiveness in the care of bipolar disorder.
Objective:To test whether two brief mania measures, the Parent General Behavior Inventory-10 Mania form (PGBI-10M) and 7-Up, retain useful psychometric properties in a large population cohort, and to evaluate whether the PGBI-10M can identify Kiddie Schedule for Affective Disorders and Schizophrenia (KSADS)-defined bipolar spectrum disorders in that setting. Method:Analyses used 11,000+ youths across late childhood and early adolescence from the Adolescent Brain Cognitive Development (ABCD) Study. For both PGBI-10M and 7-Up, we estimated descriptive statistics, internal consistency, confirmatory factor models, graded response models, and measurement-based care benchmarks (minimally important difference, reliable change, and clinical cutpoints). For the PGBI-10M, receiver operating characteristic (ROC) analyses estimated concurrent classification accuracy for bipolar diagnoses at baseline and 2-year follow-up and compared area under the curve (AUC) values with prior outpatient and community mental health samples. Results:Scores were lower than in clinical samples, but both measures remained psychometrically sound. The PGBI-10M showed alpha=.87-.88 and omega=.88; the 7-Up showed alpha=.78 and omega=.79. Longitudinal analyses indicated threshold differences across waves, likely reflecting caregiver recalibration and developmental changes, with modest impact on estimates. ABCD-based benchmarks supported meaningful and reliable change. The PGBI-10M discriminated bipolar cases (AUC=0.68 baseline; 0.77 follow-up), though performance was lower than in clinical samples. Positive predictive values were low in this population. Conclusion:The PGBI-10M and 7-Up support monitoring of manic and mixed symptoms, but the PGBI-10M alone is insufficient for universal bipolar screening. Brief mania scales are best used for targeted assessment and longitudinal monitoring within multi-informant workflows.
ObjectiveSevere behavioral outbursts characterized by reactive aggression (RA) are common and highly impairing for young children and their families. A growing body of research suggests RA may characterize a distinct group of youth with behavior problems; however, RA is not specific to a single diagnosis in the current clinical nosology. The challenges of assessing the severity of RA and elucidating its diagnostic role as a transdiagnostic symptom are compounded by a lack of validated, RA-specific assessment tools. The current work presents psychometric evidence for a novel screening tool and outcome measure: the 16-item Reactive Aggression Assessment (RAGA-16). It is the first of two manuscripts presenting evidence for the RAGA-16.MethodData comprised a nationally representative US sample (N=1,162) of parents with children ages 6 to 19 (M=11.4; SD=3.98) completed an as-yet unpublished assessment of mood and behavior problems. Latent class analysis (LCA) of expert ratings identified a pool of RA-specific items. We compared two approaches to developing the RAGA-16 from this item pool: a Data Informed approach combining item response theory (IRT) techniques and expert judgment of item characteristics; and an automated test assembly (ATA) approach that selected items based on optimizing test information. Scores from the resulting forms were evaluated based on multiple facets of reliability, differential item functioning (DIF), and associations with parent-reported health information.ResultsLCA identified 43 RA items representing two distinct but highly correlated (r=.63) content domains: Temper Loss and Aggressive Behaviors. Both the Data Informed and ATA-Derived forms showed high reliability (rxx>.80) across a wide range of symptom severity and excellent internal consistency despite containing largely different items. The Data Informed version showed less evidence of DIF across child race and sex as well as slightly higher detection of parent-report externalizing disorders (AUC=.76 vs. .73). Percentile and T-Score norms for scores on the Data Informed form are reported.ConclusionsThe Data-Informed version was selected as the final version of the RAGA-16. Having a brief, highly reliable, targeted assessment of RA with epidemiological norms allows researchers and clinicians to better characterize RA behaviors in youth. Forthcoming work will validate the RAGA-16 as a screening tool in inpatient and outpatient behavioral health settings.
Autism spectrum disorder (ASD) is a heterogeneous condition that has led some to question whether IQ scores can be meaningfully compared across neurodivergent and neurotypical populations. The purpose of this study was to evaluate the measurement invariance of the Stanford-Binet Intelligence Scales, Fifth Edition (SB-5), among 3,050 clinically referred youth ages 2–16 with a diagnosis of ASD (n = 1,329; 43.6
OBJECTIVE:Reactive aggression (RA) is a transdiagnostic symptom associated with more severe illness and worse clinical course. Despite the clinical importance of RA, few validated screening measures or outcome assessments exist for use with children under 12. The 16-item Reactive Aggression Assessment (RAGA-16) is a novel parent-report instrument with promising evidence of psychometric reliability. The current work presents clinical validation evidence for RAGA-16. METHODS:A retrospective chart review examined the records of children seeking inpatient and outpatient services. Parents completed the RAGA-16 and additional questionnaires assessing depression, hypomania, ADHD symptoms, and aggression at intake. The chart review also included clinical diagnoses, demographic information, and treatment details. Graded response models (GRMs) evaluated the RAGA-16 factor structure and associations with criterion variables evaluated the scores' validity as a screening measure assessing RA. RESULTS:Confirmatory models showed good fit to the hypothesized two-factor structure with high marginal reliability (ωt = .98). RAGA-16 scores showed large significant positive associations with measures of aggression, impulsivity, and oppositionality (rs > .70, ps < .001), lower associations with hypomanic symptoms (r = .40, p < .01), and no significant association with depression (p > .05). RAGA-16 scores statistically discriminated diagnoses associated with severely externalizing symptoms from other diagnoses. Endorsement of RA in more contexts (home, school, or elsewhere) was associated with higher aggression symptoms. CONCLUSIONS:Overall, RAGA-16 scores showed a pattern of criterion associations consistent with the construct of interest along with high score reliability. The addition of inpatient and outpatient clinical score characteristics to the existing epidemiological norms makes RAGA-16 a promising new tool for assessing youth RA.
Cognitive Disengagement Syndrome (CDS) is characterized by symptoms such as daydreaming, slowed behavior, and mental confusion, but remains understudied outside of ADHD populations. CDS has recently emerged as a potential transdiagnostic construct, yet most investigations have been limited to ADHD samples. To further clarify the clinical profile of CDS, research is needed in pediatric mood disorder populations and comorbid presentations. This study examined predictors of caregiver-reported CDS symptoms in a racially and socioeconomically diverse sample of treatment-seeking outpatient youth (N = 697), with attention to psychiatric diagnoses (ADHD and mood disorders), youth demographics, caregiver education, and number of other psychiatric diagnoses. A 5-item CBCL-based CDS index demonstrated acceptable psychometric performance and was used to capture youth CDS levels. Hierarchical regression revealed that mood disorder diagnosis was the strongest and most unique predictor of elevated CDS symptoms, followed by ADHD diagnosis, with the comorbid group (ADHD+Mood) showing the highest CDS levels. Clinical and structural factors, including mood and ADHD diagnoses, number of other psychiatric diagnoses, Black racial identity, lower caregiver education, and older age, each independently predicted CDS severity. These findings support CDS as a clinically meaningful construct extending beyond ADHD and underscore the importance of contextually informed, transdiagnostic assessment. Implications are discussed through a developmental psychopathology framework, emphasizing equifinality.
Abstract Introduction The Nationwide Quality of Life Scale (NQLS) is a brief, mental-health focused quality of life (QoL) scale with seven items that are non-overlapping with symptom scales. We developed a parent version (P-NQLS), obtained national norms, and calculated psychometric properties for the P-NQLS. Methods Parents ( N =2251) of children aged 6-18 years who were representative of the U.S. population on key demographics completed the P-NQLS along with measures of depression, suicidality, internalizing, externalizing, and attention symptoms. We assessed the P-NQLS’s factor structure through exploratory factor analysis (EFA) and evaluated its internal reliability and convergent validity. Age- and sex-specific norms were established using GAMLSS with BCPE distributions and P-spline smoothers, with percentile curves and tables (5 th -95 th ) provided. Results EFA suggested a one-factor solution for P-NQLS in the national sample. The scale showed good internal consistency (Cronbach’s α =0.85). P-NQLS total scores (M±SD=20.7±4.7, range=0-28, higher scores indicate higher QoL) were negatively correlated (all p <.0001) with depression (Pearson’s r =-0.47), suicidality ( r =-0.50), internalizing ( r =-0.43), externalizing ( r =-0.41), and attention ( r =-0.37) symptoms. P-NQLS scores declined steadily with age in both sexes, with the most pronounced decreases (3-5 points) observed at lower percentiles (5 th , 10 th ), suggesting greater age-related decline among children with lower baselines. Females scored slightly higher than males across most ages and percentile levels, though the differences were within one point. Conclusions The newly created P-NQLS, a 7-item parent-reported QoL scale with one underlying factor, demonstrates strong reliability and validity and has robust national norms for youth aged 6-18.
Pediatric bipolar disorder is challenging to diagnose accurately due to symptom heterogeneity. More standardized and data-driven approaches are needed to enhance diagnostic reliability. We evaluated a clinical decision tool (nomogram), statistical methods (logistic regression, LASSO), machine learning (support vector machine, random forest, k-nearest neighbors, extreme gradient boosting), and deep learning (multilayer perceptron) for pediatric bipolar disorder prediction across two datasets collected in academic (N = 550) and community (N = 511) clinical settings. We compared three modeling strategies: cross-dataset validation, cross-dataset with interaction terms, and pooled-dataset. We assessed model performance using discrimination, calibration, and predictor importance ranking.In the baseline cross-dataset approach, all models showed good internal discrimination in the academic dataset, but external discrimination in the community dataset substantially declined. Interaction-enhanced models slightly improved internal discrimination but not external performance or calibration. Recalibration substantially improved cross-dataset calibration. Models trained on the pooled sample showed strong performance on held-out samples drawn from the heterogeneous pooled cohort, with good calibration for most models. Across models and training strategies, PGBI-10M was consistently identified as the most important predictor.Predictive models for pediatric bipolar disorder showed strong internal performance but limited cross-setting generalizability due to dataset shift and miscalibration. Within the present study, increasing model complexity did not improve external performance, whereas training on pooled data improved performance on held-out samples from the heterogeneous pooled cohort. These findings suggest that training-data diversity may provide greater practical benefit than increasing model complexity for developing robust psychiatric prediction models, underscoring the importance of open and collaborative datasets.
BACKGROUND:Research demonstrates the effectiveness of evidence-based psychological treatment adjunctive to pharmacotherapy for reducing mood symptoms in bipolar disorder. However, access to these treatments is limited, and innovative strategies are needed to ensure that more patients with bipolar disorder receive the gold-standard treatments that may help them achieve wellness. "Stepped care" models of psychological service delivery represent one potential solution to this problem of treatment access. Under a stepped care model, patients are assigned the minimum necessary psychological treatment for symptom improvement. This typically means that patients who are experiencing more symptoms are assigned to a treatment of greater intensity (e.g., weekly individual therapy) whereas patients who are experiencing fewer symptoms are assigned to a treatment of relatively lesser intensity (e.g., biweekly group therapy). Stepped care models are dynamic, meaning that the level of treatment can be modified depending on the patient's response. Stepped care models have been explored in other clinical populations but require further exploration in bipolar disorder. METHODS:Members of the Psychological Interventions Task Force for the International Society of Bipolar Disorders conducted a narrative review of stepped care models and their application to bipolar disorder. RESULTS:We found evidence that stepped care models are useful approaches to delivering psychosocial treatments for bipolar disorder. We discuss several contextual factors in executing stepped care models in this population (i.e., cultural and pediatric applications), as well as share an example of a stepped care model-Focused Integrated Team-based Treatment for Bipolar Disorder (FITT-BD)-that is currently being evaluated in an academic medical center. CONCLUSION:Further research is warranted to develop and assess robust stepped care models to determine whether they can improve access to treatment of bipolar disorder while not sacrificing outcomes.
Introduction: The Pubertal Development Scale (PDS) is widely used for puberty assessment, yet its psychometric properties and norms are limited to research data. This study examined the psychometric properties of parent- and self-report PDS and established continuous norms in nationally representative samples. Methods: We analyzed two deidentified survey samples: a parent-report sample of children aged 6-18 (N=2000, Mage=11.37, 47.2% female, 74.9% White), and a youth self-report sample aged 12-18 (N=754, Mage=14.33, 49.6% female, 75.3% White). Both samples were representative of the U.S. population on key demographics, and the self-report sample consisted entirely of children whose parents also participated in the parent sample, thus creating parent-child dyads. Internal consistency was evaluated using Cronbach's alpha and McDonald's Omega. Cross-informant agreement was assessed with Intraclass Correlation Coefficient (ICC; two-way model, absolute agreement, single unit) and Bland-Altman plots. Age-dependent norms of each sex were established with Generalized Additive Models for Location, Scale, and Shape (GAMLSS), with 5th-95th percentile curves and reference tables provided. Results: Parent- and self-report PDS demonstrated acceptable-to-good internal consistency (Cronbach's alpha: 0.78-0.89; McDonald's omega: 0.79-0.90). Among the 754 parent-youth dyads, excellent cross-informant agreement was observed for both sexes (ICC(2,1)=0.88). Parents' and children's PDS total scores did not differ significantly for boys; for girls, parents rated pubertal development on average 0.13 points lower than children's self-report. Regardless of informants, PDS scores increased nonlinearly with age and exhibited sex-specific developmental patterns. Girls showed earlier pubertal onset, faster progression, and greater convergence toward pubertal completion by late adolescence. Discussion: The PDS demonstrated strong psychometrics in national samples, supporting its utility in the general pediatric population. The national norms provide empirical benchmarks for PDS score interpretation, strengthening its value as a broad estimation of pubertal status and a pre-screening tool for identifying early or delayed puberty. ### Competing Interest Statement Eric A. Youngstrom is the co-founder and Executive Director of Helping Give Away Psychological Science, a 501c3; he has consulted about psychological assessment with Signant Health and received royalties from the American Psychological Association and Guilford Press, and he holds equity in Joe Startup Technologies and held equity in Autism Intervention Measures. ### Funding Statement This study did not receive any funding ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Institutional Review Board of Nationwide Children's Hospital provided a Not Human Subject Research determination for analyses of deidentified survey data used in this study. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Code and sufficient data for replication available upon request from the corresponding author.
Objective:Although aggression in children and adolescents remains a leading cause of seeking mental health services, the nosology of aggression remains underdeveloped. Prior work characterized youth with aggression by using secondary analyses of large research protocol-based data sets, which revealed a profile characterized by high levels of aggression impulsive/reactive (AIR) and symptoms of hyperactivity/impulsivity. The goal of this study was to evaluate whether the AIR profile was present in a clinical sample using similar methodology with novel dimensional measures of AIR. Method:Medical records of patients 4 to 17 years of age who presented with behavior concerns to outpatient, inpatient, and emergency department settings were reviewed. Caregivers completed dimensional measures evaluating symptoms of aggression, mania, depression, and hyperactivity/impulsivity. Latent profile analysis was performed with indicators representing the same 6 content domains as in prior work: AIR, depression, mania, rule-breaking, self-harm, and hyperactivity-impulsivity. Results:A total of 430 patients completed the questionnaires and were included in the analyses. Patients' mean age was 12.52 years, and 53% were female. Four profiles were identified: predominantly AIR with hyperactivity-impulsivity (n = 83) and low mood symptoms; high AIR and high mood symptoms (n = 54); high AIR, mood, and self-harm (n = 56); and moderate overall symptom profile (n = 237). Children with high AIR and hyperactivity-impulsivity symptoms were more likely to be younger and male (mean age, 9.8 years), to have younger ages of onset of aggressive behaviors (mean age, 5.1 years), and to be less likely to have an abuse history. Conclusion:This work further validated the previously described AIR profile in a new cohort of youth presenting for clinical care. Future studies should focus on developing diagnostic criteria for children with AIR.
Objective To evaluate whether latent profile analysis (LPA) of dimensional mood symptoms can identify clinically meaningful subgroups of youth, using a measure that captures both positive (manic) and negative (depressive) valence, and to assess convergence with psychiatric diagnoses and interview-based ratings of mood symptom severity. Method We conducted a secondary analysis of data from 1369 treatment-seeking youths (ages 5–18) recruited from an urban community mental health center and an academic medical center. Consensus DSM diagnoses were derived from KSADS interviews, family history, and treatment records. LPA was applied to 20 parent-rated symptom facets from the Parent General Behavior Inventory (PGBI), modeling two continuous mood dimensions. Profiles were compared against mood diagnoses and symptom severity based on the Children's Depression Rating Scale-Revised (CDRS-R) and the Young Mania Rating Scale (YMRS), both scored via direct parent-youth interviews by trained clinicians (with PGBI scores masked). Results A 5-profile solution was supported by fit indices (BIC, entropy) and interpretability (minimum cluster size: n=123). Profile membership was strongly associated with consensus diagnoses (χ2[16]=579.39, p<.000005) and current mood state ratings (χ2[20]=583.99, p<.000005). Profiles varied significantly in rates of depressive, manic, and mixed states, and in average CDRS-R and YMRS scores, as well as global functioning and family history of bipolar disorder. Conclusion Latent profiles based on dimensional ratings of positive and negative mood symptoms aligned well with structured interview-derived diagnoses and symptom states. These profiles may offer a data-driven method to enhance mood disorder classification and support early identification efforts in youth psychiatry.
Developing accurate test norms is crucial to developmental disability practice but requires substantial resources. Continuous norms show promise relative to traditional norms, yet there is limited guidance on optimal methods and minimum sample sizes. This simulation study compared traditional norming with continuous norming models across 96 data conditions varying in sample size (500−1500), age trajectories (ages 2–22), variance patterns, and skewness. Generalized Additive Models for Location, Scale, and Shape (GAMLSS) that approximated the data generating conditions tended to be the best fitting models. Best fitting GAMLSS models showed closer fit to the data at N = 500 than traditional windowing norms at N = 1500 with the largest GAMLSS performance gains occurring by N = 750. Real-world neurobehavioral data supported the need for an iterative model selection process when developing continuous norms. Test developers can achieve high norming accuracy with smaller samples using a GAMLSS model selection process, potentially reducing costs while improving precision.
Despite encouraging evidence for the efficacy of comprehensive and intensive behavioral intervention (CIBI) programs, the majority of studies have focused on relatively narrow, deficit-focused outcomes. More specifically, although adaptive social communication and interaction (SCI) are essential for facilitative functioning, the majority of studies have utilized instruments that capture only the severity of SCI symptoms. Thus, given the importance of the comprehensive and appropriate characterization of distinct SCI adaptive skills in CIBI, in this review, based on PubMed search strategies to identify relevant published articles, we provide a critical appraisal of two of the most commonly used adaptive functioning measures—the Vineland Adaptive Behavior Scales-Third Edition (Vineland-3) and the Adaptive Behavior Assessment System-Third Edition (ABAS-3), for characterizing SCI in the behavioral intervention context. The review focused on periodic outcome and treatment planning assessment in people with autism spectrum disorder receiving CIBI programs. Instrument technical manuals were reviewed and a PubMed search was used to identify published manuscripts, with relevance to Vineland-3 and ABAS-3 development, psychometric properties, or measure interpretation. Instrument analysis begins by introducing the roles of periodic outcome assessment for CIBI programs. Next, the Vineland-3 and ABAS-3 are evaluated in terms of their development processes, psychometric characteristics, and the practical aspects of their implementation. Examination of psychometric evidence for each measure demonstrated that the evidence for several key psychometric characteristics is either unavailable or suggests less-than-desirable properties. Evaluation of practical considerations for implementation revealed weaknesses in ongoing intervention monitoring and clinical decision support. The Vineland-3 and ABAS-3 have significant strengths for cross-sectional outpatient mental health assessment, particularly as related to the identification of intellectual disability, but also substantial weaknesses relevant to their application in CIBI outcome assessment. Alternative approaches are offered, including adopting measures specifically developed for the CIBI context.
Do the shortened Positive and Negative Syndrome Scale (PANSS) (Kay et al., J Clin Psychiatry 58:538–546, 1987) versions recently developed from a National Institute of Mental Health (NIMH) pediatric dataset continue to perform well in a third independent randomized double-blind clinical trial of adolescents with schizophrenia? Secondary analysis of the double-blind, placebo-controlled aripiprazole pivotal trial data ( N = 302) found that the 10-item (and 20-item) PANSS versions on which we have previously reported (Findling et al., J Am Acad Child Adolesc Psychiatry, https://doi.org/10.1016/j.jaac.2022.07.864 , 2023) continued to provide high reliability, strong convergent correlation with expected measures, and treatment effects that equaled those found in the 30-item adult PANSS. Our shortened PANSS, derived originally from the randomized non-placebo controlled NIMH Treatment of Early Onset Schizophrenia Spectrum study (TEOSS) (Sikich et al., Am J Psychiatry 165(11):1420–1431, 2008), and independently replicated in both the placebo-controlled paliperidone pivotal trial for adolescents with schizophrenia (Youngstrom et al., PsyArxiv, https://doi.org/10.31234/osf.io/zb695 , 2023), and now the placebo-controlled aripiprazole pivotal trial for adolescents with schizophrenia, has again performed as well as the full 30 item adult-patient derived PANSS. The findings suggest it is possible to reduce the PANSS interview by 2 thirds, thus reducing burden on families and pediatric patients as well as administration and training costs, while maintaining high reliability, validity, and sensitivity to treatment equal to that of the 30-item version.
Behavioral interventions have shown substantial positive effects at the group level in improving the developmental trajectory of individuals with autism spectrum disorder (ASD), including a wide range of benefits from symptom reductions to skill development. However, there remain pronounced individual differences in the response to interventions and substantial practice variability in the choice and implementation of outcome assessments to evaluate progress for individual cases. Unfortunately, legacy outcome assessments were not specifically designed for the behavioral intervention context or for use with individuals with ASD. Furthermore, legacy instruments have been normed using traditional approaches that are often very inefficient and have limited sensitivity to divergence from neurotypical expectation. Recently, new measures, specifically designed for ASD and related neurodevelopmental conditions, have been developed and revised for use as behavioral intervention outcome assessments. To maximize the value of these measures, the present study aimed to identify optimal norming methods by comparing five distinct continuous norming models. Results indicated that more complex models that include estimation of non-linear age trends fit better and appear to provide more accurate identification of deviation from normative expectation, especially at younger ages where normative data is dense. For some symptom and skill domains, inclusion of sex-specific age-trends was necessary for best fit and most accurate performance. These findings support the use of continuous norming methods using non-linear modeling of developmental trends in the norming of outcome measures for behavioral intervention. Behavior intervention outcome assessments would benefit from implementing these norming approaches to improve the ability to detect deviation from neurotypical symptom and skill levels.