Improvements in organization, time management, and planning (OTMP) skills have downstream effects on student academic performance. Brief, practical assessment tools are needed to identify students needing OTMP intervention and monitor response to intervention. The purpose of this study was to develop and validate brief rating scales of OTMP skills, derived from the Children’s Organizational Skills Scale parent and teacher versions (COSS-P, COSS-T), for screening and progress monitoring. Three samples were used: (1) general and clinical sample (n for caregivers = 1155; n for teachers = 1139; Abikoff Gallagher, 2009); (2) sample from the randomized controlled trial (RCT) evaluating Organizational Skills Training-Tier 2 (OST-T2; n = 185; Nissley-Tsiopinis et al., 2024); (3) sample from the RCT evaluating Homework, Organization, and Planning Skills (HOPS; n = 163; Langberg et al., 2018). Based on indices of discrimination and item effect sizes, we identified 10 items from the COSS-P/T representing multiple components of OTMP interventions, in addition to one item assessing OTMP interference with performance. Confirmatory factor analyses demonstrated strong fit to a bifactor model. Only the general factor demonstrated acceptable reliability. The total score demonstrated convergent validity with measures of school and executive functioning and was excellent in differentiating children with and without substantial, impairing OTMP deficits (Area Under Curve [AUC] = .947 for COSS-P, .982 for COSS-T). Treatment sensitivity for the screeners and three, five-item formative behavior assessment measures, computed as effect sizes (Cohen’s d) based on response to OST-T2 and HOPS, all exceeded −1.20. The brief tools developed in this study have substantial utility for screening and progress monitoring.
Patients with obstructive sleep apnea (OSA) frequently seek information online, yet the comparative quality of content delivered by web search engines versus generative AI systems is unclear. This study evaluated how different digital information sources perform in answering common patient questions about OSA. Thirty high-volume, patient-facing OSA questions were identified using Google Trends. Each question was submitted verbatim to four general-purpose large language models (GPT-4, GPT-5, DeepSeek, Mistral), a medically specialized retrieval-augmented model (OpenEvidence), and Google Search. Seven otolaryngologists with clinical experience in OSA independently rated each response for accuracy, clarity, completeness, relevance, and usefulness using a five-point rubric. Composite and domain scores were analyzed using one-way analysis of variance with multiple-comparison correction; inter-rater reliability was assessed with two-way random-effects intraclass correlation coefficients. A total of 180 question–system pairs received 6295 domain-level ratings. OpenEvidence achieved the highest mean composite score (4.33), followed by a tightly clustered group of LLMs (means 4.00–4.04). Google Search scored significantly lower (3.15). Differences among systems were statistically significant across all domains (p < 0.001), with large effect sizes for comparisons of OpenEvidence and general LLMs versus Google. Composite average-rater reliability was good (ICC = 0.70). For common OSA questions, generative AI systems—particularly a retrieval-augmented medical model—produced higher-quality patient-facing information than standard web search. These findings support cautious consideration of GenAI tools to supplement patient education in OSA, while underscoring the need for ongoing evaluation across diseases, disciplines, and patient populations. Patients with obstructive sleep apnea (OSA) frequently rely on online sources such as Google Search to understand symptoms, testing, and treatment, yet the quality of patient-facing information varies widely. As generative artificial intelligence tools are increasingly used for health questions, their comparative performance for OSA education has not been systematically evaluated using blinded expert review. In this blinded comparative study, generative AI systems, particularly a retrieval-augmented medical model, provided more accurate, clear, complete, and useful answers to common OSA questions than standard web search. These findings highlight that the choice of digital information source can meaningfully influence the quality of patient education in sleep medicine and support further evaluation of AI tools within clinical practice.
PURPOSE:Interactions with police officers can be a significant source of stress and anxiety, particularly for vulnerable and/or misunderstood populations. Here, we report a parallel randomized controlled clinical trial designed to examine and compare the effects of two different intervention programs designed to support autistic individuals as they prepare to interact with police officers. METHODS:Forty-seven autistic participants, aged 12 to 60 years, were randomized to participate in either the Floreo Police Safety Module virtual reality (VR) intervention or the BeSAFE The Movie video modeling intervention. For both conditions, three 45-min intervention sessions were completed an average of 9 days apart. RESULTS:Results revealed a significant intervention type by time interaction on fidgeting behavior during live interactions with police officers, indicating a reduction from pre-training to post-training that was specific to the VR intervention condition (estimate: 0.47, SE: 0.16, t = 2.95, p = 0.005). The statistical interaction between intervention type and time was not significant for participants' responding in accordance with expectations (p = .07) or overall behavior during live interactions with police officers (p = .23), but planned follow-up tests revealed improvements for both responding in accordance with expectations (estimate: -.21, SE: .07, t = -3.14, p = .02) and overall behavior (estimate: -.29, SE: .10, t = -3.04, p = .02) in the VR intervention condition specifically. CONCLUSIONS:Overall, these findings provide empirical support for practice as a means to empower autistic people to prepare for police interactions, with evidence suggesting that practicing police interactions using VR may be especially effective for driving skill acquisition.
Artificial intelligence (AI) for surgical workflow analysis often fails to generalize because surgical actions lack a standardized, fine-grained representation. Gesture-level “tokenization” of surgery, capturing instrument–tissue interactions as the smallest intentional functional units, offers greater technical specificity than phase- or step-level labels and has demonstrated associations with proficiency and clinical outcomes. However, the field remains fragmented by heterogeneous gesture terminology, limiting dataset interoperability and model reproducibility. We conducted a SAGES-led, accelerated Delphi consensus process to establish a standardized surgical gesture taxonomy. Starting with 270 literature-derived gesture terms, we employed a novel hybrid pipeline combining large language model (LLM)-assisted semantic clustering with multi-round expert review. The process involved two Delphi surveys (open-ended, then structured agreement) with a predefined ≥ 80
The optimal timing for surgical intervention in single-suture craniosynostosis (SSC) remains debated, despite advances in minimally invasive and open reconstructive approaches. This work integrates Children’s National Medical Center (CNMC) institutional data with current literature to clarify age-related outcomes and guide timing recommendations. Published series from CMNC were analyzed and outcomes included perioperative morbidity, morphometric correction, intracranial pressure, and revision rates. Findings were contextualized with recent systematic reviews and multicenter data. Early endoscopic repair (≤ 4 months) yielded superior morphometric gains, lower blood loss, and shorter hospital stays. O’Brien et al. demonstrated a 2–4-month “sweet spot” for sagittal synostosis1, while Lajthia et al. confirmed excellent outcomes for metopic deformities in the same age range2. Open reconstruction at 9–12 months achieved durable aesthetic correction with low complication rates3. Delayed presentations were associated with elevated intracranial pressure but benefited from surgical decompression4. Meta-analyses corroborate these trends. CNMC’s experience and global evidence converge on an early-infancy window (2–4 months) as optimal for endoscopic repair. Open cranial vault reconstruction remains effective for older infants or complex anatomy. Surgical timing should balance biological potential, institutional resources, and neurodevelopmental opportunity.