This cross-sectional study examines trends in US patient portal messaging, telehealth and telephone encounters, and office visits, assessing differences in messaging by various patient sociodemographic characteristics.
Privacy is a human right that sustains patient-provider trust. Clinical notes capture a patient's private vulnerability and individuality, which are used for care coordination and research. Under HIPAA Safe Harbor, these notes are de-identified to protect patient privacy. However, Safe Harbor was designed for an era of categorical tabular data, focusing on the removal of explicit identifiers while ignoring the latent information found in correlations between identity and quasi-identifiers, which can be captured by modern LLMs. We first formalize these correlations using a causal graph, then validate it empirically through individual re-identification of patients from scrubbed notes. The paradox of de-identification is further shown through a diagnosis ablation: even when all other information is removed, the model can predict the patient's neighborhood based on diagnosis alone. This position paper raises the question of how we can act as a community to uphold patient-provider trust when de-identification is inherently imperfect. We aim to raise awareness and discuss actionable recommendations.
Background and Objectives:Are some brain regions intrinsically more vulnerable to metastatic colonization? We sought to characterize the spatial distribution of brain metastases and determine whether regional patterns vary according to primary tumor origin. Methods:We analyzed baseline MRI scans and expert tumor segmentations from 906 patients with 3,492 brain metastases treated with stereotactic radiosurgery. Lesions were normalized to MNI152 standard space and superimposed to generate probabilistic atlases of metastatic occurrence. Regional metastatic burden was quantified using anatomical and vascular atlases. Spatial distributions were additionally compared between lung cancer and melanoma metastases. Results:Metastatic burden was distributed nonuniformly throughout the brain. The cerebellum demonstrated the strongest enrichment relative to its anatomical volume (fold change 1.61, p < 0.001), accompanied by overrepresentation of the vertebrobasilar circulation (fold change 1.49, p < 0.001). Spatial distribution also varied by primary tumor type. Lung cancer metastases demonstrated greater infratentorial involvement than melanoma metastases (16.6% vs. 8.7%, p < 0.05), with a corresponding increase in cerebellar burden (14.8% vs. 6.8%, p < 0.05), whereas melanoma metastases were relatively concentrated within the frontal lobe (37.7% vs. 24.6%, p < 0.01). Infratentorial enrichment was observed across all carcinoma subgroups, with the greatest enrichment seen in gastrointestinal metastases (32.9% infratentorial). Conclusion:Brain metastases exhibit nonrandom spatial distributions, with preferential involvement of posterior and infratentorial structures. Regional patterns vary according to primary tumor origin, supporting the existence of region-specific vulnerability to metastatic disease.
BACKGROUND AND OBJECTIVES: Myxopapillary ependymomas (MPE) and intradural lumbosacral schwannomas may be challenging to distinguish based on presenting characteristics and preoperative imaging. Accurate differentiation is crucial, as MPEs carry a risk of cerebrospinal fluid dissemination and warrant earlier intervention, a more tailored surgical strategy, consideration for adjuvant radiation, and frequent surveillance. Here, we describe our institutional experience with these tumors and develop a radiomics-based machine learning model to help distinguish them on preoperative imaging. METHODS: Institutional surgical records from 2011 to 2025 were queried and clinical data were extracted for the retrospective cohort analysis. Tumors were manually segmented in ITK-Snap from T1 postcontrast images, and radiomics features were extracted using the PyRadiomics package. An ensemble of random forest, k-nearest neighbors, and naive Bayes classifiers was trained on a subset of radiomics features using nested cross-validation. RESULTS: Our cohort included 101 cases, including 32 MPEs, 61 intradural schwannomas, and 8 dumbbell schwannomas with a circumscribed intradural component. Twenty-four consecutive tumors (3 MPEs and 21 schwannomas) were used as a held-out pseudoprospective test set. No significant difference in presenting International Standards for Neurological Classification of Spinal Cord Injury grade was observed ( P = .558). Our radiomics model incorporated segmentations with inter-rater DICE scores >0.83 for all tumors. Cross-validation (area under the receiver operating characteristic = 0.895) and test (area under the receiver operating characteristic = 0.984) discriminatory performance was high and identified biologically relevant features, such as maximum 3-dimensional diameter, that were significantly increased ( P < .001) in MPEs, likely due to longitudinal tumor growth along the filum. Excluding scoliotic patients did not significantly alter discrimination, suggesting robustness to vertebral column malalignment that may coexist with intradural tumors. CONCLUSION: A radiomics-based machine learning model demonstrated excellent discriminative ability between MPE and lumbosacral schwannoma, achieving high accuracy and robustness to vertebral alignment variations. These results suggest that radiomics-based models may be developed into a useful tool for preoperative planning and patient counseling.
Clinical evaluations of large language models (LLMs) have rapidly expanded since 2022, yet their evidence base remains opaque. The overwhelming volume of studies creates challenges for manual curation and review. However, LLMs themselves offer the scalability and capability to evaluate the ever-growing evidence base. This LLM-assisted review identified 4,609 peer-reviewed studies in clinical medicine between January 2022 and September 2025, equating to roughly 3.2 papers per day. Only 1,048 studies used real-world patient data and of these only 19 were prospective randomized trials; most addressed simulated scenarios (n = 1,857) or exam-style tasks (n = 1,704). ChatGPT and related OpenAI models constitute 65.7% of evaluated models, with Gemini/Bard a distant second constituting 13.1% of evaluated models. Patient-facing communication and education comprised 17% of tasks, followed by knowledge retrieval, and education and assessment simulation. Across 1,046 head-to-head comparisons, LLMs outperformed humans in 33% of comparisons, with a strong dependency on task realism and level of training. At least 25% of studies had sample sizes less than 30. Despite the growth of LLMs in medicine, rigorous, patient-centered evidence remains scarce, underscoring the need for larger prospective trials before clinical adoption.
The ability to quickly learn and generalize is one of the brain's most impressive feats and recreating it remains a major challenge for modern artificial intelligence research. One of the most mysterious one-shot learning abilities displayed by humans is one-shot perceptual learning, whereby a single viewing experience drastically alters visual perception in a long-lasting manner. Where in the brain one-shot perceptual learning occurs and what mechanisms support it remain enigmatic. Combining psychophysics, 7 T fMRI, and intracranial recordings, we identify the high-level visual cortex as the most likely neural substrate wherein neural plasticity supports one-shot perceptual learning. We further develop a deep neural network model incorporating top-down feedback into a vision transformer, which recapitulates and predicts human behavior. The prior knowledge learnt by this model is highly similar to the neural code in the human high-level visual cortex. These results reveal the neurocomputational mechanisms underlying one-shot perceptual learning in humans.
OBJECTIVE:Spinal meningiomas (SMs) are common primary spinal tumors for which surgery is considered the first-line treatment when safe and feasible. The ability to extrapolate the tumor grade from preoperative imaging may significantly inform early patient expectation-setting regarding recurrence. Building on radiomics studies in cranial meningiomas, the authors aimed to construct a benchmark radiomics model to preoperatively identify the histological grade of SMs. METHODS:Institutional surgical records from May 2012 to November 2025 were queried for pathology-confirmed meningiomas below the foramen magnum, with preoperative contrast-enhanced imaging available for segmentation. SMs were classified as low-grade (WHO grade 1) and high-grade (WHO grade 2 tumors and grade 1 tumors with atypia). Tumors were manually segmented, and features were extracted using the PyRadiomics software package. An ensemble model of k-nearest neighbors, random forest, and support vector machine classifiers was trained using nested cross-validation on a subset of 10 features to differentiate tumor grades. Clinical data for the cohort were also extracted, and disease control in an adjunctive clinical series was assessed. RESULTS:Seventy-four patients were included in radiomics analysis, with an area under the receiver operating characteristic curve of 0.879 and a mean F1 score of 0.748. The model's top 5 features were all texture features that differed significantly (p < 0.05) across low- and high-grade SMs. These included measures of tumor textural and contrast-enhancement heterogeneity, with overlap with features reported in radiomics models for histological grading of intracranial meningiomas. Fifty-five patients with a median radiographic follow-up of 22.2 (range 1.9-86.4) months remained for clinical analysis after exclusion of patients with less than 1 month of follow-up and syndromic meningiomas. Four recurrences occurred at a median of 20.8 (range 1.8-41.8) months. High-grade tumor pathology did not significantly impact progression-free survival (p = 0.682, log-rank test; Cox regression high vs low grade hazard ratio [HR] 0.62, 95% CI 0.06-6.11, p = 0.685). Subtotal resection was associated with poorer progression-free survival than gross-total resection (p = 0.004, log-rank test; Cox regression subtotal vs gross-total resection HR 10.62, 95% CI 1.46-77.05, p = 0.019). These findings remain contextualized within a relatively limited follow-up window and small recurrence event count, suggesting a need to characterize the interplay between tumor grade and extent of resection as drivers of local disease control in SMs. CONCLUSIONS:A preoperative radiomics model can stratify high-grade SMs using open-source tools applied to single-institution data.
BACKGROUND AND OBJECTIVES:Molecular markers such as isocitrate dehydrogenase (IDH) and alpha-thalassemia/mental retardation syndrome X-linked (ATRX) status are essential for glioma classification and treatment planning, but their manual extraction from pathology reports creates significant research bottlenecks. This study evaluated 3 Natural Language Processing approaches with increasing computational complexity: deterministic Regular Expressions (RegEx), statistical Term Frequency-Inverse Document Frequency (TF-IDF) with logistic regression, and contextual deep learning Bidirectional Encoder Representations from Transformers (BERT). We address whether more intensive approaches provide sufficient performance benefits over simpler approaches in computational pathology research. METHODS:We analyzed pathology reports from 404 patients with glioma at Institution A and 197 at Institution B for external validation. IDH analysis included 399 (Institution A) and 193 (Institution B) patients; ATRX analysis included 361 and 130 patients, respectively. All approaches underwent identical preprocessing steps, including text normalization, terminology standardization, and context extraction. Performance was evaluated using standard classification metrics and memory usage benchmarks on internal and external validation data sets. RESULTS:Simpler approaches outperformed more intensive approaches on external validation. For IDH, Regex achieved near-perfect accuracy (99%, area under the curve [AUC] 1.000) and TF-IDF performed exceptionally (94.2%, AUC 0.984), while BlueBERT underperformed (85.2%, AUC 0.934). For ATRX, Regex achieved perfect accuracy (100%, AUC 1.000) and TF-IDF maintained high accuracy (98.0%, AUC 0.998), outperforming BERT-large (84.6%, AUC 0.931). BERT-based approaches required 1825-1953 MB of memory vs Regex (0.82-5.52 MB) and TF-IDF (17.27-34.89 MB). CONCLUSION:Simple Natural Language Processing approaches effectively automate molecular marker extraction from pathology reports with near-perfect accuracy while requiring minimal computational resources. This enables expanded sample sizes in retrospective studies, multi-institutional analyses of rare molecular subgroups, and accelerated biomarker research. Future work will focus on validation across larger data sets, infrastructure integration, and expansion to additional molecular markers.
Despite recent advances, we hypothesized that women remain underrepresented as authors in neurosurgical literature–a key indicator of success in academic medicine. We queried Web of Science for all contents published in the top 15 neurosurgical journals by h5-index. Articles and metadata were exported. Analyses were performed using Python packages (Gender-Guesser and Wiki-Gendersort for gender identification). 143,713 articles from 112,932 unique authors with full names were identified (62.0% (69,985/112,932) male, 24.0% (27,148/112,932) female, 14.0% (15,799/112,932) unknown). Female co-authorship was limited before 1980, with an increase in the 2000’s (7.0% in 2000 to 60.7% in 2023). However, female first and last authorship have yet to exceed 17.6% (2023) and 11.2% (2022), respectively. This is consistent across the top 3 neurosurgical journals. Compared to male last authors, female last authors are 3.45 times more likely to publish with female first authors (95% CI 3.27-3.64). Female first authors are also 3.73 times more likely than male first authors to publish with female last authors (95% CI 3.53-3.93). Of the 10,000 most productive authors, male authors have a higher h-index than female authors (p<0.001). However, female authors have a higher m-index (p<0.03), which is normalized to impact over time in the field, suggesting an emerging contingent of female authors who are academically productive and authoring impactful works. While progress has been made in female co-authorship, female first and last authorship remain below 20%. Female last authors are more likely to be associated with female first authors and vice versa. This finding suggests that women are paving the way towards a more equitable future of neurosurgery by mentoring other women, and that efforts to support female senior authors can also benefit female trainees.
Artificial intelligence makes strides in specialized diagnostics but faces challenges in complex clinical scenarios, such as rare disease diagnosis and emergency condition identification. To address these limitations, we develop Meta General Practitioner (MetaGP), a 32-billion-parameter generative foundation model trained on extensive datasets, including over 8 million electronic health records, biomedical literature, and medical textbooks. MetaGP demonstrates robust diagnostic capabilities, achieving accuracy comparable to experienced clinicians. In rare disease cases, it achieves an average diagnostic score of 1.57, surpassing GPT-4's 0.93. For emergency conditions, it improves diagnostic accuracy for junior and mid-level clinicians by 53% and 46%, respectively. MetaGP also excels in generating medical imaging reports, producing high-quality outputs for chest X-rays and computed tomography, often rated comparable to or superior to physician-authored reports. These findings highlight MetaGP's potential to transform clinical decision-making across diverse medical contexts.
The mentorship and training relationships within the field of neurosurgery are pivotal in shaping the careers and influence of neurosurgeons. Understanding these relationships through network analysis can reveal key individuals and structural characteristics that drive the field's development. A combined multi-directed and undirected network graph comprising neurosurgical training relationships was constructed, including edges between chair and trainee, program director and trainee, and peers from the same residency graduating within 5 years of one another. Graph level metrics were calculated and communities were detected via modularity maximization. Centrality metrics were calculated for a undirected subgraph of neurosurgical department chairs. The neurosurgical training network consisted of 8443 nodes and 259,445 edges, with a density of 0.0036, average degree of 61.5, 338 weakly connected components and 530 strongly connected components. We detected 394 communities (mean size 21.4 ± 59.9 members), with the 30 largest having 209.9 ± 69.1 members. Most of the members from the largest communities were from the same residency (60.28% ± 21.16). Harvey Cushing was dominant in closeness, betweenness and local-reaching centrality, highlighting his role as the quintessential neurosurgeon-influencer. Wilder Penfield had the top PageRank, demonstrating the breadth of his direct connections with other influential individuals. Our analysis reveals a sparse but highly connected network, typical of large real-world professional networks. This indicates that while there are many small groups of neurosurgeons who are closely interconnected, each is connected to a small fraction of the total. This suggests most neurosurgeons may not interact outside their immediate training hierarchy or peer group. The central positions and high influence of various chairmen suggest they have been pivotal in mentoring and connecting many neurosurgeons, contributing to the network's overall structure and dynamics. Future analyses will seek insights into the present-day importance of other key figures in neurosurgery.
Surgical site infections (SSI) can be devastating, yet optimal preparation techniques have not been clearly determined in neurosurgical patients. Furthermore, the profile and origin of pathogens that ultimately go on to cause SSI remain unclear. This is a prospective single-center study including adult patients undergoing elective procedures. Cultures are taken before and after preparation with providone iodine (PVI), chlorhexidine gluconate (CHG), or a combination based on surgeon preference, and again from bone after exposure. Primary outcome is SSI requiring return to the OR. Secondary outcomes are the effects of preparation on surgical site cultures, and the correlation between cultures and ultimate infectious pathogens. The study includes 846 patients undergoing open procedures. 692 cases used CHG, 96 PVI, and 58 both. 11 SSIs required return to the OR with an infection rate of 1.3%. 6 had used CHG (0.9%), 5 PVI (5.2%) during the index surgery. 67 additional patients underwent endoscopic endonasal procedures with PVI prep and no infections. The infection rate for PVI was significantly higher than CHG (p = 0.01), even among matched procedures (p = 0.03). Among the index surgeries with complete surgical site culture data, the causative organism was never present. We found that pathogens present in the first culture but not the second were sometimes present in the third cultures, consistent with the hypothesis that there is a bacterial reservoir in hair follicles inaccessible to skin prep. CHG reduces the rate of SSI. The pathogens that go on to cause infection are not present on the skin at the time of surgery, and thus that causative organisms are introduced into the wound at a later time. This raises important questions about the optimal care of neurosurgical wounds postoperatively.
We present Head CT Ontology Normalized Evaluation (HeadCT-ONE), a metric for evaluating head CT report generation through ontology-normalized entity and relation extraction. HeadCT-ONE enhances current information extraction derived metrics (such as RadGraph F1) by implementing entity normalization through domain-specific ontologies, addressing radiological language variability. HeadCT-ONE compares normalized entities and relations, allowing for controllable weighting of different entity types or specific entities. Through experiments on head CT reports from three health systems, we show that HeadCT-ONE's normalization and weighting approach improves the capture of semantically equivalent reports, better distinguishes between normal and abnormal reports, and aligns with radiologists' assessment of clinically significant errors, while offering flexibility to prioritize specific aspects of report content. Our results demonstrate how HeadCT-ONE enables more flexible, controllable, and granular automated evaluation of head CT reports.