Myeloid cells, including microglia and perivascular macrophages, are central to Alzheimer's disease (AD) neurobiology, yet their role remains incompletely understood. We profiled 832,505 human myeloid cells from the prefrontal cortex of 1,607 donors spanning the lifespan and showing varying degrees of AD neuropathology. We delineated six subclasses comprising 13 transcriptionally distinct subtypes and identified adaptive changes associated with aging and AD progression. Here we show that a disease-associated microglial subtype, characterized by elevated GPNMB expression and enriched for polygenic AD risk, expands with AD pathology and shows increased phagocytic activity. We identify MITF as an upstream regulator required to maintain this microglial state. Cell-cell interaction analyses prioritize APOE-SORL1 and APOE-TREM2 signaling pairs associated with disease progression. Using human and mouse models, we demonstrate that the neuroprotective effects of this microglial subtype depend on TREM2. These findings provide mechanistic insights into myeloid cell function in aging and AD, aiding therapeutic discovery.
Microglia are resident immune cells of the brain and are implicated in the etiology of Alzheimer's disease (AD) and other diseases. Yet the cellular and molecular processes regulating their function throughout the course of the disease are poorly understood. Here, we present a transcriptional analysis of primary microglia from 189 human postmortem brains, including 58 healthy aging individuals and 131 with a range of disease phenotypes, such as 63 patients representing the full clinical and pathological spectra of AD. We identified changes associated with multiple AD phenotypes, capturing the severity of dementia and neuropathological lesions. Transcript-level analyses identified additional genes with heterogeneous isoform usage and AD phenotypes. We identified changes in gene-gene coordination in AD, dysregulation of coexpression modules and disease subtypes with distinct gene expression patterns. Taken together, these data further our understanding of the key role that microglia have in AD biology and nominate candidates for therapeutic intervention.
Noncoding variants increase neuropsychiatric disease risk, but our understanding of their cell-type-specific role remains incomplete. We conducted large-scale chromatin accessibility profiling of neurons and non-neurons from 2 neocortical regions in 1,393 libraries. We observed substantial differences in neuronal chromatin accessibility between schizophrenia (SCZ) cases and controls, with upregulated open chromatin regions (OCRs) in neurons associated with SCZ risk loci. A comparison of SCZ-associated OCRs with fetal brain-specific OCRs revealed a strong correlation between upregulated changes in SCZ chromatin and openness in fetal cortical brains, linking disease-related chromatin alterations to neurodevelopment. Here we show that a prominent neuronal trans-regulatory domain containing upregulated OCRs consolidates key neurodevelopmental chromatin signatures and is enriched for immature glutamatergic neurons. These findings link altered adult cortical chromatin states to early developmental mechanisms in SCZ. This study provides a comprehensive cell-type-resolved chromatin accessibility resource for the human cortex and offers insights into the regulatory architecture underlying SCZ risk.
Neurodegenerative diseases and serious mental illnesses often exhibit overlapping characteristics, highlighting the potential for shared underlying mechanisms. To facilitate a deeper understanding of these diseases and pave the way for more effective treatments, we have generated a population-scale multi-omics dataset consisting of genotype and single-nucleus transcriptome data from the prefrontal cortex of frozen human brain specimens. Encompassing over 6.3 million nuclei from 1,494 donors, our dataset represents a diverse range of neurodegenerative and serious mental illnesses, including Alzheimer's and Parkinson's diseases, schizophrenia, bipolar disorder and diffuse Lewy body dementia, as well as neurotypical controls. Our dataset offers a unique opportunity to study disease interactions, as 21% of donors had comorbid diagnoses of two or more major brain disorders. Additionally, it includes detailed phenotypic information on neuropsychiatric symptoms, such as apathy and weight loss, which commonly accompany Alzheimer's disease and related dementias. We have performed stringent preprocessing and quality controls, ensuring the reliability and usability of the data. As a commitment to fostering collaborative research, we provide this valuable resource as an online repository, enabling widespread analyses across the scientific community.
Genetic risk variants for common diseases are predominantly located in non-coding regulatory regions and modulate gene expression. Although bulk tissue studies have elucidated shared mechanisms of regulatory and disease-associated genetics, the cellular specificity of these mechanisms remains largely unexplored. This study presents a comprehensive single-nucleus multi-ancestry atlas of genetic regulation of gene expression in the human prefrontal cortex, comprising 5.6 million nuclei from 1,384 donors of diverse ancestries. Through multi-resolution analyses spanning eight major cell classes and 27 subclasses, we identify genetic regulation for 14,258 genes, with 857 showing cell type-specific regulatory effects at the class level and 981 at the subclass level. Colocalization of genetic variants associated with gene regulation and disease traits uncovers novel cell type-specific genes implicated in Alzheimer's disease, schizophrenia, and other disorders, which were not detectable in bulk tissue analyses. Analysis of dynamic genetic regulation at the single nucleus level identifies 2,073 genes with regulatory effects that vary across developmental trajectories, inferred from a broad age range of donors. We also uncover 1,655 genes with trans-regulatory effects, revealing distal regulation of gene expression. This high-resolution atlas provides unprecedented insight into the cell type-specific regulatory architecture of the human brain, and offers novel mechanistic targets for understanding the genetic basis of neuropsychiatric and neurodegenerative diseases.
Neurodegenerative and neuropsychiatric diseases impose a significant societal and public health burden. However, our understanding of the molecular mechanisms underlying these highly complex conditions remains limited. To gain deeper insights into the etiology of different brain diseases, we used specimens from 1,494 unique donors to generate a population-scale single-cell transcriptomic atlas of the human dorsolateral prefrontal cortex (DLPFC), comprising over 6.3 million individual nuclei. The cohort includes neurotypical controls as well as donors affected by eight common and complex brain disorders: Alzheimer's disease (AD), diffuse Lewy body disease (DLBD), vascular dementia (Vas), Parkinson's disease (PD), tauopathy, frontotemporal dementia, schizophrenia, and bipolar disorder. We show that inter-individual variation accounts for a substantial portion of gene expression variation in the DLPFC. By comparing transcriptomic variation across diseases, we reveal universal signatures enriched in basic cellular functions such as mRNA splicing and protein localization. After discounting these cross-disease signatures, we show strong genetic and transcriptomic concordance among AD, DLBD, Vas, and PD, largely driven by alteration of synaptic signaling functions in neurons. Furthermore, we characterize transcriptomic variation among different AD phenotypes that were distinct from healthy aging. We uncover mitigating effects of interneurons and aggravating effects of immune and vascular cells in AD dementia. Further exploring the effect of the neuropsychiatric symptoms frequently accompanying AD, we identify a link to deep layer excitatory neurons. By constructing transcriptome trajectories that capture AD progression, we show cell-type specific responses implicated in early and late stages of AD. Our atlas provides an unprecedented perspective of the transcriptomic landscape in neurodegenerative and neuropsychiatric diseases, shedding light on shared and distinct processes involving the neuro-immune-vascular systems, and identifying potential targets for therapeutic intervention.
Neuropsychiatric and neurodegenerative disorders exhibit cell–type–specific characteristics 1–8, yet most transcriptome–wide association studies have been constrained by the use of homogenate brain tissue9–11, limiting their resolution and power. Here, we present a single–nucleus transcriptome–wide association study (snTWAS) leveraging single–nucleus RNA sequencing of over 6 million nuclei from the dorsolateral prefrontal cortex of 1,494 donors across three ancestries–European, African, and Admixed American. We constructed ancestry–specific single–nucleus–derived transcriptomic imputation models (snTIMs) including up to 27 non–overlapping cellular populations, enhancing the resolution of genetically regulated gene expression (GReX) in the brain and uncovering novel gene–trait associations across 12 neuropsychiatric and neurodegenerative traits. Our snTWAS framework revealed cell–type–specific dysregulation of GReX, identifying over 4,000 novel gene–trait associations not detected by bulk tissue approaches. By applying these snTIMs to the Million Veteran Program, we validated major findings and explored the pleiotropy of cell–type–specific GReX, revealing cross–ancestry concordance and fine–mapping causal genes. This approach enhances the discovery of biologically relevant pathways and gene targets, highlighting the importance of cell–type resolution and ancestry–specific models in understanding the genetic architecture of complex brain disorders. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research is based on data from the Million Veteran Program, Office of Research and Development, Veterans Health Administration, and was supported by award I01BX004189. This publication does not represent the views of the Department of Veteran Affairs or the United States Government. We thank the participants of the Million Veteran Program, the scientists, clinicians and supportive staff involved in the construction of this biobank, and the scientific computing staff for the expertise that they provided. We thank the computational resources and staff expertise provided by the Scientific Computing at the Icahn School of Medicine at Mount Sinai. This study was also supported by the National Institutes of Health (NIH), Bethesda, MD under award numbers R01AG067025 (PR), R01AG082185 (PR), K08MH122911 (GV), R01AG078657 (GV), BX004189 (PR), R01AG065582 (PR), R01AG067025 (PR), R01MH125246 (PR) and T32MH087004 (KT). Human tissues were obtained from the NIH NeuroBioBank at the Mount Sinai Brain Bank (MSSM; supported by NIMH-75N95019C00049), the Rush Alzheimer's Disease Center (RADC; funding: P30AG10161, P30AG72975, R01AG15819, R01AG17917, R01AG22018, U01AG46152, and U01AG61356), and NIMH-IRP Human Brain Collection Core (HBCC, project # ZIC MH002903). This work was supported in part through the computational and data resources and staff expertise provided by Scientific Computing and Data at the Icahn School of Medicine at Mount Sinai and supported by the Clinical and Translational Science Award (CTSA) grant UL1TR004419 from the National Center for Advancing Translational Sciences. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved by the VA Central Institutional Review Board (IRB), and participating studies received approval from their respective IRBs. All participants provided written informed consent. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All results are included either in the main text or provided in supplementary tables or data.
Parkinson's Disease (PD) is a debilitating neurodegenerative disorder, characterized by motor and cognitive impairments, that affects >1% of the population over the age of 60. The pathogenesis of PD is complex and remains largely unknown. Due to the cellular heterogeneity of the human brain and changes in cell type composition with disease progression, this complexity cannot be fully captured with bulk tissue studies. To address this, we generated single-nucleus RNA sequencing and whole-genome sequencing data from 100 postmortem cases and controls, carefully selected to represent the entire spectrum of PD neuropathological severity and diverse clinical symptoms. The single nucleus data were generated from five brain regions, capturing the subcortical and cortical spread of PD pathology. Rigorous preprocessing and quality control were applied to ensure data reliability. Committed to collaborative research and open science, this dataset is available on the AMP PD Knowledge Platform, offering researchers a valuable tool to explore the molecular bases of PD and accelerate advances in understanding and treating the disease.
With the advent of healthcare-based genotyped biobanks, genome-wide association studies (GWAS) leverage larger sample sizes, incorporate patients with diverse ancestries and introduce noisier phenotypic definitions. Yet the extent and impact of phenotypic misclassification on large-scale datasets is not currently well understood due to a lack of statistical methods to estimate relevant parameters from empirical data. Here, we develop a statistical method and scalable software, PheMED, Phenotypic Measurement of Effective Dilution, to quantify phenotypic misclassification across GWAS using only summary statistics. We illustrate how the parameters estimated by PheMED relate to the negative and positive predictive value of the labeled phenotype, compared to ground truth, and how misclassification of the phenotype yields diluted effect-sizes of variant-phenotype associations. Furthermore, we apply our methodology to detect multiple instances of statistically significant dilution in real-world data. We demonstrate how effective dilution biases downstream GWAS replication and heritability analyses despite utilizing current best practices, and provide a dilution-aware meta-analysis approach that outperforms existing methods. Consequently, we anticipate that PheMED will be a valuable tool for researchers to address phenotypic data quality issues both within and across cohorts.
Binge eating disorder (BED) is the most common eating disorder, yet its genetic architecture remains largely unknown. Studying BED is challenging because it is often comorbid with obesity, a common and highly polygenic trait, and it is underdiagnosed in biobank data sets. To address this limitation, we apply a supervised machine-learning approach (using 822 cases of individuals diagnosed with BED) to estimate the probability of each individual having BED based on electronic medical records from the Million Veteran Program. We perform a genome-wide association study of individuals of African ( n = 77,574) and European ( n = 285,138) ancestry while controlling for body mass index to identify three independent loci near the HFE , MCHR2 and LRP11 genes and suggest APOE as a risk gene for BED. We identify shared heritability between BED and several neuropsychiatric traits, and implicate iron metabolism in the pathophysiology of BED. Overall, our findings provide insights into the genetics underlying BED and suggest directions for future translational research.
Non-coding variants increase risk of neuropsychiatric disease. However, our understanding of the cell-type specific role of the non-coding genome in disease is incomplete. We performed population scale (N=1,393) chromatin accessibility profiling of neurons and non-neurons from two neocortical brain regions: the anterior cingulate cortex and dorsolateral prefrontal cortex. Across both regions, we observed notable differences in neuronal chromatin accessibility between schizophrenia cases and controls. A per-sample disease pseudotime was positively associated with genetic liability for schizophrenia. Organizing chromatin into cis- and trans-regulatory domains, identified a prominent neuronal trans-regulatory domain (TRD1) active in immature glutamatergic neurons during fetal development. Polygenic risk score analysis using genetic variants within chromatin accessibility of TRD1 successfully predicted susceptibility to schizophrenia in the Million Veteran Program cohort. Overall, we present the most extensive resource to date of chromatin accessibility in the human cortex, yielding insights into the cell-type specific etiology of schizophrenia.
Our understanding of the genetic basis of autism spectrum disorder (ASD) is advancing rapidly. The number of ASD-associated, linkage disequilibrium independent risk loci is constantly growing due to ongoing PGC efforts. The majority of the common variants reside within non-coding regions of the genome and, as such, the formulation of testable hypotheses to elucidate their potential function in the brain is challenging. Integration of genetic findings with large-scale genotype-tissue brain expression efforts have enabled the execution of independent transcriptome-wide association studies (TWAS); however, heritability mediated by gene expression as assayed in homogenate tissues with RNA-seq is only estimated to account for about 10% of the total heritability. Recent advances in cell type specific transcriptomic and epigenetic profiling of the human brain are highlighting the unrealized potential of integrating cell-type specific transcriptomic and epigenetic features, such as chromatin accessibility and histone modifications, towards elucidating the mechanisms of cis-regulation of brain gene expression and causal gene prioritization. Here, we are leveraging large-scale high resolution multi-omics datasets derived from brain homogenate, FANS/FACS-sorted cells/nuclei and single-nucleus data to increase our mechanistic understanding of the genetic liability for autism spectrum disorder as follows: 1) We prioritize putatively causal SNPs associated with changes in transcriptomes and chromatin accessibility. 2) We quantify heritability mediated by brain gene expression. 3) We perform a transcriptome-wide association study to identify ASD-associated transcripts and biological pathways. 4) We perform epigenome-based prioritization of relevant cell types and tissues. 5) We identify cell subtypes exhibiting high expression across ASD-associated genes. Finally, we are exploring associations of aggregated genetic liability for ASD across complex cognition domains by leveraging several genotyped cohorts.
We describe the Predicting Protein-Compound Interactions (PrePCI) database which comprises over 5 billion predicted interactions between 6.8 million chemical compounds and 19,797 human proteins. PrePCI relies on a proteome-wide database of structural models based on both traditional modeling techniques and the AlphaFold Protein Structure Database. Sequence- and structural similarity-based metrics are established between template proteins, T, in the Protein Data Bank that bind compounds, C, and query proteins in the model database, Q. When the metrics exceed threshold values, it is assumed that C also binds to Q with a likelihood ratio (LR) derived from machine learning. If the relationship is based on structural similarity, the LR is based on a scoring function that measures the extent to which C is compatible with the binding site of Q as described in the LT-scanner algorithm. For every predicted complex derived in this way, chemical similarity based on the Tanimoto coefficient identifies other small molecules that may bind to Q. An overall LR for the binding of C to Q is obtained from Naive Bayesian statistics. The PrePCI database can be queried by entering a UniProt ID or gene name for a protein to obtain a list of compounds predicted to bind to it along with associated LRs. Alternatively, entering an identifier for the compound outputs a list of proteins it is predicted to bind. Specific applications of the database to lead discovery, elucidation of drug mechanism of action, and biological function annotation are described.
Binge-eating disorder (BED) is the most common eating disorder yet its genetic architecture remains largely unknown. Studying BED is challenging because it is often comorbid with obesity, a common and highly polygenic trait, and it is underdiagnosed in biobank datasets. To address this limitation, we apply a supervised machine learning approach to estimate the probability of each individual having BED based on electronic medical records from the Million Veteran Program. We perform a genome-wide association study on individuals of African (n = 77,574) and European (n = 285,138) ancestry while controlling for body mass index to identify three independent loci near the HFE, MCHR2 and LRP11 genes, which are reproducible across three independent cohorts. We identify genetic association between BED and several neuropsychiatric traits and implicate iron metabolism in the pathophysiology of BED. Overall, our findings provide insights into the genetics underlying BED and suggest directions for future translational research.
Nanostructures generated by self-assembly of peptides yield nanomaterials that have many therapeutic applications, including drug delivery and biomedical engineering, due to their low cytotoxicity and higher uptake by targeted cells owing to their high affinity and specificity towards cell surface receptors. Despite the promising implications of this rapidly expanding field, there is no dedicated resource to study peptide nanostructures. This study endeavours to create a repository of short peptides, which may prove to be the best models to study ordered nanostructures formed by peptide self-assembly. SAPdb has a repertoire of 1049 entries of experimentally validated nanostructures formed by the self-assembly of small peptides. It consists of 328 tripeptides, 701 dipeptides, and 20 single amino acids with some conjugate partners. Each entry encompasses comprehensive information about the peptide, such as chemical modifications, the type of nanostructure formed, experimental conditions like pH, temperature, solvent required for the self-assembly, etc. Our analysis indicates that peptides containing aromatic amino acids favour the formation of self-assembling nanostructures. Additionally, we observed that these peptides form different nanostructures under different experimental conditions. SAPdb provides this comprehensive information in a hassle-free tabulated manner at a glance. User-friendly browsing, searching, and analysis modules have been integrated for easy data retrieval, data comparison, and examination of properties. We anticipate SAPdb to be a valuable repository for researchers engaged in the burgeoning arena of nanobiotechnology. It is freely available at https://webs.iiitd.edu.in/raghava/sapdb.
This paper describes a web server developed for designing therapeutic peptides with desired half-life in blood. In this study, we used 163 natural and 98 modified peptides whose half-life has been determined experimentally in mammalian blood, for developing in silico models. Firstly, models have been developed on 261 peptides containing natural and modified residues, using different chemical descriptors. The best model using 43 PaDEL descriptors got a maximum correlation of 0.692 between the predicted and the actual half-life peptides. Secondly, models were developed on 163 natural peptides using amino acid composition feature of peptides and achieved a maximum correlation of 0.643. Thirdly, models were developed on 163 natural peptides using chemical descriptors and attained a maximum correlation of 0.743 using 45 selected PaDEL descriptors. In order to assist researchers in the prediction and designing of half-life of peptides, the models developed have been integrated into PlifePred web server.
TopicalPdb (http://crdd.osdd.net/raghava/topicalpdb/) is a repository of experimentally verified topically delivered peptides. Data was manually collected from research articles. The current release of TopicalPdb consists of 657 entries, which includes peptides delivered through the skin (462 entries), eye (173 entries), and nose (22 entries). Each entry provides comprehensive information related to these peptides like the source of origin, nature of peptide, length, N- and C-terminal modifications, mechanism of penetration, type of assays, cargo and biological properties of peptides, etc. In addition to natural peptides, TopicalPdb contains information of peptides having non-natural, chemically modified residues and D-amino acids. Besides this primary information, TopicalPdb stores predicted tertiary structures as well as peptide sequences in SMILE format. Tertiary structures of peptides were predicted using state-of-art method PEPstrMod. In order to assist users, a number of web-based tools have been integrated that includes keyword search, data browsing, similarity search and structural similarity. We believe that TopicalPdb is a unique database of its kind and it will be very useful in designing peptides for non-invasive topical delivery.
Numerous therapeutic peptides do not enter the clinical trials just because of their high hemolytic activity. Recently, we developed a database, Hemolytik, for maintaining experimentally validated hemolytic and non-hemolytic peptides. The present study describes a web server and mobile app developed for predicting, and screening of peptides having hemolytic potency. Firstly, we generated a dataset HemoPI-1 that contains 552 hemolytic peptides extracted from Hemolytik database and 552 random non-hemolytic peptides (from Swiss-Prot). The sequence analysis of these peptides revealed that certain residues (e.g., L, K, F, W) and motifs (e.g., "FKK", "LKL", "KKLL", "KWK", "VLK", "CYCR", "CRR", "RFC", "RRR", "LKKL") are more abundant in hemolytic peptides. Therefore, we developed models for discriminating hemolytic and non-hemolytic peptides using various machine learning techniques and achieved more than 95% accuracy. We also developed models for discriminating peptides having high and low hemolytic potential on different datasets called HemoPI-2 and HemoPI-3. In order to serve the scientific community, we developed a web server, mobile app and JAVA-based standalone software (http://crdd.osdd.net/raghava/hemopi/).