ABSTRACTLimited ancestral diversity has impaired our ability to detect risk variants more prevalent in non-European ancestry groups in genome-wide association studies (GWAS). We constructed and analyzed a multi-ancestry GWAS dataset in the Alzheimer’s Disease (AD) Genetics Consortium (ADGC) to test for novel shared and ancestry-specific AD susceptibility loci and evaluate underlying genetic architecture in 37,382 non-Hispanic White (NHW), 6,728 African American, 8,899 Hispanic (HIS), and 3,232 East Asian individuals, performing within-ancestry fixed-effects meta-analysis followed by a cross-ancestry random-effects meta-analysis. We identified 13 loci with cross-ancestry associations including known loci at/nearCR1,BIN1,TREM2,CD2AP,PTK2B,CLU,SHARPIN,MS4A6A,PICALM,ABCA7,APOEand two novel loci not previously reported at 11p12 (LRRC4C) and 12q24.13 (LHX5-AS1). Reflecting the power of diverse ancestry in GWAS, we observed theSHARPINlocus using 7.1% the sample size of the original discovering single-ancestry GWAS (n=788,989). We additionally identified three GWS ancestry-specific loci at/near (PTPRK(P=2.4×10-8) andGRB14(P=1.7×10-8) in HIS), andKIAA0825(P=2.9×10-8in NHW). Pathway analysis implicated multiple amyloid regulation pathways (strongest withPadjusted=1.6×10-4) and the classical complement pathway (Padjusted=1.3×10-3). Genes at/near our novel loci have known roles in neuronal development (LRRC4C, LHX5-AS1, andPTPRK) and insulin receptor activity regulation (GRB14). These findings provide compelling support for using traditionally-underrepresented populations for gene discovery, even with smaller sample sizes.
This chapter describes the mathematical properties that have been developed since to support the GraphBLAS. It describes the key mathematical concepts of the GraphBLAS and presents preliminary results that show the overhead of the GraphBLAS is minimal. Matrix multiplication is the most important matrix operation and can be used to implement a wide range of graph algorithms. One of the most common uses of matrix multiplication is to construct an adjacency matrix from an incidence matrix representation of a graph. The GraphBLAS performance of sparse matrix sparse vector multiplication is similar toGunrock BFS performance. The similarity in performance indicates that the GraphBLAS is not introducing a high overhead. The GraphBLAS allows these matrix properties to be readily applied to graphs in a low-overhead manner.
Combinatorial algorithms such as those that arise in graph analysis, modeling of discrete systems, bioinformatics, and chemistry, are often hard to parallelize. The Combinatorial BLAS library implements key computational primitives for rapid development of combinatorial algorithms in distributed-memory systems. During the decade since its first introduction, the Combinatorial BLAS library has evolved and expanded significantly. This article details many of the key technical features of Combinatorial BLAS version 2.0, such as communication avoidance, hierarchical parallelism via in-node multithreading, accelerator support via GPU kernels, generalized semiring support, implementations of key data structures and functions, and scalable distributed I/O operations for human-readable files. Our article also presents several rules of thumb for choosing the right data structures and functions in Combinatorial BLAS 2.0, under various common application scenarios.
Each additional copy of the apolipoprotein E4 (APOE4) allele is associated with a higher risk of Alzheimer's dementia, while the APOE2 allele is associated with a lower risk of Alzheimer's dementia, it is not yet known whether APOE2 homozygotes have a particularly low risk. We generated Alzheimer's dementia odds ratios and other findings in more than 5,000 clinically characterized and neuropathologically characterized Alzheimer's dementia cases and controls. APOE2/2 was associated with a low Alzheimer's dementia odds ratios compared to APOE2/3 and 3/3, and an exceptionally low odds ratio compared to APOE4/4, and the impact of APOE2 and APOE4 gene dose was significantly greater in the neuropathologically confirmed group than in more than 24,000 neuropathologically unconfirmed cases and controls. Finding and targeting the factors by which APOE and its variants influence Alzheimer's disease could have a major impact on the understanding, treatment and prevention of the disease.
Whole exome sequencing and copy‐number variant analysis was performed on a family with three brothers diagnosed with autism. Each of the siblings shares an alteration in the nuclear receptor subfamily 3 group C member 2 (NR3C2) gene that is predicted to result in a stop‐gain mutation (p.Q919X) in the mineralocorticoid receptor (MR) protein. This variant was maternally inherited and provides further evidence for a connection between the NR3C2 and autism. Interestingly, the NR3C2 gene encodes the MR protein, a steroid hormone‐regulated transcription factor that acts in the hypothalamic–pituitary–adrenal axis and has been connected to stress and anxiety, both of which are features often seen in individuals with autism. Autism Res 2020, 13: 523–531. © 2020 International Society for Autism Research, Wiley Periodicals, Inc.Lay SummaryGiven the complexity of the genetics underlying autism, each gene contributes to risk in a relatively small number of individuals, typically less than 1% of all autism cases. Whole exome sequencing of three brothers with autism identified a rare variant in the nuclear receptor subfamily 3 group C member 2 gene that is predicted to strongly interfere with its normal function. This gene encodes the mineralocorticoid receptor protein, which plays a role in how the body responds to stress and anxiety, features that are often elevated in people diagnosed with autism. This study adds further support to the relevance of this gene as a risk factor for autism.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Numerical computational science dominated the first half century of high- performance computing; graph theory served numerical linear algebra by enabling efficient sparse matrix methods. Turnabout is fair play: Nowadays more and more computational problems concern graphs in their own right, and sparse matrix methods are often a good way to look at algorithms on graphs. This has led via a long path to the Graph BLAS and its reference implementations, which are a significant milestone. But there’s a lot left to do. What happens now?
Alterations of the gamma-aminobutyric acid (GABA) signaling system has been strongly linked to the pathophysiology of autism spectrum disorder (ASD). Genetic associations of common variants in GABA receptor subunits, in particular GABRA4 on chromosome 4p12, with ASD have been replicated by several studies. Moreover, molecular investigations have identified altered transcriptional and translational levels of this gene and protein in brains of ASD individuals. Since the genotyped common variants are likely not the functional variants contributing to the molecular consequences or underlying ASD phenotype, this study aims to examine rare sequence variants in GABRA4, including those outside the protein coding regions of the gene. We comprehensively re-sequenced the entire protein coding and noncoding portions of the gene and putative regulatory sequences in 82 ASD individuals and 55 developmentally typical pediatric controls, all homozygous for the most significant previously associated ASD risk allele (G/G at rs1912960). We identified only a single common, coding variant, and no association of any single marker or set of variants with ASD. Functional annotation of noncoding variants identified several rare variants in putative regulatory sites. Finally, a rare variant unique to ASD cases, in an evolutionary conserved site of the 3'UTR, shows a trend toward decreasing gene expression. Hence, GABRA4 rare variants in noncoding DNA may be variants of modest physiological effects in ASD etiology.
We seek to discover Laplacian linear systems that stress the ability of existing Laplacian solver packages to solve them efficiently. We employ a genetic algorithm to explore the problem space of graphs of fixed size and edge density. The goal is to measure the gap between theoretical and existing Laplacian solvers, by trying to find worst case example graphs for existing solvers. These problems may have little use inside any real world application, but they give great insight into solver behavior. We report performance results of our genetic algorithm, and explore the properties of the evolved graphs.
Sven van der Lee, Julie Williams, Gerard Schellenberg and colleagues identify rare coding variants in PLCG2, ABI3 and TREM2 associated with Alzheimer's disease. These genes are highly expressed in microglia and provide additional evidence that the microglia-mediated immune response contributes to the development of Alzheimer's disease. We identified rare coding variants associated with Alzheimer's disease in a three-stage case–control study of 85,133 subjects. In stage 1, we genotyped 34,174 samples using a whole-exome microarray. In stage 2, we tested associated variants (P < 1 × 10−4) in 35,962 independent samples using de novo genotyping and imputed genotypes. In stage 3, we used an additional 14,997 samples to test the most significant stage 2 associations (P < 5 × 10−8) using imputed genotypes. We observed three new genome-wide significant nonsynonymous variants associated with Alzheimer's disease: a protective variant in PLCG2 (rs72824905: p.Pro522Arg, P = 5.38 × 10−10, odds ratio (OR) = 0.68, minor allele frequency (MAF)cases = 0.0059, MAFcontrols = 0.0093), a risk variant in ABI3 (rs616338: p.Ser209Phe, P = 4.56 × 10−10, OR = 1.43, MAFcases = 0.011, MAFcontrols = 0.008), and a new genome-wide significant variant in TREM2 (rs143332484: p.Arg62His, P = 1.55 × 10−14, OR = 1.67, MAFcases = 0.0143, MAFcontrols = 0.0089), a known susceptibility gene for Alzheimer's disease. These protein-altering changes are in genes highly expressed in microglia and highlight an immune-related protein–protein interaction network enriched for previously identified risk genes in Alzheimer's disease. These genetic findings provide additional evidence that the microglia-mediated innate immune response contributes directly to the development of Alzheimer's disease.
Importance Mutations in APP, PSEN1, and PSEN2 lead to early-onset Alzheimer disease (EOAD) but account for only approximately 11% of EOAD overall, leaving most of the genetic risk for the most severe form of Alzheimer disease unexplained. This extreme phenotype likely harbors highly penetrant risk variants, making it primed for discovery of novel risk genes and pathways for AD. Objective To search for rare variants contributing to the risk for EOAD. Design, Setting, and Participants In this case-control study, whole-exome sequencing (WES) was performed in 51 non-Hispanic white (NHW) patients with EOAD (age at onset <65 years) and 19 Caribbean Hispanic families previously screened as negative for established APP, PSEN1, and PSEN2 causal variants. Participants were recruited from John P. Hussman Institute for Human Genomics, Case Western Reserve University, and Columbia University. Rare, deleterious, nonsynonymous, or loss-of-function variants were filtered to identify variants in known and suspected AD genes, variants in multiple unrelated NHW patients, variants present in 19 Hispanic EOAD WES families, and genes with variants in multiple unrelated NHW patients. These variants/genes were tested for association in an independent cohort of 1524 patients with EOAD, 7046 patients with late-onset AD (LOAD), and 7001 cognitively intact controls (age at examination, >65 years) from the Alzheimer’s Disease Genetics Consortium. The study was conducted from January 21, 2013, to October 13, 2016. Main Outcomes and Measures Alzheimer disease diagnosed according to standard National Institute of Neurological and Communicative Disorders and Stroke and the Alzheimer Disease and Related Disorders Association criteria. Association between Alzheimer disease and genetic variants and genes was measured using logistic regression and sequence kernel association test–optimal gene tests, respectively. Results Of the 1524 NHW patients with EOAD, 765 (50.2%) were women and mean (SD) age was 60.0 (4.9) years; of the 7046 NHW patients with LOAD, 4171 (59.2%) were women and mean (SD) age was 77.4 (8.6) years; and of the 7001 NHW controls, 4215 (60.2%) were women and mean (SD) age was 77.4 (8.6) years. The gene PSD2, for which multiple unrelated NHW cases had rare missense variants, was significantly associated with EOAD (P = 2.05 × 10−6; Bonferroni-corrected P value [BP] = 1.3 × 10−3) and LOAD (P = 6.22 × 10−6; BP = 4.1 × 10−3). A missense variant in TCIRG1, present in a NHW patient and segregating in 3 cases of a Hispanic family, was more frequent in EOAD cases (odds ratio [OR], 2.13; 95% CI, 0.99-4.55; P = .06; BP = 0.413), and significantly associated with LOAD (OR, 2.23; 95% CI, 1.37-3.62; P = 7.2 × 10−4; BP = 5.0 × 10−3). A missense variant in the LOAD risk gene RIN3 showed suggestive evidence of association with EOAD after Bonferroni correction (OR, 4.56; 95% CI, 1.26-16.48; P = .02, BP = 0.091). In addition, a missense variant in RUFY1 identified in 2 NHW EOAD cases showed suggestive evidence of an association with EOAD as well (OR, 18.63; 95% CI, 1.62-213.45; P = .003; BP = 0.129). Conclusions and Relevance The genes PSD2, TCIRG1, RIN3, and RUFY1 all may be involved in endolysosomal transport—a process known to be important to development of AD. Furthermore, this study identified shared risk genes between EOAD and LOAD similar to previously reported genes, such as SORL1, PSEN2, and TREM2.
Solving Laplacian linear systems is an important task in a variety of practical and theoretical applications. This problem is known to have solutions that perform in linear times polylogarithmic work in theory, but these algorithms are difficult to implement in practice. We examine existing solution techniques in order to determine the best methods currently available and for which types of problems are they useful. We perform timing experiments using a variety of solvers on a variety of problems and present our results. We discover differing solver behavior between web graphs and a class of synthetic graphs designed to model them.
We study the performance of linear solvers for graph Laplacians based on the combinatorial cycle adjustment methodology proposed by [Kelner-Orecchia-Sidford-Zhu STOC-13]. The approach finds a dual flow solution to this linear system through a sequence of flow adjustments along cycles. We study both data structure oriented and recursive methods for handling these adjustments. The primary difficulty faced by this approach, updating and querying long cycles, motivated us to study an important special case: instances where all cycles are formed by fundamental cycles on a length $n$ path. Our methods demonstrate significant speedups over previous implementations, and are competitive with standard numerical routines.
Text: Background: Mutations in APP, PSEN1 and PSEN2lead to familial early-onset Alzheimer disease (EOAD). These mutations account for ~11% of EOAD overall, leaving the majority of genetic risk for the most severe form of AD unexplained. Methods: We performed whole-exome sequencing (WES) on 53 Non-Hispanic White EOAD cases screened negative for APP, PSEN1, and PSEN2 to search for rare EOAD risk variants. Variant filtering for missense and loss-of-function (LOF) rare variants (MAF<0.1%) was performed, and filtered variants present on the Illumina exome chip array were tested in a cohort of 1,292 EOAD cases (Age-at-onset<65) and 5,625 controls (Age≥65) from the Alzheimer’s Disease Genetics Consortium (ADGC). As no LOF filtered variants were on the chip, we assessed these variants by using both the lowest variant P-value per gene and gene-based SKAT-O results. Rare LOF variants were prioritized for testing if they were in 2+ cases and in haploinsufficient genes (Haploinsufficient Score≥50). Results: 1,803 rare missense variants identified in the WES sample were available for testing in in the exome chip study. 70 of these variants were nominally associated with EOAD, with the strongest signals in the neuropeptide NPPB (OR=11.24, P=0.003), and NDST2 (OR=14.95, P=0.001), a processor of heparin sulfate, a molecule with potential importance in Aβ formation. Assessment of the 70 nominally significant genes containing these variants for consistent dysregulation (3+ expression studies) in AD using AlzBase prioritized 11 genes with strong potential for involvement in EOAD, including the APP interector ADRA1A (OR=3.60, P=0.032). One variant in the LOAD GWAS gene RIN3 (OR=6.94, P=0.009) also was of interest. Testing of 30 LOF genes revealed several associations that approached significance including a splicing variant in CACNA1G (OR=4.93, P=0.002), a gene associated with cognitive decline and age-related production of Aβ, a stopgain in CENPF (OR=11.46, P=0.004), a gene involved in endolysosomal transport and amyloid plaque generation, and a frameshift in the gene TDP2 (Gene P=0.006), for which homozygous LOF mutations cause neurodegeneration with epilepsy. Conclusions: Testing of candidate WES EOAD risk variants and genes in the ADGC EOAD exome chip study identified several genes with potential roles in AD pathogenesis.
A new method for solving Laplacian linear systems proposed by Kelner et al. involves the random sampling and update of fundamental cycles in a graph. Kelner et al. proved asymptotic bounds on the complexity of this method but did not report experimental results. We seek to both evaluate the performance of this approach and to explore improvements to it in practice. We compare the performance of this method to other Laplacian solvers on a variety of real world graphs. We consider different ways to improve the performance of this method by exploring different ways of choosing the set of cycles and the sequence of updates, with the goal of providing more flexibility and potential parallelism. We propose a parallel model of the Kelner et al. method, for evaluating potential parallelism in terms of the span of edges updated at each iteration. We provide experimental results comparing the potential parallelism of the fundamental cycle basis and our extended cycle set. Our preliminary experiments show that choosing a non-fundamental set of cycles can save significant work compared to a fundamental cycle basis.
BACKGROUND:Essential tremor is a neurological condition characterized by tremor during voluntary movement. To date, 3 loci linked to familial essential tremor have been identified. METHODS:We examined 48 essential tremor patients in 5 large essential tremor pedigrees in our data set for genetic linkage using an Affymetrix Axiom array. Linkage analysis was performed using an affecteds-only dominant model in SIMWALK2. To incorporate all genotype information, GERMLINE was used to identify genome segments shared identical-by-descent in pairs of affecteds. Exome sequencing was performed in pedigrees showing evidence of linkage. RESULTS:For one family, chromosomes 5 and 18 showed genome-wide significant linkage to essential tremor. Shared segment analysis excluded the 18p11 candidate region and reduced the 5q35 region by 1 megabase. Exome sequencing did not identify a potential causative variant in this region. CONCLUSION:A locus on chromosome 5 is linked to essential tremor. Further research is needed to identify a causative variant. © 2016 International Parkinson and Movement Disorder Society.
Objective: To identify a causative variant(s) that may contribute to Alzheimer disease (AD) in African Americans (AA) in the ATP-binding cassette, subfamily A (ABC1), member 7 (ABCA7) gene, a known risk factor for late-onset AD. Methods: Custom capture sequencing was performed on ∼150 kb encompassing ABCA7 in 40 AA cases and 37 AA controls carrying the AA risk allele (rs115550680). Association testing was performed for an ABCA7 deletion identified in large AA data sets (discovery n = 1,068; replication n = 1,749) and whole exome sequencing of Caribbean Hispanic (CH) AD families. Results: A 44-base pair deletion (rs142076058) was identified in all 77 risk genotype carriers, which shows that the deletion is in high linkage disequilibrium with the risk allele. The deletion was assessed in a large data set (531 cases and 527 controls) and, after adjustments for age, sex, and APOE status, was significantly associated with disease (p = 0.0002, odds ratio [OR] = 2.13 [95% confidence interval (CI): 1.42–3.20]). An independent data set replicated the association (447 cases and 880 controls, p = 0.0117, OR = 1.65 [95% CI: 1.12–2.44]), and joint analysis increased the significance (p = 1.414 × 10−5, OR = 1.81 [95% CI: 1.38–2.37]). The deletion is common in AA cases (15.2%) and AA controls (9.74%), but in only 0.12% of our non-Hispanic white cohort. Whole exome sequencing of multiplex, CH families identified the deletion cosegregating with disease in a large sibship. The deleted allele produces a stable, detectable RNA strand and is predicted to result in a frameshift mutation (p.Arg578Alafs) that could interfere with protein function. Conclusions: This common ABCA7 deletion could represent an ethnic-specific pathogenic alteration in AD.
Recommender system data presents unique challenges to the data mining, machine learning, and algorithms communities. The high missing data rate, in combination with the large scale and high dimensionality that is typical of recommender systems data, requires new tools and methods for efficient data analysis. Here, we address the challenge of evaluating similarity between two users in a recommender system, where for each user only a small set of ratings is available. We present a new similarity score, that we call LiRa, based on a statistical model of user similarity, for large-scale, discrete valued data with many missing values. We show that this score, based on a ratio of likelihoods, is more effective at identifying similar users than traditional similarity scores in user-based collaborative filtering, such as the Pearson correlation coefficient. We argue that our approach has significant potential to improve both accuracy and scalability in collaborative filtering.
The GraphBLAS standard (GraphBlas.org) is being developed to bring the potential of matrix based graph algorithms to the broadest possible audience. Mathematically the Graph- BLAS defines a core set of matrix-based graph operations that can be used to implement a wide class of graph algorithms in a wide range of programming environments. This paper provides an introduction to the mathematics of the GraphBLAS. Graphs represent connections between vertices with edges. Matrices can represent a wide range of graphs using adjacency matrices or incidence matrices. Adjacency matrices are often easier to analyze while incidence matrices are often better for representing data. Fortunately, the two are easily connected by matrix mul- tiplication. A key feature of matrix mathematics is that a very small number of matrix operations can be used to manipulate a very wide range of graphs. This composability of small number of operations is the foundation of the GraphBLAS. A standard such as the GraphBLAS can only be effective if it has low performance overhead. Performance measurements of prototype GraphBLAS implementations indicate that the overhead is low.
Arthur B. Maccabe合作论文数Oak Ridge National Laboratory18