Impaired lung function in early life is associated with the subsequent development of chronic respiratory disease. Most genetic associations with lung function have been identified in adults of European descent and therefore may not represent those most relevant to pediatric populations and populations of different ancestries. In this study, we performed genome-wide association analyses of lung function in a multiethnic cohort of children (n = 1,035) living in low-income urban neighborhoods. We identified one novel locus at the TDRD9 gene in chromosome 14q32.33 associated with percent predicted forced expiratory volume in one second (FEV 1 ) (p = 2.4x10 -9 ; β z = -0.31, 95% CI = -0.41- -0.21). Mendelian randomization and mediation analyses revealed that this genetic effect on FEV 1 was partially mediated by DNA methylation levels at this locus in airway epithelial cells, which were also associated with environmental tobacco smoke exposure (p = 0.015). Promoter-enhancer interactions in airway epithelial cells revealed chromatin interaction loops between FEV 1 -associated variants in TDRD9 and the promoter region of the PPP1R13B gene, a stimulator of p53-mediated apoptosis. Expression of PPP1R13B in airway epithelial cells was significantly associated the FEV 1 risk alleles (p = 1.3x10 -5 ; β = 0.12, 95% CI = 0.06–0.17). These combined results highlight a potential novel mechanism for reduced lung function in urban youth resulting from both genetics and smoking exposure.
BACKGROUND:Viruses may drive immune mechanisms responsible for chronic rhinosinusitis with nasal polyposis (CRSwNP), but little is known about the underlying molecular mechanisms. OBJECTIVES:To identify epigenetic and transcriptional responses to a common upper respiratory pathogen, rhinovirus (RV), that are specific to patients with CRSwNP using a primary sinonasal epithelial cell culture model. METHODS:Airway epithelial cells were collected at surgery from patients with CRSwNP (cases) and from controls without sinus disease, cultured, and then exposed to RV or vehicle for 48 h. Differential gene expression and DNA methylation (DNAm) between cases and controls in response to RV were determined using linear mixed models. Weighted gene co-expression analysis (WGCNA) was used to identify (a) co-regulated gene expression and DNAm signatures, and (b) genes, pathways, and regulatory mechanisms specific to CRSwNP. RESULTS:We identified 5585 differential transcriptional and 261 DNAm responses (FDR <0.10) to RV between CRSwNP cases and controls. These differential responses formed three co-expression/co-methylation modules that were related to CRSwNP and three that were related to RV (Bonferroni corrected p < .01). Most (95%) of the differentially methylated CpGs (DMCs) were in modules related to CRSwNP, whereas the differentially expressed genes (DEGs) were more equally distributed between the CRSwNP- and RV-related modules. Genes in the CRSwNP-related modules were enriched in known CRS and/or viral response immune pathways. CONCLUSION:RV activates specific epigenetic programs and correlated transcriptional networks in the sinonasal epithelium of individuals with CRSwNP. These novel observations suggest epigenetic signatures specific to patients with CRSwNP modulate response to viral pathogens at the mucosal environmental interface. Determining how viral response pathways are involved in epithelial inflammation in CRSwNP could lead to therapeutic targets for this burdensome airway disorder.
PDF file - 197K, Table S1. Discovery set study populations. Table S2. Distribution of discovery set subjects by site, case-control status and chip. Table S3. Discovery set quality control summary. Table S4. SNPs with score-based P-values < 4 X 10-8 among 58016 SNPs for which Plink aborted the logistic regressions. Table S5. Replication set study populations. Table S6. Distribution of replication set subjects by site, case-control status and
We study the change point problem that considers alterations in the conditional distribution of an inferential target on a set of covariates. This paired data scenario is in contrast to the standard setting where a sequentially observed variable is analyzed for potential changes in the marginal distribution. We propose new methodology for solving this problem, by starting from a simpler task that analyzes changes in conditional expectation, and generalizing the tools developed for that task to conditional distributions. Large sample properties of the proposed statistics are derived. In empirical studies, we illustrate the performance of the proposed method against baselines adapted from existing tools. Two real data applications are presented to demonstrate its potential.
Background Genome-wide association studies of asthma have revealed robust associations with variation across the human leukocyte antigen (HLA) complex with independent associations in the HLA class I and class II regions for both childhood-onset asthma (COA) and adult-onset asthma (AOA). However, the specific variants and genes contributing to risk are unknown. Methods We used Bayesian approaches to perform genetic fine-mapping for COA and AOA ( n =9432 and 21,556, respectively; n =318,167 shared controls) in White British individuals from the UK Biobank and to perform expression quantitative trait locus (eQTL) fine-mapping in immune (lymphoblastoid cell lines, n =398; peripheral blood mononuclear cells, n =132) and airway (nasal epithelial cells, n =188) cells from ethnically diverse individuals. We also examined putatively causal protein coding variation from protein crystal structures and conducted replication studies in independent multi-ethnic cohorts from the UK Biobank (COA n =1686; AOA n =3666; controls n =56,063). Results Genetic fine-mapping revealed both shared and distinct causal variation between COA and AOA in the class I region but only distinct causal variation in the class II region. Both gene expression levels and amino acid variation contributed to risk. Our results from eQTL fine-mapping and amino acid visualization suggested that the HLA-DQA1 *03:01 allele and variation associated with expression of the nonclassical HLA-DQA2 and HLA-DQB2 genes accounted entirely for the most significant association with AOA in GWAS. Our studies also suggested a potentially prominent role for HLA-C protein coding variation in the class I region in COA. We replicated putatively causal variant associations in a multi-ethnic cohort. Conclusions We highlight roles for both gene expression and protein coding variation in asthma risk and identified putatively causal variation and genes in the HLA region. A convergence of genomic, transcriptional, and protein coding evidence implicates the HLA-DQA2 and HLA-DQB2 genes and HLA-DQA1 *03:01 allele in AOA.
We introduce a multi-modes tensor clustering method that implements a fused version of the alternating least squares algorithm (Fused-Orth-ALS) for simultaneous tensor factorization and clustering. The statistical convergence rates of recovery and clustering are established when the data are a noise contaminated tensor with a latent low rank CP decomposition structure. Furthermore, we show that a modified alternating least squares algorithm can provably recover the true latent low rank factorization structure when the data form an asymmetric tensor with perturbation. Clustering consistency is also established. Finally, we illustrate the accuracy and computational efficient implementation of the Fused-Orth-ALS algorithm by using both simulations and real datasets.
Review international efforts to build a global public health initiative focused on toxoplasmosis with spillover benefits to save lives, sight, cognition and motor function benefiting maternal and child health. Multiple countries’ efforts to eliminate toxoplasmosis demonstrate progress and context for this review and new work. Problems with potential solutions proposed include accessibility of accurate, inexpensive diagnostic testing, pre-natal screening and facilitating tools, missed and delayed neonatal diagnosis, restricted access, high costs, delays in obtaining medicines emergently, delayed insurance pre-approvals and high medicare copays taking considerable physician time and effort, harmful shortcuts being taken in methods to prepare medicines in settings where access is restricted, reluctance to perform ventriculoperitoneal shunts promptly when needed without recognition of potential benefit, access to resources for care, especially for marginalized populations, and limited use of recent advances in management of neurologic and retinal disease which can lead to good outcomes.
Purpose of Review Review work to create and evaluate educational materials that could serve as a primary prevention strategy to help both providers and patients in Panama, Colombia, and the USA reduce disease burden of Toxoplasma infections. Recent Findings Educational programs had not been evaluated for efficacy in Panama, USA, or Colombia. Summary Educational programs for high school students, pregnant women, medical students and professionals, scientists, and lay personnel were created. In most settings, short-term effects were evaluated. In Panama, Colombia, and USA, all materials showed short-term utility in transmitting information to learners. These educational materials can serve as a component of larger public health programs to lower disease burden from congenital toxoplasmosis. Future priorities include conducting robust longitudinal studies of whether education correlates with reduced adverse disease outcomes, modifying educational materials as new information regarding region-specific risk factors is discovered, and ensuring materials are widely accessible.
Review building of programs to eliminate Toxoplasma infections. Morbidity and mortality from toxoplasmosis led to programs in USA, Panama, and Colombia to facilitate understanding, treatment, prevention, and regional resources, incorporating student work. Studies foundational for building recent, regional approaches/programs are reviewed. Introduction provides an overview/review of programs in Panamá, the United States, and other countries. High prevalence/risk of exposure led to laws mandating testing in gestation, reporting, and development of broad-based teaching materials about Toxoplasma. These were tested for efficacy as learning tools for high-school students, pregnant women, medical students, physicians, scientists, public health officials and general public. Digitized, free, smart phone application effectively taught pregnant women about toxoplasmosis prevention. Perinatal infection care programs, identifying true regional risk factors, and point-of-care gestational screening facilitate prevention and care. When implemented fully across all demographics, such programs present opportunities to save lives, sight, and cognition with considerable spillover benefits for individuals and societies.
Review comprehensive data on rates of toxoplasmosis in Panama and Colombia. Samples and data sets from Panama and Colombia, that facilitated estimates regarding seroprevalence of antibodies to Toxoplasma and risk factors, were reviewed. Screening maps, seroprevalence maps, and risk factor mathematical models were devised based on these data. Studies in Ciudad de Panamá estimated seroprevalence at between 22 and 44%. Consistent relationships were found between higher prevalence rates and factors such as poverty and proximity to water sources. Prenatal screening rates for anti-Toxoplasma antibodies were variable, despite existence of a screening law. Heat maps showed a correlation between proximity to bodies of water and overall Toxoplasma seroprevalence. Spatial epidemiological maps and mathematical models identify specific regions that could most benefit from comprehensive, preventive healthcare campaigns related to congenital toxoplasmosis and Toxoplasma infection.
Motivated by the rising abundance of observational data with continuous treatments, we investigate the problem of estimating the average dose-response curve (ADRF). Available parametric methods are limited in their model space, and previous attempts in leveraging neural network to enhance model expressiveness relied on partitioning continuous treatment into blocks and using separate heads for each block; this however produces in practice discontinuous ADRFs. Therefore, the question of how to adapt the structure and training of neural network to estimate ADRF remains open. This paper makes two important contributions. First, we propose a novel varying coefficient neural network (VCNet) that improves model expressiveness while preserving continuity of the estimated ADRF. Second, to improve finite sample performance, we generalize targeted regularization to obtain a doubly robust estimator of the whole ADRF curve.
We consider the detection and localization of change points in the distribution of an offline sequence of observations. Based on a nonparametric framework that uses a similarity graph among observations, we propose new test statistics when at most one change point occurs and generalize them to multiple change points settings. The proposed statistics leverage edge weight information in the graphs, exhibiting substantial improvements in testing power and localization accuracy in simulations. We derive the null limiting distribution, provide accurate analytic approximations to control type I error, and establish theoretical guarantees on the power consistency under contiguous alternatives for the one change point setting, as well as the minimax localization rate. In the multiple change points setting, the asymptotic correctness of the number and location of change points are also guaranteed. The methods are illustrated on the MIT proximity network data.
Motivated by the rising abundance of observational data with continuous treatments, we investigate the problem of estimating the average dose-response curve (ADRF). Available parametric methods are limited in their model space, and previous attempts in leveraging neural network to enhance model expressiveness relied on partitioning continuous treatment into blocks and using separate heads for each block; this however produces in practice discontinuous ADRFs. Therefore, the question of how to adapt the structure and training of neural network to estimate ADRFs remains open. This paper makes two important contributions. First, we propose a novel varying coefficient neural network (VCNet) that improves model expressiveness while preserving continuity of the estimated ADRF. Second, to improve finite sample performance, we generalize targeted regularization to obtain a doubly robust estimator of the whole ADRF curve.
Additional file 10: Tables showing enrichment and colocalization results from this study. Table S1: Interaction model results for genotype x atopy and genotype x steroid use for eQTLs and meQTLs. Table S2: Enrichment estimates of eQTLs for TAGC asthma GWAS SNPs from six tissues. No P-values were significant after FDR correction. Table S3: Enrichment estimates of eQTLs for adult onset asthma GWAS SNPs from six tissues. No P-values were significant after FDR correction. Table S4: moloc results for molecular QTL-GWAS pairs and triplets. Table S5: Asthma GWAS risk allele effects on gene expression and DNA methylation.
We consider the detection and localization of gradual changes in the distribution of a sequence of time-ordered observations. Existing literature focuses mostly on the simpler abrupt setting which assumes a discontinuity jump in distribution, and is unrealistic for some applied settings. We propose a general method for detecting and localizing gradual changes that does not require any specific data generating model, any particular data type, or any prior knowledge about which features of the distribution are subject to change. Despite relaxed assumptions, the proposed method possesses proven theoretical guarantees for both detection and localization.
Rationale: Birth cohort studies have identified several temporal patterns of wheezing, only some of which are associated with asthma. Whether 17q12-21 genetic variants, which are closely associated with asthma, are also associated with childhood wheezing phenotypes remains poorly explored. Objectives: To determine whether wheezing phenotypes, defined by latent class analysis (LCA), are associated with nine 17q12-21 SNPs and if so, whether these relationships differ by race/ancestry. Methods: Data from seven U.S. birth cohorts (n = 3,786) from the CREW (Children's Respiratory Research and Environment Workgroup) were harmonized to represent whether subjects wheezed in each year of life from birth until age 11 years. LCA was then performed to identify wheeze phenotypes. Genetic associations between SNPs and wheeze phenotypes were assessed separately in European American (EA) (n = 1,308) and, for the first time, in African American (AA) (n = 620) children. Measurements and Main Results: The LCA best supported four latent classes of wheeze: infrequent, transient, late-onset, and persistent. Odds of belonging to any of the three wheezing classes (vs. infrequent) increased with the risk alleles for multiple SNPs in EA children. Only one SNP, rs2305480, showed increased odds of belonging to any wheezing class in both AA and EA children. Conclusions: These results indicate that 17q12-21 is a "wheezing locus," and this association may reflect an early life susceptibility to respiratory virusescommon to all wheezing children. Which children will have their symptoms remit or reoccur during childhood may be independent of the influence of rs2305480.
Background Genome-wide association studies (GWASs) have identified thousands of variants associated with asthma and other complex diseases. However, the functional effects of most of these variants are unknown. Moreover, GWASs do not provide context-specific information on cell types or environmental factors that affect specific disease risks and outcomes. To address these limitations, we used an upper airway epithelial cell (AEC) culture model to assess transcriptional and epigenetic responses to rhinovirus (RV), an asthma-promoting pathogen, and provide context-specific functional annotations to variants discovered in GWASs of asthma. Methods Genome-wide genetic, gene expression, and DNA methylation data in vehicle- and RV-treated upper AECs were collected from 104 individuals who had a diagnosis of airway disease (n=66) or were healthy participants (n=38). We mapped cis expression and methylation quantitative trait loci (cis-eQTLs and cis-meQTLs, respectively) in each treatment condition (RV and vehicle) in AECs from these individuals. A Bayesian test for colocalization between AEC molecular QTLs and adult onset asthma and childhood onset asthma GWAS SNPs, and a multi-ethnic GWAS of asthma, was used to assign the function to variants associated with asthma. We used Mendelian randomization to demonstrate DNA methylation effects on gene expression at asthma colocalized loci. Results Asthma and allergic disease-associated GWAS SNPs were specifically enriched among molecular QTLs in AECs, but not in GWASs from non-immune diseases, and in AEC eQTLs, but not among eQTLs from other tissues. Colocalization analyses of AEC QTLs with asthma GWAS variants revealed potential molecular mechanisms of asthma, including QTLs at the TSLP locus that were common to both the RV and vehicle treatments and to both childhood onset and adult onset asthma, as well as QTLs at the 17q12-21 asthma locus that were specific to RV exposure and childhood onset asthma, consistent with clinical and epidemiological studies of these loci. Conclusions This study provides evidence of functional effects for asthma risk variants in AECs and insight into RV-mediated transcriptional and epigenetic response mechanisms that modulate genetic effects in the airway and risk for asthma.
In many transcriptomic studies, the correlation of genes might fluctuate with quantitative factors such as genetic ancestry. We propose a method that models the covariance between two variables to vary against a continuous covariate. For the bivariate case, the proposed score test statistic is computationally simple and robust to model misspecification of the covariance term. Subsequently, the method is expanded to test relationships between one highly connected gene, such as a transcription factor, and several other genes for a more global investigation of the dynamic of the coexpression network. Simulations show that the proposed method has higher statistical power than alternatives, can be used in more diverse scenarios, and is computationally cheaper. We apply this method to African American subjects from GTEx to analyze the dynamic behavior of their gene coexpression against genetic ancestry and to identify transcription factors whose coexpression with their target genes change with the genetic ancestry. The proposed method can be applied to a wide array of problems that require covariance modeling.
This paper outlines a framework for quantifying the prior's contribution to posterior inference in the presence of prior-likelihood discordance, a broader concept than the usual notion of prior-likelihood conflict. We achieve this dual purpose by extending the classic notion of prior sample size, M, in three directions: (I) estimating M beyond conjugate families; (II) formulating M as a relative notion that is as a function of the likelihood sample size k, M(k), which also leads naturally to a graphical diagnosis; and (III) permitting negative M, as a measure of prior-likelihood conflict, that is, harmful discordance. Our asymptotic regime permits the prior sample size to grow with the likelihood data size, hence making asymptotic arguments meaningful for investigating the impact of the prior relative to that of likelihood. It leads to a simple asymptotic formula for quantifying the impact of a proper prior that only involves computing a centrality and a spread measure of the prior and the posterior. We use simulated and real data to illustrate the potential of the proposed framework, including quantifying how weak is a 'weakly informative' prior adopted in a study of lupus nephritis. Whereas we take a pragmatic perspective in assessing the impact of a prior on a given inference problem under a specific evaluative metric, we also touch upon conceptual and theoretical issues such as using improper priors and permitting priors with asymptotically non-vanishing influence.