Recent advances in genome sequencing have improved variant calling in complex regions of the human genome. However, it is difficult to quantify variant calling performance because existing standards often focus on specificity, neglecting completeness in difficult-to-analyze regions. To create a more comprehensive truth set, we used Mendelian inheritance in a large pedigree (CEPH-1463) to filter variants across PacBio high-fidelity (HiFi), Illumina and Oxford Nanopore Technologies platforms. This generated a variant map with over 4.7 million single-nucleotide variants, 767,795 insertions and deletions (indels), 537,486 tandem repeats and 24,315 structural variants, covering 2.77 Gb of the GRCh38 genome. This work adds 200 Mb of high-confidence regions, including 8
Each human genome has tens of thousands of rare genetic variants; however, identifying impactful rare variants remains a major challenge. We demonstrate how use of personal multi-omics can enable identification of impactful rare variants by using the Multi-Ethnic Study of Atherosclerosis, which included several hundred individuals, with whole-genome sequencing, transcriptomes, methylomes, and proteomes collected across two time points, 10 years apart. We evaluated each multi-omics phenotype's ability to separately and jointly inform functional rare variation. By combining expression and protein data, we observed rare stop variants 62 times and rare frameshift variants 216 times as frequently as controls, compared to 13-27 times as frequently for expression or protein effects alone. We extended a Bayesian hierarchical model, "Watershed," to prioritize specific rare variants underlying multi-omics signals across the regulatory cascade. With this approach, we identified rare variants that exhibited large effect sizes on multiple complex traits including height, schizophrenia, and Alzheimer's disease.
Integrative approaches that simultaneously model multi-omics data have gained increasing popularity because they provide holistic system biology views of multiple or all components in a biological system of interest. Canonical correlation analysis (CCA) is a correlation-based integrative method designed to extract latent features shared between multiple assays by finding the linear combinations of features–referred to as canonical variables (CVs)–within each assay that achieve maximal across-assay correlation. Although widely acknowledged as a powerful approach for multi-omics data, CCA has not been systematically applied to multi-omics data in large cohort studies, which has only recently become available. Here, we adapted sparse multiple CCA (SMCCA), a widely-used derivative of CCA, to proteomics and methylomics data from the Multi-Ethnic Study of Atherosclerosis (MESA) and Jackson Heart Study (JHS). To tackle challenges encountered when applying SMCCA to MESA and JHS, our adaptations include the incorporation of the Gram-Schmidt (GS) algorithm with SMCCA to improve orthogonality among CVs, and the development of Sparse Supervised Multiple CCA (SSMCCA) to allow supervised integration analysis for more than two assays. Effective application of SMCCA to the two real datasets reveals important findings. Applying our SMCCA-GS to MESA and JHS, we identified strong associations between blood cell counts and protein abundance, suggesting that adjustment of blood cell composition should be considered in protein-based association studies. Importantly, CVs obtained from two independent cohorts also demonstrate transferability across the cohorts. For example, proteomic CVs learned from JHS, when transferred to MESA, explain similar amounts of blood cell count phenotypic variance in MESA, explaining 39.0% ~ 50.0% variation in JHS and 38.9% ~ 49.1% in MESA. Similar transferability was observed for other omics-CV-trait pairs. This suggests that biologically meaningful and cohort-agnostic variation is captured by CVs. We anticipate that applying our SMCCA-GS and SSMCCA on various cohorts would help identify cohort-agnostic biologically meaningful relationships between multi-omics data and phenotypic traits.
BACKGROUND:Chronic obstructive pulmonary disease (COPD) varies significantly in symptomatic and physiologic presentation. Identifying disease subtypes from molecular data, collected from easily accessible blood samples, can help stratify patients and guide disease management and treatment. METHODS:Blood gene expression measured by RNA-sequencing in the COPDGene Study was analyzed using a network perturbation analysis method. Each COPD sample was compared against a learned reference gene network to determine the part that is deregulated. Gene deregulation values were used to cluster the disease samples. RESULTS:The discovery set included 617 former smokers from COPDGene. Four distinct gene network subtypes are identified with significant differences in symptoms, exercise capacity and mortality. These clusters do not necessarily correspond with the levels of lung function impairment and are independently validated in two external cohorts: 769 former smokers from COPDGene and 431 former smokers in the Multi-Ethnic Study of Atherosclerosis (MESA). Additionally, we identify several genes that are significantly deregulated across these subtypes, including DSP and GSTM1, which have been previously associated with COPD through genome-wide association study (GWAS). CONCLUSIONS:The identified subtypes differ in mortality and in their clinical and functional characteristics, underlining the need for multi-dimensional assessment potentially supplemented by selected markers of gene expression. The subtypes were consistent across cohorts and could be used for new patient stratification and disease prognosis.
Exome sequencing of genes associated with heritable thoracic aortic disease (HTAD) failed to identify a pathogenic variant in a large family with Marfan syndrome (MFS). A genome-wide linkage analysis for thoracic aortic disease identified a peak at 15q21.1, and genome sequencing identified a novel deep intronic FBN1 variant that segregated with thoracic aortic disease in the family (LOD score 2.7) and was predicted to alter splicing. RT-PCR and bulk RNA sequencing of RNA harvested from fibroblasts explanted from the affected proband revealed an insertion of a pseudoexon between exons 13 and 14 of the FBN1 transcript, predicted to lead to nonsense mediated decay (NMD). Treating the fibroblasts with an NMD inhibitor, cycloheximide, greatly improved the detection of the pseudoexon-containing transcript. Family members with the FBN1 variant had later onset aortic events and fewer MFS systemic features than typical for individuals with haploinsufficiency of FBN1. Variable penetrance of the phenotype and negative genetic testing in MFS families should raise the possibility of deep intronic FBN1 variants and the need for additional molecular studies.
Cold snare polypectomy (CSP) is the preferred resection technique for small (6–9 mm) polyps due to lower rate of incomplete resection compared to cold forceps polypectomy (CFP) and improved safety profile over hot snare polypectomy (HSP). To describe resection techniques for small (6–9 mm) polyps and determine factors associated with sub-optimal technique. This was retrospective cohort study of colonoscopies performed by gastroenterological and surgical endoscopists from 2012 to 2019 where at least one 6–9 mm polyp was removed. Patient, provider, and procedure characteristics were collected. Univariate and multivariate regression analyses were performed to determine factors associated with sub-optimal technique. In total, 773 colonoscopies where 1,360 6–9 mm polyps removed by 21 endoscopists were included. CSP was used for 1,122 (82.5
Despite the growing number of genome-wide association studies (GWASs), it remains unclear to what extent gene-by-gene and gene-by-environment interactions influence complex traits in humans. The magnitude of genetic interactions in complex traits has been difficult to quantify because GWASs are generally underpowered to detect individual interactions of small effect. Here, we develop a method to test for genetic interactions that aggregates information across all trait-associated loci. Specifically, we test whether SNPs in regions of European ancestry shared between European American and admixed African American individuals have the same causal effect sizes. We hypothesize that in African Americans, the presence of genetic interactions will drive the causal effect sizes of SNPs in regions of European ancestry to be more similar to those of SNPs in regions of African ancestry. We apply our method to two traits: gene expression in 296 African Americans and 482 European Americans in the Multi-Ethnic Study of Atherosclerosis (MESA) and low-density lipoprotein cholesterol (LDL-C) in 74K African Americans and 296K European Americans in the Million Veteran Program (MVP). We find significant evidence for genetic interactions in our analysis of gene expression; for LDL-C, we observe a similar point estimate, although this is not significant, most likely due to lower statistical power. These results suggest that gene-by-gene or gene-by-environment interactions modify the effect sizes of causal variants in human complex traits.
Integrative approaches that simultaneously model multi-omics data have gained increasing popularity because they provide holistic system biology views of multiple or all components in a biological system of interest. Canonical correlation analysis (CCA) is a correlation-based integrative method. It was initially designed to extract latent features shared between two assays by finding the linear combinations of features – referred to as canonical vectors (CVs) – within each assay that achieve maximal across-assay correlation. Sparse multiple CCA (SMCCA), a widely-used derivative of CCA, allows more than two assays but can result in non-orthogonal CVs when applied to high-dimensional data. Here, we incorporated a variation of the Gram-Schmidt (GS) algorithm with SMCCA to improve orthogonality among CVs. Applying our SMCCA-GS method to proteomics and methylomics data from the Multi-Ethnic Study of Atherosclerosis (MESA) and Jackson Heart Study (JHS), we identified strong associations between blood cell counts and protein abundance. This finding suggests that adjustment of blood cell composition should be considered in protein-based association studies. Importantly, CVs obtained from two independent cohorts demonstrate transferability across the cohorts. For example, proteomic CVs learned from JHS explain similar amounts of blood cell count phenotypic variance in MESA, explaining 39.0% ~ 50.0% variation in JHS and 38.9% ~ 49.1% in MESA, similar transferability was observed for other omics-CV-trait pairs. This suggests that biologically meaningful and cohort-agnostic variation is captured by CVs. We further developed Sparse Supervised Multiple CCA (SSMCCA) to allow supervised integration analysis for more than two assays. We anticipate that applying our SMCCA-GS and SSMCCA on various cohorts would help identify cohort-agnostic biologically meaningful relationships between multi-omics data and phenotypic traits. Author Summary Comprehensive understanding of human complex traits may benefit from incorporation of molecular features from multiple biological layers such as genome, epigenome, transcriptome, proteome, and metabolome. CCA is a correlation-based method for multi-omics data which reduces the dimension of each omic assay to several orthogonal components – commonly referred to as canonical vectors (CVs). The widely-used SMCCA method allows effective dimension reduction and integration of multi-omics data, but suffers from potentially highly correlated CVs when applied to high-dimensional omics data. Here, we improve the statistical independence among the CVs by adopting a variation of the GS algorithm. We applied our SMCCA-GS method to proteomic and methylomic data from two cohort studies, MESA and JHS. Our results reveal a pronounced effect of blood cell counts on protein abundance, strongly suggesting blood cell composition adjustment in protein-based association studies may be necessary. Finally, we present SSMCCA which allows supervised CCA analysis for the association between one phenotype of interest and more than two assays. We anticipate that SMCCA-GS would help reveal meaningful system-level factors from biological processes involving features from multiple assays; and SSMCCA would further empower interrogation of these factors for phenotypic traits related to health and diseases.
Protein bound uremic toxins (PBUTs), a series of chemicals that remain a challenge for removal strategies used on patients suffering with chronic kidney disease, could be strong candidates for MD study in order to better understand the interactions and time scales associated with binding mode transitions. Currently, traditional dialysis methods cannot satisfactorily remove PBUTs from the bloodstream. This is at least partly due to these toxin's high level of affinity for protein binding sites, particularly the prominent human serum albumin (HSA) and two of its drug binding sites (Sudlow site I and II). We investigate the dynamics of binding site transitions and interactions by MD simulations targeting four well-known toxins: indoxyl sulfate (IS), p-cresyl sulfate (PCS), indole-3-acetic acid (IAA), and hippurate acid (HIP). Long-time scale dynamics are obtained by the use of time-structure independent component analysis (tICA) for dimensionality reduction followed by spectral analysis of a Markov state model (MSM) scored using the generalized matrix Rayleigh quotient (GMRQ). Our results add new insights to prior findings related to the key role of charge-pairing in governing toxin-protein interactions. We find that IAA, the bulkiest hydrophobic toxin studied, observes the slowest process of at least 3 times slower than the smaller, less hydrophobic toxins. In general, we find that the processes slower than 15 ns are correlated with a transition from dominantly hydrophobic interactions deep in the binding pocket to a gain in hydrogen bonding partners near the mouth of the pocket. Our results indicate that aromatic residues such as PHE play a part in a type of toxin stabilization akin to π-stacking. In conclusion, this work presents mechanistic descriptions of interactions/transitions for a set of important PBUTs that bind Sudlow site II on time scales relevant to the underlying binding kinetics of most interest.
Protein bound uremic toxins (PBUTs) are known to bind strongly with the primary drug carrying sites of human serum albumin (HSA), Sudlow site I and Sudlow site II. A detailed energetic and structural description of PBUT interactions with these binding sites would provide useful insight into the design of materials that specifically displace and capture PBUTs. In this work, we used molecular dynamics (MD) simulations to study in atomistic detail 4 PBUTs bound in Sudlow site II. Specifically, we used the experimentally resolved X-ray structure of simulated indoxyl sulfate (IS) bound to Sudlow site II (PBD ID: 2BXH) to generate initial binding poses for p-cresyl sulfate (pCS), indole-3-acetic acid (IAA), and hippuric acid (HA). We calculated the interaction energy between toxin and protein in MD simulations and performed mean shift clustering on the collection of molecular structures from MD to identify the primary binding modes of each toxin. We find that all 4 toxins are primarily stabilized by electrostatic interactions between their anionic moiety and the hydrophilic residues in Sudlow site II. We observed transience in the strongest toxin-protein interaction, a charge-pairing with the positively charged R410 residue. We confirm the finding that the primary binding pose of IS in Sudlow site II is stabilized by a hydrogen bond with the carbonyl oxygen of L430, and find that this is also true for IAA. We provide insight into the chemical functional groups that might be incorporated to improve the specificity of synthetic materials for PBUT capture. This work represents a next step toward the de novo design of solutions to the problem of PBUT management in CKD patients. Significance Statement In spite of their implication in poor clinical outcomes, surprisingly little information is available about the structure and mechanisms that govern the binding of protein bound uremic toxins to their primary carrier human serum albumin. To date, only the structure of indoxyl sulfate has been determined by experiment. This paper describes a comprehensive characterization of four toxins that are known to bind Sudlow site II using molecular dynamics simulations. Based on the experimental structure of indoxyl sulfate bound to HSA, the binding mode within Sudlow site II of three additional PBUTs was determined. The structures, energetic and mechanistic analysis provide substantial new information for the nephrology community about these toxins as well as new protocols to aid future studies of PBUTs.
The therapeutic potential of protein drugs has been hindered by difficulties with long-term stability and rapid clearance from the body. Recombinant fusion proteins provide a scalable platform for engineered biologics, whereby a polypeptide domain is appended to alter the physical characteristics of a therapeutic protein and enhance its pharmaceutical viability. Two simple design principles for recombinant fusion proteins, based on the physical properties of the polypeptide domain, have been separately applied to address issues with the stability and delivery of biologics. "Conformationally disordered" peptides, exemplified by the homo amino acid peptide polyG, have been shown to increase the circulation half-life and bioactivity of protein therapeutics in vivo. Superhydrophilic peptides, exemplified by the alternating-charge peptide poly(EK), have been shown to increase the thermostability of proteins in vitro. The combination of superhydrophilicity and conformational disorder in a single fusion peptide could simultaneously address concerns regarding the stability and therapeutic lifetime of biologics. In the current work, we use enhanced sampling molecular dynamics (MD) simulations to investigate the conformational ensemble of poly(EK) and glycine-substituted poly(EK) variants and validate our structural predictions with circular dichroism (CD). We find the (EK)15 peptide exhibits a high propensity for forming antiparallel β-strand secondary structures, which are stabilized by extensive salt bridging of the positive and negative side chains. MD simulations predict that limited glycine substitutions effectively disrupt the secondary structure and promote disordered conformations at physiologically relevant temperatures. We conclude that the conformational disorder of alternating-charge peptides should be taken into account to improve their suitability for drug delivery applications. We also contribute a computational approach to quantify conformational disorder in polypeptides, which should facilitate the de novo design of effective fusion proteins.
Enzymes play a critical role in many applications in biology and medicine as potential therapeutics. One specific area of interest is enzyme encapsulation in polymer nanostructures, which have applications in drug delivery and catalysis. A detailed understanding of the mechanisms governing protein/polymer interactions is crucial for optimizing the performance of these complex systems for different applications. Using a combined computational and experimental approach, this study aims to quantify the relative importance of molecular and mesoscale driving forces to protein release from polymeric nanoparticles. Classical molecular dynamics (MD) simulations have been performed on bovine serum albumin (BSA) in aqueous solutions with oligomeric surrogates of poly(lactic-co-glycolic acid) copolymer, poly(styrene)-poly(lactic acid) copolymer, and poly(lactic acid). The simulated strength and location of polymer surrogate binding to the surface of BSA have been compared to experimental BSA release rates from nanoparticles formulated with these same polymers. Results indicate that the self-interaction tendencies of the polymer surrogates and other macroscale properties may play governing roles in protein release. Additional MD simulations of BSA in solution with poly(styrene)-acrylate copolymer reveal the possibility of enhanced control over the enzyme encapsulation process by tuning polymer self-interaction. Last, the authors find consistent protein surface binding preferences across simulations performed with polymer surrogates of varying lengths, demonstrating that protein/polymer interactions can be understood in part by studying the interactions and affinity of proteins with small polymer surrogates in solution.
To clarify the molecular changes of sublesional muscle in the acute phase of spinal cord injury (SCI), a moderately severe injury (40 g cm) was induced in the spinal cord (T10 vertebral level) of adult male Sprague–Dawley rats (injury) and compared with sham (laminectomy only). Rats were sacrificed at 48 h (acute) post injury, and gastrocnemius muscles were excised. Morphological examination revealed no significant changes in the muscle fiber diameter between the sham and injury rats. Western blot analyses performed on the visibly red, central portion of the gastrocnemius muscle showed significantly higher expression of muscle specific E3 ubiquitin ligases (muscle ring finger-1 and muscle atrophy f-box) and significantly lower expression of phosphorylated Akt-1/2/3 in the injury group compared to the sham group. Cyclooxygenase 2, tumor necrosis factor alpha (TNF-α), and caspase-1, also had a significantly higher expression in the injury group; although, the mRNA levels of TNF-α and IL-6 did not show any significant difference between the sham and injury groups. These results suggest activation of protein degradation, deactivation of protein synthesis, and development of inflammatory reaction occurring in the sublesional muscles in the acute phase of SCI before overt muscle atrophy is seen.
BackgroundIchthyoses are clinically characterized by scaling or hyperkeratosis of the skin or both. It can be an isolated condition limited to the skin or appear secondarily with involvement of other cutaneous or systemic abnormalities.MethodsThe present study investigated clinical and molecular characterization of three consanguineous families (A, B, C) segregating two different forms of autosomal recessive congenital ichthyosis (ARCI). Linkage in three consanguineous families (A, B, C) segregating two different forms of ARCI was searched by typing microsatellite and single nucleotide polymorphism marker analysis. Sequencing of the two genes TGM1 and ALOXE3 was performed by the dideoxy chain termination method.ResultsGenome-wide linkage analysis established linkage in family A to TGM1 gene on chromosome 14q11 and in families B and C to ALOXE3 gene on chromosome 17p13. Subsequently, sequencing of these genes using samples from affected family members led to the identification of three novel mutations: a missense variant p.Trp455Arg in TGM1 (family A); a nonsense variant p.Arg140* in ALOXE3 (family B); and a complex rearrangement in ALOXE3 (family C).ConclusionThe present study further extends the spectrum of mutations in the two genes involved in causing ARCI. Characterizing the clinical spectrum resulting from mutations in the TGM1 and ALOXE3 genes will improve diagnosis and may direct clinical care of the family members.
Estrogen (EST) is a steroid hormone that exhibits several important physiological roles in the human body. During the last few decades, EST has been well recognized as an important neuroprotective agent in a variety of neurological disorders in the central nervous system (CNS), such as spinal cord injury (SCI), traumatic brain injury (TBI), Alzheimer's disease, and multiple sclerosis. The exact molecular mechanisms of EST-mediated neuro-protection in the CNS remain unclear due to heterogeneity of cell populations that express EST receptors (ERs) in the CNS as well as in the innate and adaptive immune system. Recent investigations suggest that EST protects the CNS from injury by suppressing pro-inflammatory pathways, oxidative stress, and cell death, while promoting neurogenesis, angiogenesis, and neurotrophic support. In this review, we have described the currently known molecular mechanisms of EST-mediated neuroprotection and neuroregeneration in SCI and TBI. At the same time, we have emphasized on the recent in vitro and in vivo findings from our and other laboratories, implying potential clinical benefits of EST in the treatment of SCI and TBI.