The widespread application of high-throughput Next Generation Sequencing (NGS) technologies has made microbiome research an emerging field in public health and biomedical sciences. However, there are still many challenges that need to be addressed in this field. Pipelines available to generate microbiome data across cohorts are diverse, and sources of variation to be recorded and evaluated during microbiome profiling have not been standardized. Moreover, meticulous quality control of the microbiome data processing, from collection to computational quantification is still challenging, especially in large population studies. Innovative approaches are required to handle samples and to minimize the potential bias introduced by logistic hurdles in biobanking. In this paper, we describe the methodological steps surrounding the optimization of the 16s rRNA gut microbiome profiling in two large prospective cohorts the Generation R Study (mean age 9.83 years, SD:0.32 years) and the Rotterdam Study (mean age 62.67,SD:5.66 years). This paper also highlights potential solutions to sample mislabeling in large-scale microbiome analysis. To summarize, our study addresses common problems in human microbiome research. It aims to improve the research quality and reliability by integrating more stringent quality control standards into microbiome research. ### Competing Interest Statement The authors have declared no competing interest.
When profiling the human gut microbiome, technical biases introduced by analytical approaches impede translational research, reducing data reliability and study comparability. Here, through a global study involving 23 labs, we analyzed a wide range of sequencing and bioinformatic approaches for the taxonomic profiling of two well-defined DNA reference reagents (RRs) comprised of 20 common gut bacteria. Through both shotgun and 16S rRNA gene amplicon sequencing, we aimed to isolate sources of bias and understand their impact on microbiome profiling accuracy. Importantly, minimum quality criteria (MQC) were established and are used to evaluate profiling performance. We found that the variability of shotgun sequencing data sets was greater than that of 16S rRNA gene amplicon sequencing and isolated sources of bias in wet and dry lab steps, such as sequencing depth, primer and database choices, rarefaction, and 16S copy number adjustment. This study presents well-defined RRs and MQC to combat technical bias, paving the way for reliable and comparable microbiome research.IMPORTANCEThis benchmark paper highlights the true level of variability in microbiome data across the world and across sectors, underscoring the critical need for the use of WHO International DNA Gut Reference Reagents (RRs) to elevate the quality of data in microbiome research. This global study is the first of its kind, revealing the reality of the bias in the field, comprehensively testing methodologies used by leading laboratories across the world, but also providing avenues for workflow optimization, to accelerate innovation and translational research and move the field forward.
A common feature of human aging is the acquisition of somatic mutations, and mitochondria are particularly prone to mutation, leading to a state of mitochondrial DNA heteroplasmy. Cross-sectional studies have demonstrated that detection of heteroplasmy increases with participant age, a phenomenon that has been attributed to genetic drift. In this large-scale longitudinal study, we measured heteroplasmy in two prospective cohorts (combined n = 1404) at two time points (mean time between visits, 8.6 years), demonstrating that deleterious heteroplasmies were more likely to increase in variant allele fraction (VAF). We further demonstrated that increase in VAF was associated with increased risk of overall mortality. These results challenge the claim that somatic mtDNA mutations arise mainly due to genetic drift, instead suggesting a role for positive selection for a subset of predicted deleterious mutations at the cellular level, despite a negative impact of these mutations on overall mortality.
A common feature of human aging is the acquisition of somatic mutations, and mitochondria are particularly prone to mutation due to their inefficient DNA repair and close proximity to reactive oxygen species, leading to a state of mitochondrial DNA heteroplasmy1,2. Cross-sectional studies have demonstrated that detection of heteroplasmy increases with participant age3, a phenomenon that has been attributed to genetic drift4-7. In this first large-scale longitudinal study, we measured heteroplasmy in two prospective cohorts (combined n=1405) at two timepoints (mean time between visits, 8.6 years), demonstrating that deleterious heteroplasmies were more likely to increase in variant allele fraction (VAF). We further demonstrated that increase in VAF was associated with increased risk of overall mortality. These results challenge the claim that somatic mtDNA mutations arise mainly due to genetic drift, instead demonstrating positive selection for predicted deleterious mutations at the cellular level, despite an negative impact on overall mortality.
Introduction Pulmonary fibrosis is a severe disease which can be familial. A genetic cause can only be found in ∼40% of families. Searching for shared novel genetic variants may aid the discovery of new genetic causes of disease. Methods Whole-exome sequencing was performed in 152 unrelated patients with a suspected genetic cause of pulmonary fibrosis from the St Antonius interstitial lung disease biobank. Variants of interest were selected by filtering for novel, potentially deleterious variants that were present in at least three unrelated pulmonary fibrosis patients. Results The novel c.586G>A p.(E196K) variant in the ZCCHC8 gene was observed in three unrelated patients: two familial patients and one sporadic patient, who was later genealogically linked to one of the families. The variant was identified in nine additional relatives with pulmonary fibrosis and other telomere-related phenotypes, such as pulmonary arterial venous malformations, emphysema, myelodysplastic syndrome, acute myeloid leukaemia and dyskeratosis congenita. One family showed incomplete segregation, with absence of the variant in one pulmonary fibrosis patient who carried a PARN variant. The majority of ZCCHC8 variant carriers showed short telomeres in blood. ZCCHC8 protein was located in different lung cell types, including alveolar type 2 (AT2) pneumocytes, the culprit cells in pulmonary fibrosis. AT2 cells showed telomere shortening and increased DNA damage, which was comparable to patients with sporadic pulmonary fibrosis and those with pulmonary fibrosis carrying a telomere-related gene variant, respectively. Discussion The ZCCHC8 c.586G>A variant confirms the involvement of ZCCHC8 in pulmonary fibrosis and short-telomere syndromes and underlines the importance of including the ZCCHC8 gene in diagnostic gene panels for these diseases.
Desmosomes are dynamic complex protein structures involved in cellular adhesion. Disruption of these structures by loss-of-function variants in desmosomal genes leads to a variety of skin- and heart-related phenotypes. In this study, we report TUFT1 as a desmosome-associated protein, implicated in epidermal integrity. In two siblings with mild skin fragility, woolly hair, and mild palmoplantar keratoderma but without a cardiac phenotype, we identified a homozygous splice-site variant in the TUFT1 gene, leading to aberrant mRNA splicing and loss of TUFT1 protein. Patients' skin and keratinocytes showed acantholysis, perinuclear retraction of intermediate filaments, and reduced mechanical stress resistance. Immunolabeling and transfection studies showed that TUFT1 is positioned within the desmosome and that its location is dependent on the presence of the desmoplakin carboxy-terminal tail. A Tuft1-knockout mouse model mimicked the patients' phenotypes. Altogether, this study reveals TUFT1 as a desmosome-associated protein, whose absence causes skin fragility, woolly hair, and palmoplantar keratoderma.
Single nucleotide polymorphism (SNP) data generated with microarray technologies have been used to solve murder cases via investigative leads obtained from identifying relatives of the unknown perpetrator included in accessible genomic databases, referred to as investigative genetic genealogy (IGG). However, SNP microarrays were developed for relatively high input DNA quantity and quality, while SNP microarray data from compromised DNA typically obtainable from crime scene stains are largely missing. By applying the Illumina Global Screening Array (GSA) to 264 DNA samples with systematically altered quantity and quality, we empirically tested the impact of SNP microarray analysis of deprecated DNA on kinship classification success, as relevant in IGG. Reference data from manufacturer-recommended input DNA quality and quantity were used to estimate genotype accuracy in the compromised DNA samples and for simulating data of different degree relatives. Although stepwise decrease of input DNA amount from 200 nanogram to 6.25 picogram led to decreased SNP call rates and increased genotyping errors, kinship classification success did not decrease down to 250 picogram for siblings and 1st cousins, 1 nanogram for 2nd cousins, while at 25 picogram and below kinship classification success was zero. Stepwise decrease of input DNA quality via increased DNA fragmentation resulted in the decrease of genotyping accuracy as well as kinship classification success, which went down to zero at the average DNA fragment size of 150 base pairs. Combining decreased DNA quantity and quality in mock casework and skeletal samples further highlighted possibilities and limitations. Overall, GSA analysis achieved maximal kinship classification success from 800-200 times lower input DNA quantities than manufacturer-recommended, although DNA quality plays a key role too, while compromised DNA produced false negative kinship classifications rather than false positive ones. Author Summary Investigative genetic genealogy (IGG), i.e., identifying unknown perpetrators of crime via genomic database-tracing of their relatives by means of microarray-based single nucleotide polymorphism (SNP) data, is a recently emerging field. However, SNP microarrays were developed for much higher DNA quantity and quality than typically available from crime scenes, while SNP microarray data on quality and quantity compromised DNA are largely missing. As first attempt to investigate how SNP microarray analysis of quantity and quality compromised DNA impacts kinship classification success in the context of IGG, we performed systematic SNP microarray analyses on DNA samples below the manufacturer-recommended quantity and quality as well as on mock casework samples and on skeletal remains. In addition to IGG, our results are also relevant for any SNP microarray analysis of compromised DNA, such as for the DNA prediction of appearance and biogeographic ancestry in forensics and anthropology and for other purposes.
The aetiology of late-onset neurodegenerative diseases is largely unknown. Here we investigated whether de novo somatic variants for semantic dementia can be detected, thereby arguing for a more general role of somatic variants in neurodegenerative disease. Semantic dementia is characterized by a non-familial occurrence, early onset (<65 years), focal temporal atrophy and TDP-43 pathology. To test whether somatic variants in neural progenitor cells during brain development might lead to semantic dementia, we compared deep exome sequencing data of DNA derived from brain and blood of 16 semantic dementia cases. Somatic variants observed in brain tissue and absent in blood were validated using amplicon sequencing and digital PCR. We identified two variants in exon one of the TARDBP gene (L41F and R42H) at low level (1-3%) in cortical regions and in dentate gyrus in two semantic dementia brains, respectively. The pathogenicity of both variants is supported by demonstrating impaired splicing regulation of TDP-43 and by altered subcellular localization of the mutant TDP-43 protein. These findings indicate that somatic variants may cause semantic dementia as a non-hereditary neurodegenerative disease, which might be exemplary for other late-onset neurodegenerative disorders.
Purpose We studied the penetrance of pathogenically classified variants in an elderly Dutch population from the Rotterdam Study, for which deep phenotyping is available. We screened the 59 actionable genes for which reporting of known pathogenic variants was recommended by the American College of Medical Genetics and Genomics (ACMG), and demonstrate that determining what constitutes a known pathogenic variant can be quite challenging. Methods We defined “known pathogenic” as classified pathogenic by both ClinVar and the Human Gene Mutation Database (HGMD). In 2628 individuals, we performed exome sequencing and identified known pathogenic variants. We investigated the clinical records of carriers and evaluated clinical events during 25 years of follow-up for evidence of variant pathogenicity. Results Of 3815 variants detected in the 59 ACMG genes, 17 variants were considered known pathogenic. For 14/17 variants the ClinVar classification had changed over time. Of 24 confirmed carriers of these variants, we observed at least one clinical event possibly caused by the variant in only three participants (13%). Conclusion We show that the definition of “known pathogenic” is often unclear and should be approached carefully. Additionally variants marked as known pathogenic do not always have clinical impact on their carriers. Definition and classification of true (individual) expected pathogenic impact should be defined carefully.
Macrophage-mediated inflammation is thought to have a causal role in osteoarthritis-related pain and severity, and has been suggested to be triggered by endotoxins produced by the gastrointestinal microbiome. Here we investigate the relationship between joint pain and the gastrointestinal microbiome composition, and osteoarthritis-related knee pain in the Rotterdam Study; a large population based cohort study. We show that abundance of Streptococcus species is associated with increased knee pain, which we validate by absolute quantification of Streptococcus species. In addition, we replicate these results in 867 Caucasian adults of the Lifelines-DEEP study. Finally we show evidence that this association is driven by local inflammation in the knee joint. Our results indicate the microbiome is a possible therapeutic target for osteoarthritis-related knee pain.
Disease incidences increase with age, but the molecular characteristics of ageing that lead to increased disease susceptibility remain inadequately understood. Here we perform a whole-blood gene expression meta-analysis in 14,983 individuals of European ancestry (including replication) and identify 1,497 genes that are differentially expressed with chronological age. The age-associated genes do not harbor more age-associated CpG-methylation sites than other genes, but are instead enriched for the presence of potentially functional CpG-methylation sites in enhancer and insulator regions that associate with both chronological age and gene expression levels. We further used the gene expression profiles to calculate the 'transcriptomic age' of an individual, and show that differences between transcriptomic age and chronological age are associated with biological features linked to ageing, such as blood pressure, cholesterol levels, fasting glucose, and body mass index. The transcriptomic prediction model adds biological relevance and complements existing epigenetic prediction models, and can be used by others to calculate transcriptomic age in external cohorts.
Small insertions and deletions (indels) and large structural variations (SVs) are major contributors to human genetic diversity and disease. However, mutation rates and characteristics of de novo indels and SVs in the general population have remained largely unexplored. We report 332 validated de novo structural changes identified in whole genomes of 250 families, including complex indels, retrotransposon insertions, and interchromosomal events. These data indicate a mutation rate of 2.94 indels (1-20 bp) and 0.16 SVs (>20 bp) per generation. De novo structural changes affect on average 4.1 kbp of genomic sequence and 29 coding bases per generation, which is 91 and 52 times more nucleotides than de novo substitutions, respectively. This contrasts with the equal genomic footprint of inherited SVs and substitutions. An excess of structural changes originated on paternal haplotypes. Additionally, we observed a nonuniform distribution of de novo SVs across offspring. These results reveal the importance of different mutational mechanisms to changes in human genome structure across generations.
Background Osteoporosis is a systemic skeletal disease characterised by reduced bone mineral density and increased susceptibility to fracture; these traits are highly heritable. Both common and rare copy number variants (CNVs) potentially affect the function of genes and may influence disease risk.Aim To identify CNVs associated with osteoporotic bone fracture risk.Method We performed a genome-wide CNV association study in 5178 individuals from a prospective cohort in the Netherlands, including 809 osteoporotic fracture cases, and performed in silico lookups and de novo genotyping to replicate in several independent studies.Results A rare (population prevalence 0.14%, 95% CI 0.03% to 0.24%) 210kb deletion located on chromosome 6p25.1 was associated with the risk of fracture (OR 32.58, 95% CI 3.95 to 1488.89; p=8.69x10(-5)). We performed an in silico meta-analysis in four studies with CNV microarray data and the association with fracture risk was replicated (OR 3.11, 95% CI 1.01 to 8.22; p=0.02). The prevalence of this deletion showed geographic diversity, being absent in additional samples from Australia, Canada, Poland, Iceland, Denmark, and Sweden, but present in the Netherlands (0.34%), Spain (0.33%), USA (0.23%), England (0.15%), Scotland (0.10%), and Ireland (0.06%), with insufficient evidence for association with fracture risk.Conclusions These results suggest that deletions in the 6p25.1 locus may predispose to higher risk of fracture in a subset of populations of European origin; larger and geographically restricted studies will be needed to confirm this regional association. This is a first step towards the evaluation of the role of rare CNVs in osteoporosis.
Hanneke J.M. Kerkhof ⁎, Rik J. Lories , Ingrid Meulenbelt , Ingileif Jonsdottir , Ana M. Valdes , Pascal Arp , Thorvaldur Ingvarsson , Mila Jhamai , Helgi Jonsson , Lisette Stolk , Gudmar Thorleifsson , Guangju Zhai , Feng Zhang , Yanyan Zhu , Ruud van der Breggen , Andrew Carr , Michael Doherty , Sally Doherty , David T. Felson , Antonio Gonzalez , Bjarni V. Halldorsson , Deborah J. Hart , Valdimar B. Hauksson , Albert Hofman , John P.A. Ioannidis , Margreet Kloppenburg , Nancy E. Lane , John Loughlin , Frank P. Luyten , Michael C. Nevitt , Neeta Parimi , Huibert A.P. Pols , Tom van de Putte , Fernando Rivadeneira , Eline P. Slagboom , Unnur Styrkarsdottir , Aspasia Tsezou , Joseph Zmuda , Tim D. Spector , Kari Stefansson , Andre G. Uitterlinden , Joyce B.J. van Meurs a,d
OBJECTIVE:To identify novel genes involved in osteoarthritis (OA), by means of a genome-wide association study.METHODS:We tested 500,510 single-nucleotide polymorphisms (SNPs) in 1,341 Dutch Caucasian OA cases and 3,496 Dutch Caucasian controls. SNPs associated with at least 2 OA phenotypes were analyzed in 14,938 OA cases and approximately 39,000 controls. Meta-analyses were performed using the program Comprehensive Meta-analysis, with P values <1 x 10(-7) considered genome-wide significant.RESULTS:The C allele of rs3815148 on chromosome 7q22 (minor allele frequency 23%; intron 12 of the COG5 gene) was associated with a 1.14-fold increased risk (95% confidence interval 1.09-1.19) of knee and/or hand OA (P = 8 x 10(-8)) and also with a 30% increased risk of knee OA progression (95% confidence interval 1.03-1.64) (P = 0.03). This SNP is in almost complete linkage disequilibrium with rs3757713 (68 kb upstream of GPR22), which is associated with GPR22 expression levels in lymphoblast cell lines (P = 4 x 10(-12)). Immunohistochemistry experiments revealed that G protein-coupled receptor protein 22 (GPR22) was absent in normal mouse articular cartilage or synovium. However, GPR22-positive chondrocytes were found in the upper layers of the articular cartilage of mouse knee joints that were challenged with in vivo papain treatment or methylated bovine serum albumin treatment. GPR22-positive chondrocyte-like cells were also found in osteophytes in instability-induced OA.CONCLUSION:Our findings identify a novel common variant on chromosome 7q22 that influences susceptibility to prevalence and progression of OA. Since the GPR22 gene encodes a G protein-coupled receptor, this is potentially an interesting therapeutic target.
Hanneke J. M. Kerkhof, Rik J. Lories, Ingrid Meulenbelt, Ingileif Jonsdottir, Ana M. Valdes, Pascal Arp, Thorvaldur Ingvarsson, Mila Jhamai, Helgi Jonsson, Lisette Stolk, Gudmar Thorleifsson, Guangju Zhai, Feng Zhang, Yanyan Zhu, Ruud van der Breggen, Andrew Carr, Michael Doherty, Sally Doherty, David T. Felson, Antonio Gonzalez, Bjarni V. Halldorsson, Deborah J. Hart, Valdimar B. Hauksson, Albert Hofman, John P. A. Ioannidis, Margreet Kloppenburg, Nancy E. Lane, John Loughlin, Frank P. Luyten, Michael C. Nevitt, Neeta Parimi, Huibert A. P. Pols, Fernando Rivadeneira, Eline P. Slagboom, Unnur Styrkársdóttir, Aspasia Tsezou, Tom van de Putte, Joseph Zmuda, Tim D. Spector, Kari Stefansson, André G. Uitterlinden, and Joyce B. J. van Meurs
Kerkhof, Hanneke JM ; Lories, Rik J. ; Meulenbelt, Ingrid ; Jonsdottir, Ingileif ; Valdes, Ana M. ; Arp, Pascal ; Ingvarsson, Thorvaldur ; Jhamai, Mila ; Jonsson, Helgi ; Stolk, Lisette ; Thorleifsson, Gudmar ; Zhai, Guangju ; Zhang, Feng ; Zhu, Yanyan ; van der Breggen, Ruud ; Carr, Andrew ; Doherty, Michael ; Doherty, Sally ; Felson, David T. ; Gonzalez, Antonio ; Halldorsson, Bjarni V. ; Hart, Deborah J. ; Hauksson, Valdimar B. ; Hofman, Albert ; Ioannidis, John PA ; Kloppenburg, Margreet ; Lane, Nancy E ; Loughlin, John ; Luyten, Frank P. ; Nevitt, Michael C. ; Parimi, Neeta ; Pols, Huibert AP ; van de Putte, Tom ; Rivadeneira, Fernando ; Slagboom, Eline P. ; Styrkársdóttir, Unnur ; Tsezou, Aspasia ; Zmuda, Joseph ; Spector, Tim D. ; Stefansson, Kari ; Uitterlinden, André G. ; van Meurs, Joyce B J …
AIMS:Contradictory reports exist regarding the influence of the exon 3 deleted (d3)/full-length (fl) growth hormone receptor (GHR) polymorphism on responsiveness to recombinant human growth-hormone therapy in idiopathic short stature, small for gestational age and GH-deficient children, Turner syndrome patients and GH-deficient adults. In some of these studies, the d3 allele was associated with increased responsiveness to GH. The aim of this study was to test this association in a group of GH-deficient adult patients receiving recombinant GH treatment. MATERIALS & METHODS:Patients were derived from the prospective German Pfizer International Metabolic Study (KIMS) Pharmacogenetics Study. The GHRd3/fl polymorphism was determined in 133 German adult patients (66 men and 67 women; mean age: 45.4 years +/- 13.1 standard deviation; majority Caucasian) with a GH-deficiency of different origin. Patients received GH treatment for 12 months with a finished dose-titration of GH and standardized insulin-like growth factor (IGF)-1 measurements in one central laboratory. GH dose after 1 year of treatment, IGF-1 serum concentrations, IGF-1 standard deviation score (SDS) values and anthropometric data were analyzed by GHRd3/fl genotypes. RESULTS:After 1 year of GH treatment, the individually required GH dose was significantly lower in GH-deficient patients carrying one or two d3 alleles, compared with patients with the full-length receptor (p = 0.04). Genotype groups (d3-allele carriers vs noncarriers) showed no significant differences in IGF-1 serum concentrations (p = 0.51), IGF-1 SDS (p = 0.36) nor in gender (p = 0.53), age (p = 0.28), weight (p = 0.13), height (p = 0.53) or BMI (p = 0.15). CONCLUSION:The d3-allele carriers required approximately 25% less exogenous GH compared with the homozygous fl-allele carriers, which may express an increased responsiveness to exogenous GH. Variability of the individually required GH dose in adult GH-deficient patients may therefore be partly due to the GHRd3/fl polymorphism. Further studies are required to confirm these results.
Osteoporosis is a bone disease leading to an increased fracture risk. It is considered a complex multifactorial genetic disorder with interaction of environmental and genetic factors. As a candidate gene for osteoporosis, we studied vitamin D binding protein (DBP, or group-specific component, Gc), which binds to and transports vitamin D to target tissues to maintain calcium homeostasis through the vitamin D endocrine system. DBP can also be converted to DBP-macrophage activating factor (DBP-MAF), which mediates bone resorption by directly activating osteoclasts. We summarized the genetic linkage structure of the DBP gene. We genotyped two single-nucleotide polymorphisms (SNPs, rs7041 = Glu416Asp and rs4588 = Thr420Lys) in 6,181 elderly Caucasians and investigated interactions of the DBP genotype with vitamin D receptor (VDR) genotype and dietary calcium intake in relation to fracture risk. Haplotypes of the DBP SNPs correspond to protein variations referred to as Gc1s (haplotype 1), Gc2 (haplotype 2), and Gc1f (haplotype3). In a subgroup of 1,312 subjects, DBP genotype was found to be associated with increased and decreased serum 25-(OH)D(3) for haplotype 1 (P = 3 x 10(-4)) and haplotype 2 (P = 3 x 10(-6)), respectively. Similar associations were observed for 1,25-(OH)(2)D(3). The DBP genotype was not significantly associated with fracture risk in the entire study population. Yet, we observed interaction between DBP and VDR haplotypes in determining fracture risk. In the DBP haplotype 1-carrier group, subjects of homozygous VDR block 5-haplotype 1 had 33% increased fracture risk compared to noncarriers (P = 0.005). In a subgroup with dietary calcium intake <1.09 g/day, the hazard ratio (95% confidence interval) for fracture risk of DBP hap1-homozygote versus noncarrier was 1.47 (1.06-2.05). All associations were independent of age and gender. Our study demonstrated that the genetic effect of the DBP gene on fracture risk appears only in combination with other genetic and environmental risk factors for bone metabolism.