Recently, a number of studies have looked at the problem of privacy and data-sharing restrictions in the context of missing-genotype imputation servers. This relates to the most typical imputation pipelines which involve a whole-genome sequenced haplotype reference panel being compared to genotyped study individuals (who have missing data to be imputed). Hence, involving two datasets from separate sources coming together in one informatic environment, where relatively complicated statistical models are applied: specifically, hidden Markov modelling. We give a short review of the current literature in this domain, observing three prevalent strategies: complicated data encryption, technical solutions to secure computation environments, and rearrangements of haplotype data to provide anonymisation. We embarked on a thought experiment to provide a potential fourth type of solution involving federating the different internal tasks within the statistical methods used for imputation. This idea is relevant considering there is currently motivation for federated analyses platforms in Europe for making combined inference across multiple genomic data resources. Our solution allows for very simple manipulations to protect sensitive individual level data, which enable imputation algorithms to complete on simple plain-text files. We provide here an illustration of how such a federated imputation server could be put in place, along with associated code, including a simple implementation of the Li-Stephens haplotype mosaic model to achieve the imputation of missing genotypes. We name our general framework ANONYMP for anonymised imputation. A demonstration of the concept is given involving simulated data generated with msprime. We show that dividing different parts of the required calculations for statistical imputation between several sites is a valuable new avenue in the field of privacy-preserving imputation server development.
INTRODUCTION:Next-generation sequencing (NGS) data analysis has become an integral part of clinical genetic diagnosis, raising the question of variant prioritization. The Population Sampling Probability (PSAP) method has been developed to tackle the issue of variant prioritization in the exome of a single patient, by leveraging allele frequencies from population databases and a variant pathogenicity score. METHODS:Here, we present Easy-PSAP, a completely new implementation of the PSAP method comprising two user-friendly and highly adaptable pipelines. Easy-PSAP allows the gene-based recalibration of any in silico pathogenicity prediction score compared to scores of variants seen in the general population, including popular scores like CADD or AlphaMissense. Easy-PSAP can evaluate genetic variants at the scale of a whole exome or a whole genome using information from the latest population and annotation databases. RESULTS:Through simulations on synthetic disease exomes, we show that Easy-PSAP is able to rank more than 50% of causal pathogenic variants in the top 10 variants for an autosomal dominant model of transmission and in top 1 for an autosomal recessive model of transmission. DISCUSSION:These findings, along with the accessibility of the pipeline to both researchers and clinicians, make Easy-PSAP a state-of-the-art tool for variant prioritization in NGS data that can continue to evolve as new frameworks and databases become available. Easy-PSAP is implemented in R and bash within an open-source Snakemake framework. It is available on GitHub alongside conda environments containing the required dependencies (https://github.com/msogloblinsky/Easy-PSAP).
Non-canonical small open reading frames (sORFs) are among the main regulators of gene expression. The most studied of these are upstream ORFs (upORFs) located in the 5'-untranslated region (UTR) of coding genes. Internal ORFs (intORFs) in the coding sequence and downstream ORFs (dORFs) in the 3'UTR have received less attention. Different bioinformatics tools permit the prediction of single nucleotide variants (SNVs) altering upORFs, mainly those creating AUGs or deleting stop codons, but no tool predicts variants altering non-canonical translation initiation sites and those altering intORFs or dORFs. We propose an upgrade of our MORFEE bioinformatics tool to identify SNVs that may alter all types of sORFs in coding transcripts from a VCF file. Moreover, we generate an exhaustive catalog, named MORFEEdb, reporting all possible SNVs altering existing upORFs or creating new ones in human transcripts, and provide an R script for visualizing the results. MORFEEdb has been implemented in the public platform Mobidetails. Finally, the annotation of ClinVar variants with MORFEE reveals that > 45% of UTR-SNVs can alter upORFs or dORFs. In conclusion, MORFEE and MORFEEdb have the potential to improve the molecular diagnosis of rare human diseases and to facilitate the identification of functional variants from genome-wide association studies of complex traits.
PURPOSE:Survivors of childhood cancers are at risk of second malignant neoplasms (SMNs). Exposure to radiation therapy and different chemotherapy agents are known risk factors for SMN development. However, there remains interindividual variability that could be attributed to genetic variations. Our study aims to identify rare genetic variants associated with SMN risk and assess the influence of treatment on the associated risk of these variants. MATERIALS AND METHODS:We conducted a whole-exome sequencing in a nested case-control study within the French Childhood Cancer Survivors Study for 450 survivors including 163 cases and 287 controls. Rare variants in genes within replication, recombination, and repair pathways were selected. Gene-based association tests were conducted and a logistic regression analysis was used to estimate the effect of selected genes on SMN risk, adjusting for relevant clinical variables. Specific analysis restricted to breast and thyroid SMNs were also performed. Stratified analyses on the basis of the radiation doses received by the bone marrow and specific organs were also conducted. RESULTS:We have identified several genetic associations: the RNASEL and APOBEC3F genes (odds ratio [OR], 5.43 [95% CI, 1.76 to 20.44]; P = .0055 and OR, 4.8 [95% CI, 1.28 to 22.95]; P = .027, respectively) are associated with the risk of SMN, while the FANCM gene (OR, 4.65 [95% CI, 1.06 to 21.56]; P = .041) is associated with the risk of developing breast SMN. CONCLUSION:This study provides new evidence regarding the involvement of genetics in the development of SMN. Additional research is needed to confirm these results and uncover the underlying mechanisms behind these associations. If replicated, these findings will help identify survivors at higher risk of developing SMN, enabling tailored treatment and follow-up strategies.
The demographical history of France remains largely understudied despite its central role toward understanding modern population structure across Western Europe. Here, by exploring publicly available Europe-wide genotype datasets together with the genomes of 3234 present-day and six newly sequenced medieval individuals from Northern France, we found extensive fine-scale population structure across Brittany and the downstream Loire basin and increased population differentiation between the northern and southern sides of the river Loire, associated with higher proportions of steppe vs. Neolithic-related ancestry. We also found increased allele sharing between individuals from Western Brittany and those associated with the Bell Beaker complex. Our results emphasise the need for investigating local populations to better understand the distribution of rare (putatively deleterious) variants across space and the importance of common genetic legacy in understanding the sharing of disease-related alleles between Brittany and people from western Britain and Ireland. Here, the authors look at present day and medieval genome samples from Northern France to piece together demographic history and population structure across Brittany and the Loire basin.
Background and aims: PCSK9 is a key regulator of LDL-cholesterol levels. PCSK9 gain of function variants (GOFVs) cause autosomal dominant hypercholesterolemia (ADH). The first described PCSK9-GOFV, p.Ser127Arg, almost exclusively reported in France, represents 67 % of the PCSK9 French GOFVs due to a founder effect. Few other carriers are reported in South Africa and Norway. This study aims to estimate when the common ancestor lived and to describe a cohort of p.Ser127Arg carriers. Methods: Eight families and 14 p.Ser127Arg carriers were genotyped and phenotyped. Haplotypes were constructed using 11 microsatellites around PCSK9 and 6 intragenic single nucleotide polymorphisms (SNPs). To add to the biological analysis, eight additional p.Ser127Arg carriers, 12 carriers of other PCSK9-GOFVs, 93 LDLR loss of function variant (LOFV) carriers and 49 non-carriers subjects were phenotyped. Results: The most common ancestor of p.Ser127Arg was estimated to have lived 775 years ago [95 % CI: 5751075]. French Protestants exiled after the revocation of the Edict of Nantes in 1685 AD likely brought the variant to South Africa and Norway. As expected for ADH subjects, carriers of LDLR-LOFV, the p.Ser127Arg, or other PCSK9-GOFVs showed significantly higher LDL-C levels than that of the non-carriers. Interestingly, LDL-C levels are higher for LDLR-LOFVs and for the reduced secreted p.Ser127Arg than for secreted PCSK9-GOFVs, suggesting a greater effect of the p.Ser127Arg. Conversely, HDL-C was significantly lower for LDLR-LOFV and p.Ser127Arg carriers. Conclusions: This first report from a large cohort of PCSK9-p.Ser127Arg carriers provides observations suggesting a stronger hypercholesterolemic potential of the mutated pro-PCSK9 compared with the secreted mature protein. This work also provides additional data to support the association between PCSK9 and HDL metabolism, and molecular evidence that this variant appeared in France around 1248 AD (Graphical Abstract = Fig. 1).
Achondrite meteorites record the earliest evidence of volcanic activity in the Solar System.The 182 Hf-182 W records of noncarbonaceous (NC) iron meteorites provide model ages for core formation as early as 0.3 Myr after calcium-aluminium-rich inclusions (CAIs), supporting even earlier accretion times for their parent bodies [1].Stony achondrite meteorites provide evidence of early melting of chondritic [2] or differentiated [3] sources.These planetesimals were predominantly heated by the decay of short-lived radionuclide 26 Al (half-life ~0.7 Myr), which was subsequently enriched in crustal reservoirs during partial melting.Large uncertainties remain for modelling the accretion timescales owing to complex histories of magmatic differentiation and secondary processes (e.g., metamorphism, metasomatism, impact).The oldest known lava is an andesitic meteorite Erg Chech 002, which crystallized within ~2 Myr from CAIs [2, 4, 5].Using its chronological records, we model a best-fit planetesimal with a radius of 20-30 km that formed at ~0.1 Myr after CAIs [6].Angrites are sub-grouped either fine-grained volcanic angrites (formed 3-4 Myr after CAIs) or coarse-grained plutonic angrites (~10 Myr after CAIs).We obtained petrological and thermochronological records using U-Pb dating of mineral separates for an olivine-bearing plutonic angrite NWA 10463 [7] and of phosphates for NWA 10463 and the unique angritic dunite Northwest Africa (NWA) 8535.Phosphate SIMS U-Pb ages range from 4552±5 Myr to 4512±18 Myr (2s).In the absence of evidence of shock effects, such extended heating within the crust supports a relatively larger angrite parent body [8].
Rare variant association tests (RVAT) have been developed to study the contribution of rare variants widely accessible through high-throughput sequencing technologies. RVAT require to aggregate rare variants in testing units and to filter variants to retain only the most likely causal ones. In the exome, genes are natural testing units and variants are usually filtered based on their functional consequences. However, when dealing with whole-genome sequence (WGS) data, both steps are challenging. No natural biological unit is available for aggregating rare variants. Sliding windows procedures have been proposed to circumvent this difficulty, however they are blind to biological information and result in a large number of tests. We propose a new strategy to perform RVAT on WGS data: “RAVA-FIRST” (RAre Variant Association using Functionally-InfoRmed STeps) comprising three steps. (1) New testing units are defined genome-wide based on functionally-adjusted Combined Annotation Dependent Depletion (CADD) scores of variants observed in the gnomAD populations, which are referred to as “CADD regions”. (2) A region-dependent filtering of rare variants is applied in each CADD region. (3) A functionally-informed burden test is performed with sub-scores computed for each genomic category within each CADD region. Both on simulations and real data, RAVA-FIRST was found to outperform other WGS-based RVAT. Applied to a WGS dataset of venous thromboembolism patients, we identified an intergenic region on chromosome 18 enriched for rare variants in early-onset patients. This region that was missed by standard sliding windows procedures is included in a TAD region that contains a strong candidate gene. RAVA-FIRST enables new investigations of rare non-coding variants in complex diseases, facilitated by its implementation in the R package Ravages.
Next‐generation sequencing technologies have opened up the possibility to sequence large samples of cases and controls to test for association with rare variants. To limit cost and increase sample sizes, data from controls could be used in multiple studies and might thus be generated on different sequencing platforms. This could pose some problems of comparability between cases and controls due to batch effects that could be confounding factors, leading to false‐positive association signals. To limit batch effects and ensure comparability of datasets, stringent quality controls are required. We propose an integrative five‐steps pipeline, RAVAQ, that (a) performs a specific three‐step quality control taking into account the case–control status to ensure data comparability, (b) selects qualifying variants as defined by the user, and (c) performs rare variant association tests per genomic region. The RAVAQ pipeline is wrapped in an R package. It is user‐friendly and flexible in its arguments to adapt to the specificity of each research project. We provide examples showing how RAVAQ improves rare variant association tests. The default RAVAQ quality control outperformed the widely used Variant Quality Score Recalibration method, removing inflation due to spurious signals. RAVAQ is open source and freely available at https://gitlab.com/gmarenne/ravaq .
Research question: Are there genetic determinants shared by unrelated women with unexplained recurrent early miscarriage (REM)? Design: Thirty REM cases and 30 controls were selected with extreme phenotype among women from Eastern Brittany (France), previously enrolled in an incident case-control study on thrombophilic mutations. Cases and controls were selected based on the number of early miscarriages or live births, respectively. Peripheral blood was collected for DNA extraction at initial visit. The burden of low-frequency variants in the coding part of the genes was compared using whole exome sequencing (WES). Results: Cases had 3 to 17 early miscarriages (20 cases: >_5 previous losses). Controls had 1 to 4 live births (20 controls: >_3 previous live births) and no miscarriages. WES data were available for 29 cases and 30 controls. A total of 209,387 variants were found (mean variant per patient: 59,073.05) with no difference between groups (P = 0.68). The top five most significantly associated genes were ABCA4, NFAM1, TCN2, AL078585.1 and EPS15. Previous studies suggest the involvement of vitamin B12 deficiency in REM. TCN2 encodes for vitamin B12 transporter into cells. Therefore, holotranscobalamin (active vitamin B12) was measured for both cases and controls (81.2 +/- 32.1 versus 92.9 +/- 34.3 pmol/l, respectively, P = 0.186). Five cases but no controls were below 50 pmol/l (P = 0.052). Conclusions: This study highlights four new genes of interest in REM, some of which belong to known networks of genes involved in embryonic development (clathrin-mediated endocytosis and ciliary pathway). The study also confirms the involvement of TCN2 (vitamin B12 pathway) in the early first trimester of pregnancy.
ObjectiveThe majority of patients with a familial cerebral small vessel disease (CSVD) referred for molecular screening do not show pathogenic variants in known genes. In this study, we aimed to identify novel CSVD causal genes.MethodsWe performed a gene‐based collapsing test of rare protein‐truncating variants identified in exome data of 258 unrelated CSVD patients of an ethnically matched control cohort and of 2 publicly available large‐scale databases, gnomAD and TOPMed. Western blotting was used to investigate the functional consequences of variants. Clinical and magnetic resonance imaging features of mutated patients were characterized.ResultsWe showed that LAMB1 truncating variants escaping nonsense‐mediated messenger RNA decay are strongly overrepresented in CSVD patients, reaching genome‐wide significance (p < 5 × 10−8). Using 2 antibodies recognizing the N‐ and C‐terminal parts of LAMB1, we showed that truncated forms of LAMB1 are expressed in the endogenous fibroblasts of patients and trapped in the cytosol. These variants are associated with a novel phenotype characterized by the association of a hippocampal type episodic memory defect and a diffuse vascular leukoencephalopathy.InterpretationThese findings are important for diagnosis and clinical care, to avoid unnecessary and sometimes invasive investigations, and also from a mechanistic point of view to understand the role of extracellular matrix proteins in neuronal homeostasis. ANN NEUROL 2021;90:962–975
This paper proposes a new privacy-preserving framework to perform rare variant case-control association tests with information provided by two parties: a Genomic Research Unit (GRU) with sequencing data from individuals affected by a disease D (cases); a Genomic Research Center (GRC) with sequencing data from healthy individuals (controls). To identify genes with rare variants involved in D, GRU needs to compare cases against controls using association tests (genome-wide association study). The main originality of our proposal is twofold. First, it positions GRC as a proxy between GRU and the server. Doing so makes it possible to use classical cryptographic tools to securely conduct association tests with no computation complexity increase, contrarily to actual state of the art proposals which are of very high complexity being based on homomorphic encryption, for instance. In particular, we show how sensitive data confidentiality can be ensured with secret key based cryptographic hashing with no need to modify statistical algorithms. In our protocol the server simply conducts statistical analyses on partially hashed data. Secondly, we introduce a novel privacy constraint: GRU’s identity should remain unknown to the server as this knowledge can give it clues about GRU’s data (e.g., diseases and genes of interest). We exhibit how Pretty Good Privacy (PGP) can be used to solve this problem. We illustrate our protocol in the case of one rare variant association test, the Weighted-Sum Statistic (WSS) algorithm, carried out on real genetic data. This secure WSS achieves the same accuracy as its nonsecure version with no increase of complexity. Furthermore, we establish that our protocol can be extended to the different rare variant association tests available in the literature.
Summary: When analyzing sequence data, genetic variants are considered one by one, taking no account of whether or not they are found in the same individual. However, variant combinations might be key players in some diseases as variants that are neutral on their own can become deleterious when associated together. GEMPROT is a new analysis tool that allows, from a phased vcf file, to visualize the consequences of the genetic variants on the protein. At the level of an individual, the program shows the variants on each of the two protein sequences and the Pfam functional protein domains. When data on several individuals are available, GEMPROT lists the haplotypes found in the sample and can compare the haplotype distributions between different sub-groups of individuals. By offering a global visualization of the gene with the genetic variants present, GEMPROT makes it possible to better understand the impact of combinations of genetic variants on the protein sequence. Availability and implementation: GEMPROT is freely available at https://github.com/ TaniaCuppens/GEMPROT. An on-line version is also available at http://med-laennec.univ-brest.fr/ GEMPROT/. Contact: tania.cuppens@inserm.fr or emmanuelle.genin@inserm.fr Supplementary information: Supplementary data are available at Bioinformatics online.
Genetic association studies have provided new insights into the genetic variability of human complex traits with a focus mainly on continuous or binary traits. Methods have been proposed to take into account disease heterogeneity between subgroups of patients when studying common variants but none was specifically designed for rare variants. Because rare variants are expected to have stronger effects and to be more heterogeneously distributed among cases than common ones, subgroup analyses might be particularly attractive in this context. To address this issue, we propose an extension of burden tests by using a multinomial regression model, which enables association tests between rare variants and multicategory phenotypes. We evaluated the type I error and the power of two burden tests, CAST and WSS, by simulating data under different scenarios. In the case of genetic heterogeneity between case subgroups, we showed an advantage of multinomial regression over logistic regression, which considers all the cases against the controls. We replicated these results on real data from Moyamoya disease where the burden tests performed better when cases were stratified according to age-of-onset. We implemented the functions for association tests in the R package "Ravages" available on Github.
SUMMARY:When analyzing sequence data, genetic variants are considered one by one, taking no account of whether or not they are found in the same individual. However, variant combinations might be key players in some diseases as variants that are neutral on their own can become deleterious when associated together. GEMPROT is a new analysis tool that allows, from a phased vcf file, to visualize the consequences of the genetic variants on the protein. At the level of an individual, the program shows the variants on each of the two protein sequences and the Pfam functional protein domains. When data on several individuals are available, GEMPROT lists the haplotypes found in the sample and can compare the haplotype distributions between different sub-groups of individuals. By offering a global visualization of the gene with the genetic variants present, GEMPROT makes it possible to better understand the impact of combinations of genetic variants on the protein sequence.AVAILABILITY AND IMPLEMENTATION:GEMPROT is freely available at https://github.com/TaniaCuppens/GEMPROT. An on-line version is also available at http://med-laennec.univ-brest.fr/GEMPROT/.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Summary: Assemble is an intuitive graphical interface to analyze, manipulate and build complex 3D RNA architectures. It provides several advanced and unique features within the framework of a semi-automated modeling process that can be performed by homology and ab initio with or without electron density maps. Those include the interactive editing of a secondary structure and a searchable, embedded library of annotated tertiary structures. Assemble helps users with performing recurrent and otherwise tedious tasks in structural RNA research. Availability and Implementation: Assemble is released under an open-source license (MIT license) and is freely available at http://bioinformatics.org/assemble. It is implemented in the Java language and runs on MacOSX, Linux and Windows operating systems. Contact: f. jossinet@ibmc-cnrs. unistra. fr