RNA sequencing (RNA-Seq) is increasingly used alongside exome and genome sequencing to identify causal variants underlying rare Mendelian disorders. We present short-read RNA-Seq data from 5,412 individuals with a diverse range of rare disorders recruited to Genomics England's 100,000 Genomes Project. We show that the proportion of genes from gene panels applied to different disorders which are well captured (transcripts per million (TPM) ≥ 5) from blood RNA varies widely, highlighting differences in applicability across disorder types. Using OUTRIDER and FRASER2 to identify gene expression and splicing outliers respectively, we identify at least one outlier event in a disorder relevant gene in 20% of the cohort. To prioritise likely diagnostic candidates, we apply multiple strategies including focussing on outlier events in known haploinsufficient genes (n=78), integrating outliers with structural variant calls (n=19), and using strategies integrating phenotypic presentation (Exomiser, n=39). We present a series of candidate diagnoses involving diverse variant types and disease mechanisms, demonstrating the broad utility of RNA-Seq in identifying and prioritising diagnostic candidates in individuals with a variety of different rare conditions and no known genetic diagnosis. Our findings demonstrate that blood-based RNA-Seq can deliver clinically relevant findings across a broad range of rare disorders.
Despite the introduction of genome sequencing (GS) for rare disease diagnostics, a genetic cause is not identified in most patients. Here, we explored the potential of proteomics to improve the diagnostic yield in 424 patients with rare diseases from the 100,000 Genomes Project (100kGP) without a genetic diagnosis. Serum proteomic profiling was performed using the Olink Explore 1536 assay ( N = 1463 proteins). For 13 patients without genetic diagnoses, detection of lower serum protein “outliers” ( z -score < −2) led to confirmed genetic diagnoses by resolving variants of uncertain significance or prioritizing genes for targeted GS reanalysis. For 23 additional patients without genetic diagnoses (64% of findings), we identified candidate gene-disease links and variants through convergent evidence from lower protein outliers and variants ranked through the variant prioritization tool Exomiser. For example, we identified a candidate heterozygous missense variant [Genome Aggregation Database (gnomAD) minor allele frequency = 0.006%] in tyrosine kinase with immunoglobulin-like and epidermal growth factor homology domains 1 ( TIE1 ) that was only present in a patient with lower TIE1 serum abundance ( z -score = −5.12) and their father, both of whom were affected by the same monogenic cardiac disorder, but in no other individuals from the 100kGP. Missense (52.5%) and splice region (27.5%) variants accounted for most diagnostic or candidate variants prioritized. This proof-of-principle study demonstrated that serum proteomics can support rare disease diagnosis and identify disease-causing genes in patients undiagnosed after GS, although successful implementation will likely depend on tissue specificity of protein expression, detectability in blood, proteomic platform coverage, and sensitivity.
AIMS:Whole-genome sequencing (WGS) is beginning to be applied to cancer samples in the clinical setting. This ideally requires high-quality, minimally degraded DNA of high tumour cell content, while retaining sufficient tissue with excellent morphology for histopathological diagnosis and immunohistochemistry. The aim of this study was to investigate alternative ways of handling cancer samples to fulfil both diagnostic and molecular requirements. METHODS:Ex vivo biopsies were taken to investigate the feasibility of using cancer cells 'shaken' from the surface of a biopsy for WGS, while maintaining the tissue biopsy for histological diagnosis. WGS from the shaken cells was compared with the gold standard of a fresh-frozen (FF) biopsy. The procedure was piloted in the real-world setting for breast cancer samples. RESULTS:Cells shaken from ex vivo biopsies can yield DNA of sufficient quantity and quality for WGS, while having no discernible impact on quality of tissue morphology. WGS data showed good coverage, comparable variant calls and generally higher tumour content in shaken cell samples compared with the control FF samples. For real-world biopsies, DNA yields were lower, but WGS data were of excellent quality for the cases analysed. CONCLUSIONS:Shaken biopsy sampling allows genomic sequencing from patients with cancer who may otherwise not receive a genome sequence due to limited sample availability. It represents a way of overcoming the logistics of obtaining and storing FF tissue making it a suitable technique for wider scale implementation in the clinical setting.
Accurate detection of somatic structural variants (SVs) and somatic copy number aberrations (SCNAs) is critical to study the mutational processes underpinning cancer evolution. Here we describe SAVANA, an algorithm designed to detect somatic SVs and SCNAs at single-haplotype resolution and estimate tumor purity and ploidy using long-read sequencing data with or without a germline control sample. We also establish best practices for benchmarking SV detection algorithms across the entire genome in a data-driven manner using replication and read-backed phasing analysis. Through the analysis of matched Illumina and nanopore whole-genome sequencing data for 99 human tumor-normal pairs, we show that SAVANA has significantly higher sensitivity and 13- and 82-times-higher specificity than the second and third-best performing algorithms. Moreover, SVs reported by SAVANA are highly consistent with those detected using short-read sequencing. In summary, SAVANA enables the application of long-read sequencing to detect SVs and SCNAs reliably.
In susceptible patients, COVID-19 causes life-threatening disease driven by immune-mediated inflammatory lung injury. We have previously shown that multiple common host genetic variants are significantly associated with susceptibility to critical Covid-19, and in one case, we demonstrated that such variants can inform development of new, effective drug treatment. Here we report an association analysis of whole-genome sequences (WGS) from 11,423 cases from the GenOMICC study and 60,628 controls, together with meta-analyses with available genome-wide data. We identify a rare association signal at SLC50A1, primarily driven by a missense variant rs147850817 (1:155138217:G:T, Arg201Leu) that may interfere with transport function, and we identify four common association signals near ARF1, ZNF462, KLF13 and MVP genes. Finally, we build a WGS-derived polygenic risk score (PRS) for critical Covid-19, which offers only marginal improvement in risk estimation for the general population but may provide clinically-valuable discrimination for extreme susceptibility. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Protocols ### Funding Statement GenOMICC was funded by Sepsis Research (the Fiona Elizabeth Agnew Trust), the Intensive Care Society, a Wellcome Trust Senior Research Fellowship (J.K.Baillie, 223164/Z/21/Z), the Department of Health and Social Care (DHSC), Illumina, LifeArc, the Medical Research Council, UKRI, a BBSRC Institute Strategic Program Support Grant to the Roslin Institute (BBS/E/D/20002172, BBS/E/D/10002070 and BBS/E/D/30002275) and UKRI grants MC PC 20004, MC PC 19025, MC PC 1905, and MRNO2995X/1. ADB acknowledges funding from the Wellcome PhD training fellowship for clinicians (204979/Z/16/Z), the Edinburgh Clinical Academic Track (ECAT) programme. This research is supported in part by the Data and Connectivity National Core Study, led by Health Data Research UK in partnership with the Office for National Statistics and funded by UK Research and Innovation (grant ref MC PC 20029). This study owes a great deal to the National Institute for Healthcare Research Clinical Research Network (NIHR CRN) and the Chief Scientist's Office (Scotland), who facilitate recruitment into research studies in NHS hospitals, and to the global ISARIC and InFACT consortia. This work forms part of the translational research portfolio of the National Institute for Health and Care Research Barts Biomedical Research Centre. T.M. is supported by Cancer Research UK grant DRCRPG-May23/100002 to C. Siebold. Genomics England: This research was made possible through access to data in the National Genomic Research Library, which is managed by Genomics England Limited (a wholly owned company of the Department of Health and Social Care). The National Genomic Research Library (\url{https://www.genomicsengland.co.uk/research}) holds data provided by patients and collected by the NHS as part of their care and data collected as part of their participation in research. The National Genomic Research Library is funded by the National Institute for Health Research and NHS England. The Wellcome Trust, Cancer Research UK and the Medical Research Council have also funded research infrastructure. REACT: National Institute for Health and Care Research (NIHR) and UK Research and Innovation (UKRI) - REACT-Genomics England (REACT-GE) (MR/V030841/1) and REACT-Long COVID (REACT-LC) (COV-LT-0040). The REACT study was funded by the UK Department of Health and Social Care with supplemental funding from the Huo Family Foundation. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: GenOMICC was approved by the following research ethics committees: Scotland A Research Ethics Committee (15/SS/0110) and Coventry and Warwickshire Research Ethics Committee (England, Wales and Northern Ireland) (19/WM/0247). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All other data produced in the present study are available upon reasonable request to the authors
Osteosarcoma is the most common primary cancer of the bone, with a peak incidence in children and young adults. Using multi-region whole-genome sequencing, we find that chromothripsis is an ongoing mutational process, occurring subclonally in 74% of osteosarcomas. Chromothripsis generates highly unstable derivative chromosomes, the ongoing evolution of which drives the acquisition of oncogenic mutations, clonal diversification, and intra-tumor heterogeneity across diverse sarcomas and carcinomas. In addition, we characterize a new mechanism, termed loss-translocation-amplification (LTA) chromothripsis, which mediates punctuated evolution in about half of pediatric and adult high-grade osteosarcomas. LTA chromothripsis occurs when a single double-strand break triggers concomitant TP53 inactivation and oncogene amplification through breakage-fusion-bridge cycles. It is particularly prevalent in osteosarcoma and is not detected in other cancers driven by TP53 mutation. Finally, we identify the level of genome-wide loss of heterozygosity as a strong prognostic indicator for high-grade osteosarcoma.
The identification of structural variants (SVs) in genomic data represents an ongoing challenge because of difficulties in reliable SV calling leading to reduced sensitivity and specificity. We prepared high-quality DNA from 9 parent–child trios, who had previously undergone short-read whole-genome sequencing (Illumina platform) as part of the Genomics England 100,000 Genomes Project. We reanalysed the genomes using both Bionano optical genome mapping (OGM; 8 probands and one trio) and Nanopore long-read sequencing (Oxford Nanopore Technologies [ONT] platform; all samples). To establish a “truth” dataset, we asked whether rare proband SV calls (n = 234) made by the Bionano Access (version 1.6.1)/Solve software (version 3.6.1_11162020) could be verified by individual visualisation using the Integrative Genomics Viewer with either or both of the Illumina and ONT raw sequence. Of these, 222 calls were verified, indicating that Bionano OGM calls have high precision (positive predictive value 95%). We then asked what proportion of the 222 true Bionano SVs had been identified by SV callers in the other two datasets. In the Illumina dataset, sensitivity varied according to variant type, being high for deletions (115/134; 86%) but poor for insertions (13/58; 22%). In the ONT dataset, sensitivity was generally poor using the original Sniffles variant caller (48% overall) but improved substantially with use of Sniffles2 (36/40; 90% and 17/23; 74% for deletions and insertions, respectively). In summary, we show that the precision of OGM is very high. In addition, when applying the Sniffles2 caller, the sensitivity of SV calling using ONT long-read sequence data outperforms Illumina sequencing for most SV types.
Whole genome sequencing (WGS) provides comprehensive, individualised cancer genomic information. However, routine tumour biopsies are formalin-fixed and paraffin-embedded (FFPE), damaging DNA, historically limiting their use in WGS. Here we analyse FFPE cancer WGS datasets from England's 100,000 Genomes Project, comparing 578 FFPE samples with 11,014 fresh frozen (FF) samples across multiple tumour types. We use an approach that characterises rather than discards artefacts. We identify three artefactual signatures, including one known (SBS57) and two previously uncharacterised (SBS FFPE, ID FFPE), and develop an "FFPEImpact" score that quantifies sample artefacts. Despite inferior sequencing quality, FFPE-derived data identifies clinically-actionable variants, mutational signatures and permits algorithmic stratification. Matched FF/FFPE validation cohorts shows good concordance while acknowledging SBS, ID and copy-number artefacts. While FF-derived WGS data remains the gold standard, FFPE-samples can be used for WGS if required, using analytical advancements developed here, potentially democratising whole cancer genomics to many. Formalin fixation is commonly used in tissue storage; however, this process has traditionally limited downstream whole genome sequencing usage. Here, the authors identify artefactual signatures in FFPE-derived sequencing data and demonstrate the preservation of clinical utility, thus enabling FFPE whole genome sequencing when required.
Abstract Despite the recent advances in genomic analysis, causative variants cannot be found for a sizeable proportion of patients with suspected genetic disorders. Many of these disorders involve genes in difficult-to-align genomic regions which are recalcitrant to short read approaches. Structural variants in these regions can be particularly hard to detect or define with short reads, yet may account for a significant number of cases. Long read sequencing can overcome these difficulties and is providing new hope for diagnosis and patient care. Here, we present a case of unusually complex, severe fatigue where a potentially relevant structural variant was indicated but could not be resolved by short-read sequencing. We use nanopore sequencing to identify and fully characterise a large inversion in a highly homologous region spanning the AKR1C gene locus, along with serum steroid analysis to investigate the functional consequences. The DNA inversion appears to increase the expression of AKR1C2 while limiting AKR1C1 activity, resulting in a relative increase of inhibitory neurosteroids and impaired progesterone metabolism. This study provides an example of where long read sequencing may supplement the use of more traditional sequencing methods in clinical care to increase diagnostic yield for rare disease, and highlights some of the challenges that arise in sequencing complex regions containing tandem arrays of genes. It also proposes a novel gene associated with a specific disease aetiology that may be an underlying cause of unexplained severe fatigue.
BACKGROUND:Causative genetic variants cannot yet be found for many disorders with a clear heritable component, including chronic fatigue disorders like myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS). These conditions may involve genes in difficult-to-align genomic regions that are refractory to short read approaches. Structural variants in these regions can be particularly hard to detect or define with short reads, yet may account for a significant number of cases. Long read sequencing can overcome these difficulties but so far little data is available regarding the specific analytical challenges inherent in such regions, which need to be taken into account to ensure that variants are correctly identified. Research into chronic fatigue disorders faces the additional challenge that the heterogeneous patient populations likely encompass multiple aetiologies with overlapping symptoms, rather than a single disease entity, such that each individual abnormality may lack statistical significance within a larger sample. Better delineation of patient subgroups is needed to target research and treatment.METHODS:We use nanopore sequencing in a case of unexplained severe fatigue to identify and fully characterise a large inversion in a highly homologous region spanning the AKR1C gene locus, which was indicated but could not be resolved by short-read sequencing. We then use GC-MS/MS serum steroid analysis to investigate the functional consequences.RESULTS:Several commonly used bioinformatics tools are confounded by the homology but a combined approach including visual inspection allows the variant to be accurately resolved. The DNA inversion appears to increase the expression of AKR1C2 while limiting AKR1C1 activity, resulting in a relative increase of inhibitory GABAergic neurosteroids and impaired progesterone metabolism which could suppress neuronal activity and interfere with cellular function in a wide range of tissues.CONCLUSIONS:This study provides an example of how long read sequencing can improve diagnostic yield in research and clinical care, and highlights some of the analytical challenges presented by regions containing tandem arrays of genes. It also proposes a novel gene associated with a novel disease aetiology that may be an underlying cause of complex chronic fatigue. It reveals biomarkers that could now be assessed in a larger cohort, potentially identifying a subset of patients who might respond to treatments suggested by the aetiology.
Abstract Despite the recent advances in genomic analysis, causative variants cannot be found for a sizeable proportion of patients with suspected genetic disorders. Many of these disorders involve genes in difficult-to-align genomic regions which are recalcitrant to short read approaches. Structural variants in these regions can be particularly hard to detect or define with short reads, yet may account for a significant number of cases. Long read sequencing can overcome these difficulties and is providing new hope for diagnosis and patient care. Here, we present a case of unusually complex, severe fatigue where a potentially relevant structural variant was indicated but could not be resolved by short-read sequencing. We use nanopore sequencing to identify and fully characterise a large inversion in a highly homologous region spanning the AKR1C gene locus, along with serum steroid analysis to investigate the functional consequences. The DNA inversion appears to increase the expression of AKR1C2 while limiting AKR1C1 activity, resulting in a relative increase of inhibitory neurosteroids and impaired progesterone metabolism. This study provides an example of where long read sequencing may supplement the use of more traditional sequencing methods in clinical care to increase diagnostic yield for rare disease, and highlights some of the challenges that arise in sequencing complex regions containing tandem arrays of genes. It also proposes a novel gene associated with a specific disease aetiology that may be an underlying cause of unexplained severe fatigue.
Osteosarcoma is the most common primary cancer of bone with a peak incidence in children and young adults. Despite progress, the genomic aberrations underpinning osteosarcoma evolution remain poorly understood. Using multi-region whole-genome sequencing, we find that chromothripsis is an ongoing mutational process, occurring subclonally in 74% of tumours. Chromothripsis drives the acquisition of oncogenic mutations and generates highly unstable derivative chromosomes, the evolution of which drives clonal diversification and intra-tumour heterogeneity. In addition, we report a novel mechanism, loss-translocation-amplification (LTA) chromothripsis, which mediates rapid malignant transformation and punctuated evolution in about half of paediatric and adult high-grade osteosarcomas. Specifically, a single double-strand break triggers concomitant TP53 inactivation and segmental amplifications, often amplifying oncogenes to high copy numbers in extrachromosomal circular DNA elements through breakage-fusion-bridge cycles involving multiple chromosomes. LTA chromothripsis is detected at low frequency in soft-tissue sarcomas, but not in epithelial cancers, including those driven by TP53 mutation. Finally, we identify genome-wide loss of heterozygosity as a strong prognostic indicator for high-grade osteosarcoma.
DNA transfer from cytoplasmic organelles to the cell nucleus is a legacy of the endosymbiotic event-the majority of nuclear-mitochondrial segments (NUMTs) are thought to be ancient, preceding human speciation1-3. Here we analyse whole-genome sequences from 66,083 people-including 12,509 people with cancer-and demonstrate the ongoing transfer of mitochondrial DNA into the nucleus, contributing to a complex NUMT landscape. More than 99% of individuals had at least one of 1,637 different NUMTs, with 1 in 8 individuals having an ultra-rare NUMT that is present in less than 0.1% of the population. More than 90% of the extant NUMTs that we evaluated inserted into the nuclear genome after humans diverged from apes. Once embedded, the sequences were no longer under the evolutionary constraint seen within the mitochondrion, and NUMT-specific mutations had a different mutational signature to mitochondrial DNA. De novo NUMTs were observed in the germline once in every 104 births and once in every 103 cancers. NUMTs preferentially involved non-coding mitochondrial DNA, linking transcription and replication to their origin, with nuclear insertion involving multiple mechanisms including double-strand break repair associated with PR domain zinc-finger protein 9 (PRDM9) binding. The frequency of tumour-specific NUMTs differed between cancers, including a probably causal insertion in a myxoid liposarcoma. We found evidence of selection against NUMTs on the basis of size and genomic location, shaping a highly heterogenous and dynamic human NUMT landscape.
Critical COVID-19 is caused by immune-mediated inflammatory lung injury. Host genetic variation influences the development of illness requiring critical care1 or hospitalization2-4 after infection with SARS-CoV-2. The GenOMICC (Genetics of Mortality in Critical Care) study enables the comparison of genomes from individuals who are critically ill with those of population controls to find underlying disease mechanisms. Here we use whole-genome sequencing in 7,491 critically ill individuals compared with 48,400 controls to discover and replicate 23 independent variants that significantly predispose to critical COVID-19. We identify 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2). Using transcriptome-wide association and colocalization to infer the effect of gene expression on disease severity, we find evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease. Mendelian randomization provides evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets. Our results are broadly consistent with a multi-component model of COVID-19 pathophysiology, in which at least two distinct mechanisms can predispose to life-threatening disease: failure to control viral replication; or an enhanced tendency towards pulmonary inflammation and intravascular coagulation. We show that comparison between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Titins, giant sarcomere proteins with major mechanical/signaling functions, are expressed in 2 main isoform classes in the mammalian heart: N2B (3000 kDa) and N2BA (>3200 kDa). A dramatic isoform switch occurs during cardiac development, from fetal N2BA titin (3700 kDa) expressed before birth to a mix of smaller N2BA/N2B isoforms found postnatally; adult rat hearts almost exclusively have N2B titin. The isoform switch, which can be reversed in chronic human heart failure, alters myocardial distensibility and mechanosignaling. Here we determined factors regulating this switch using, as a model system, primary cardiomyocyte cultures prepared from embryonic rats. In standard culture, the mean N2B percentage initially was 14% and increased by ≈60% within 1 week, resembling the in vivo switching. The titin isoform transition was independent of endothelin-1–induced myocyte hypertrophy and was not altered by pacing, contractile arrest, or cell stretch; however, it was modestly impaired by decreasing substrate rigidity and strongly dependent on serum components. Angiotensin II significantly promoted the transition. The mean N2B proportion in 1-week-old cultures dropped 20% to 25% in hormone-reduced medium, but addition of 3,5,3′-triiodo-l-thyronine (T3) nearly restored the proportion to that found in standard culture. This T3 effect was not prevented by bisphenol A, a specific inhibitor of the classic genomic pathway of T3 action. In contrast, the titin switch could be stalled by the phosphatidylinositol 3-kinase inhibitor LY294002, which decreased the proportion of N2B mRNA transcripts within hours and suppressed a rapid T3-induced increase in Akt phosphorylation. Also, angiotensin II, but not endothelin-1 or cell stretch, enhanced Akt phosphorylation. Thus, although matrix stiffness modulates developmental titin isoform transitions, these transitions are mainly regulated through phosphatidylinositol 3-kinase/Akt-dependent signaling triggered particularly by T3 via a rapid action pathway.