The D4Z4 locus is a macrosatellite array on Chromosome 4q normally comprising 8 to >100 3.3-kb repeat units. Its size and repetitiveness render it refractory to most sequencing technologies; consequently, its genetic and epigenetic architectures remain incompletely understood despite their relevance to facioscapulohumeral muscular dystrophy (FSHD). Current FSHD molecular testing relies on complex, multistep and low-resolution assays, which aim to identify contractions on permissive haplotypes (FSHD type 1) or epigenetic reactivation due to pathogenic variants in the epigenetic machinery, most often in SMCHD1 (FSHD type 2). Recent guideline updates highlight the need for more accurate and comprehensive diagnostic approaches. Here, we leverage ultra-long whole-genome and Cas9-targeted sequencing to develop a fast and accurate workflow, D4Z4End2End, for comprehensive genetic and methylation analysis of D4Z4 alleles. We apply it to samples from two controls, four FSHD1 patients, four FSHD2 patients, and two patients with Bosma arhinia microphthalmia syndrome (BAMS) caused by SMCHD1 variants, as well as publicly available data from 30 B-lymphoblastoid cell lines from the 1000 Genomes Project and Human Pangenome Reference Consortium. We attain high-depth sequencing of full-length D4Z4 arrays of up to 40 repeat units (∼132 kb), accurately capture contracted arrays, genetic mosaicism, and pathogenic SMCHD1 variants, and generate consensus sequences of all D4Z4 alleles. We identify new allelic variants, analyze complex D4Z4 rearrangements including in-cis duplications, and reveal length- and SMCHD1-dependent methylation patterns across the D4Z4 array. Our findings offer insights into D4Z4 genetics and epigenetics, and demonstrate the potential of long-read nanopore sequencing to accelerate FSHD research and diagnostics.
Long-read RNA sequencing enables full-length transcript profiling and improved isoform resolution, but variable platforms and evolving chemistries demand careful benchmarking for reliable application. We present LongBench , a matched, multi-platform reference dataset spanning bulk, single-cell, and single-nucleus transcriptomics across eight human lung cancer cell lines with synthetic spike-in controls. LongBench incorporates three state-of-the-art long-read protocols alongside Illumina short reads: Oxford Nanopore Technologies (ONT) PCR-cDNA, ONT direct RNA, and PacBio Kinnex. We systematically evaluate transcript capture, quantification accuracy, differential expression, isoform usage, variant detection, and allele-specific analyses. Our results show high concordance in gene-level differential analyses across protocols, but reduced consistency for transcript-level and isoform analyses due to lengthand platform-dependent biases. Single-cell long-read data are highly concordant with bulk for high-confidence features, though single-nuclei data show reduced feature detection. LongBench provides one of the largest publicly available long-read benchmarking resources, enabling rigorous cross-platform evaluation and guiding technology selection for transcriptomic research. ### Competing Interest Statement The authors acknowledge the support of both PacBio and Oxford Nanopore Technologies who provided sequencing reagents used in the single-cell / single-nuclei arm of this study. Y.Y, K.Z, M.B.C. and Q.G. received travel support from Oxford Nanopore Technologies to attend conferences. Q.G. received travel support from PacBio to attend a conference. The authors have no other competing interests, and the Funders had no involvement in study design, data analysis, interpretation and writing of the article, or the decision to submit the work for publication. Australian National Health and Medical Research Council Medical Research Future Fund Researcher Exchange and Development in Industry Fellowship Research Foundation-Flanders Dutch MS Research Foundation KU Leuven (BOF-FKO, Bijzonder Onderzoeksfonds – Fundamenteel Klinisch Onderzoeker) Victorian State Government Operational Infrastructure Support Australian Cancer Research Foundation
Long-read sequencing technologies have transformed the field of epigenetics by enabling direct, single-base resolution detection of DNA modifications, such as methylation. This produces novel opportunities for studying the role of DNA methylation in gene regulation, imprinting, and disease. However, the unique characteristics of long-read data, including the modBAM format and extended read lengths, necessitate the development of specialised software tools for effective analysis. The NanoMethViz package provides a suite of tools for loading in long-read methylation data, visualising data at various data resolutions. It can convert the data for use with other Bioconductor software such as bsseq, DSS, dmrseq and edgeR to discover differentially methylated regions (DMRs). In this workflow article, we demonstrate the process of converting modBAM files into formats suitable for comprehensive downstream analysis. We leverage NanoMethViz to conduct an exploratory analysis, visually summarizing differences between samples, examining aggregate methylation profiles across gene and CpG islands, and investigating methylation patterns within specific regions at the single-read level. Additionally, we illustrate the use of dmrseq for identifying DMRs and show how to integrate these findings into gene-level visualization plots. Our analysis is applied to a triplicate dataset of haplotyped long-read methylation data from mouse neural stem cells, allowing us to visualize and compare the characteristics of the parental alleles on chromosome 7. By applying DMR analysis, we recover DMRs associated with known imprinted genes and visualise the methylation patterns of these genes summarised at single-read resolution. Through DMR analysis, we identify DMRs associated with known imprinted genes and visualize their methylation patterns at single-read resolution. This streamlined workflow is adaptable to common experimental designs and offers flexibility in the choice of upstream data sources and downstream statistical analysis tools.
Spatial transcriptomics technology has developed rapidly in recent years, with various sequencing-based platforms such as 10× Visium, Slide-seq, and Stereo-seq becoming widely used by researchers. Each platform brings its own set of protocols and customized data analysis pipelines, which presents challenges when the goal is to obtain uniformly preprocessed data that is conveniently formatted for downstream analysis. To address the need for simpler, open-source solutions that deal with sequencing-based spatial transcriptomics (sST) data from different platforms, we present stPipe, a comprehensive and modular preprocessing pipeline for all current mainstream sST platforms. stPipe is implemented as an R/Bioconductor package that handles various analysis steps, including (i) data processing from raw FASTQ files to create a spatially resolved gene count matrix; (ii) the collation of relevant quality control metrics to ensure unwanted artifacts can be filtered; and (iii) the adoption of standardized data storage containers to allow results to be easily passed on to a wide range of downstream analysis packages. A key use case for stPipe is in methods benchmarking, and we demonstrate how the uniform processing of sST data collected on reference tissue samples from the cadasSTre and SpatialBenchVisium projects is made easier, allowing comparisons between different technology platforms and downstream analysis tools.
Long-read sequencing technologies have transformed the field of epigenetics by enabling direct, single-base resolution detection of DNA modifications, such as methylation. This produces novel opportunities for studying the role of DNA methylation in gene regulation, imprinting, and disease. However, the unique characteristics of long-read data, including the modBAM format and extended read lengths, necessitate the development of specialised software tools for effective analysis. The NanoMethViz package provides a suite of tools for loading in long-read methylation data, visualising data at various data resolutions. It can convert the data for use with other Bioconductor software such as bsseq, DSS, dmrseq and edgeR to discover differentially methylated regions (DMRs). In this workflow article, we demonstrate the process of converting modBAM files into formats suitable for comprehensive downstream analysis. We leverage NanoMethViz to conduct an exploratory analysis, visually summarizing differences between samples, examining aggregate methylation profiles across gene and CpG islands, and investigating methylation patterns within specific regions at the single-read level. Additionally, we illustrate the use of dmrseq for identifying DMRs and show how to integrate these findings into gene-level visualization plots. Our analysis is applied to a triplicate dataset of haplotyped long-read methylation data from mouse neural stem cells, allowing us to visualize and compare the characteristics of the parental alleles on chromosome 7. By applying DMR analysis, we recover DMRs associated with known imprinted genes and visualise the methylation patterns of these genes summarised at single-read resolution. Through DMR analysis, we identify DMRs associated with known imprinted genes and visualize their methylation patterns at single-read resolution. This streamlined workflow is adaptable to common experimental designs and offers flexibility in the choice of upstream data sources and downstream statistical analysis tools.
X-linked genetic disorders typically affect females less severely than males due to the presence of a second X chromosome not carrying the deleterious variant. However, the phenotypic expression in females is highly variable, which may be explained by an allelic skew in X chromosome inactivation. Accurate measurement of X inactivation skew is crucial to understand and predict disease phenotype in carrier females, with prediction especially relevant for degenerative conditions.We propose a novel approach using nanopore sequencing to quantify skewed X inactivation accurately. By phasing sequence variants and methylation patterns, this single assay reveals the disease variant, X inactivation skew, its directionality, and is applicable to all patients and X-linked variants. Enrichment of X-chromosome reads through adaptive sampling enhances cost-efficiency. Our study includes a cohort of 16 X-linked variant carrier females affected by two X-linked inherited retinal diseases: choroideremia andRPGR-associated retinitis pigmen-tosa. As retinal DNA cannot be readily obtained, we instead determine the skew from peripheral samples (blood, saliva and buccal mucosa), and correlate it to phenotypic outcomes. This revealed a strong correlation between X inactivation skew and disease presentation, confirming the value in performing this assay and its potential as a way to prioritise patients for early intervention, such as gene therapy currently in clinical trials for these conditions.Our method of assessing skewed X inactivation is applicable to all long-read genomic datasets, providing insights into disease risk and severity and aiding in the development of individualised strategies for X-linked variant carrier females.
SUMMARY Skeletal muscle contains a resident population of somatic stem cells capable of both self-renewal and differentiation. The signals that regulate this important decision have yet to be fully elucidated. Here we use metabolomics and mass spectrometry imaging (MSI) to identity a state of localized hyperglycaemia following skeletal muscle injury. We show that committed muscle progenitor cells exhibit an enrichment of glycolytic and TCA cycle genes and that extracellular monosaccharide availability regulates intracellular citrate levels and global histone acetylation. Muscle stem cells exposed to a reduced (or altered) monosaccharide environment demonstrate reduced global histone acetylation and transcription of myogenic determination factors (including myod1 ). Importantly, reduced monosaccharide availability was linked directly to increased rates of asymmetric division and muscle stem cell self-renewal in regenerating skeletal muscle. Our results reveal an important role for the extracellular metabolic environment in the decision to undergo self-renewal or myogenic commitment during skeletal muscle regeneration.
scPipe is a flexible R/Bioconductor package originally developed to analyse platform-independent single-cell RNA-Seq data. To expand its preprocessing capability to accommodate new single-cell technologies, we further developed scPipe to handle single-cell ATAC-Seq and multi-modal (RNA-Seq and ATAC-Seq) data. After executing multiple data cleaning steps to remove duplicated reads, low abundance features and cells of poor quality, a SingleCellExperiment object is created that contains a sparse count matrix with features of interest in the rows and cells in the columns. Quality control information (e.g. counts per cell, features per cell, total number of fragments, fraction of fragments per peak) and any relevant feature annotations are stored as metadata. We demonstrate that scPipe can efficiently identify “true” cells and provides flexibility for the user to fine-tune the quality control thresholds using various feature and cell-based metrics collected during data preprocessing. Researchers can then take advantage of various downstream single-cell tools available in Bioconductor for further analysis of scATAC-Seq data such as dimensionality reduction, clustering, motif enrichment, differential accessibility and cis-regulatory network analysis. The scPipe package enables a complete beginning-to-end pipeline for single-cell ATAC-Seq and RNA-Seq data analysis in R.
Abstract Schwann Cell Precursors (SCPs) are multipotent precursor cells and express SOX10, but relatively little is known about the specification and molecular make-up of human SCPs. To address this, we subjected a human SOX10 knock-in reporter iPSC line to a one step cranial Schwann cell precursor differentiation protocol. We show that SOX10-expressing cells acquire the morphology, motility and gene expression of SCP-like cells, and we exemplify that these cells migrate to multiple developing craniofacial structures following injection into early mouse embryos. We next defined the bulk and single-cell transcriptomes, metabolomes, and proteomes of SOX10 expressing human SCP-like cells, revealing gene expression heterogeneity, shifts in splicing events, a reduced dependence on glycolysis, and changes in the expression of cell adhesion proteins, miRNAs and LncRNAs that accompany SCP specification. Discovery proteomics identifies MCAM (CD146) as a cell surface marker that permits the isolation of pure SOX10 expressing human cranial SCP-like cells from human pluripotent stem cell lines subjected to SCP differentiation. We further show that this permits isolation of Down syndrome cranial SCP-like cells that display defects in proliferation. Collectively these data provide a detailed map of the molecular make-up of in vitro generated human cranial Schwann cell precursor-like cells and exemplify the utility of MCAM-sorted cranial SCP-like cells for modelling human neurocristopathies.
Differential expression (DE) analysis is one of the most common forms of analysis done using RNAseq. It performs statistical tests to determine the set of genes that have significant differences between experimental groups, linking the epxression of genes to disease or treatment effects. The classic presentation of results involves static plots thousands of points and large tables, with little ability for a non-computational scientist to associate points on plots to the data in the table. The original Glimma R/Bioconductor package provided the ability to produce interactive plots with linked interactions between the plots and table of results, allowing non-computational end-users to explore points on the plot and explore the gene and associated statistic, as well as search the genes in the table and locate its position in the plots. It produced standalone HTML pages with data files that can be opened with any modern web-browser, making it high accessible for all researchers. Glimma V2 is a re-write of the original Glimma library using more modern Javascript libraries and the R htmlwidgets package, enabling new features to be included along with integration into Rmarkdown reports for better reproducibility and portability.
Pooled short hairpin RNA sequencing (shRNA-seq) screens are becoming increasingly popular in functional genomics research, and there is a need to establish optimal analysis tools to handle such data. Our open-source shRNA processing pipeline in edgeR provides a complete analysis solution for shRNA-seq screen data, that begins with the raw sequence reads and ends with a ranked lists of candidate shRNAs for downstream biological validation. We first summarize the raw data contained in a fastq file into a matrix of counts (samples in the columns, hairpins in the rows) with options for allowing mismatches and small shifts in hairpin position. Diagnostic plots, normalization and differential representation analysis can then be performed using established methods to prioritize results in a statistically rigorous way, with the choice of either the classic exact testing methodology or a generalized linear modelling that can handle complex experimental designs. A detailed users’ guide that demonstrates how to analyze screen data in edgeR along with a pointand-click implementation of this workflow in Galaxy are also provided. The edgeR package is freely available from http://www.bioconductor.org. This article is included in the Bioconductor gateway. Open Peer Review
A key benefit of long-read nanopore sequencing technology is the ability to detect modified DNA bases, such as 5-methylcytosine. The lack of R/Bioconductor tools for the effective visualization of nanopore methylation profiles between samples from different experimental groups led us to develop the NanoMethViz R package. Our software can handle methylation output generated from a range of different methylation callers and manages large datasets using a compressed data format. To fully explore the methylation patterns in a dataset, NanoMethViz allows plotting of data at various resolutions. At the sample-level, we use dimensionality reduction to look at the relationships between methylation profiles in an unsupervised way. We visualize methylation profiles of classes of features such as genes or CpG islands by scaling them to relative positions and aggregating their profiles. At the finest resolution, we visualize methylation patterns across individual reads along the genome using the spaghetti plotand heatmaps, allowing users to explore particular genes or genomic regions of interest. In summary, our software makes the handling of methylation signal more convenient, expands upon the visualization options for nanopore data and works seamlessly with existing methylation analysis tools available in the Bioconductor project. Our software is available at https://bioconductor.org/packages/NanoMethViz.
CD1c presents lipid-based antigens to CD1c-restricted T cells, which are thought to be a major component of the human T cell pool. However, the study of CD1c-restricted T cells is hampered by the presence of an abundantly expressed, non-T cell receptor (TCR) ligand for CD1c on blood cells, confounding analysis of TCR-mediated CD1c tetramer staining. Here, we identified the CD36 family (CD36, SR-B1, and LIMP-2) as ligands for CD1c, CD1b, and CD1d proteins and showed that CD36 is the receptor responsible for non-TCR-mediated CD1c tetramer staining of blood cells. Moreover, CD36 blockade clarified tetramer-based identification of CD1c-restricted T cells and improved identification of CD1b- and CD1d-restricted T cells. We used this technique to characterize CD1c-restricted T cells ex vivo and showed diverse phenotypic features, TCR repertoire, and antigen-specific subsets. Accordingly, this work will enable further studies into the biology of CD1 and human CD1-restricted T cells.
Application of Oxford Nanopore Technologies’ long-read sequencing platform to transcriptomic analysis is increasing in popularity. However, such analysis can be challenging due to small library sizes and high sequence error, which decreases quantification accuracy and reduces power for statistical testing. Here, we report the analysis of two nanopore sequencing RNA-seq datasets with the goal of obtaining gene-level and isoform-level differential expression information. A dataset of synthetic, spliced, spike-in RNAs (“sequins”) as well as a mouse neural stem cell dataset from samples with a null mutation of the epigenetic regulator Smchd1 were analysed using a mix of long-read specific tools for preprocessing together with established short-read RNA-seq methods. We used limma-voom to perform differential gene expression analysis, and the novel FLAMES pipeline to perform isoform identification and quantification, followed by DRIMSeq and limma-diffSplice (with stageR ) to perform differential transcript usage analysis. We compared results from the sequins dataset to the ground truth, and results of the mouse dataset to a previous short-read study on equivalent samples. Overall, our work shows that transcriptomic analysis of long-read nanopore data using short-read software and methods that are already in wide use can yield meaningful results.
Despite advances in single-cell multi-omics, a single stem or progenitor cell can only be tested once. We developed clonal multi-omics, in which daughters of a clone act as surrogates of the founder, thereby allowing multiple independent assays per clone. With SIS-seq, clonal siblings in parallel "sister" assays are examined either for gene expression by RNA sequencing (RNA-seq) or for fate in culture. We identified, and then validated using CRISPR, genes that controlled fate bias for different dendritic cell (DC) subtypes. This included Bcor as a suppressor of plasmacytoid DC (pDC) and conventional DC type 2 (cDC2) numbers during Flt3 ligand-mediated emergency DC development. We then developed SIS-skew to examine development of wild-type and Bcor-deficient siblings of the same clone in parallel. We found Bcor restricted clonal expansion, especially for cDC2s, and suppressed clonal fate potential, especially for pDCs. Therefore, SIS-seq and SIS-skew can reveal the molecular and cellular mechanisms governing clonal fate.
Glimma 1.0 introduced intuitive, point-and-click interactive graphics for differential gene expression analysis. Here, we present a major update to Glimma which brings improved inter-activity and reproducibility using high-level visualisation frame-works for R and JavaScript. Glimma 2.0 plots are now readily embeddable in R Markdown, thus allowing users to create reproducible reports containing interactive graphics. The revamped multidimensional scaling plot features dashboard-style controls allowing the user to dynamically change the colour, shape and size of sample points according to different experimental conditions. Interactivity was enhanced in the MA-style plot for comparing differences to average expression, which now supports selecting multiple genes, export options to PNG, SVG or CSV formats and includes a new volcano plot function. Feature-rich and user-friendly, Glimma makes exploring data for gene expression analysis more accessible and intuitive and is available on Bioconductor and GitHub.
RNA-seq datasets can contain millions of intron reads per sequenced library that are typically removed from downstream analysis. Only reads overlapping annotated exons are considered to be informative since mature mRNA is assumed to be the major component sequenced, especially when examining poly(A) RNA samples. In this paper, we demonstrate that intron reads are informative and that pre-mRNA is the major source of intron signal. Making use of pre-mRNA signal, our index method combines differential expression analyses from intron and exon counts to categorise changes observed in each count set, giving additional genes with evidence of transcriptional changes when compared to a classic approach. Considering the importance of intron retention in some biological systems, another novel method, superintronic , looks for evidence of intron retention after accounting for the presence of pre-mRNA signal. The results presented here overcomes deficiencies and biases in previous works related to intron reads by exploring multiple sources for intron reads simultaneously using a data-driven approach, and provides a broad overview into how intron reads can be utilised in relation to multiple aspects of transcriptional biology.
Archetypal human pluripotent stem cells (hPSC) are widely considered to be equivalent in developmental status to mouse epiblast stem cells, which correspond to pluripotent cells at a late post-implantation stage of embryogenesis. Heterogeneity within hPSC cultures complicates this interspecies comparison. Here we show that a subpopulation of archetypal hPSC enriched for high self-renewal capacity (ESR) has distinct properties relative to the bulk of the population, including a cell cycle with a very low G1 fraction and a metabolomic profile that reflects a combination of oxidative phosphorylation and glycolysis. ESR cells are pluripotent and capable of differentiation into primordial germ cell-like cells. Global DNA methylation levels in the ESR subpopulation are lower than those in mouse epiblast stem cells. Chromatin accessibility analysis revealed a unique set of open chromatin sites in ESR cells. RNA-seq at the subpopulation and single cell levels shows that, unlike mouse epiblast stem cells, the ESR subset of hPSC displays no lineage priming, and that it can be clearly distinguished from gastrulating and extraembryonic cell populations in the primate embryo. ESR hPSC correspond to an earlier stage of post-implantation development than mouse epiblast stem cells.
Long-read technologies are overcoming early limitations in accuracy and throughput, broadening their application domains in genomics. Dedicated analysis tools that take into account the characteristics of long-read data are thus required, but the fast pace of development of such tools can be overwhelming. To assist in the design and analysis of long-read sequencing projects, we review the current landscape of available tools and present an online interactive database, long-read-tools.org, to facilitate their browsing. We further focus on the principles of error correction, base modification detection, and long-read transcriptomics analysis and highlight the challenges that remain.