Abstract Tandem repeat expansion disorders can be difficult to diagnose when expansions exceed 200 repeats, as standard methods (for example, Southern blot and modified PCR) often fail. We present a Cas9-targeted nanopore sequencing workflow and an automated analysis pipeline, RepeatLab, for accurate repeat-length estimation, structure assessment, and high-resolution methylation profiling. Validated on 13 myotonic dystrophy type 1 samples, 4 healthy controls, and 4 cell lines, this approach demonstrates improved sensitivity and accuracy for large expansions. Key refinements include an alternative basecalling strategy for extended repeats and a repeat-length calling algorithm that remains robust at lower sequencing throughput. The platform also automatically reports methylation near the DMPK repeat region, including five CpG site groups that could inform more nuanced clinical evaluations. This integrated workflow offers a rapid, cost-effective diagnostic solution with a turnaround time under 24 h and costs comparable to standard assays. Its compatibility with readily available computational resources enhances accessibility and scalability.
Abstract mRNA vaccine efficacy depends on sequence optimization, but designing optimal sequences is challenging due to complex cellular RNA regulatory mechanisms. Here we present VaxLab, an open-source web platform that provides the complete design-to-synthesis workflow for mRNA vaccines in a unified interface. VaxLab incorporates four codon optimization algorithms based on distinct approaches: codon usage matching, secondary structure design and deep generative models. 5′ and 3′ untranslated regions can be designed using provided generative models or selected from a predefined sequence library. The designed sequences are then evaluated and reported for ten predictive metrics for RNA stability, protein expression and manufacturing risks. An integrated sequence editor with annotation tools enables further adjustments, and final sequences are exported in formats compatible with commercial DNA synthesis services. We validated VaxLab using influenza hemagglutinin sequences, completing the full design workflow in under 5 min per variant. In HCT116 cells, optimized variants showed up to 9.5-fold differences in protein expression, with most achieving at least 2.9-fold higher expression than the wild-type sequence. VaxLab provides an integrated optimization workbench that accelerates mRNA vaccine design by enabling principled selection among multiple optimization techniques informed by current understanding of RNA regulation.
The limited stability of mRNA in vivo remains a major challenge for vaccines and therapeutics. While alternative RNA formats such as circular RNA or self-amplifying RNA offer greater durability, these modalities often suffer from low translation, modification incompatibility and difficult manufacturing. To overcome these limitations, we screen 196,277 viral sequences and identify eleven elements that strongly enhance mRNA stability and translation. Mechanistically, they recruit TENT4 to extend the poly(A) tail, preventing deadenylation. Five of them are compatible with N1-methylpseudouridine, which improves mRNA efficacy and reduces immunogenicity. An element named A7 demonstrates particularly robust performance across cell types, delivery methods, modifications and coding sequences, making linear mRNA as stable as circular RNA while achieving higher translation efficiency. In mouse liver, A7-containing linear mRNA exhibits substantially higher protein levels than circular RNA, with sustained expression lasting for over 2 weeks. These RNA stability enhancers enable robust linear mRNA platforms that combine high and durable expression, low immunogenicity and simple manufacturing. A systematic screen identifies RNA elements that enhance the stability and translation of base-modified mRNA.
In vitro-transcribed (IVT) mRNA therapeutics are promising for preventing and treating diseases, including infectious diseases and cancer, by delivering nucleic acid sequences. Assessing the stability of mRNA sequences and poly(A) tail length is crucial to minimize adverse effects and ensure drug efficacy. However, accurately measuring long homopolymeric nucleotides remains technically challenging, and a specialized method for IVT mRNAs is lacking. We introduce 3AIM-seq, a technique optimized experimentally and analytically for sequencing 3' poly(A) tails in IVT mRNAs. To generate high-quality data, we used a ligation-free adapter followed by 3' end amplification to prepare high-throughput sequencing libraries. The 3AIM-seq algorithm employed a sliding window approach to base quality scores, providing higher resolution and accuracy than the base-calling method and distinguishing poly(A) length differences as small as ±5 bp. Using in silico synthetic standard spike-in experiments, the method estimated the homogeneity of poly(A) tail lengths at the individual DNA molecule level in IVT mRNAs with various poly(A) tail structures. Although poly(A) tails of up to 70 bases were accurately measured, stretches exceeding 100 bases exhibited high heterogeneity for length, highlighting the importance of thorough assessments. Therefore, 3AIM-seq is a reliable method for evaluating the structural integrity of IVT mRNA products before clinical use.
Spinocerebellar ataxias (SCAs) represent a diverse group of neurodegenerative disorders characterized by progressive cerebellar ataxia. In South Korea, diagnostic laboratories typically focus on common SCA subtypes, leaving the prevalence of rare SCAs uncertain. This study aimed to explore the frequency of rarer forms of SCA, including SCA10, 12, 31, and 36 utilizing molecular techniques including long-read sequencing (LRS). Patients from ataxia cohorts who remained undiagnosed after testing for common genetic ataxias (SCA1, 2, 3, 6, 7, 8 17, and dentatorubral-pallidoluysian atrophy) were analyzed, along with unselected ataxia patients referred for screening of common SCAs. Expanded alleles for SCA10, 12, 31, and 36 were investigated through allele-length PCR, repeat-primed PCR, and LRS. Among 78 patients from 67 families with undiagnosed cerebellar ataxia, SCA36 was identified in 8 families (11.9%), while SCA10, 12, or 31 were not found. In unselected ataxia, SCA36 was present in 1.0% (1/99). Korean SCA36 patients exhibited clinical characteristics similar to global reports, with a higher incidence of hyperreflexia. The haplotype of expanded alleles identified in LRS was consistent among SCA36 patients. The findings indicate that SCA36 accounts for 11.9% of diagnoses after excluding common SCAs and 1.0% in unselected ataxia patients. The study underscores the prevalence of SCA36 in South Korea and emphasizes the potential of LRS as a diagnostic tool for this condition. Integrating LRS into diagnostic protocol could enhance diagnostic efficacy, particularly in populations with a high prevalence of SCA36 like South Korea. Further research is necessary to standardize LRS for routine clinical application.
The 3’ terminal oligo-uridylation, a post-transcriptional mRNA modification, is conserved among eukaryotes and drives mRNA degradation, thereby affecting several key biological processes such as animal development and viral infection. Our TAIL-seq experiment of mouse liver mRNA collected from six zeitgeber times reveals transcripts with rhythmic poly(A) tail lengths and demonstrates that overall 3’ terminal uridylation frequencies at mRNA poly(A) tail very-ends undergo rhythmic change. Consistently, major terminal uridylyl transferases, TUT4 and TUT7, have cycling protein expression in mouse liver corresponding to 3’ terminal uridylation rhythms, indicating that the cycling expression of TUTases correlates with the rhythmic pattern of uridylation. Furthermore, the double knockdown of TUT4 and TUT7 in U2OS cells lengthens the circadian period and decreases the rhythmic amplitude of clock gene expression. Our work thoroughly profiles the dynamic changes in poly(A) tail lengths and terminal modifications and uncovers uridylation as a post-transcriptional modulator in the mammalian circadian clock.
Translational regulation in tissue environments during in vivo viral pathogenesis has rarely been studied due to the lack of translatomes from virus-infected tissues, although a series of translatome studies using in vitro cultured cells with viral infection have been reported. In this study, we exploited tissue-optimized ribosome profiling (Ribo-seq) and severe-COVID-19 model mice to establish the first temporal translation profiles of virus and host genes in the lungs during SARS-CoV-2 pathogenesis. Our datasets revealed not only previously unknown targets of translation regulation in infected tissues but also hitherto unreported molecular signatures that contribute to tissue pathology after SARS-CoV-2 infection. Specifically, we observed gradual increases in pseudoribosomal ribonucleoprotein (RNP) interactions that partially overlapped the trails of ribosomes, being likely involved in impeding translation elongation. Contemporaneously developed ribosome heterogeneity with predominantly dysregulated 5 S rRNP association supported the malfunction of elongating ribosomes. Analyses of canonical Ribo-seq reads (ribosome footprints) highlighted two obstructive characteristics to host gene expression: ribosome stalling on codons within transmembrane domaincoding regions and compromised translation of immunity- and metabolism-related genes with upregulated transcription. Our findings collectively demonstrate that the abrogation of translation integrity may be one of the most critical factors contributing to pathogenesis after SARS-CoV-2 infection of tissues.
Small, compact genomes confer a selective advantage to viruses, yet human cytomegalovirus (HCMV) expresses the long non-coding RNAs (lncRNAs); RNA1.2, RNA2.7, RNA4.9, and RNA5.0. Little is known about the function of these lncRNAs in the virus life cycle. Here, we dissected the functional and molecular landscape of HCMV lncRNAs. We found that HCMV lncRNAs occupy ~ 30% and 50–60% of total and poly(A)+viral transcriptome, respectively, throughout virus life cycle. RNA1.2, RNA2.7, and RNA4.9, the three abundantly expressed lncRNAs, appear to be essential in all infection states. Among these three lncRNAs, depletion of RNA2.7 and RNA4.9 results in the greatest defect in maintaining latent reservoir and promoting lytic replication, respectively. Moreover, we delineated the global post-transcriptional nature of HCMV lncRNAs by nanopore direct RNA sequencing and interactome analysis. We revealed that the lncRNAs are modified with N 6 -methyladenosine (m 6 A) and interact with m 6 A readers in all infection states. In-depth analysis demonstrated that m 6 A machineries stabilize HCMV lncRNAs, which could account for the overwhelming abundance of viral lncRNAs. Our study lays the groundwork for understanding the viral lncRNA–mediated regulation of host-virus interaction throughout the HCMV life cycle.
RNA modifications are a common occurrence across all domains of life. Several chemical modifications, including N6-methyladenosine, have also been found in viral transcripts and viral RNA genomes. Some of the modifications increase the viral replication efficiency while also helping the virus to evade the host immune system. Nonetheless, there are numerous examples in which the host's RNA modification enzymes function as antiviral factors. Although established methods like MeRIP-seq and miCLIP can provide a transcriptome- wide overview of how viral RNA is modified, it is difficult to distinguish between the complex overlapping viral transcript isoforms using the short read-based techniques. Nanopore direct RNA sequencing (DRS) provides both long reads and direct signal readings, which may carry information about the modifications. Here, we describe a refined protocol for analyzing the RNA modifications in viral transcriptomes using nanopore technology.
The accuracy of methods for assembling transcripts from short-read RNA sequencing data is limited by the lack of long-range information. Here we introduce Ladder-seq, an approach that separates transcripts according to their lengths before sequencing and uses the additional information to improve the quantification and assembly of transcripts. Using simulated data, we show that a kallisto algorithm extended to process Ladder-seq data quantifies transcripts of complex genes with substantially higher accuracy than conventional kallisto. For reference-based assembly, a tailored scheme based on the StringTie2 algorithm reconstructs a single transcript with 30.8% higher precision than its conventional counterpart and is more than 30% more sensitive for complex genes. For de novo assembly, a similar scheme based on the Trinity algorithm correctly assembles 78% more transcripts than conventional Trinity while improving precision by 78%. In experimental data, Ladder-seq reveals 40% more genes harboring isoform switches compared to conventional RNA sequencing and unveils widespread changes in isoform usage upon m6A depletion by Mettl14 knockout.
TENT4 enzymes generate ‘mixed tails’ of diverse nucleotides at 3′ ends of RNAs via nontemplated nucleotide addition to protect messenger RNAs from deadenylation. Here we discover extensive mixed tailing in transcripts of hepatitis B virus (HBV) and human cytomegalovirus (HCMV), generated via a similar mechanism exploiting the TENT4–ZCCHC14 complex. TAIL-seq on HBV and HCMV RNAs revealed that TENT4A and TENT4B are responsible for mixed tailing and protection of viral poly(A) tails. We find that the HBV post-transcriptional regulatory element (PRE), specifically the CNGGN-type pentaloop, is critical for TENT4-dependent regulation. HCMV uses a similar pentaloop, an interesting example of convergent evolution. This pentaloop is recognized by the sterile alpha motif domain–containing ZCCHC14 protein, which in turn recruits TENT4. Overall, our study reveals the mechanism of action of PRE, which has been widely used to enhance gene expression, and identifies the TENT4–ZCCHC14 complex as a potential target for antiviral therapeutics.
We would like to correct author’s affiliations, add an author and an edit one sentence as shown below 1) The corrected affiliations (switch affiliation 1 and 2) and added author are marked by bold and underlines Joungha Won1,2, Solji Lee2, Myungsun Park2, Tai Young Kim2, Mingu Gordon Park2,3, Byung Yoon Choi4, Dongwan Kim5,6, Hyeshik Chang5,6, Won Do Heo1, V Narry Kim5,6 and C Justin Lee2* 1Department of Biological Sciences, Korea Advanced Institute of Science and Technology (KAIST), Daejeon 34141, 2Center for Cognition and Sociality, Cognitive Glioscience Group, Institute for Basic Science, Daejeon 34126, 3KU-KIST Graduate School of Converging Science and Technology, Korea University, Seoul 02841, 4Department of Otorhinolaryngology, Seoul National University Bundang Hospital, Seongnam 13620, 5Center for RNA Research, Institute for Basic Science, Seoul 08826, 6School of Biological Sciences, Seoul National University, Seoul 08826, Korea 2) In Quantitative rtPCR section (Page 110, in material and methods) we would like to correct the following sentence from;“In brief, each reaction buffer consisted of a total volume of 20 µl containing 8 µl of 100 µM forward and reverse primers (4 µl for each primer), 2 µl of cDNA, and 10 µl power SYBR Green PCR Master Mix ” to;“In brief, each reaction buffer consisted of a total volume of 20 µl containing 2 µl of 10 µM forward and reverse primers (1 µl for each primer), 2 µl of cDNA, and 10 µl power SYBR Green PCR Master Mix” © 2020 Korean Society for Neurodegenerative Disease All rights reserved
Joungha Won, Solji Lee, Myungsun Park, Tai Young Kim, Mingu Gordon Park, Byung Yoon Choi, Dongwan Kim, Hyeshik Chang, Won Do Heo, V. Narry Kim and C. Justin Lee. Exp Neurobiol 2020;29:402. https://doi.org/10.5607/en20009e1
The severe acute respiratory coronavirus 2 (SARS-CoV-2), which emerged in December 2019 in Wuhan, China, has spread rapidly to over a dozen countries. Especially, the spike of case numbers in South Korea sparks pandemic worries. This virus is reported to spread mainly through person-to-person contact via respiratory droplets generated by coughing and sneezing, or possibly through surface contaminated by people coughing or sneezing on them. More critically, there have been reports about the possibility of this virus to transmit even before a virus-carrying person to show symptoms. Therefore, a low-cost, easy-access protocol for early detection of this virus is desperately needed. Here, we have established a real-time reverse-transcription PCR (rtPCR)-based assay protocol composed of easy specimen self-collection from a subject via pharyngeal swab, Trizol-based RNA purification, and SYBR Green-based rtPCR. This protocol shows an accuracy and sensitivity limit of 1-10 virus particles as we tested with a known lentivirus. The cost for each sample is estimated to be less than 15 US dollars. Overall time it takes for an entire protocol is estimated to be less than 4 hours. We propose a cost-effective, quick-and-easy method for early detection of SARS-CoV-2 at any conventional Biosafety Level II laboratories that are equipped with a rtPCR machine. Our newly developed protocol should be helpful for a first-hand screening of the asymptomatic virus-carriers for further prevention of transmission and early intervention and treatment for the rapidly propagating virus.
SARS-CoV-2 is a betacoronavirus responsible for the COVID-19 pandemic. Although the SARS-CoV-2 genome was reported recently, its transcriptomic architecture is unknown. Utilizing two complementary sequencing techniques, we present a high-resolution map of the SARS-CoV-2 transcriptome and epitranscriptome. DNA nanoball sequencing shows that the transcriptome is highly complex owing to numerous discontinuous transcription events. In addition to the canonical genomic and 9 subgenomic RNAs, SARS-CoV-2 produces transcripts encoding unknown ORFs with fusion, deletion, and/or frameshift. Using nanopore direct RNA sequencing, we further find at least 41 RNA modification sites on viral transcripts, with the most frequent motif, AAGAA. Modified RNAs have shorter poly(A) tails than unmodified RNAs, suggesting a link between the modification and the 3' tail. Functional investigation of the unknown transcripts and RNA modifications discovered in this study will open new directions to our understanding of the life cycle and pathogenicity of SARS-CoV-2.
This dataset contains the raw sequencing data from a TAIL-seq run for Xenopus laevis embryos. The cluster intensities of fluorescence signals are repacked as an HDF5 formatted file, then split into multiple parts to fit in the dataset size limitation of the Zenodo.
RNA tails play integral roles in the regulation of messenger RNA (mRNA) translation and decay. Guanylation of the poly(A) tail was discovered recently, yet the enzymology and function remain obscure. Here we identify TENT4A (PAPD7) and TENT4B (PAPD5) as the enzymes responsible for mRNA guanylation. Purified TENT4 proteins generate amixed poly(A) tail with intermittent non-adenosine residues, the most common of which is guanosine. A single guanosine residue is sufficient to impede the deadenylase CCR4-NOT complex, which trims the tail and exposes guanosine at the 3' end. Consistently, depletion of TENT4A and TENT4B leads to a decrease in mRNA half-life and abundance in cells. Thus, TENT4A and TENT4B produce a mixed tail that shields mRNA from rapid deadenylation. Our study unveils the role of mixed tailing and expands the complexity of posttranscriptional gene regulation.