Virus long noncoding RNAs (vlncRNAs) play crucial roles in viral infections, yet their identification and characterization remain limited. This study identified 5,053 novel vlncRNAs across 25 viral species using third-generation sequencing, with two from Influenza A virus and Vesicular stomatitis virus validated by RT-qPCR. Most vlncRNAs originated from dsDNA viruses. Only ~1% of vlncRNAs have annotated RNA families, suggesting many novel RNA structures. Interestingly, a total of 772 vlncRNAs from 15 human viruses structurally mimicked human lncRNAs (hlncRNAs), while only seven vlncRNAs shared sequence similarities with hlncRNAs. These vlncRNA and hlncRNAs bound to similar miRNAs, potentially acting as miRNA sponges to promote essential life processes. Splicing analysis showed vlncRNAs had a prevalence of alternative first exon. Finally, we developed vlncRNAbase (http://computationalbiology.cn/vlncRNAbase/#/) to store and organize the newly identified and known vlncRNAs. Overall, the study provides a valuable resource for further investigation into vlncRNAs and deepens our understanding of the diversity, structure, and function of the molecule.
Extracellular vesicle DNAs (evDNAs) hold significant diagnostic value for various diseases and facilitate transcellular transfer of genetic material. Our study identifies transcription factor FOXM1 as a mediator for directing chromatin genes or DNA fragments (termed FOXM1-chDNAs) to extracellular vesicles (EVs). FOXM1 binds to MAP1LC3/LC3 in the nucleus, and FOXM1-chDNAs, such as the DUX4 gene and telomere DNA, are designated by FOXM1 binding and translocated to the cytoplasm before being released to EVs through the secretory autophagy during lysosome inhibition (SALI) process involving LC3. Disrupting FOXM1 expression or the SALI process impairs FOXM1-chDNAs incorporation into EVs. FOXM1-chDNAs can be transmitted to recipient cells via EVs and expressed in recipient cells when they carry functional genes. This finding provides an example of how chromatin DNA fragments are specified to EVs by transcription factor FOXM1, revealing its contribution to the formation of evDNAs from nuclear chromatin. It provides a basis for further exploration of the roles of evDNAs in biological processes, such as horizontal gene transfer. Abbreviation: ATG5: autophagy related 5; CCFs: cytoplasmic chromatin fragments; ChIP: chromatin immunoprecipitation; cytoDNA: cytoplasmic DNA; CQ: chloroquine; FOXM1-DBD: FOXM1 DNA binding domain; DUX4:double homeobox 4; EVs: extracellular vesicles; evDNAs: extracellular vesicle DNAs; FOXM1: forkhead box M1; FOXM1-chDNAs: chromatin DNA fragments directed by FOXM1 to EVs; HGT: horizontal gene transfer; LC3-II: lipid modified LC3; LMNB1: lamin B1; LIR: LC3-interacting region; MAP1LC3/LC3: microtubule associated protein 1 light chain 3; MVBs: multivesicular bodies; M1-binding DNA: a linear DNA containing 72× FOXM1 binding sites; SALI: secretory autophagy during lysosome inhibition; siRNA: small interfering RNA; TetO-DUX4: TetO array-containing DUX4 DNA; TetO: tet operator; TetR: tet repressor
Identification of viruses and further assembly of viral genomes from the next-generation-sequencing data are essential steps in virome studies. This study presented a one-stop tool named VIGA (available at https://github.com/viralInformatics/VIGA) for eukaryotic virus identification and genome assembly from NGS data. It was composed of four modules, namely, identification, taxonomic annotation, assembly and novel virus discovery, which integrated several third-party tools such as BLAST, Trinity, MetaCompass and RagTag. Evaluation on multiple simulated and real virome datasets showed that VIGA assembled more complete virus genomes than its competitors on both the metatranscriptomic and metagenomic data and performed well in assembling virus genomes at the strain level. Finally, VIGA was used to investigate the virome in metatranscriptomic data from the Human Microbiome Project and revealed different composition and positive rate of viromes in diseases of prediabetes, Crohn's disease and ulcerative colitis. Overall, VIGA would help much in identification and characterization of viromes, especially the known viruses, in future studies.
Virus circular RNAs (circRNA) have been reported to be extensively expressed and play important roles in viral infections. Previously we build the first database of virus circRNAs named VirusCircBase which has been widely used in the field. This study significantly improved the database on both the data quantity and database functionality: the number of virus circRNAs, virus species, host organisms was increased from 46440, 23, 9 to 60859, 43, 22, respectively, and 1902 full-length virus circRNAs were newly added; new functions were added such as visualization of the expression level of virus circRNAs and visualization of virus circRNAs in the Genome Browser. Analysis of the expression of virus circRNAs showed that they had low expression levels in most cells or tissues and showed strong expression heterogeneity. Analysis of the splicing of virus circRNAs showed that they used a much higher proportion of non-canonical back-splicing signals compared to those in animals and plants, and mainly used the A5SS (alternative 5’ splice site) in alternative-splicing. Most virus circRNAs have no more than two isoforms. Finally, human genes associated with the virus circRNA production were investigated and more than 1000 human genes exhibited moderate correlations with the expression of virus circRNAs. Most of them showed negative correlations including 42 genes encoding RNA-binding proteins. They were significantly enriched in biological processes related to cell cycle and RNA processing. Overall, the study provides a valuable resource for further studies of virus circRNAs and also provides new insights into the biogenesis mechanisms of virus circRNAs.
Virus-encoded small RNAs (vsRNAs) have been reported to play an important role in viral infections. Unfortunately, there is still a lack of a systematic characterization and resource of vsRNAs. Herein, we identified a total of 19 734 high-confidence vsRNAs including 2746 microRNAs (miRNAs) in 64 viral species from more than 800 samples of public small RNA-Seq data. The number of vsRNAs identified in viruses varied from 1 to 2489 with a median of 170. The length distribution of vsRNAs peaked at 21 and 22 nt. Plant viruses were found to express larger number and higher levels of vsRNAs than those of animal viruses. Besides, the number of vsRNAs identified increased as the viral infection persisted. Interestingly, the vsRNA showed strong expression specificity as little overlap was observed among vsRNAs identified in different strains of a virus, or in different hosts, cells, or tissues infected by the same virus. Little conservation was observed among vsRNAs of different viruses. The viral miRNAs were found to interact with host genes involved in multiple biological processes related to organization, development, action potential, polarity establishment, methylation, immune response, gene regulation, localization, and so on. To facilitate the usage of vsRNAs, a database named vsRNAdb was built for organizing and storing vsRNAs which is available at . Overall, the study deepens our understanding about the diversity and complexity of vsRNAs and provides a rich resource for further studies of vsRNAs.
Parkinson's disease (PD) is a kind of neurodegenerative disease that causes a huge burden to society. Previous studies have suggested the association between PD and multiple viruses. However, there is still a lack of a virome study about PD. This study systematically identified viruses from the public RNA‐sequencing data of more than 700 samples from both PD patients and the control group (most were healthy people). Only nine viruses such as human betaherpesvirus 5 and Merkel cell polyomavirus have been detected in several human brain tissues of the central nervous system, the appendix, and blood of PD patients, and all of these viruses were also detected in the control group. Most viruses were observed to have low abundance in no more than three tissues. No statistically significant differences were observed between the virus abundance in the PD patients and the control group for all viruses. The positive rates of most viruses in PD patients were higher or similar to that in the control group, although those were less than 5% for most viruses. Overall, this is the first study to systematically investigate the virome in PD patients, and provides new insights into the association between viruses and PD.
Plant viruses cause huge damage to commercial crops, yet the studies towards plant viruses are limited and the diversity of plant viruses are under-estimated yet. This study built an up-to-date atlas of plant viruses by computationally identifying viruses from the RNA-seq data in the One Thousand Plant Transcriptomes Initiative (1KP) and by integrating plant viruses from public databases, and further built the Plant Virus Database (PVD, freely available at http://47.90.94.155/PlantVirusBase/#/home ) to store and organize these viruses. The PVD contained 3,206 virus species and 9,604 virus-plant host interactions which were more than twice that reported in previous plant virus databases. The plant viruses were observed to infect only a few plant hosts and vice versa. Analysis and comparison of the viromes in the Monocots and Eudicots, and those in the plants in tropical and temperate regions showed significant differences in the virome composition. Finally, several factors including the viral group (DNA and RNA viruses), enveloped or not, and the transmission mode of viruses, were found to have no or weak associations with the host range of plant viruses. Overall, the study not only provides a valuable resource for further studies of plant viruses, but also deepens our understanding towards the genetic diversity of plant viruses and the virus-host interactions.
Virus-encoded small RNAs (vsRNA) have been reported to play an important role in viral infection. Unfortunately, there is still a lack of an effective method for vsRNA identification. Herein, we presented vsRNAfinder, a de novo method for identifying high-confidence vsRNAs from small RNA-Seq (sRNA-Seq) data based on peak calling and Poisson distribution and is publicly available at https://github.com/ZenaCai/vsRNAfinder. vsRNAfinder outperformed two widely used methods namely miRDeep2 and ShortStack in identifying viral miRNAs with a significantly improved sensitivity. It can also be used to identify sRNAs in animals and plants with similar performance to miRDeep2 and ShortStack. vsRNAfinder would greatly facilitate effective identification of vsRNAs from sRNA-Seq data.
African swine fever virus (ASFV) is a large DNA virus that infects domestic pigs with high morbidity and mortality rates. Repeat sequences, which are DNA sequence elements that are repeated more than twice in the genome, play an important role in the ASFV genome. The majority of repeat sequences, however, have not been identified and characterized in a systematic manner. In this study, three types of repeat sequences, including microsatellites, minisatellites and short interspersed nuclear elements (SINEs), were identified in the ASFV genome, and their distribution, structure, function, and evolutionary history were investigated. Most repeat sequences were observed in noncoding regions and at the 5' end of the genome. Noncoding repeat sequences tended to form enhancers, whereas coding repeat sequences had a lower ratio of alpha-helix and beta-sheet and a higher ratio of loop structure and surface amino acids than nonrepeat sequences. In addition, the repeat sequences tended to encode penetrating and antimicrobial peptides. Further analysis of the evolution of repeat sequences revealed that the pan-repeat sequences presented an open state, showing the diversity of repeat sequences. Finally, CpG islands were observed to be negatively correlated with repeat sequence occurrences, suggesting that they may affect the generation of repeat sequences. Overall, this study emphasizes the importance of repeat sequences in ASFVs, and these results can aid in understanding the virus's function and evolution.