OBJECTIVES:Gut microbiota develops dynamically during infancy in parallel with early growth processes. This study aimed to assess the pattern of gut microbiota colonization in full-term SGA infants with catch-up weight gain in the 1st year of life. METHODS:This longitudinal cohort study included 19 full-term SGA and 46 full-term appropriate-for-gestational-age (AGA) infants. Stool samples and body mass measurements were collected at multiple time points during the 1st year of life. Gut microbiota composition was analyzed using 16S rRNA gene sequencing. Alpha diversity, beta diversity, and taxa abundances were used to evaluate microbial composition and diversity across developmental stages. Associations between the rate of weight gain and the pace of gut microbiota maturation were examined. RESULTS:SGA infants exhibited higher alpha diversity than AGA children at most time points. In this group, the Shannon index, reflecting the level of gut microbiota maturation, was positively associated with the rate of body weight gain over time (p = 0.015), an association that was not observed in AGA infants. Characteristic genera associated with SGA included Citrobacter, Staphylococcus, Blautia, Veillonella, Klebsiella, and Clostridium XIVa. CONCLUSIONS:SGA children had a distinct gut microbiota with higher alpha diversity than AGA peers. In this group, more mature microbiota was linked to faster weight gain and an increased abundance of short-chain fatty acid-producing and obesity-associated bacteria, suggesting that early microbial development may affect the risk of overweight and obesity later in life.
Background: The gut microbiota undergoes dynamic changes during infancy, a period that coincides with intensive early growth and metabolic development. Objective: This study aimed to assess gut microbiota colonization patterns in full-term small for gestational age (SGA) infants with catch-up weight gain during the first year of life. Subjects and methods: The longitudinal cohort study included 19 full-term SGA infants and 46 full-term infants born appropriate for gestational age (AGA). Stool samples and body weight measurements were collected at several points throughout the first year of life. Gut microbiota composition was assessed using 16S rRNA gene sequencing. Microbial composition and diversity across developmental stages were evaluated using alpha diversity, beta diversity, and taxa abundance analyses. The relationship between the rate of weight gain and the pace of gut microbiota maturation was also examined. Results: SGA infants demonstrated higher alpha diversity than AGA infants at most time points. In the SGA group, the Shannon index, reflecting gut microbiota maturation, was positively associated with the rate of body weight gain over time (p=0.015), whereas no such association was observed in AGA infants. Genera characteristic of SGA group included Citrobacter, Staphylococcus, Blautia, Veillonella, Klebsiella and Clostridium XIVa. Conclusions: Overall, SGA infants exhibited a distinct gut microbiota profile with higher alpha diversity compared with AGA peers. In this group, a more mature microbiota was associated with faster weight gain and a greater abundance of short-chain fatty acid-producing and obesity-associated bacteria, suggesting that early microbial development may contribute to the risk of overweight and obesity later in life.
PURPOSE:The microbiome of the saliva can be influenced by various factors, including systemic diseases and chemotherapy. Oral dysbiosis manifests as altered bacterial composition and abundance, which often correlates with increased local and systemic inflammation. The aim of the study was to investigate the dysbiosis in the saliva of breast cancer (BC) patients before and during neoadjuvant chemotherapy (NAC). METHODS:Saliva samples were collected from 50 breast cancer patients at three timepoints (before, during, and after NAC). Saliva from 10 healthy women was used as control samples. Full-length gene 16S rRNA sequencing and analysis were performed using the Microbiome Analyst platform, R and JADBIO AutomatedML platform to compare the abundances of bacterial taxa. RESULTS:Alpha and beta diversity measures differed between breast cancer patients and healthy controls. In addition, eight bacterial genera differed significantly between breast cancer patients and controls, including Porphyromonas, Campylobacter, Oribacterium, Veillonella, and Alloprevotella. Longitudinal analysis revealed significant decrease of bacterial diversity in the course of neoadjuvant chemotherapy as well as significant change in the prevalence of a few low-abundant genera. CONCLUSIONS:The obtained results confirm BC-related and NAC-related dysbiosis in saliva, which emphasizes the potential of saliva as a diagnostic and prognostic tool in patients with breast cancer.
Microbiome studies aim to answer the following questions: which organisms are in the sample and what is their impact on the patient or the environment? To answer these questions, investigators have to perform comparative analyses on their classified sequences based on the collected metadata, such as treatment, condition of the patient, or the environment. The integrity of sequences, classifications, and metadata is paramount for the success of such studies. Still, the area of data management for the preliminary study results appears to be neglected. Here, we present the development of MetagenomicsDB (http://github.com/IOB-Muenster/MetagenomicsDB; accessed 2024/12/18), a central data management system for the study of the gut microbiome in children who are small for their gestational age (SGA). Our system provided more flexibility to conduct study-specific analyses and to integrate specific external resources than existing and necessarily more generic solutions. It supports short or long read data produced by virtually any sequencing instrument targeting (parts of) popular marker genes, such as the 16S rRNA gene and its variable regions. Classifications of these reads from the MetaG and Kraken 2 software are supported. The main goals of the system are to store the pre-computed study data securely under concurrent load and to make downstream analyses accessible to all researchers, regardless of programming proficiency. Thus, after initial plausibility checks on the input data to reduce human error, data are stored in a relational database and can be continuously updated over the whole life time of the study. We used a modular approach for MetagenomicsDB with comprehensive tests verifying the expected behavior and extensively described the underlying rational which allows users to adapt the system to their needs. We advocate the use of MetagenomicsDB as the backend for a graphical web interface. We showcase the potential of this approach at the example of our study on SGA children (http://www.bioinformatics.uni-muenster.de/tools/metagenomicsDB; accessed 2024/12/02). Without restrictions caused by the level of programming proficiency, our team members could explore the study data and optionally filter them using the graphical interface, before exporting the data in a format directly suitable for external normalization of read counts and statistical analyses. Study results could be conveniently and transparently shared with the public, as demonstrated here. Links to external resources facilitated literature search with regard to the SGA condition and assessments of the potential pathogenicity of taxa. Since different users will have different demands regarding features, data security, and web environments, we provide our implementation of the web interface as a visual example. By providing users with the MetagenomicsDB backend which constitutes the major part of the system, we ensure that custom development can be finished in a reasonable amount of time. We report our endeavors in order to motivate the application of data management systems at the scale of single studies in microbiome research.
Although the existence of overlapping protein-coding genes in eukaryotic genomes is known for decades, their role in regulating expression remains far from fully understood. Here, the mechanism regulating the expression of head-to-head overlapping genes, a pair of INO80E and HIRIP3 genes is presented. Based on a series of experiments, we show that the expression of these genes is strongly dependent on sense/antisense interactions. The overlapping transcripts form an RNA:RNA duplex that has a stabilizing effect on the mRNAs involved, and this stabilization may be mediated by the ELAVL1 protein. We also show that the transcription factor RARG is important for the transcription of both genes studied. In addition, we demonstrate that the overlapping isoform of INO80E forms an R-loop that may positively regulate HIRIP3 isoforms. We propose that both structures, dsRNA and R-loops, help to keep the DNA loop open to allow the transcription of the remaining variants of both genes. However, experiments suggest that RNA:RNA duplex formation plays a major role, while R-loops play only a complementary one. The absence of this dsRNA structure leads to the loss of a stable DNA opening and consequently to transcriptional interference.
Microbiome studies aim to answer the following questions: which organisms are in the sample and what is their impact on the patient or the environment? To answer these questions, investigators have to perform comparative analyses on their classified sequences based on the collected metadata, such as treatment, condition of the patient, or the environment. The integrity of sequences, classifications, and metadata is paramount for the success of such studies. Still, the area of data management for the preliminary study results appears to be neglected. Here, we present the development of a central data management system with an accessible web interface for the study of the gut microbiome in children who are small for their gestational age. We have called this system MetagenomicsDB. The web interface allows users, regardless of their bioinformatics expertise, to operate the system. The interface contains links to external resources in order to facilitate literature search, statistical analyses, and assessments of the potential pathogenicity of taxa. Preliminary study results are automatically quality-controlled and subsequently imported into a relational database. After exploration and optional filtering by the user, data are exported in a format directly suitable for follow-up analyses. Compared to a more conventional approach of storing the data in plain files, the automated quality control and database storage offered by MetagenomicsDB provides an enhanced quality assurance of the produced study results. This is especially true in a collaborative setting. Also, the automation of the data transfer from the format of the input data to the format needed for downstream analyses makes basic statistical and bioinformatic analyses accessible. The web interface not only allows us to perform our internal analyses, but will also facilitate transparent sharing of the complete study results at publication time with reviewers and the general public. We demonstrated the viability of this approach by making a subset of our preliminary study results already publicly accessible: In the context of our study, our system provided more flexibility to conduct study- specific analyses and to integrate specific external resources than existing and necessarily more generic solutions. Our system is not yet ready to be widely applicable out-of-the-box for any microbiome study. However, we expect that due to the modular concept, the tests, and the extensive description of the underlying rationale, large parts of our implementation can be adapted for future projects in a fraction of the time needed to develop a new data management system from scratch. Thus, we report our endeavors in order to motivate the application of data management systems at the scale of single studies in microbiome research. ### Competing Interest Statement The authors have declared no competing interest.
U7 snRNA is part of the U7 snRNP complex, required for the 3' end processing of replication-dependent histone pre-mRNAs in S phase of the cell cycle. Here, we show that U7 snRNA plays another function in inhibiting the expression of a subset of long terminal repeats of human endogenous retroviruses (HERV1/LTR12s) and LTR12-containing long intergenic noncoding RNAs (lincRNAs), both bearing sequence motifs that perfectly match the 5' end of U7 snRNA. We demonstrate that U7 snRNA inhibits LTR12 and lincRNA transcription and propose a mechanism in which U7 snRNA hampers the binding/activity of the NF-Y transcription factor to CCAAT motifs within LTR12 elements. Thereby, U7 snRNA plays a protective role in maintaining the silencing of deleterious genetic elements in selected types of cells.
Retrotransposition is one of the main factors responsible for gene duplication and thus genome evolution. However, the sequences that undergo this process are not only an excellent source of biological diversity, but in certain cases also pose a threat to the integrity of the DNA. One of the mechanisms that protects against the incorporation of mobile elements is the HUSH complex, which is responsible for silencing long, intronless, transcriptionally active transposed sequences that are rich in adenine on the sense strand. In this study, broad sets of human and porcine retrocopies were analysed with respect to the above factors, taking into account evolution of these molecules. Analysis of expression pattern, genomic structure, transcript length, and nucleotide substitution frequency showed the strong relationship between the expression level and exon length as well as the protective nature of introns. The results of the studies also showed that there is no direct correlation between the expression level and adenine content. However, protein-coding retrocopies, which have a lower adenine content, have a significantly higher expression level than the adenine-rich non-coding but expressed retrocopies. Therefore, although the mechanism of HUSH silencing may be an important part of the regulation of retrocopy expression, it is one component of a more complex molecular network that remains to be elucidated.
Retroposed protein-coding genes are commonly considered to be nonfunctional duplicates. However, they often gain transcriptional capa-bility and have important roles. Amici et al. recently identified novel functions of a retroposed gene. HAPSTR2, a retrocopy of HAPSTR1, encodes a protein that stabilizes the HAPSTR1 protein and functionally buffers its loss.
ABSTRACT U7 snRNA is part of U7 snRNP, a complex required for the 3’end processing of replication-dependent histone pre-mRNAs in the S phase of the cell cycle. During this maturation event, the 5’ region of U7 snRNA hybridizes with the highly complementary sequence present in the 3’UTR of histone pre-mRNAs, called histone downstream element, HDE. This base-pair interaction triggers subsequent reactions that eventually result in cleavage and release of mature histone transcripts. Intriguingly, U7 snRNP is constitutively expressed throughout the cell cycle and in nondividing cells, suggesting another function of U7 snRNA/snRNP in cells. Here, we show that several human endogenous retroviruses (HERVs) are significantly upregulated in HEK293T cells with U7 snRNA knockdown. They predominantly belong to the LTR12 class. Interestingly, some of them are located within long intergenic noncoding RNAs (lincRNAs), which in turn are upregulated in U7 snRNA knockdown cells as well. Significantly, both these HERV1/LTR12s and lincRNAs contain two or more sequence motifs that perfectly match the 5’ end of U7 snRNA, which we called HDE-like motifs. We confirmed that mutations within the HDE-like motifs abrogate U7 snRNA regulatory function and stimulate the expression of selected lincRNAs. Furthermore, we demonstrate that U7 snRNA inhibits HERV1/LTR12 and lincRNA expression at the transcription level. We propose a mechanism in which U7 snRNA hampers binding/activity of NF-Y transcription factor to CCAAT motifs that are frequently found in LTRs as well as in a close proximity to HDE-like motifs. The expression of many HERV1/LTR12s and lincRNAs regulated by U7 snRNA seems to be tissue specific, therefore, we suggest that U7 snRNA plays a protective role in keeping deleterious genetic elements in silence in selected types of cells.
Head and neck squamous cell carcinoma is one of the most common and fatal cancers worldwide. Lack of appropriate preventive screening tests, late detection, and high heterogeneity of these tumors are the main reasons for the unsatisfactory effects of therapy and, consequently, unfavorable outcomes for patients. An opportunity to improve the quality of diagnostics and treatment of this group of cancers are microRNAs (miRNAs) - molecules with a great potential both as biomarkers and therapeutic targets. This review aims to present the characteristics of these short non-coding RNAs (ncRNAs) and summarize the current reports on their use in oncology focused on medical strategies tailored to patients' needs.
The INO80E gene encodes the protein involved in the chromatin remodeling processes as a part of the multi-subunit INO80-chromatin remodeling complex. The INO80E gene is located on chromosome 16 and overlaps head-to-head with the HIRIP3 gene encoding protein, which binds H2B and H3 core histones and HIRA protein and regulates the chromatin and histone metabolism. Antisense transcription of head-to-head overlapping PC genes may have several consequences and none of them have been comprehensively investigated and explained. Here, we determined that INO80E-201, which overlaps the HIRIP3 gene transcripts, forms an R-loop at its 5’ end. We also demonstrated that overlapping transcripts of INO80E and HIRIP3 form an RNA:RNA duplex that has the stabilizing effect on the involved mRNAs. Our results confirmed that this stabilization could be mediated by the ELAVL1 protein. We additionally introduced de novo methylation using the CRISPR/Cas system into the promoter sequence of INO80E gene. As a result of the introduced changes, reduced expression of HIRIP3 and INO80E gene transcripts was observed. It was determined that methylated cytosines were located in the binding sites for four transcription factors including RARG, which was further confirmed to be important in the transcription of both studied genes. Our results strongly suggest that the formation of an RNA:RNA duplex is necessary for stable simultaneous expression of both genes. Lack of this dsRNA structure results in a loss of a wider DNA opening and in consequence transcriptional interference. We also concluded that forming R-loops probably plays only a supplementary role and is not required for proper expression of HIRIP3 and INO80E gene transcripts.### Competing Interest StatementThe authors have declared no competing interest.
As it is well known, messenger RNA has many regulatory regions along its sequence length. One of them is the 5′ untranslated region (5’UTR), which itself contains many regulatory elements such as upstream ORFs (uORFs), internal ribosome entry sites (IRESs), microRNA binding sites, and structural components involved in the regulation of mRNA stability, pre-mRNA splicing, and translation initiation. Activation of the alternative, more upstream transcription start site leads to an extension of 5′UTR. One of the consequences of 5′UTRs extension may be head-to-head gene overlap. This review describes elements in 5′UTR of protein-coding transcripts and the functional significance of protein-coding genes 5′ overlap with implications for transcription, translation, and disease.
Gut microbiota succession overlaps with intensive growth in infancy and early childhood. The multitude of functions performed by intestinal microbes, including participation in metabolic, hormonal, and immune pathways, makes the gut bacterial community an important player in cross-talk between intestinal processes and growth. Long-term disturbances in the colonization pattern may affect the growth trajectory, resulting in stunting or wasting. In this review, we summarize the evidence on the mediating role of gut microbiota in the mechanisms controlling the growth of children.
BACKGROUND:Long noncoding RNAs represent a large class of transcripts with two common features: they exceed an arbitrary length threshold of 200 nt and are assumed to not encode proteins. Although a growing body of evidence indicates that the vast majority of lncRNAs are potentially nonfunctional, hundreds of them have already been revealed to perform essential gene regulatory functions or to be linked to a number of cellular processes, including those associated with the etiology of human diseases. To better understand the biology of lncRNAs, it is essential to perform a more in-depth study of their evolution. In contrast to protein-encoding transcripts, however, they do not show the strong sequence conservation that usually results from purifying selection; therefore, software that is typically used to resolve the evolutionary relationships of protein-encoding genes and transcripts is not applicable to the study of lncRNAs.RESULTS:To tackle this issue, we developed lncEvo, a computational pipeline that consists of three modules: (1) transcriptome assembly from RNA-Seq data, (2) prediction of lncRNAs, and (3) conservation study-a genome-wide comparison of lncRNA transcriptomes between two species of interest, including search for orthologs. Importantly, one can choose to apply lncEvo solely for transcriptome assembly or lncRNA prediction, without calling the conservation-related part.CONCLUSIONS:lncEvo is an all-in-one tool built with the Nextflow framework, utilizing state-of-the-art software and algorithms with customizable trade-offs between speed and sensitivity, ease of use and built-in reporting functionalities. The source code of the pipeline is freely available for academic and nonacademic use under the MIT license at https://gitlab.com/spirit678/lncrna_conservation_nf .
Long noncoding RNAs (lncRNAs) have emerged as prominent regulators of gene expression in eukaryotes. The identification of lncRNA orthologs is essential in efforts to decipher their roles across model organisms, as homologous genes tend to have similar molecular and biological functions. The relatively high sequence plasticity of lncRNA genes compared with protein-coding genes, makes the identification of their orthologs a challenging task. This is why comparative genomics of lncRNAs requires the development of specific and, sometimes, complex approaches. Here, we briefly review current advancements and challenges associated with four levels of lncRNA conservation: genomic sequences, splicing signals, secondary structures and syntenic transcription.
Retroposition is RNA-based gene duplication leading to the creation of single exon nonfunctional copies. Nevertheless, over time, many of these duplicates acquire transcriptional capabilities. In human in most cases, these so-called retrogenes do not code for proteins but function as regulatory long noncoding RNAs (lncRNAs). The mechanisms by which they can regulate other genes include microRNA sponging, modulation of alternative splicing, epigenetic regulation and competition for stabilizing factors, among others. Here, we summarize recent findings related to lncRNAs originating from retrocopies that are involved in human diseases such as cancer and neurodegenerative, mental or cardiovascular disorders. Special attention is given to retrocopies that regulate their progenitors or host genes. Presented evidence from the literature and our bioinformatics analyses demonstrates that these retrocopies, often described as unimportant pseudogenes, are significant players in the cell’s molecular machinery.
Despite the number of studies focused on sense-antisense transcription, the key question of whether such organization evolved as a regulator of gene expression or if this is only a byproduct of other regulatory processes has not been elucidated to date. In this study, protein-coding sense-antisense gene pairs were analyzed with a particular focus on pairs overlapping at their 5’ ends. Analyses were performed in 73 human transcription start site libraries. The results of our studies showed that the overlap between genes is not a stable feature and depends on which TSSs are utilized in a given cell type. An analysis of gene expression did not confirm that overlap between genes causes downregulation of their expression. This observation contradicts earlier findings. In addition, we showed that the switch from one promoter to another, leading to genes overlap, may occur in response to changing environment of a cell or tissue. We also demonstrated that in transfected and cancerous cells genes overlap is observed more often in comparison with normal tissues. Moreover, utilization of overlapping promoters depends on particular state of a cell and, at least in some groups of genes, is not merely coincidental.
Waterlogging (WL), excess water in the soil, is a phenomenon often occurring during plant cultivation causing low oxygen levels (hypoxia) in the soil. The aim of this study was to identify candidate genes involved in long-term waterlogging tolerance in cucumber using RNA sequencing. Here, we also determined how waterlogging pre-treatment (priming) influenced long-term memory in WL tolerant (WL-T) and WL sensitive (WL-S) i.e., DH2 and DH4 accessions, respectively. This work uncovered various differentially expressed genes (DEGs) activated in the long-term recovery in both accessions. De novo assembly generated 36,712 transcripts with an average length of 2236 bp. The results revealed that long-term waterlogging had divergent impacts on gene expression in WL-T DH2 and WL-S DH4 cucumber accessions: after 7 days of waterlogging, more DEGs in comparison to control conditions were identified in WL-S DH4 (8927) than in WL-T DH2 (5957). Additionally, 11,619 and 5007 DEGs were identified after a second waterlogging treatment in the WL-S and WL-T accessions, respectively. We identified genes associated with WL in cucumber that were especially related to enhanced glycolysis, adventitious roots development, and amino acid metabolism. qRT-PCR assay for hypoxia marker genes i.e., alcohol dehydrogenase (adh), 1-aminocyclopropane-1-carboxylate oxidase (aco) and long chain acyl-CoA synthetase 6 (lacs6) confirmed differences in response to waterlogging stress between sensitive and tolerant cucumbers and effectiveness of priming to enhance stress tolerance.