Viruses make extensive use of host cell machinery, however, most systematic studies of virus-host interactions focused on proteins with less attention to nucleic acids. RNA plays important roles in storing, conveying, and regulating genetic information. Our understanding of functional interactions between viral and host RNA is dominated by interactions with micro-RNAs (miRNAs), such as the interaction between hepatitis C virus (HCV) and the liver specific miR-122, critical for viral replication. Methodological developments, however, allow for broader exploration of the RNA interaction landscape beyond that of miRNAs. We here set out to identify virus-host RNA interactions by optimizing RNA antisense purification to systematically map RNA-RNA interactions (RAP-RNA) for viral RNA. After using the HCV/miR-122 interaction for validation, we applied RAP-RNA to determine the RNA interactomes for three important human pathogens; HCV, yellow fever virus (YFV) and chikungunya virus (CHIKV). Comparing virus-host RNA interactomes, we observed patterns of mRNAs encoding factors involved in translation and the proteasome, mitochondrial (mt)RNAs, small nucleolar (sno)RNAs and small nuclear (sn)RNAs, thereby providing a more comprehensive understanding of cellular RNA interactions for RNA viruses. Experimental targeting of selected HCV interactors did not lead to significant impact on viral infection, suggesting that the majority of virus-host RNA interactions are not critical for the virus. In contrast, transcription data were consistent with a possible role in viral stabilization of host RNA interactors. These findings may guide future research directions, e.g., for the role of snoRNAs and snRNAs in viral RNA regulation, with potential to provide further insight to viral exploitation of host factors.
CO$_2$ concentrations above air level (0.04%) are beneficial for the growth of cyanobacteria. However, very high CO$_2$ levels inhibit growth, limiting the usability of cyanobacteria for carbon capture and biotechnological applications. The transcriptomic changes in the cyanobacterium Synechococcus sp. PCC 7002 were investigated at varying growth rates governed by different gas-phase CO$_2$ concentrations, ranging from limiting (0.04%) to optimal (4% and 8%) to inhibitory high (30%). Compared to optimal CO$_2$ concentrations, large differences in the transcriptome were observed in limiting and inhibiting CO$_2$. At 30% CO$_2$, genes encoding CO$_2$ uptake mechanisms, photosynthetic electron transfer, and light-harvesting antennae proteins were down-regulated compared to lower CO$_2$. Genes involved in the ribosomal machinery and biosynthetic pathways were down-regulated at 30% and 0.04% CO$_2$, consistent with the observed reduced growth. The genes most strongly up-regulated at 30% CO$_2$ were primarily associated with stress responses, but did not closely resemble the transcriptomic changes at other stress conditions previously described. The small RNA PsrR1 was strongly up-regulated at 30% CO$_2$ and is likely to be involved in the regulation of growth. These transcriptomic insights are essential for the engineering of fast-growing cyanobacteria at high CO$_2$, which can be applied in carbon capture from industrial point sources.
Cultivation of photosynthetic microbes under elevated CO2 is widely used in biotechnology to enhance growth. However, very high CO2 levels inhibit growth for reasons that are not always clear. We established a method that allowed cultivation under different CO2 concentrations and controlled pH to study the cyanobacterium Synechococcus sp. PCC 7002. Growth rates, cell composition, and transcriptomic profiles were investigated under limiting (0.04%), optimal (1-4%), and inhibitory high (15-30%) gas phase CO2 contents, while maintaining pH similar to 7.5. At intermediate light (200 mu mol m(-2) s(-1)), the growth rate at 30% CO2 was similar to 45% of that under optimal conditions. Under low light (<100 mu mol m(-2) s(-1)), growth was less inhibited by high CO2 than at higher light intensities. Pigmentation and photosynthesis likewise decreased under high CO2, except that carotenoids indicative of high-light stress increased. Transcriptome sequencing confirmed that the photosynthesis machinery was downregulated and stress responses were upregulated under high CO2. Genetic knockdown of CO2 uptake (ndhF3 and ndhF4) or overexpression of a heterologous Na+/H+ antiporter (nhaS3) improved tolerance to high CO2, while knockdown of the CcmR transcriptional regulator reduced it. We conclude that the inhibiting species under very high inorganic carbon conditions was CO2, and not HCO3-, and that reduced light intensity or genetically suppressing active CO2 uptake improved the tolerance of Synechococcus to very high CO2 conditions. This work provides a basis for cultivation under a very wide range of CO2 concentrations, thereby broadening the scope for metabolic manipulation and biotechnological applications of cyanobacteria.
Abstract Spatial navigation is a fundamental mammalian ability, supported by the entorhinal cortex (EC), a structurally conserved yet functionally diverse region across mammalian species. However, how molecular signaling underlies both shared and species-specific navigational strategies remains unclear. Here, we present a cross-species single-cell atlas of the EC from human, Hamadryas baboon, mouse, and Egyptian fruit bat - species spanning distinct evolutionary lineages and navigational demands, including true 3D navigation in bats. Using this resource, we identify conserved principal neuron populations as well as species-specific innovations, including mixed-layer or functional identities and fruit bat-specific subtypes. GABAergic interneurons neurons show strong conservation of somatostatin (SST) and parvalbumin (PV) families, while VIP GABAergic neurons exhibit pronounced species-specific divergence, with an expanded repertoire in primates. Integration with whole-brain diffusion tensor imaging reveals conserved and species-specific connectivity between the EC, hippocampus, and sensory cortices. Major species-specific cellular innovations were further validated using orthogonal histological approaches, confirming their anatomical and laminar organization. Overall, this atlas provides a comparative framework available for the research community to dissect the molecular, cellular, and circuit principles underlying conserved and specialized spatial navigation across mammals.
Lysine is an essential amino acid often limited in plant-based foods, making its enrichment in fermented foods a key nutritional goal. Bacillus and Priestia species are natural lysine producers, but their biosynthetic capacity is constrained by riboswitch regulation of the aspartokinase gene lysC. In this study, we employed mutagenesis using the lysine analogue aminoethylcysteine (AEC) to generate lysine-overproducing strains. AEC-resistant mutants carried distinct mutations in the lysC riboswitch and in lysine biosynthetic and transporter genes. Several mutants showed significantly increased lysine production in defined media. One riboswitch mutant produced over 150-fold more free lysine than its wild type in defined media and threefold during oat fermentation, establishing proof of concept for lysine enrichment in food matrices. Structural predictions revealed that lysC riboswitch mutations occurred near the lysine binding pocket, suggesting altered ligand-binding dynamics. These findings highlight riboswitch-targeted mutagenesis as a promising non-recombinant strategy to enhance lysine production in food fermentations.
Cyanobacteria are one of the oldest and most abundant groups of prokaryotes and are crucial for research in climate, ecology, medicine, and agriculture. Despite intensive efforts in metabolic engineering of cyanobacteria, the mechanisms of gene regulation, particularly through regulatory RNA structures, are often ignored. We computationally searched 202 cyanobacterial genomes for putative conserved RNA structures (CRSs) in the upstream and downstream intergenic regions of 931 orthologous gene groups with the comparative genomics tool CMfinder. The predicted structures were scored according to their local phylogeny and filtered for a maximal false discovery rate of 10%. The screen identified 402 CRSs that match known RNA families (Rfam and Rho-independent bacterial terminators) and 409 novel CRSs. The structures are not limited to either low or high nucleotide conservation, and about half have a high level of significant covariation. The majority of novel CRSs are supported by transcription in at least one species in public RNA-seq data. The regulatory associations of CRSs are discussed in different metabolic pathways, such as photosynthesis, nitrogen fixation, and CO$_2$ metabolism. This resource will support future research on the regulatory mechanisms of RNA in cyanobacteria.
Topologically associating domains (TADs) are generally considered as a homogeneous basic units of genome folding, which is critical for transcriptional regulation. However, recent studies indicate that both the TAD domain structures and the boundaries between them are not as homogeneous as originally recognized. Here, we address the heterogeneity of the TAD boundaries in the human genome at a large scale, which varies between active and inactive chromatin and across cell lines and tissues. To address this, based on the well-annotated TAD boundaries extracted from multiple cell lines and tissues, we examine their nucleotide content, resulting in two main clusters, one GC-rich and one AT-rich, which are mainly distributed in active and inactive chromatin, respectively. Also, they contain different types of repetitive sequences and have different epigenetic patterns, with more CTCF binding motifs in the GC-rich cluster. Hence, our observations of the TAD boundary content provide novel insights into TAD genomic architecture. In addition, we find that cell- or tissue-specific boundaries are less evolutionarily conserved than other boundaries. We highlight the importance of TAD boundary diversity in different functional contexts and discuss the importance of the different types of repetitive sequences and epigenetic patterns in the two main types of boundaries.
A fundamental understanding of genome organization relies on accurately annotating topologically associating domains (TADs) and their boundaries. This is crucial for understanding how cis-regulatory elements regulate gene expression. To go beyond calling TADs and boundaries from Hi-C data, several machine learning-based methods have been proposed to go the step further and predict TAD boundaries from genomic sequences. As the growing evidence of TADs and their boundaries, TADs have been proved exhibiting diverse properties, such as differences in replication timing and epigenetic patterns. However, existing methods do not take this heterogeneity into account. To address this, we propose a method called TADBpred for TAD boundary prediction in a large genomic context in humans. TADBpred focuses on TAD boundaries in active and inactive chromatin across cell-lines and tissues, which are GC-rich and AT-rich, respectively. By integrating genomic elements and sequence composition, we designed two models for GC-rich and AT-rich boundaries, respectively. When testing the performance on respective independent held-out datasets, we obtain AUC scores of 0.91 and 0.80. Our results indicate that TADBpred excels in TAD boundary prediction. Additionally, feature importance analysis highlights the essential features for different classes of TAD boundaries, thereby enhancing our understanding of these TAD boundaries.
Leishmania parasites alternate between hosts, facing environmental changes that demand rapid gene expression adaptation. Lacking canonical RNA polymerase II promoters, transcription in these eukaryotes is polycistronic, with gene regulation occurring post-transcriptionally. Although non-coding RNAs (ncRNAs) have been identified in Leishmania transcriptomes, their functions remain unclear. Recognizing RNA structure's importance, we performed a genome-wide alignment of L. braziliensis and related species, identifying conserved RNA structures, 38 of which overlap with known ncRNAs. One such ncRNA, lncRNA45, was functionally characterized. Using a knockout cell line, we demonstrated that lncRNA45 is crucial for parasite fitness. Reintroducing the wild type lncRNA45 restored fitness, while a version with a single nucleotide substitution in the structured region did not. This mutation also altered RNA-protein interactions. These findings suggest that lncRNA45's regulatory role and protein interactions rely on its secondary structure. This study highlights the significance of structured lncRNAs in Leishmania biology and their potential as therapeutic targets. Further research into these ncRNAs could uncover new parasite regulation mechanisms and inspire novel treatment strategies.
Multiplexed solid-phase polymerase chain reaction (SP-PCR) has emerged as an indispensable modality for concurrent amplification of multiple genetic loci within a singular reaction vessel, facilitating efficient molecular diagnostics. Nevertheless, SP-PCR has seldom been integrated into point-of-care diagnostic devices due to several technical challenges, such as bubble formation during PCR, long reaction time, and low fluorescence signals generated from the PCR products on a solid surface. To circumvent these constraints, we engineered a microfluidic chip comprising SP-PCR and nanophotonic enhancement to enable highly sensitive, high-throughput, and cost-efficient molecular diagnostics. The chip's vertical orientation integrates preloaded reagent chambers for sequential lysis, washing, elution, and amplification, driven by a synchronized stepper motor and air vacuum, achieving robust nucleic acid purification and reverse transcription-PCR, and enabling bubble-free, gravity-assisted fluid dynamics during the PCR thermocycling. Thermal cycling is expedited through a dual-heater configuration alternating at subsecond intervals, obviating active cooling and shortening the reaction time. All-dielectric nanostructured metasurface was incorporated beneath the PCR chamber, allowing for the facile immobilization of DNA arrays to conduct SP-PCR. Taking advantage of guided-mode resonance supported by the metasurface and the SP-PCR approaches permits multiplexed detection and achieves a detection limit of 10 copies/reaction, highlighting the platform's potential for point-of-care diagnostics, personalized medicine, and high-throughput pathogen surveillance. Facile fabrication and automation emphasize scalability for mass production and deployment and collectively represent an advancement in point-of-care diagnostics.
Design of guide RNA (gRNA) with high efficiency and specificity is vital for successful application of the CRISPR gene editing technology. Although many machine learning (ML) and deep learning (DL)-based tools have been developed to predict gRNA activities, a systematic and unbiased evaluation of their predictive performance is still needed. Here, we provide a brief overview of in silico tools for CRISPR design and assess the CRISPR datasets and statistical metrics used for evaluating model performance. We benchmark seven ML and DL-based CRISPR-Cas9 editing efficiency prediction tools across nine CRISPR datasets covering six cell types and three species. The DL models CRISPRon and DeepHF outperform the other models exhibiting greater accuracy and higher Spearman correlation coefficient across multiple datasets. We compile all CRISPR datasets and in silico prediction tools into a GuideNet resource web portal, aiming to facilitate and streamline the sharing of CRISPR datasets. Furthermore, we summarize features affecting CRISPR gene editing activity, providing important insights into model performance and the further development of more accurate CRISPR prediction models.
Megabase-sized extrachromosomal circular DNA with intact oncogenes (ecDNA) plays crucial roles in cancer. However, the impact of smaller (100 bp-1 Mb), more ubiquitous extrachromosomal circular DNA (eccDNA) on tumor pathology is unclear. We analyze eccDNA from 122 renal tumors and adjacent tissues, finding increased eccDNA in late-stage cancers, correlating with heightened patient mortality and copy-number amplification. Large eccDNAs are rare, and smaller eccDNAs predominate. The microRNA (miRNA) genes MIR107, MIR196a, MIR495, and MIR519 are recurrently found on eccDNA, almost exclusively in tumors, and patients with any one of these exhibited shorter progression-free survival. Synthetic eccDNAs with these MIR genes increase cancer cell proliferation in 786-O cells, suggesting oncogenic potential. Additionally, we found eccDNA carrying full protein-coding genes that are overexpressed in transcriptional analyses. Our findings suggest that small eccDNAs with miRNA genes are functional and contribute to the tumorigenesis of renal cell carcinoma.
Leishmania parasites alternate between hosts, facing environmental changes that demand rapid gene expression adaptation. Lacking canonical RNA polymerase II promoters, transcription in these eukaryotes is polycistronic, with gene regulation occurring post-transcriptionally. Although non-coding RNAs (ncRNAs) have been identified in Leishmania transcriptomes, their functions remain unclear. Recognizing RNA structure's importance, we performed a genome-wide alignment of L. braziliensis and related species, identifying conserved RNA structures, 38 of which overlap with known ncRNAs. One such ncRNA, lncRNA45, was functionally characterized. Using a knockout cell line, we demonstrated that lncRNA45 is crucial for parasite fitness. Reintroducing the wild-type lncRNA45 restored fitness, while a version with a single nucleotide substitution in the structured region did not. This mutation also altered RNA-protein interactions. These findings suggest that lncRNA45's regulatory role and protein interactions rely on its secondary structure. This study highlights the significance of structured lncRNAs in Leishmania biology and their potential as therapeutic targets. Further research into these ncRNAs could uncover new parasite regulation mechanisms and inspire novel treatment strategies. ### Competing Interest Statement The authors have declared no competing interest.
Abstract One strategy for CO2 mitigation is using photosynthetic microorganisms to sequester CO2 under high concentrations, such as in flue gases. While elevated CO2 levels generally promote growth, excessively high levels inhibit growth through uncertain mechanisms. This study investigated the physiology of the cyanobacterium Synechocystis sp. PCC 6803 under very high CO2 concentrations and yet stable pH around 7.5. The growth rate of the wild type (WT) at 200 µmol photons m−2 s−1 and a gas phase containing 30% CO2 was 2.7-fold lower compared to 4% CO2. Using a CRISPR interference mutant library, we identified genes that, when repressed, either enhanced or impaired growth under 30% or 4% CO2. Repression of genes involved in light harvesting (cpc and apc), photochemical electron transfer (cytM, psbJ, and petE), and several genes with little or unknown functions promoted growth under 30% CO2, while repression of key regulators of photosynthesis (pmgA) and CO2 capture and fixation (ccmR, cp12, and yfr1) increased growth inhibition under 30% CO2. Experiments confirmed that WT cells were more susceptible to light inhibition under 30% than under 4% CO2 and that a light-harvesting-impaired ΔcpcG mutant showed improved growth under 30% CO2 compared to the WT. These findings suggest that enhanced fitness under very high CO2 involves modifications in light harvesting, electron transfer, and carbon metabolism, and that the native regulatory machinery is insufficient, and in some cases obstructive, for optimal growth under 30% CO2. This genetic profiling provides potential targets for engineering cyanobacteria with improved photosynthetic efficiency and stress resilience for biotechnological applications. Key points • Synechocystis growth was inhibited under very high CO 2 . • Inhibition of growth under very high CO 2 was light dependent. • Repression of photosynthesis genes improved growth under very high CO 2 . Graphical Abstract
The current nucleic acid signal amplification methods for SARS-CoV-2 RNA detection heavily rely on the functions of biological enzymes which imposes stringent transportation and storage conditions, high cost and global supply shortages. Here, a non-enzymatic whole genome detection method based on a simple isothermal signal amplification approach is developed for rapid detection of SARS-CoV-2 RNA and potentially any types of nucleic acids regardless of their size. The assay, termed non-enzymatic isothermal strand displacement and amplification (NISDA), is able to quantify 10 RNA copies.µL −1 . In 164 clinical oropharyngeal RNA samples, NISDA assay is 100 % specific, and it is 96.77% and 100% sensitive when setting up in the laboratory and hospital, respectively. The NISDA assay does not require RNA reverse-transcription step and is fast (<30 min), affordable, highly robust at room temperature (>1 month), isothermal (42 °C) and user-friendly, making it an excellent assay for broad-based testing.
Targeted nucleases, primarily CRISPR-Cas-based systems, have revolutionized genome editing by enabling precise modification of target genes or transcripts. Many pre-clinical and clinical studies leverage this technology to develop treatments for human diseases; however, substantial off-target genotoxicity concerns delay its clinical translation. Despite the development of a wide array of tools, assays, and technologies aimed at identifying and quantifying off-target effects, the absence of standardized guidelines leads to inconsistent practices across studies. This review highlights the key challenges and potential solutions in ensuring the safety of gene editing studies for therapeutic applications, focusing on gRNA design, off-target sites prediction, and off-target activity measurement.
CRISPR-derived base editors (BE) enable precise single nucleotide substitution without introducing double-stranded DNA breaks. Apart from the base editing enzymes, efficient base editing strongly depends on both the CRISPR guide RNA (gRNA) efficiency and the edited position. Here, we show that the accuracy of BE gRNA design can be significantly improved by generating more data and by introducing deep neural networks trained on multiple different datasets simultaneously. Generating ~20,000 gRNAs for A•T to G•C and C•G to T•A conversions, we present such deep learning models, which also allow users to do dataset-aware predictions. The methods are available online and as stand-alone software.
Long intergenic non-coding RNAs (lincRNAs) and single nucleotide polymorphisms (SNPs) have been associated with cancers for years, yet the molecular mechanism is mostly unclear. The secondary structure of lincRNAs is often crucial for their biological function but can be disrupted by SNPs. However, only a small number of studies investigated how lincRNAs function through secondary structure. Given that the vast majority of cancer-related SNPs are located mainly in non-coding regions, there is a large potential for associating SNPs to disrupted structure in lincRNAs. To estimate the structural impacts of cancer-associated SNPs on lincRNAs, we predicted local secondary structures for lincRNAs and computed structural distances between structural ensembles of the wild-type and mutant sequences. Manual literature curation was performed to study the function of lincRNAs that are structurally disrupted by cancer-associated SNPs. By integrating with RBP binding sites annotation, we estimated the impacts of structural changes from cancer-associated SNPs on the protein binding of lincRNAs. We predict 559 SNPs to cause significant structural disruption in 231 lincRNAs using RNAsnp (P-value < 0.1). In addition, we find that these disrupted regions have the potential to alter the binding ability of lincRNAs with RBPs. An example is the structural change in the lincRNA small nucleolar RNA host gene 25 (SNHG25), which overlaps the binding site of protein insulin like growth factor 2 mRNA binding protein 2 (IGF2BP2) in glioblastoma multiforme. The results show the importance of the lincRNA secondary structure in understanding their biological function, especially the structural changes from SNPs in cancer. The predicted structural change in lincRNA SNHG25 holds a potential insight into the mechanism of protein IGF2BP2 in recognizing RNA methylation signals in glioblastoma multiforme.
Clustered regularly interspaced short palindromic repeats-based editing is inefficient at over two-thirds of genetic targets. A primary cause is ribonucleic acid (RNA) misfolding that can occur between the spacer and scaffold regions of the gRNA, which hinders the formation of functional Cas9 ribonucleoprotein (RNP) complexes. Here, we uncover hundreds of highly efficient gRNA variant scaffolds for Staphylococcus aureus (Sa)Cas9 utilizing an innovative binding and ligand activation driven enrichment (BLADE) methodology, which leverages asymmetrical product dissociation over rounds of evolution. SaBLADE-derived gRNA scaffolds contain 7%-42% of nucleotide variation relative to wild type. gRNA variants are able to improve gene editing efficiency at all targets tested, and they achieve their highest levels of editing improvement (>400%) at the most challenging DNA target sites for the wild-type SaCas9 gRNA. This arsenal of SaBLADE-derived gRNA variants showcases the power and flexibility of combinatorial chemistry and directed evolution to enable efficient gene editing at challenging, or previously intractable, genomic sites.
Background The global rise in obesity prevalence highlights an urgent need to understand its underlying pathophysiology, which ranges from preclinical obesity (excess body fat without overt disease) to clinical obesity with impairment of organ function. To study obesity, we fed obesity-prone Ossabaw pigs a medium cholesterol (0.5 weight %), high-fat (50% of energy) and high-fructose (21% of energy) diet (MC-HFD) for relatively short-term (11 weeks). Morphometric parameters, adipocyte area, and blood biomarkers of metabolic syndrome and inflammation were evaluated. Furthermore, RNA sequencing was adopted to characterize the transcriptome-wide response to this obesogenic challenge and identify key genes and molecular processes affected by the attained obesity state in visceral adipose tissue (VAT) and liver. Results Pigs fed the MC-HFD diet exhibited significant increases in body weight, body size, and adipocyte area compared to the control group. However, these pigs remained metabolically healthy, with only a minor increase in low-density lipoprotein (LDL), and normal C-reactive protein, triglyceride, glucose, and insulin levels. In VAT, 666 differentially expressed transcripts (DETs) were identified, while only 40 were found in liver tissue. Notably, FASN was the only transcript that was regulated similarly (significantly downregulated) in both tissues of MC-HFD pigs, suggesting tissue-specific transcriptional signatures. In VAT, transcripts related to the mitochondrial respiratory chain complex, protein synthesis, spliceosome and certain extracellular matrix components were upregulated, while collagen, GPCR signalling and fatty acid metabolism were generally downregulated. Compared to VAT, the liver of MC-HFD fed pigs exhibited a greater enrichment in transcripts commonly linked to obesity. Transcripts related to fatty acid biosynthesis and amino acid catabolism were suppressed, while hormone and lipoprotein metabolism associated transcripts were upregulated. Conclusions Taken together the data presented here suggest that MC-HFD fed Ossabaw pigs attained a pre-metabolic syndrome state of obesity characterized by an ‘obesity tolerant’ transcriptional response in VAT in contrast to the liver response resembling an early canonical obesity response.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark14