Plant systems show dynamic responses, such as changes in architecture and physiology, to adjust their growth in changing environments. The reconfiguration of network modules underlies these responses. There is multi-scale regulation acting on these networks that can be measured as changes in mRNA synthesis, stability, and decay; and in protein translation, activity, affinity, and decay, among others. The regulatory linkages across biological scales can be constitutive, tunable, or switchable under changing environments. The ultimate goal of breeding efforts is to create novel ideotypes with desired traits, which maintain high yield despite challenging environments. Gain of beneficial or adaptive phenotypic traits in commercial crops is often due to ideal coupling/uncoupling of the network modules with tunable linkages from parental lines. Regulatory genomic variation of influential nodes in network modules perturbs network properties, such as hubs, topology, and clustering, and serves as a source of variation for novel traits. These critical variants that perturb network modules to create novel phenotypes can be discovered using natural diversity in wild relatives of cultivated crops, thereby aiding breeding programs for target discovery. Predictive modeling and quantitatively characterized synthetic modules provide detailed understanding on predictable and heritable behavior of complex regulatory and signaling networks, but implementation of these models in crop plants lags behind. Future efforts to incorporate multi-scale layers of information to predict systems level behavior of crop plant networks and their dynamics in changing environments are an exciting area.
The genetic engineering of value-added traits such as the accumulation of bioproducts in high biomass C4 grass stems is one promising strategy to make plant-derived biofuels more economical for industrial use. A first step toward achieving this goal is to identify stem-specific promoters that can drive the expression of genes of interest with good temporal and spatial specificity. However, a comprehensive characterization of the spatial-temporal regulatory elements of stem tissue-specific promoters for C4 grasses has not been reported. Therefore, we performed an in-silico analysis on Sorghum bicolor cv BTx623 transcriptomes from multiple tissues over development to identify stem-expressed genes. The analysis identified 10 genes that are “Always-On-Stem-Specific,” 59 genes that are “Temporally-Stem-Specific during early development,” and 21 genes that are “Temporally-Stem-Specific during late development.” Promoter analysis revealed common and/or unique cis-regulatory elements in promoters of genes within each of the three categories. Subsequent gene regulatory network (GRN) analysis revealed that different transcriptional regulatory programs are responsible for the temporal activation of the stem-expressed genes. The analysis of temporal stem GRNs between sweet ( cv Della ) and grain ( cv BTx623 ) sorghum varieties revealed genetic variation that could influence the regulatory landscape. This study provides new insights about sorghum stem biology, and information for future genetic engineering efforts to fine-tune the spatial-temporal expression of transgenes in C4 grass stems.
The distributions of secondary structural elements appear to differ between coding regions (CDS) of mRNAs compared to the untranslated regions (UTRs), presumably as a mechanism to fine-tune gene expression, including efficiency of translation. However, a systematic and comprehensive analysis of secondary structure avoidance because of potential bias in codon usage is difficult as some of the common secondary structures, such as, hairpins can be formed by numerous sequence combinations. Using G-quadruplex (GQ) as the model secondary structure we studied the impact of codon bias on GQs within the CDS. Because GQs can be predicted using specific consensus sequence motifs, they provide an excellent platform for investigation of the selectivity of such putative structures at the codon level. Using a bioinformatics approach, we calculated the frequencies of putative GQs within the CDS of a variety of species. Our results suggest that the most stable GQs appear to be significantly underrepresented within the CDS, through the use of specific synonymous codon combinations. Furthermore, we identified many peptide sequence motifs in which silent mutations can potentially alter translation via stable GQ formation. This work not only provides a comprehensive analysis on how stable secondary structures appear to be avoided within the CDS of mRNA, but also broadens the current understanding of synonymous codon usage as they relate to the structure-function relationship of RNA.
The piwi-interacting RNAs (piRNAs) are small non-coding RNAs, mostly 24-32 nucleotides in length. The piRNAs are not known to have any conserved secondary structure or sequence motifs. Using bioinformatics analysis, we discovered the presence of putative G-quadruplex (GQ) forming sequences in human piRNAs. We studied human piR-48164/piR-GQ containing a potential GQ forming sequence and using biochemical and biophysical techniques confirmed its ability to form a GQ. Using EMSA, we discovered that the formation of GQ structure led to inhibition of the piRNA binding to the HIWI-PAZ domain as well as the complementary base pairing to a target RNA. The inability of the piR-GQ to interact with the PIWI protein might be detrimental to the function of the piRNA. To investigate if the formation of a GQ structure in piRNA prevents its target gene silencing in vivo, we used a reporter assay. The piR-GQ failed to inhibit the reporter gene expression while a mutated version that lacked the ability to form GQ inhibited reporter gene expression indicating that the presence of GQ in piRNA is detrimental to its function. These studies unraveled the dependence of a piRNA's functionality on an RNA secondary structure and added a new layer of regulation to their function. (C) 2018 Elsevier B.V. and Societe Francaise de Biochimie et Biologie Moleculaire (SFBBM). All rights reserved.
Unicellular flagellates that make up the class Kinetoplastida include multiple parasites responsible for public health concerns, including Trypanosoma brucei and T. cruzi (agents of African sleeping sickness and Chagas disease, respectively), and various Leishmania species, which cause leishmaniasis. These diseases are generally difficult to eradicate, with treatments often having lethal side effects and/or being effective only during the acute phase of the diseases, when most patients are still asymptomatic. Phospholipid signaling and metabolism are important in the different life stages of Trypanosoma, including playing a role in transitions between stages and in immune system evasion, thus, making the responsible enzymes into potential therapeutic targets. However, relatively little is understood about how the pathways function in these pathogens. Thus, in this study we examined evolutionary history of proteins from one such signaling pathway, namely phospholipase D (PLD) homologs. PLD is an enzyme responsible for synthesizing phosphatidic acid (PA) from membrane phospholipids. PA is not only utilized for phospholipid synthesis, but is also involved in many other signaling pathways, including biotic and abiotic stress response. 37 different representative Kinetoplastida genomes were used for an exhaustive search to identify putative PLD homologs. The genome of Bodo saltans was the only one of surveyed Kinetoplastida genomes that encoded a protein that clustered with plant PLDs. The representatives from other Kinetoplastida species clustered together in two different clades, thought to be homologous to the PLD superfamily, but with shared sequence similarity with cardiolipin synthases (CLS), and phosphatidylserine synthases (PSS). The protein structure predictions showed that most Kinetoplastida sequences resemble CLS and PSS, with the exception of 5 sequences from Bodo saltans that shared significant structural similarities with the PLD sequences, suggesting the loss of PLD-like sequences during the evolution of parasitism in kinetoplastids. On the other hand, diacylglycerol kinase (DGK) homologs were identified for all species examined in this study, indicating that DGK could be the only pathway for the synthesis of PA involved in lipid signaling in these organisms due to genome streamlining during transition to parasitic lifestyle. Our findings offer insights for development of potential therapeutic and/or intervention approaches, particularly those focused on using PA, PLD and/or DGK related pathways, against trypanosomiasis, leishmaniasis, and Chagas disease.
Recognition of viral epitopes by the host immune system plays a pivotal role in controlling viral infection, while at the same time exerting selective pressure for escape mutations, which in turn leads to many epitopes evolving faster than the adjacent regions. Some epitope regions appear to be highly conserved despite the strong immune pressure due to functional and/or structural constraints acting on them. Yet, we still have relatively little understanding of the nature of protein structural and functional constraints operating on epitope regions. Here, we identify coevolving epitope regions that have high functional and/or structural constraints. We examined patterns of coevolution of protein segment pairs between Integrase-Reverse Transcriptase, Integrase-Vpr, and Reverse Transcriptase-Vpr protein interactions pairs in HIV-1 pre-integration complex. Our coevolutionary and structural analysis shows that protein regions with strong multiple coevolutionary constraints with few regions are located in structurally conserved regions (i.e., those with well-defined secondary structures). Meanwhile, protein regions with strong multiple coevolutionary constraints with many regions are clustered in the structurally flexible regions (regions with high B-factor) of the proteins and regions with high conformational diversity (relative mobility in the collective dynamics). On the other hand, protein regions with weak coevolutionary signals are clustered in structurally disordered regions. Identifying viral epitope regions that harbor strong constraints against escape mutations is important in development of vaccines that can induce immune responses against multiple conserved epitopes. This analysis also offers important insights into molecular evolution of protein interactions.
The replication of human immunodeficiency virus-1 (HIV-1) requires reverse transcription of the viral RNA genome and integration of newly synthesized pro-viral DNA into the host genome. This is mediated by the viral proteins reverse transcriptase (RT) and integrase (IN). The formation and stabilization of the pre-integration complex (PIC), which is an essential step for reverse transcription, nuclear import, chromatin targeting, and subsequent integration, involves direct and indirect modes of interaction between RT and IN proteins. While epitope-based treatments targeting IN–viral DNA and IN–RT complexes appear to be a promising combination for an anti-HIV treatment, the mechanisms of IN-RT interactions within the PIC are not well understood due to the transient nature of the protein complex and the intrinsic flexibility of its components. Here, we identify potentially interacting regions between the IN and RT proteins within the PIC through the coevolutionary analysis of amino acid sequences of the two proteins. Our results show that specific regions in the two proteins have strong coevolutionary signatures, suggesting that these regions either experience direct and prolonged interactions between them that require high affinity and/or specificity or that the regions are involved in interactions mediated by dynamic conformational changes and, hence, may involve both direct and indirect interactions. Other regions were found to exhibit weak, but positive correlations, implying interactions that are likely transient and/or have low affinity. We identified a series of specific regions of potential interactions between the IN and RT proteins (e.g., specific peptide regions within the C-terminal domain of IN were identified as potentially interacting with the Connection domain of RT). Coevolutionary analysis can serve as an important step in predicting potential interactions, thus informing experimental studies. These studies can be integrated with structural data to gain a better understanding of the mechanisms of HIV protein interactions.