Pennycress ( Thlaspi arvense ) is a promising intermediate oilseed crop, producing oil suitable for conversion to biofuels—including aviation fuels. While domestication efforts are ongoing, a deeper understanding of the genetic architecture of traits is crucial for informing future breeding efforts. Here, we conducted the largest genomic and phenotypic survey of pennycress to date, analyzing 739 accessions collected across four continents. Leveraging whole-genome sequencing and field-collected phenotypes, we characterized the standing genetic variation underlying key agronomic traits and climate resilience. Our findings revealed multiple independent migration events to North America, with substantial genetic admixture. We identified homologs of Arabidopsis thaliana flowering-time genes that contribute to adaptation and demonstrated the agronomic benefits of winter-type pennycress. Furthermore, through multi-season field trials, we identified a genomic region containing a cluster of mTERF genes strongly associated with green canopy coverage, a critical trait for biomass retention and yield stability. These insights provide a genomic roadmap for accelerating pennycress domestication and improving its resilience to climate variability. ### Competing Interest Statement The authors have declared no competing interest.
Microbiome assembly, structure, and dynamics significantly influence plant health. Secreted microbial signaling molecules initiate and mediate symbiosis by binding to structurally compatible plant receptors. For example, lipo-chitooligosaccharides (LCOs), produced by nitrogen-fixing rhizobial bacteria and various fungi, are recognized by plant lysin motif receptor-like kinases (LysM-RLKs), which activate the common symbiotic pathway. Accurately predicting these molecular interactions could reveal complementary signatures underlying the initial stages of endosymbiosis. Despite the breakthrough in protein-ligand structure prediction with deep learning-based tools, such as AlphaFold3, the large size and highly flexible nature of signaling compounds like LCOs present major challenges for detailed structural characterization and binding-affinity prediction. Typical structure-/physics-based methods of ligand virtual screening are designed for small, drug-like molecules, often rely on high-resolution, experimentally determined structures of the protein receptors, and rarely achieve sufficient sampling to obtain converged thermodynamic quantities with large ligands. In this study, we developed a hybrid molecular dynamics/machine learning (MD/ML) approach capable of predicting binding affinity rankings with high accuracy in systems involving large, flexible ligands, despite limited experimental structural information. Using coarse initial structural models, the predictions using the MD/ML workflow achieved strong alignment with experimental trends, particularly in the top-affinity tier for four legume LysM-RLKs (LYR3) binding to LCOs and a chitooligosaccharide. Furthermore, the MD-based conformation selection protocol provided critical structural insights into substrate specificity and binding mechanisms. This study demonstrates a powerful method to screen for challenging cognate ligand-receptors and advance our understanding of the molecular basis of microbial colonization in plants.
Field pennycress (Thlaspi arvense) is a new biofuel winter annual crop with extreme cold hardiness and a short life cycle, enabling off-season integration into corn and soybean rotations across the U.S. Midwest. Pennycress fields are susceptible to winter snow melt and spring rainfall, leading to waterlogged soils. The objective of this research was to determine the extent to which waterlogging during the reproductive stage affected gene expression, morphology, physiology, recovery, and yield between two pennycress lines (SP32-10 and MN106). In a controlled environment, total pod number, shoot/root dry weight, and total seed count/weight were significantly reduced in SP32-10 in response to waterlogging, whereas primary branch number, shoot dry weight, and single seed weight were significantly reduced in MN106. This indicated waterlogging had a greater negative impact on seed yield in SP32-10 than MN106. We compared the transcriptomic response of SP32-10 and MN106 to determine the gene expression patterns underlying these different responses to seven days of waterlogging. The number of differentially expressed genes (DEGs) between waterlogged and control roots were doubled in MN106 (3,424) compared to SP32-10 (1,767). Functional enrichment analysis of upregulated DEGs revealed Gene Ontology (GO) terms associated with hypoxia and decreased oxygen, with genes in these categories encoding proteins involved in alcoholic fermentation and glycolysis. Additionally, downregulated DEGs revealed GO terms associated with cell wall biogenesis and suberin biosynthesis, indicating suppressed growth and energy conservation. Interestingly, MN106 waterlogged roots exhibited significant stronger regulation of these genes than SP32-10, displaying a more robust transcriptomic response overall. Together, these results reveal the reconfiguration of cellular and metabolic processes in response to the severe energy crisis invoked by waterlogging in pennycress.
Herbicide-resistant weeds are increasingly a problem in crop fields when exposed to similar chemistry over time. To avoid future yield losses, identifying herbicidal chemistry needs to be accelerated. We screened 50,000 small molecules using a liquid-handling robot and light microscopy focusing on pre-emergent herbicides in the family of cellulose biosynthesis inhibitors. Through phenotypic, chemical, genetic, and in silico methods we uncovered 6-{[4-(2-fluorophenyl)-1-piperazinyl]methyl}-N-(2-methoxy-5-methylphenyl)-1,3,5-triazine-2,4-diamine (fluopipamine). Symptomologies support fluopipamine as a putative antagonist of cellulose synthase enzyme 1 (CESA1) from Arabidopsis (Arabidopsis thaliana). Ectopic lignification, inhibition of etiolation, phenotypes including loss of anisotropic cellular expansion, swollen roots, and live cell imaging link fluopipamine to cellulose biosynthesis inhibition. Radiolabeled glucose incorporation of cellulose decreased in short-duration experiments when seedlings were incubated in fluopipamine. To elucidate the mechanism, ethylmethanesulfonate mutagenized M2 seedlings were screened for fluopipamine resistance. Two loci of genetic resistance were linked to CESA1. In silico docking of fluopipamine, quinoxyphen, and flupoxam against various CESA1 mutations suggests that an alternative binding site at the interface between CESA proteins is necessary to preserve cellulose polymerization in compound presence. These data uncovered potential fundamental mechanisms of cellulose biosynthesis in plants along with feasible leads for herbicidal uses.
While the proliferation of data-driven omics technologies has continued to accelerate, methods of identifying relationships among large-scale changes from omics experiments have stagnated. It is therefore imperative to develop methods that can identify key mechanisms among one or more omics experiments in order to advance biological discovery. To solve this problem, here we describe the network-based algorithm MENTOR - Multiplex Embedding of Networks for Team-Based Omics Research. We demonstrate MENTOR's utility as a supervised learning approach to successfully partition a gene set containing multiple ontological functions into their respective functions. Subsequently, we used MENTOR as an unsupervised learning approach to identify important biological functions pertaining to the host genetic architectures in Populus trichocarpa associated with microbial abundance of multiple taxa. Moreover, as open source software designed with scientific teams in mind, we demonstrate the ability to use the output of MENTOR to facilitate distributed interpretation of omics experiments.
For plants, distinguishing between mutualistic and pathogenic microbes is a matter of survival. All microbes contain microbe-associated molecular patterns (MAMPs) that are perceived by plant pattern recognition receptors (PRRs). Lysin motif receptor-like kinases (LysM-RLKs) are PRRs attuned for binding and triggering a response to specific MAMPs, including chitin oligomers (COs) in fungi, lipo-chitooligosaccharides (LCOs), which are produced by mycorrhizal fungi and nitrogen-fixing rhizobial bacteria, and peptidoglycan in bacteria. The identification and characterization of LysM-RLKs in candidate bioenergy crops including Populus are limited compared to other model plant species, thus inhibiting our ability to both understand and engineer microbe-mediated gains in plant productivity. As such, we performed a sequence analysis of LysM-RLKs in the Populus genome and predicted their function based on phylogenetic analysis with known LysM-RLKs. Then, using predictive models, molecular dynamics simulations, and comparative structural analysis with previously characterized CO and LCO plant receptors, we identified probable ligand-binding sites in Populus LysM-RLKs. Using several machine learning models, we predicted remarkably consistent binding affinity rankings of Populus proteins to CO. In addition, we used a modified Random Walk with Restart network-topology based approach to identify a subset of Populus LysM-RLKs that are functionally related and propose a corresponding signal transduction cascade. Our findings provide the first look into the role of LysM-RLKs in Populus-microbe interactions and establish a crucial jumping-off point for future research efforts to understand specificity and redundancy in microbial perception mechanisms.
The unprecedented scientific achievements in combating the COVID-19 pandemic reflect a global response informed by unprecedented access to data. We now have the ability to rapidly generate a diversity of information on an emerging pathogen and, by using high-performance computing and a systems biology approach, we can mine this wealth of information to understand the complexities of viral pathogenesis and contagion like never before. These efforts will aid in the development of vaccines, antiviral medications, and inform policymakers and clinicians. Here we detail computational protocols developed as SARS-CoV-2 began to spread across the globe. They include pathogen detection, comparative structural proteomics, evolutionary adaptation analysis via network and artificial intelligence methodologies, and multiomic integration. These protocols constitute a core framework on which to build a systems-level infrastructure that can be quickly brought to bear on future pathogens before they evolve into pandemic proportions.
We developed Distilled Graph Attention Policy Network (DGAPN), a reinforcement learning model to generate novel graph-structured chemical representations that optimize user-defined objectives by efficiently navigating a physically constrained domain. The framework is examined on the task of generating molecules that are designed to bind, noncovalently, to functional sites of SARS-CoV-2 proteins. We present a spatial Graph Attention (sGAT) mechanism that leverages self-attention over both node and edge attributes as well as encoding the spatial structure --- this capability is of considerable interest in synthetic biology and drug discovery. An attentional policy network is introduced to learn the decision rules for a dynamic, fragment-based chemical environment, and state-of-the-art policy gradient techniques are employed to train the network with stability. Exploration is driven by the stochasticity of the action space design and the innovation reward bonuses learned and proposed by random network distillation. In experiments, our framework achieved outstanding results compared to state-of-the-art algorithms, while reducing the complexity of paths to chemical synthesis.
Gene-to-gene networks, such as Gene Regulatory Networks (GRN) and Predictive Expression Networks (PEN) capture relationships between genes and are beneficial for use in downstream biological analyses. There exists multiple network inference tools to produce these gene-to-gene networks from matrices of gene expression data. Random Forest-Leave One Out Prediction (RF-LOOP) is a method that has been shown to be efficient at producing these gene-to-gene networks, frequently known as GEne Network Inference with Ensemble of trees (GENIE3). Random Forest can be replaced in this process by iterative Random Forest (iRF), which performs variable selection and boosting. Here we validate that iterative Random Forest-Leave One Out Prediction (iRF-LOOP) produces higher quality networks than GENIE3 (RF-LOOP). We use both synthetic and empirical networks from the Dialogue for Reverse Engineering Assessment and Methods (DREAM) Challenges by Sage Bionetworks, as well as two additional empirical networks created from Arabidopsis thaliana and Populus trichocarpa expression data.
Abstract Despite SARS-CoV and SARS-CoV-2 being equipped with highly similar protein arsenals, the corresponding zoonoses have spread among humans at extremely different rates. The specific characteristics of these viruses that led to such distinct outcomes remain unclear. Here, we apply proteome-wide comparative structural analysis aiming to identify the unique molecular elements in the SARS-CoV-2 proteome that may explain the differing consequences. By combining protein modeling and molecular dynamics simulations, we suggest nonconservative substitutions in functional regions of the spike glycoprotein (S), nsp1, and nsp3 that are contributing to differences in virulence. Particularly, we explain why the substitutions at the receptor-binding domain of S affect the structure–dynamics behavior in complexes with putative host receptors. Conservation of functional protein regions within the two taxa is also noteworthy. We suggest that the highly conserved main protease, nsp5, of SARS-CoV and SARS-CoV-2 is part of their mechanism of circumventing the host interferon antiviral response. Overall, most substitutions occur on the protein surfaces and may be modulating their antigenic properties and interactions with other macromolecules. Our results imply that the striking difference in the pervasiveness of SARS-CoV-2 and SARS-CoV among humans seems to significantly derive from molecular features that modulate the efficiency of viral particles in entering the host cells and blocking the host immune response.
Using a Systems Biology approach, we integrated genomic, transcriptomic, proteomic, and molecular structure information to provide a holistic understanding of the COVID-19 pandemic. The expression data analysis of the Renin Angiotensin System indicates mild nasal, oral or throat infections are likely and that the gastrointestinal tissues are a common primary target of SARS-CoV-2. Extreme symptoms in the lower respiratory system likely result from a secondary-infection possibly by a comorbidity-driven upregulation of ACE2 in the lung. The remarkable differences in expression of other RAS elements, the elimination of macrophages and the activation of cytokines in COVID-19 bronchoalveolar samples suggest that a functional immune deficiency is a critical outcome of COVID-19. We posit that using a non-respiratory system as a major pathway of infection is likely determining the unprecedented global spread of this coronavirus. One Sentence Summary A Systems Approach Indicates Non-respiratory Pathways of Infection as Key for the COVID-19 Pandemic
Background: The magnitude and severity of the COVID-19 pandemic cannot be overstated. Although the mortality rate is less than SARS and MERS, the global outbreak has already resulted in orders of magnitude more deaths. In order to tackle the complexities of this disease, a Systems Biology approach can provide insights into the biology of the virus and mechanisms of disease. Methods: Using a Systems Biology approach, we have integrated genomic, transcriptomic, proteomic, and molecular evolution data layers to understand its impact on host cells. We overlay these analyses with high-resolution structural models and atomistic molecular dynamics simulations conducted on the Summit supercomputer at the Oak Ridge National Laboratory. Findings: Transcriptomic and proteomic data indicate little to no expression of ACE2 in lung tissue. Molecular modeling simulations support ACE2 as the receptor for SARS-CoV-2, but ACE may also act as a receptor for the virus and may be important for entry of SARS-CoV-1. Gene expression data from bronchoalveolar lavage samples from COVID-19 patients identify upregulation of renin, angiotensin, and the angiotensin 1-7 receptor MAS as well as a cellular landscape consistent with large-scale dissolution of lung parenchyma tissues, likely comprised of all lung epithelial cell types as well as lymphatic endothelial cells, but an absence of cells, such as macrophages, normally essential for host defense. Interpretation: Our analyses indicate that the commonly accepted view that SARS-CoV-2 enters host cells via ACE2 expressed in the lung is unlikely because ACE2 is undetectable there. Instead, given the greater target space of ACE2-positive nasal, oral, and gastrointestinal tissues, a more likely scenario suggests initial infection in those tissues is followed by a secondary infection via migration through the lymphatic system and bloodstream to the lung microvasculature. The elimination of macrophages and complete lack of activated cytokine signature in COVID-19 lung samples suggest that a major component of SARS-CoV-29s virulence is its net effect of causing a functional immune deficiency syndrome. Our structural analysis of the SARS-CoV-2 proteome suggests involvement of the highly conserved nsp5 protein as part of a major mechanism that suppresses the nuclear factor transcription factor kappa B (NF-κB) pathway, eliminating the host cell9s interferon-based antiviral response.
We demonstrate a selection of network and machine learning techniques useful in the analysis of complex datasets, including 2-way similarity networks, Markov clustering, enrichment statistical networks, FCROS differential analysis, and random forests. We demonstrate each of these techniques on the Populus trichocarpa gene expression atlas.
Background A mechanistic understanding of the spread of SARS-CoV-2 and diligent tracking of ongoing mutagenesis are of key importance to plan robust strategies for confining its transmission. Large numbers of available sequences and their dates of transmission provide an unprecedented opportunity to analyze evolutionary adaptation in novel ways. Addition of high-resolution structural information can reveal the functional basis of these processes at the molecular level. Integrated systems biology-directed analyses of these data layers afford valuable insights to build a global understanding of the COVID-19 pandemic. Results Here we identify globally distributed haplotypes from 15,789 SARS-CoV-2 genomes and model their success based on their duration, dispersal, and frequency in the host population. Our models identify mutations that are likely compensatory adaptive changes that allowed for rapid expansion of the virus. Functional predictions from structural analyses indicate that, contrary to previous reports, the Asp 614 Gly mutation in the spike glycoprotein (S) likely reduced transmission and the subsequent Pro 323 Leu mutation in the RNA-dependent RNA polymerase led to the precipitous spread of the virus. Our model also suggests that two mutations in the nsp13 helicase allowed for the adaptation of the virus to the Pacific Northwest of the USA. Finally, our explainable artificial intelligence algorithm identified a mutational hotspot in the sequence of S that also displays a signature of positive selection and may have implications for tissue or cell-specific expression of the virus. Conclusions These results provide valuable insights for the development of drugs and surveillance strategies to combat the current and future pandemics.
The objective of the present study was to isolate and identify polyphenol degrading microorganism and to optimize the culture conditions for better yield and increased mass production. Waste water samples collected from the effluents of tanneries were filtered and serially diluted. The microbes were grown on nutrient agar medium for 24 hours. The isolated colonies were then transferred onto a mineral salt medium containing various concentration of polyphenol (tannic acid), followed by incubation at 370C for 72 hrs in CO2 incubator. The survival of microorganisms at various increasing concentration of polyphenol (50 ppm to 250 ppm) was performed. The plate which showed highest number of colonies was selected and the isolated colony was subjected to morphological, biochemical and molecular identification. The results showed that the isolate belonged to Bacillus genus and the resulting bacterial strain isolate was found to be Bacillus subtilis and the GenBank accession number was obtained which was MK760577. After identification, the culture conditions required for maximum enzyme production were optimized.
Various 'omics data types have been generated for Populus trichocarpa, each providing a layer of information which can be represented as a density signal across a chromosome. We make use of genome sequence data, variants data across a population as well as methylation data across 10 different tissues, combined with wavelet-based signal processing to perform a comprehensive analysis of the signature of the centromere in these different data signals, and successfully identify putative centromeric regions in P. trichocarpa from these signals. Furthermore, using SNP (single nucleotide polymorphism) correlations across a natural population of P. trichocarpa, we find evidence for the co-evolution of the centromeric histone CENH3 with the sequence of the newly identified centromeric regions, and identify a new CENH3 candidate in P. trichocarpa.
Various patterns of multi-phenotype associations (MPAs) exist in the results of Genome Wide Association Studies (GWAS) involving different topologies of single nucleotide polymorphism (SNP)-phenotype associations. These can provide interesting information about the different impacts of a gene on closely related phenotypes or disparate phenotypes (pleiotropy). In this work we present MPA Decomposition, a new network-based approach which decomposes the results of a multi-phenotype GWAS study into three bipartite networks, which, when used together, unravel the multi-phenotype signatures of genes on a genome-wide scale. The decomposition involves the construction of a phenotype powerset space, and subsequent mapping of genes into this new space. Clustering of genes in this powerset space groups genes based on their detailed MPA signatures. We show that this method allows us to find multiple different MPA and pleiotropic signatures within individual genes and to classify and cluster genes based on these SNP-phenotype association topologies. We demonstrate the use of this approach on a GWAS analysis of a large population of 882 Populus trichocarpa genotypes using untargeted metabolomics phenotypes. This method should prove invaluable in the interpretation of large GWAS datasets and aid in future synthetic biology efforts designed to optimize phenotypes of interest.