Natural product discovery is increasingly driven by the ability to analyze microbial genomes for biosynthetic gene clusters (BGCs) that encode secondary metabolites. While existing approaches have successfully linked BGCs to broad classes of chemical products, they typically operate in a single modality (genomic or chemical) limiting the scope of bidirectional prediction. In this work, we propose a multimodal framework that integrates genomic and chemical information by projecting embeddings derived from pretrained language models into a common representation space. We embed genomic sequences using a BGC foundation model and represent molecules through a chemical language model, then use a metric learning model to co-embed BGCs and their associated chemical structures. This co-embedding space allows us to quantify the similarity between BGCs and compounds using similarity measures, enabling both efficient forward and inverse retrieval tasks. Our approach consistently outperforms the non-alignment approach and represents a generalizable, scalable strategy to bridge biological and chemical modalities in natural product discovery.
The fungal kingdom represents a greatly untapped resource to produce a wide range of bioactive secondary metabolites, including antibiotics, anticancer agents, industrially significant dyes and enzymes. To-date, it is estimated only less than 5% of all fungi have been characterised, a deficit that is especially pronounced in tropical regions like Singapore, where fungal diversity remains underexplored compared to northern hemisphere counterparts. This underlines the urgency and importance of our research which motivated the creation of our curated dataset, aiming to address this gap and contribute to understanding the broader ecosystem. We developed a generalisable cultivation workflow that enables systematic strain preparation, supports high-resolution imaging, and yields sufficient fungal biomass amenable for genomic analyses. This resulted in a diverse collection of 518 phylogenetically and ecologically varied fungal strains from both terrestrial and marine environments in biodiverse Singapore. The curated dataset from this project captures both taxonomic identity and colony-level morphological traits serving as a foundation for visual phenotype to taxonomy mapping through the integration of computer vision.
Mouse lemurs (Microcebus spp.) are an emerging primate model organism, but their genetics, cellular and molecular biology remain largely unexplored. In an accompanying paper1, we performed large-scale single-cell RNA sequencing of 27 organs from mouse lemurs. We identified more than 750 molecular cell types, characterized their transcriptomic profiles and provided insight into primate evolution of cell types. Here we use the generated atlas to characterize mouse lemur genes, physiology, disease and mutations. We uncover thousands of previously unidentified lemur genes and hundreds of thousands of new splice junctions including over 85,000 primate splice junctions missing in mice. We systematically explore the lemur immune system by comparing global expression profiles of key immune genes in health and disease, and by mapping immune cell development, trafficking and activation. We characterize primate-specific and lemur-specific physiology and disease, including molecular features of the immune program, lemur adipocytes and metastatic endometrial cancer that resembles the human malignancy. We present expression patterns of more than 400 primate genes missing in mice, many with similar expression patterns to humans and some implicated in human disease. Finally, we provide an experimental framework for reverse genetic analysis by identifying naturally occurring nonsense mutations in three primate immune genes missing in mice and by analysing their transcriptional phenotypes. This work establishes a foundation for molecular and genetic analyses of mouse lemurs and prioritizes primate genes, isoforms, physiology and disease for future study.
A lack of tests to assess fetal development impacts decision making around antenatal steroid use in women at risk of preterm birth. We analyzed the expression of 21 cfRNA targets related to human fetal lung maturation. Discovery studies were performed using maternal and fetal sheep plasma, with results compared to fetal lung mRNA expression. These findings were then validated in first, second, and third trimester human maternal plasma samples. Discovery studies utilized a preterm sheep model of pregnancy. Date mated ewes received saline (control n = 6), or antenatal steroids (dexamethasone n = 12) (betamethasone n = 11) prior to delivery and ventilation. We analyzed the expression of 21 human cfRNA targets related to lung maturation in maternal and fetal sheep plasma and compared this to mRNA expression in fetal lung tissue. Findings were first validated in a separate cohort of sheep exposed to betamethasone (n = 8), intraamniotic LPS endotoxin for lung maturation (n = 6), or untreated term animals (n = 6). Findings were further validated in maternal plasma from a human cohort of uncomplicated term pregnancies (n = 10). Delivery and ventilation data were analyzed with ANOVA, Tukey HSD, and Dunnett T3 tests. A Random Forest algorithm identified genes that separated mature from immature fetal lung subgroups and determined AUC values for maternal and fetal cell-free RNA (cfRNA) feature sets to predict fetal lung maturation. We demonstrate that the analysis of 21 human cfRNA targets in maternal plasma is highly predictive of fetal lung maturation status across antenatal steroid induced (Dexamethasone AUC = 0.93; Betamethasone AUC = 1) and physiological (AUC = 1) lung development models. Maternal plasma cfRNA expression in the dexamethasone antenatal steroid group closely resembled direct fetal lung tissue mRNA expression. These findings were then validated in human maternal plasma samples (1st vs. 3rd trimester AUC = 0.96; 2nd vs. 3rd trimester AUC = 1). Further development of this technology may provide a rapid, minimally invasive, and cost-effective clinical tool to optimize patient selection for initial and repeat courses of antenatal steroids, along with insights into the molecular mechanisms underlying fetal lung development. Graphical abstract illustrating the potential future application of maternal plasma cell-free RNA analysis. Created in https://BioRender.com .
The management and prevention of key inflammatory-associated pregnancy complications such as chorioamnionitis and pre-eclampsia is hampered by a lack of early gestation risk screening tools. In a proof-of-principle study we used targeted cell-free RNA analyses of maternal plasma samples from large animal (sheep) and human pregnancy cohorts to develop a minimally invasive screening test for inflammatory markers. This study utilised a preterm sheep model of sterile and bacterial chorioamnionitis. Date-mated ewes received either intraamniotic Saline Control (n = 10) or E.Coli LPS (Sterile chorioamnionitis) with 2 days (n = 9) or 8 days exposure(n = 6). Preterm lambs were delivered at 124 ± 1d gestation. Findings were validated in a bacterial model of chorioamnionitis where ewes were exposed to 7 days of intraamniotic M.Hominis with delivery at 98 d gestion(n = 8) or 128d gestation(n = 8). Maternal blood was collected prior to intervention and at delivery in each group. Random Forest algorithm was used to analyse 8 cell-free RNA(cfRNA) targets related to inflammation in maternal plasma at baseline and delivery, identifying genes that separated animals with or without intrauterine inflammation. Plasma cfRNA data was compared to mRNA expression in placental tissue. Haematological and placental mRNA comparisons were analysed with ANOVA/Tukey HSD/Dunnett T3 tests. Maternal plasma cfRNA findings of intrauterine inflammation were then validated in human plasma samples from a cohort of patients with late onset pre-eclampsia (n = 10) or uncomplicated pregnancies (n = 10). We present data showing that targeted maternal cfRNA assays can accurately identify chorioamnionitis of sterile (AUC 1.0) and infectious (AUC 0.84) origin in a sheep model of pregnancy. Findings were then validated in human maternal plasma samples from patients with late-onset pre-eclampsia in the 1st (AUC = 0.85), 2nd (AUC = 0.90) and 3rd (AUC = 0.82) trimesters. In both sheep and human model systems, cfRNA tests offered high levels of sensitivity and specificity in the absence of overt clinical symptoms. We suggest that further development of this technology may serve as a scalable, rapidly deployed and cost-effective means for predicting major inflammatory conditions in pregnancy.
Natural products (NP) are instrumental in drug development, but their discovery and validation remain challenging and laborious despite advances in both genomic and analytical technologies. In this study, we demonstrate the use of an integrated multi-modal characterization of a microbial strain library for enhanced natural product discovery. This characterization utilizes language- and transformer-based models, integrated through a cross-validate and rank approach to search a mass spectrometry (MS)-genome multi-modal dataset with high confidence. MS data are analysed using an in-house developed tandem mass spectral MS/MS to structural elucidation workflow (WISE) that features a combination of molecular language and transformer-based models to predict corresponding molecular structures. Simultaneously, the related genomic data is pre-processed using the protein language model (ESM2) to extract meaningful embeddings. As a proof of concept, these models and pre-processed linked MS-genome datasets were applied and validated for the rapid identification of microbial strains capable of producing three diverse natural product compounds with precision ranging from 75-100%. Our findings demonstrate the transformative potential of linked MS-genome datasets at the strain-level to accelerate natural product discovery. This approach can expand the range of biotechnological innovations beyond what is currently known and curated, while also greatly reduce the resources and effort needed for discovery. ### Competing Interest Statement The authors have declared no competing interest.
Natural products possess significant therapeutic potential but remain underutilized despite advances in genomics and bioinformatics. While there are approaches to activate and upregulate natural product biosynthesis in both native and heterologous microbial strains, a comprehensive strategy to elicit production of natural products as well as a generalizable and efficient method to interrogate diverse native strains collection, remains lacking. Here, we explore a flexible and robust integrase-mediated multi-pronged activation approach to reliably perturb and globally trigger antibiotics production in actinobacteria. Across 54 actinobacterial strains, our approach yielded 124 distinct activator-strain combinations which consistently outperform wild type. Our approach expands accessible metabolite space by nearly two-fold and increases selected metabolite yields by up to >200-fold, enabling discovery of Gram-negative bioactivity in tetramic acid analogs. We envision these findings as a gateway towards a more streamlined, accelerated, and scalable strategy to unlock the full potential of Nature’s chemical repertoire.
ABSTRACTLow recycling rates have resulted in the alarming rate of accumulation of a widely used plastic material, polyethylene terephthalate (PET). With the build-up of plastics in our environment, there is an urgent need to source for more sustainable solutions to process them. Biological methods such as enzyme-catalyzed PET recycling or bioprocessing are seen as a potential solution to this problem. Actinobacteria, known for producing enzymes involved in the degradation of complex organic molecules, are of particular interest due to their potential to produce PET degrading enzymes. The highly thermostable enzyme, leaf-branch compost cutinase (LCC) found in Actinobacteria is one such example. This work expands on the discovery and characterization of new PET degrading enzymes fromMicrobispora, Nonomuraea, andMicromonosporagenus. Within this genus, we analyzed enzymes from the polyesterase-lipase-cutinase family, which have ∼60% similarity to LCC, where one of the enzymes was found to be capable of breaking down PET and BHET at 45-50 °C. Moreover, we were able to enhance the enzyme’s depolymerization rate through further engineering, resulting in a two-fold increase in activity.IMPORTANCEThe proliferation of PET plastic waste poses a significant threat to human and environmental health, making it an issue of increasing concern. In response to this challenge, scientists are investigating eco-friendly approaches, such as bioprocessing and microbial factories, to sustainably manage the growing quantity of plastic waste in our ecosystem. Despite the existence of enzymes capable of degrading PET, their scarcity in nature limits their applicability. The objective of this study is to enhance our understanding of this group of enzymes by identifying and characterizing novel ones that can facilitate the breakdown of PET waste. This data will expand the enzymatic repertoire and provide valuable insights into the prerequisites for successful PET degradation.
With growing concerns over the health impact of sugar, brazzein offers a viable alternative due to its sweetness, thermostability, and low risk profile. Here, we demonstrated the ability of protein language models to design new brazzein homologs with improved thermostability and potentially higher sweetness, resulting in new diverse optimized amino acid sequences that improve structural and functional features beyond what conventional methods could achieve. This innovative approach resulted in the identification of unexpected mutations, thereby generating new possibilities for protein engineering. To facilitate the characterization of the brazzein mutants, a simplified procedure was developed for expressing and analyzing related proteins. This process involved an efficient purification method using Lactococcus lactis (L. lactis), a generally recognized as safe (GRAS) bacterium, as well as taste receptor assays to evaluate sweetness. The study successfully demonstrated the potential of computational design in producing a more heat-resistant and potentially more palatable brazzein variant, V23.
Reducing sugar intake lowers the risk of obesity and associated metabolic disorders. Currently, this is achieved using artificial non-nutritive sweeteners, where their safety is widely debated and their contributions in various diseases is controversial. Emerging research suggests that these sweeteners may even increase the risk of cancer and cardiovascular problems, and some people experience gastrointestinal issues as a result of using them. A safer alternative to artificial sweeteners could be sweet-tasting proteins, such as brazzein, which do not appear to have any adverse health effects. In this study, protein language models were explored as a new method for protein design of brazzein. This innovative approach resulted in the identification of unexpected mutations, which opened up new possibilities for engineering thermostable and potentially sweeter versions of brazzein. To facilitate the characterization of the brazzein mutants, a simplified procedure was developed for expressing and analyzing related proteins. This process involved an efficient purification method using Lactococcus lactis (L. lactis), a generally recognized as safe (GRAS) bacterium, as well as taste receptor assays to evaluate sweetness. The study successfully demonstrated the potential of computational design in producing a more heat-resistant and potentially more palatable brazzein variant, V23.
High expression of OPCML in HER2-positive Breast Cancer samples is associated with better response to lapatinib
Low recycling rates coupled with increased production have resulted in an alarming rapid accumulation of a widely used plastic material, polyethylene terephthalate (PET). With the buildup of plastics in our environment, there is an urgent need for more sustainable solutions to process them. Biological methods such as enzyme-catalyzed PET recycling or bioprocessing are potential solutions to this problem. Actinobacteria, known for producing enzymes involved in the degradation of complex organic molecules, are of particular interest due to their potential to produce enzymes that may assimilate plastics. This work expands on the discovery and characterization of new PET-degrading enzymes from the Microbispora, Nonomuraea, and Micromonospora genera. Within the Micromonospora genus, we analyzed enzymes from the polyesterase-lipase-cutinase family with similar to 60% similarity to leaf-branch compost cutinase (LCC), in which one of the enzymes was found to be capable of breaking down PET and bis(2-hydroxyethyl) terephthalate (BHET) at 45 degrees C-50 degrees C. Moreover, we were able to enhance the enzyme's depolymerization rate through further engineering, resulting in a two-fold increase in activity. IMPORTANCE Mismanagement of PET plastic waste significantly threatens human and environmental health. Together with the relentless increase in plastic production, plastic pollution is an issue of rising concern. In response to this challenge, scientists are investigating eco-friendly approaches, such as bioprocessing and microbial factories, to sustainably manage the growing quantity of plastic waste in our ecosystem. Industrial applicability of enzymes capable of degrading PET is limited by numerous factors, including their scarcity in nature. The objective of this study is to enhance our understanding of this group of enzymes by identifying and characterizing novel enzymes that can facilitate the breakdown of PET waste. This data will expand the enzymatic repertoire and provide valuable insights into the prerequisites for successful PET degradation.
Vaccine hesitancy still threatens global efforts to end the COVID-19 pandemic caused by SARS-CoV-2 and its emerging variants.Social media-driven "conspiracy theories" cast doubts on vaccine safety for reproductive health, 1 including concerns that vaccine-induced SARS-CoV-2-neutralising antibodies (NAb) cross-react with human syncytin-1-a protein involved in gamete fertilisation and normal placental development 2resulting in infertility or pregnancy loss.Protein sequence similarities between syncytin-1 and SARS-CoV-2 spike protein S2 domain raise this possibility. 3oncerns over BNT162B2 mRNA persistence may prompt affected women to defer vaccination due to a perceived lack of reproductive toxicity or breastfeeding safety data, though COVID-19 vaccination is not contraindicated pre-conception, in pregnancy, or for breastfeeding. 4Given the enduring susceptibility of pregnant women to COVID-19 complications, and the risk when facing variants of higher transmission potential, such vaccine hesitancy is worrying. 5o address these issues, we performed an observational study of a convenience sample of female at-risk frontline workers receiving BNT162B2 to investigate post-vaccination presence of vaccine mRNA and antisyncytin-1 antibodies.We analysed 42 plasma and 30 breast milk samples from 15 consented female frontline staff at the National University Hospital, Singapore, collected before and after the first dose of BNT162B2, in this institutional review board-approved study (DSRB2012/00917).The participants included 5 breastfeeding mothers and 2 women inadvertently vaccinated in early pregnancy.The study was approved under National Healthcare Group Domain-Specific Review Board (Domain D) DSRB2012/00917 and the methods were conducted in accordance with the Declaration of Helsinki.All study participants provided written informed consent.Plasma was collected at Day 0 (pre-vaccination), 1-4 days and 4-7 weeks.Breast milk samples were collected daily for the first week, the timing designed to capture the presence of mRNA in plasma and breast milk (degraded within days), and peak neutralising activity (about 3 weeks post-vaccination) to answer our research questions. 6Blood and breast milk-derived total RNA
Reaction-diffusion (Turing) systems are fundamental to the formation of spatial patterns in nature and engineering. These systems are governed by a set of non-linear partial differential equations containing parameters that determine the rate of constituent diffusion and reaction. Critically, these parameters, such as diffusion coefficient, heavily influence the mode and type of the final pattern, and quantitative characterization and knowledge of these parameters can aid in bio-mimetic design or understanding of real-world systems. However, the use of numerical methods to infer these parameters can be difficult and computationally expensive. Typically, adjoint solvers may be used, but they are frequently unstable for very non-linear systems. Alternatively, massive amounts of iterative forward simulations are used to find the best match, but this is extremely effortful. Recently, physics-informed neural networks have been proposed as a means for data-driven discovery of partial differential equations, and have seen success in various applications. Thus, we investigate the use of physics-informed neural networks as a tool to infer key parameters in reaction─diffusion systems in the steady-state for scientific discovery or design. Our proof-of-concept results show that the method is able to infer parameters for different pattern modes and types with errors of less than 10%. In addition, the stochastic nature of this method can be exploited to provide multiple parameter alternatives to the desired pattern, highlighting the versatility of this method for bio-mimetic design. This work thus demonstrates the utility of physics-informed neural networks for inverse parameter inference of reaction-diffusion systems to enhance scientific discovery and design.
ABSTRACT Mouse lemurs ( Microcebus spp.) are an emerging primate model organism. However, little is known about their genetics or cellular and molecular biology. In the accompanying paper, we used large-scale single cell RNA-sequencing of 27 organs and tissues to identify over 750 molecular cell types, characterize their full transcriptomic profiles, and study evolution of primate cell types. Here we use the atlas to characterize mouse lemur genes, mutations, physiology, and disease. We uncover thousands of previously unidentified lemur genes and hundreds of thousands of new splice junctions that globally define lemur gene structures and reveal over 85,000 primate splice junctions missing in mice. We systematically explore the lemur immune system, comparing the global expression profiles of key immune genes in health and disease, and molecular mapping of immune cell development, trafficking, and their local and global activation to infection. We characterize primate/lemur-specific physiology and disease including molecular features of the immune program, of lemur adipocytes that exhibit dramatic seasonal rhythms, and of metastatic endometrial cancer that resembles the human malignancy. We identify and describe the expression patterns of over 400 primate genes missing in mice, many with similar expression patterns in human and lemur and some implicated in human disease. Finally, we provide an experimental framework for reverse genetic analysis by identifying naturally-occurring nonsense (null) mutations in three primate genes missing in mice and analyzing their transcriptional phenotypes. This work establishes mouse lemur as a tractable primate model organism for genetic and molecular analysis, and it prioritizes primate genes, splice junctions, physiology, and disease for future study.
Alzheimer's disease (AD) is a progressive neurodegenerative disease observed with aging that represents the most common form of dementia. To date, therapies targeting end-stage disease plaques, tangles, or inflammation have limited efficacy. Therefore, we set out to identify a potential earlier targetable phenotype. Utilizing a mouse model of AD and human fetal cells harboring mutant amyloid precursor protein, we show cell intrinsic neural precursor cell (NPC) dysfunction precedes widespread inflammation and amyloid plaque pathology, making it the earliest defect in the evolution of the disease. We demonstrate that reversing impaired NPC self-renewal via genetic reduction of USP16, a histone modifier and critical physiological antagonist of the Polycomb Repressor Complex 1, can prevent downstream cognitive defects and decrease astrogliosis in vivo. Reduction of USP16 led to decreased expression of senescence gene Cdkn2a and mitigated aberrant regulation of the Bone Morphogenetic Signaling (BMP) pathway, a previously unknown function of USP16. Thus, we reveal USP16 as a novel target in an AD model that can both ameliorate the NPC defect and rescue memory and learning through its regulation of both Cdkn2a and BMP signaling.