Microbial enzyme production and catalysis systems are crucial aspect of biotechnological research. However, building them from trustworthy published experimental data presents a major obstacle for both manual and automated techniques. Here, we introduce MEPAM (Microbial Enzyme Production and Catalytic Activity based on LLM), a question-answering system designed to accurately address inquiries related to enzyme production and catalytic reactions. Specifically, by training three machine learning models with >0.98 accuracy, we identified 11,068 high-quality, relevant articles from the Web of Science. Leveraging DeepSeek-V3 with zero-shot learning, we developed an ontology-driven knowledge representation that extracted 12,434 entities and 35,918 relations with 0.78 extraction accuracy and constructed a structured knowledge graph. Compared to few-shot learning and other machine learning methods, our framework achieved significantly higher extraction accuracy. Using this framework, we developed MEPAM based on retrieval-augmented generation and prompt engineering. Finally, using MEPAM, we extracted a comprehensive network involving the expression profiles, precise culture conditions, and substrate preferences for cellulase, demonstrating the strong utility of this tool. Compared with traditional LLMs, particularly GPT-4o, MEPAM exhibited superior performance, achieving significantly higher answer accuracy (0.86 vs. 0.52) and nearly eliminating hallucinations. MEPAM is available at http://180.76.108.212. This framework provides context-rich, verifiable insights, thus bridging predictive modeling with experimental validation to facilitate the exploration of microbial enzymatic systems.
Proteases can cleave peptide bonds of target substrate proteins. Their controlled proteolysis is vital for protein degradation, recycling, and physiological processes. Understanding the hydrolytic mechanisms of proteases is crucial, particularly for identifying their specific substrates and cleavage sites. Bioinformatics approaches can predict novel protease-substrate cleavage events with high accuracy using sequence and structural information. However, existing tools for cleavage site prediction face several limitations, including restricted accuracy due to limited data and cumbersome training processes that impede timely updates. To address these challenges, we developed MPCutter, which was created by fine-tuning a general-purpose protein sequence language model. This method combined the extensive knowledge of the general model with the targeted optimization of fine-tuning, providing a powerful tool for protease-substrate cleavage prediction. MPCutter offers optimized cleavage site prediction models with enhanced performance and broader coverage across proteases, encompassing four major protease families including 62 distinct proteases. Benchmarking experiments using independent test datasets demonstrated that MPCutter outperformed existing generic tools. In our case study and experiments, MPCutter precisely recognized the majority of cleavage sites and validated five caspase-3 cleavage sites crucial for cellular physiology. Notably, its application to the 10,260-protein human proteome and specific cancer pathways revealed potential new target substrates and provided insights into key biochemical behaviors of proteases. MPCutter is expected to serve as a powerful tool for high-throughput prediction of protease-specific substrates and to facilitate hypothesis-driven exploration of protease proteolytic events. The MPCutter code and associated data are freely available at https://github.com/2053798680wang/MPCutter.git.
Microorganism culturing is essential in microbiological research, with the selection of suitable culture media being critical for successful microbial growth. Traditionally, this selection has relied on empirical knowledge or trial and error, often resulting in inefficiency. In this study, we analysed nutrient compositions from the MediaDive database to construct a dataset of 2369 media types. Leveraging this dataset and microbial 16S rRNA sequences, we developed 45 binary classification models using the XGBoost algorithm. These models demonstrated strong predictive performance, achieving accuracies ranging from 76% to 99.3%, with the top-performing models for J386, J50 and J66 media reaching 99.3%, 98.9% and 98.8%, respectively. The models effectively predicted growth conditions for various human gut microbes, confirming their practical utility. This research improves the efficiency of microbial cultivation and highlights the potential of machine learning to optimise culture media selection and advance microbiological studies.
Beef and draft cattle have distinct rumen microbiota that can influence their metabolic processes and body composition. However, traditional metagenomic sequencing methods only provide broad surveys of the rumen microbial genomic contents. In this study, we utilized high-throughput single-cell genome sequencing to investigate these differences at the strain level. Following quality control and contig assembly, we obtained 97 bacterial genomes, 17 archaeal genomes, and 241 subspecies genomes from the rumen samples of Angus and Wuling cattle. Our analysis revealed a higher bacterial abundance in Angus rumen, characterized by an enrichment of the Succiniclasticum and Limivicinus genera. In contrast, the rumen of Wuling cattle exhibited a higher archaeal abundance. Additionally, we observed variations in the types and abundance of microbial-derived enzymes responsible for plant fiber degradation and volatile fatty acid (VFA) production between the two cattle breeds. The Angus rumen was found to harbor a higher diversity and abundance of cellulases and hemicellulases, particularly from the Ruminococcus unknown_0 genus. Furthermore, genera such as Succiniclasticum, Butyrivibrio, Limivicinus, UBA2868, and Prevotella were identified as key contributors to VFA production. Our findings suggest that the Angus rumen may have a stronger VFA production capacity due to the higher abundance of acidogenic genera. Interestingly, we also observed a greater abundance of Methanobrevibacter_A methanogens, which play a crucial role in energy flow in the rumen ecosystem, in Wuling cattle compared to Angus cattle. Our study highlights differences in the rumen microbiome of Angus and Wuling cattle. This difference could, at least partially, account for the variation in fat content that ultimately results in the superior meat quality of Angus cattle and the sustained muscle activity required by draft cattle. Overall, single-cell genome sequencing reveals distinct microbial composition and metabolic pathways between the two breeds, providing insights into their unique physiological and metabolic needs.
High soluble protein expression in heterologous hosts is crucial for various research and applications. Despite considerable research on the impact of codon usage on expression levels, the relationship between protein sequence and expression is often overlooked. In this study, a novel connection between protein expression and sequence is uncovered, leading to the development of SRAB (Strength of Relative Amino Acid Bias) based on AEI (Amino Acid Expression Index). The AEI served as an objective measure of this correlation, with higher AEI values enhancing soluble expression. Subsequently, the pre-trained protein model MP-TRANS (MindSpore Protein Transformer) is developed and fine-tuned using transfer learning techniques to create 88 prediction models (MPB-EXP) for predicting heterologous expression levels across 88 species. This approach achieved an average accuracy of 0.78, surpassing conventional machine learning methods. Additionally, a mutant generation model, MPB-MUT, is devised and utilized to enhance expression levels in specific hosts. Experimental validation demonstrated that the top 3 mutants of xylanase (previously not expressed in Escherichia coli) successfully achieved high-level soluble expression in E. coli. These findings highlight the efficacy of the developed model in predicting and optimizing gene expression based on protein sequences.
Deep learning models show promise in accelerating the design and optimization of antimicrobial peptides (AMPs), but current methods face challenges, such as low success rates, or large virtual library scales. In this study, we introduce DLFea4AMPGen, a bioactive peptide design strategy that leverages deep learning models to identify and extract key features associated with antimicrobial peptide activity. This approach enables the generation of peptide sequences with potential bioactivities. Using the SHapley Additive exPlanations (SHAP) method, we quantify the contribution of each amino acid in multifunctional peptides with potential antibacterial, antifungal, and antioxidant activities. Key feature fragments (KFFs) with the highest average contributions are extracted and classified into four subfamilies based on amino acid frequency. These high-frequency amino acids are systematically arranged to generate a plausible sequence subspace for candidate peptides, from which 16 representative sequences were selected for experimental validation. The results show that 75% (12/16) of the sequences exhibited at least two types of activity. Notably, D1 exhibits broad-spectrum antimicrobial activity, including efficacy against multidrug-resistant clinical pathogenic isolates both in vitro and in vivo. This proof-of-concept study underscores the potential of the DLFea4AMPGen platform for efficient design and screening of bioactive peptides, showcasing its value in AMP research.
Generative models have transformed protein design by enabling the generation of extensive datasets. However, accurate identification of biologically active sequences with specific functions within such data remains a significant challenge. In this study, we present a novel pipeline that integrates models for sequence generation, ranking, and selection to engineer proteins with enhanced properties. Our Omni-Directional Multipoint Mutagenesis (ODM) generation model was developed by refining a pre-trained protein BERT model to produce 100,000 mutant proteins. To evaluate the effects of mutations on protein activity, we utilized the lowest probability prediction across all masked positions as an indicator to rank the mutant sequences. Furthermore, we developed thermostability models to identify protease mutants with improved thermostability and utilized biological indicators to enhance lysozyme activity by introducing additional basic residues. Through two iterative design cycles, we observed that 62.5% of protease mutants exhibited enhanced thermostability, while 50% of lysozyme mutants displayed increased bacteriolytic activity.
The potential of seed endophytic microbes to enhance plant growth and resilience is well recognized, yet their role in alleviating cold stress in rice remains underexplored due to the complexity of these microbial communities. In this study, we investigated the diversity of seed endophytic microbes in two rice varieties, the coldsensitive CB9 and the cold-tolerant JG117. Our results revealed significant differences in the abundance of Microbacteriaceae, with JG117 exhibiting a higher abundance under both cold stress and room temperature conditions compared to CB9. Further analysis led to the identification of a specific cold-tolerant microbe, Microbacterium testaceum M15, in JG117 seeds. M15-inoculated CB9 plants showed enhanced growth and cold tolerance, with a germination rate increase from 40 % to 56.67 % at 14 degrees C and a survival rate under cold stress (4 degrees C) doubling from 22.67 % to 66.67 %. Additionally, M15 significantly boosted chlorophyll content by over 30 %, increased total protein by 16.31 %, reduced malondialdehyde (MDA) levels by 37.76 %, and increased catalase activity by 26.15 %. Overall, our study highlights the potential of beneficial endophytic microbes like M. testaceum M15 in improving cold tolerance in rice, which could have implications for sustainable agricultural practices and increased crop productivity in cold-prone regions.
Thermophilic endo-chitinases are essential for production of highly polymerized chitooligosaccharides, which are advantageous for plant immunity, animal nutrition and health. However, thermophilic endo-chitinases are scarce and the transformation from exo- to endo-activity of chitinases is still a challenging problem. In this study, to enhance the endo-activity of the thermophilic chitinase Chi304, we proposed two approaches for rational design based on comprehensive structural and evolutionary analyses. Four effective single-point mutants were identified among 28 designed mutations. The ratio of (GlcNAc)3 to (GlcNAc)2 quantity (DP3/2) in the hydrolysates of the four single-point mutants undertaking colloidal chitin degradation were 1.89, 1.65, 1.24, and 1.38 times that of Chi304, respectively. When combining to double-point mutants, the DP3/2 proportions produced by F79A/W140R, F79A/M264L, F79A/W272R, and M264L/W272R were 2.06, 1.67, 1.82, and 1.86 times that of Chi304 and all four double-point mutants exhibited enhanced endo-activity. When applied to produce chitooligosaccharides (DP ≥ 3), F79A/W140R accumulated the most (GlcNAc)4, while M264L/W272R was the best to produce (GlcNAc)3, which was 2.28 times that of Chi304. The two mutants had exposed shallower substrate-binding pockets and stronger binding abilities to shape the substrate. Overall, this research offers a practical approach to altering the cutting pattern of a chitinase to generate functional chitooligosaccharides.
Polyethylene terephthalate (PET) biodegradation is hindered by the intermediates bis (2-hydroxyethyl) terephthalate (BHET) and mono (2-hydroxyethyl) terephthalate (MHET). BMHETase, a thermophilic hydrolase identified from the UniParc database, exhibits degradation activity towards both BHET and MHET. BMHETase showed higher activity on BHET than LCCICCG and FASTPETase at temperatures ranging from 50 to 70℃. To enhance its activity in degrading MHET, BMHETase was engineered to mimic Ideonella sakaiensis MHETase. The resulting 6-point mutant's activities on MHET and BHET were 8 and 2 times those of the WT, with both optimal temperatures increased by 5℃. This enhancement may be attributed to the BMHETase6M's intensified binding ability with MHET and enlarged binding pocket. When combined with LCCICCG, BMHETase6M achieved complete degradation of MHET in PET films to terephthalic acid, indicating broad application potential. These findings suggest that BMHETase6M holds promise as a candidate for enhancing PET biodegradation efficiency and plastic waste management.
There are binding sites and hydrolytic active sites in chitinase, and the binding of key amino acids to substrate chitin can appropriately regulate the hydrolytic activity of the enzyme. The thermophilic chitinase Chi304 was as experimental material. Swiss-Model and analysis of its advanced structure revealed the presence of two tryptophans(W140 and W272) near its substrate binding pocket. And these two tryptophans were site-directed mutated to alanine.High performance liquid chromatography was used to detect the hydrolysis products of the enzyme. The ratio of product triacetyl chitosaccharide to diacetyl chitosaccharide(DP 3 /DP 2 ) was used to evaluate the effect of the mutants. The mutants(W140A, W272A and W140/272A) hydrolyzed colloidal chitin, and the proportions of product(DP 3 /DP 2 ) were increased by 23.3%, 45.7% and 80.0% compared with that of wild type, respectively. The results showed that W140 and W272 were the key amino acids affecting the binding of enzyme to substrate, and the mutation of alanine to Chi304 increased the endogenous activity and decreased the exogenous activity.
Microbial bioremediation of heavy metal-polluted soil is a promising technique for reducing heavy metal accumulation in crops. In a previous study, we isolated Bacillus vietnamensis strain 151-6 with a high cadmium (Cd) accumulation ability and low Cd resistance. However, the key gene responsible for the Cd absorption and bioremediation potential of this strain remains unclear. In this study, genes related to Cd absorption in B. vietnamensis 151-6 were overexpressed. A thiol-disulfide oxidoreductase gene (orf4108) and a cytochrome C biogenesis protein gene (orf4109) were found to play major roles in Cd absorption. In addition, the plant growth-promoting (PGP) traits of the strain were detected, which enabled phosphorus and potassium solubilization and indole-3-acetic acid (IAA) production. Bacillus vietnamensis 151-6 was used for the bioremediation of Cd-polluted paddy soil, and its effects on growth and Cd accumulation in rice were explored. The strain increased the panicle number (114.82%) and decreased the Cd content in rice rachises (23.87%) and grains (52.05%) under Cd stress, compared with non-inoculated rice in pot experiments. For field trials, compared with the non-inoculated control, the Cd content of grains inoculated with B. vietnamensis 151-6 was effectively decreased in two cultivars (low Cd-accumulating cultivar: 24.77%; high Cd-accumulating cultivar: 48.85%) of late rice. Bacillus vietnamensis 151-6 encoded key genes that confer the ability to bind Cd and reduce Cd stress in rice. Thus, B. vietnamensis 151-6 exhibits great application potential for Cd bioremediation.
The demand for high efficiency glycoside hydrolases (GHs) is on the rise due to their various industrial applications. However, improving the catalytic efficiency of an enzyme remains a challenge. This investigation showcases the capability of a deep neural network and method for enhancing the catalytic efficiency (MECE) platform to predict mutations that improve catalytic activity in GHs. The MECE platform includes DeepGH, a deep learning model that is able to identify GH families and functional residues. This model was developed utilizing 119 GH family protein sequences obtained from the Carbohydrate-Active enZYmes (CAZy) database. After undergoing ten-fold cross-validation, the DeepGH models exhibited a predictive accuracy of 96.73%. The utilization of gradient-weighted class activation mapping (Grad-CAM) was used to aid us in comprehending the classification features, which in turn facilitated the creation of enzyme mutants. As a result, the MECE platform was validated with the development of CHIS1754-MUT7, a mutant that boasts seven amino acid substitutions. The kcat/Km of CHIS1754-MUT7 was found to be 23.53 times greater than that of the wild type CHIS1754. Due to its high computational efficiency and low experimental cost, this method offers significant advantages and presents a novel approach for the intelligent design of enzyme catalytic efficiency. As a result, it holds great promise for a wide range of applications.
Rice, which feeds more than half of the world's population, confronts significant challenges due to environmental and climatic changes. Abiotic stressors such as extreme temperatures, drought, heavy metals, organic pollutants, and salinity disrupt its cellular balance, impair photosynthetic efficiency, and degrade grain quality. Beneficial microorganisms from rice and soil microbiomes have emerged as crucial in enhancing rice's tolerance to these stresses. This review delves into the multifaceted impacts of these abiotic stressors on rice growth, exploring the origins of the interacting microorganisms and the intricate dynamics between rice-associated and soil microbiomes. We highlight their synergistic roles in mitigating rice's abiotic stresses and outline rice's strategies for recruiting these microorganisms under various environmental conditions, including the development of techniques to maximize their benefits. Through an in-depth analysis, we shed light on the multifarious mechanisms through which microorganisms fortify rice resilience, such as modulation of antioxidant enzymes, enhanced nutrient uptake, plant hormone adjustments, exopolysaccharide secretion, and strategic gene expression regulation, emphasizing the objective of leveraging microorganisms to boost rice's stress tolerance. The review also recognizes the growing prominence of microbial inoculants in modern rice cultivation for their eco-friendliness and sustainability. We discuss ongoing efforts to optimize these inoculants, providing insights into the rigorous processes involved in their formulation and strategic deployment. In conclusion, this review emphasizes the importance of microbial interventions in bolstering rice agriculture and ensuring its resilience in the face of rising environmental challenges.
The advanced language models have enabled us to recognize protein-protein interactions (PPIs) and interaction sites using protein sequences or structures. Here, we trained the MindSpore ProteinBERT (MP-BERT) model, a Bidirectional Encoder Representation from Transformers, using protein pairs as inputs, making it suitable for identifying PPIs and their respective interaction sites. The pretrained model (MP-BERT) was fine-tuned as MPB-PPI (MP-BERT on PPI) and demonstrated its superiority over the state-of-the-art models on diverse benchmark datasets for predicting PPIs. Moreover, the model's capability to recognize PPIs among various organisms was evaluated on multiple organisms. An amalgamated organism model was designed, exhibiting a high level of generalization across the majority of organisms and attaining an accuracy of 92.65%. The model was also customized to predict interaction site propensity by fine-tuning it with PPI site data as MPB-PPISP. Our method facilitates the prediction of both PPIs and their interaction sites, thereby illustrating the potency of transfer learning in dealing with the protein pair task.
Polyethylene terephthalate (PET)-degrading enzymes represent a promising solution to the plastic pollution. However, PET-degrading enzymes, even thermophilic PETase, can effectively degrade low-crystallinity (similar to 8%) PETs, but exhibit weak depolymerization of more common, high-crystallinity (30-50%) PETs. Here, based on the thermophilic PETase, LCCICCG, we proposed two strategies for rational redesign of LCCICCG using the machine learning tool, Preoptem, combined with evolutionary analysis. Six single-point mutants (S32L, D18T, S98R, T157P, E173Q, N213P) were obtained that exhibit higher catalytic efficiency towards PET powder than wildtype LCCICCG at 75 degrees C. Additionally, the optimal temperature for degrading 39.07% crystalline PET increased from 65 degrees C in the wild-type LCCICCG to between 75 and 80 degrees C in the LCCICCG_I6M mutant that carries all six single-point mutations. Especially, the LCCICCG_I6M mutant has a significantly higher degradation effect on some commonly used bottle-grade plastic powders at 75-80 degrees C than that of wild type. The enzymatic digestion of ground 31.30% crystalline PET water bottles by LCCICCG_I6M yielded 31.91 +/- 0.99 mM soluble products in 24 h, which was 3.64 times that of LCCICCG (8.77 +/- 1.52 mM). Overall, this study provides a feasible route for engineering thermostable enzymes that can degrade high-crystallinity PET plastic.
Chitin is abundant in nature and its degradation products are highly valuable for numerous applications. Thermophilic chitinases are increasingly appreciated for their capacity to biodegrade chitin at high temperatures and prolonged enzyme stability. Here, using deep learning approaches, we developed a prediction tool, Preoptem, to screen thermophilic proteins. A novel thermophilic chitinase, Chi304, was mined directly from the marine metagenome. Chi304 showed maximum activity at 85 ℃, its Tm reached 89.65 ± 0.22℃, and exhibited excellent thermal stability at 80 and 90 °C. Chi304 had both endo- and exo-chitinase activities, and the (GlcNAc)2 was the main hydrolysis product of chitin-related substrates. The product yields of colloidal chitin degradation reached 97% within 80 min, and 20% over 4 days of reaction with crude chitin powder. This study thus provides a method to mine the novel thermophilic chitinase for efficient chitin biodegradation.
Background Advances in DNA sequencing technologies have transformed our capacity to perform life science research, decipher the dynamics of complex soil microbial communities and exploit them for plant disease management. However, soil is a complex conglomerate, which makes functional metagenomics studies very challenging. Results Metagenomes were assembled by long-read (PacBio, PB), short-read (Illumina, IL), and mixture of PB and IL (PI) sequencing of soil DNA samples were compared. Ortholog analyses and functional annotation revealed that the PI approach significantly increased the contig length of the metagenomic sequences compared to IL and enlarged the gene pool compared to PB. The PI approach also offered comparable or higher species abundance than either PB or IL alone, and showed significant advantages for studying natural product biosynthetic genes in the soil microbiomes. Conclusion Our results provide an effective strategy for combining long and short-read DNA sequencing data to explore and distill the maximum information out of soil metagenomics.
The expression of proteins in Escherichia coli is often essential for their characterization, modification, and subsequent application. Gene sequence is the major factor contributing expression. In this study, we used the expression data from 6438 heterologous proteins under the same expression condition in E. coli to construct a deep learning classifier for screening high- and low-expression proteins. In conjunction with conserved residue analysis to minimize functional disruption, a mutation predictor for enhanced protein expression (MPEPE) was proposed to identify mutations conducive to protein expression. MPEPE identified mutation sites in laccase 13B22 and the glucose dehydrogenase FAD-AtGDH, that significantly increased both expression levels and activity of these proteins. Additionally, a significant correlation of 0.46 between the predicted high level expression propensity with the constructed models and the protein abundance of endogenous genes in E. coli was also been detected. Therefore, the study provides foundational insights into the relationship between specific amino acid usage, codon usage, and protein expression, and is essential for research and industrial applications.
漆酶(EC 1.10.3.2)是一种氧化还原酶,在有毒和致癌化合物的氧化降解方面具有应用价值.通过序列分析,从UniParc数据库中筛选到耐热的漆酶基因ba4,其全长1860 bp,编码620个氨基酸.通过最适反应温度回归预测模型(PMT)预测出BA4是耐热的漆酶,并在NCBI蛋白数据库中进行比对分析,其与来源于Klebsiella michiganensis的铜抗性系统多铜氧化酶(STW26195.1)相似性为58.75%,证明漆酶ba4是Copper_res_A超家族新的漆酶基因.将其全序列合成并在大肠杆菌BL21(DE3)中异源表达并纯化,性质测定结果表明该酶在温度45-65℃之间均有较高的酶活,最适温度为50℃,最适pH为5.5.以ABTS为底物测定米氏常数(Km)(2144.5±358.5)μmol/L,kcat为(44.06±3.14)min-1,最大反应速率(Vmax)为623.2μmol/(min·g),kcat/Km为(0.0209±0.002)L/(μmol·min).漆酶BA4(60-70 U/L)在50℃条件下与玉米赤霉烯酮(0.1 mg/mL)反应2 h,降解率达到了90% 以上;漆酶BA4(70-80 U/L)在40℃和50℃条件下与棉酚反应(1 mg/mL)1 h,降解率均为30%.漆酶BA4良好的酶学性质以及对玉米赤霉烯酮和棉酚有效降解为酶的应用奠定了良好的基础.