Single-cell perturbation sequencing technologies (e.g., Perturb-seq, CROP-seq), which integrate CRISPR-based gene editing with single-cell transcriptome profiling, have revolutionized the analysis of transcriptomic changes induced by genetic perturbations at single-cell resolution. These technologies serve as a powerful tool for identifying key genes that inhibit tumor growth or reverse cancer cell phenotypes. However, they face two major challenges: data explosion with high experimental costs, and data complexity characterized by high dimensionality, noise, sparsity, and heterogeneity. To address these challenges, we developed the single-cell Rank-based Genetic Perturbation predictor (scRGP), the first deep learning framework leveraging gene expression rank-order information for this task. scRGP demonstrates superior performance in terms of robustness, cross-cell-line perturbation prediction, and high-throughput screening. Specifically, scRGP achieves an approximately 10-16 percentage points improvement in Pearson correlation coefficient (PCC) over state-of-the-art methods (e.g., GEARS and scFoundation) for single- and double-gene perturbation predictions, while also extending prediction capability to triple-gene perturbations. Furthermore, it outperforms these methods by approximately 5-9 percentage points in cross-cell-line predictions. These advancements promise to shift the paradigm of single-cell perturbation studies from experiment-driven to computation-driven approaches, providing new support for functional genomics and precision medicine. ### Competing Interest Statement The authors have declared no competing interest. National Natural Science Foundation of China, ZX2200521 Anti-tumor New Drug Rapid Translation Public Service Platform of Jiangsu Province, BM2023002
Subcellular localization of microRNAs (miRNAs) is an important reflection of their biological functions. Considering the spatio-temporal specificity of miRNA subcellular localization, experimental detection techniques are expensive and time-consuming, which strongly motivates an efficient and economical computational method to predict miRNA subcellular localization. In this paper, we describe a computational framework, MiRLoc, to predict the subcellular localization of miRNAs. In contrast to existing methods, MiRLoc uses the functional similarity between miRNAs instead of sequence features and incorporates information about the subcellular localization of the corresponding target mRNAs. The results show that miRNA functional similarity data can be effectively used to predict miRNA subcellular localization, and that inclusion of subcellular localization information of target mRNAs greatly improves prediction performance.
Small molecular networks within complex pathways are defined as subpathways. The identification of patient-specific subpathways can reveal the etiology of cancer and guide the development of personalized therapeutic strategies. The dysfunction of subpathways has been associated with the occurrence and development of cancer. Here, we propose a strategy to identify aberrant subpathways at the individual level by calculating the edge score and using the Gene Set Enrichment Analysis (GSEA) method. This provides a novel approach to subpathway analysis. We applied this method to the expression data of a lung adenocarcinoma (LUAD) dataset from The Cancer Genome Atlas (TCGA) database. We validated the effectiveness of this method in identifying LUAD-relevant subpathways and demonstrated its reliability using an independent Gene Expression Omnibus dataset (GEO). Additionally, survival analysis was applied to illustrate the clinical application value of the genes and edges in subpathways that were associated with the prognosis of patients and cancer immunity, which could be potential biomarkers. With these analyses, we show that our method could help uncover subpathways underlying lung adenocarcinoma.
Long non-coding RNA (lncRNA)–microRNA (miRNA) interactions are quickly emerging as important mechanisms underlying the functions of non-coding RNAs. Accordingly, predicting lncRNA–miRNA interactions provides an important basis for understanding the mechanisms of action of ncRNAs. However, the accuracy of the established prediction methods is still limited. In this study, we used structural consistency to measure the predictability of interactive links based on a bilayer network by integrating information for known lncRNA–miRNA interactions, an lncRNA similarity network, and an miRNA similarity network. In particular, by using the structural perturbation method, we proposed a framework called SPMLMI to predict potential lncRNA–miRNA interactions based on the bilayer network. We found that the structural consistency of the bilayer network was higher than that of any single network, supporting the utility of bilayer network construction for the prediction of lncRNA–miRNA interactions. Applying SPMLMI to three real datasets, we obtained areas under the curves of 0.9512 ± 0.0034, 0.8767 ± 0.0033, and 0.8653 ± 0.0021 based on 5-fold cross-validation, suggesting good model performance. In addition, the generalizability of SPMLMI was better than that of the previously established methods. Case studies of two lncRNAs (i.e., SNHG14 and MALAT1) further demonstrated the feasibility and effectiveness of the method. Therefore, SPMLMI is a feasible approach to identify novel lncRNA–miRNA interactions underlying complex biological processes.
Abstract Background Currently, large-scale gene expression profiling has been successfully applied to the discovery of functional connections among diseases, genetic perturbation, and drug action. To address the cost of an ever-expanding gene expression profile, a new, low-cost, high-throughput reduced representation expression profiling method called L1000 was proposed, with which one million profiles were produced. Although a set of ~ 1000 carefully chosen landmark genes that can capture ~ 80% of information from the whole genome has been identified for use in L1000, the robustness of using these landmark genes to infer target genes is not satisfactory. Therefore, more efficient computational methods are still needed to deep mine the influential genes in the genome. Results Here, we propose a computational framework based on deep learning to mine a subset of genes that can cover more genomic information. Specifically, an AutoEncoder framework is first constructed to learn the non-linear relationship between genes, and then DeepLIFT is applied to calculate gene importance scores. Using this data-driven approach, we have re-obtained a landmark gene set. The result shows that our landmark genes can predict target genes more accurately and robustly than that of L1000 based on two metrics [mean absolute error (MAE) and Pearson correlation coefficient (PCC)]. This reveals that the landmark genes detected by our method contain more genomic information. Conclusions We believe that our proposed framework is very suitable for the analysis of biological big data to reveal the mysteries of life. Furthermore, the landmark genes inferred from this study can be used for the explosive amplification of gene expression profiles to facilitate research into functional connections.
In genome-wide association studies, detecting high-order epistasis is important for analyzing the occurrence of complex human diseases and explaining missing heritability. However, there are various challenges in the actual high-order epistasis detection process due to the large amount of data, “small sample size problem”, diversity of disease models, etc. This paper proposes a multi-objective genetic algorithm (EpiMOGA) for single nucleotide polymorphism (SNP) epistasis detection. The K2 score based on the Bayesian network criterion and the Gini index of the diversity of the binary classification problem were used to guide the search process of the genetic algorithm. Experiments were performed on 26 simulated datasets of different models and a real Alzheimer’s disease dataset. The results indicated that EpiMOGA was obviously superior to other related and competitive methods in both detection efficiency and accuracy, especially for small-sample-size datasets, and the performance of EpiMOGA remained stable across datasets of different disease models. At the same time, a number of SNP loci and 2-order epistasis associated with Alzheimer’s disease were identified by the EpiMOGA method, indicating that this method is capable of identifying high-order epistasis from genome-wide data and can be applied in the study of complex diseases.
Ditch-buried straw return (DB-SR) is a novel soil tillage and fertility-building practice, which may allow for the formation of a straw layer structure in rice-wheat rotation systems. However, it is presently unknown how the straw layer affects soil physicochemical and microbial processes in the interface between straw layer and bulk soils. Thus, we conducted a field experiment from 2008 to 2014 to determine whether there were changes in soil structure, temperature, available nitrogen, and microbial properties in the soil layer close to the straw layers (+/- 5 cm). Three treatments were included: control with no-till and straw removal (CK), ditch-buried straw return to a depth of 20 cm (DB-SR-20), and ditch-buried straw return to a depth of 40 cm (DB-SR-40). Our results showed that soil macroaggregate formation was increased with the duration of straw return in the soils 5 cm below the straw layers under DB-SR-20 and above the straw layers under DB-SR-40. The straw layers could significantly increase soil temperature below the straw layers, but a larger increase was observed for DB-SR-40 than DB-SR-20. We also found that DB-SR-20 largely increased the NH4+ concentration in both the soil layers above and below the straw layers, but DB-SR-40 only increased NO3- in the soils above the straw layers. In addition, the catabolic profile of soil microbial communities was largely affected by the straw layers, but the effects decreased with the duration of the treatment. These findings suggest that the straw layer largely alters the spatio-temporal processes of physicochemical and microbial properties in the interface soils after long-term implementation of ditch-buried straw return. Some of the positive alterations may cascade to facilitate crop growth and enhance the productivity of rice-wheat rotation systems.
Identifying perturbed pathways at an individual level is important to discover the causes of cancer and develop individualized custom therapeutic strategies. Though prognostic gene lists have had success in prognosis prediction, using single genes that are related to the relevant system or specific network cannot fully reveal the process of tumorigenesis. We hypothesize that in individual samples, the disruption of transcription homeostasis can influence the occurrence, development, and metastasis of tumors and has implications for patient survival outcomes. Here, we introduced the individual-level pathway score, which can measure the correlation perturbation of the pathways in a single sample well. We applied this method to the expression data of 16 different cancer types from The Cancer Genome Atlas (TCGA) database. Our results indicate that different cancer types as well as their tumor-adjacent tissues can be clearly distinguished by the individual-level pathway score. Additionally, we found that there was strong heterogeneity among different cancer types and the percentage of perturbed pathways as well as the perturbation proportions of tumor samples in each pathway were significantly different. Finally, the prognosis-related pathways of different cancer types were obtained by survival analysis. We demonstrated that the individual-level pathway score (iPS) is capable of classifying cancer types and identifying some key prognosis-related pathways.
为了深入了解和探索lincRNA的调控机制,建立了lincRNA高效识别模型,有助于为后续研究提供数据源.依据最小自由能(minimum free energy,MFE)和信噪比(signal-noise ratio,SNR)等特征,并通过特征贡献度大小剔除冗余特征,构建随机森林(random forest,RF)分类模型,有效地识别lincRNAs.经检验,模型的灵敏度、特异性和精确度分别达到94.1%、93.2%和93.7%,高于现有PhyloCSF、LncRNA-ID和CPC方法的各项识别指标.模型在识别过程中表现出较好的鲁棒性,可准确识别lincRNA.
LncRNAs are regulatory noncoding RNAs that play crucial roles in many biological processes. The dysregulation of lncRNA is thought to be involved in many complex diseases; lncRNAs are often the targets of miRNAs in the indirect regulation of gene expression. Numerous studies have indicated that miRNA-lncRNA interactions are closely related to the occurrence and development of cancers. Thus, it is important to develop an effective method for the identification of cancer-related miRNA-lncRNA interactions. In this study, we compiled 155653 experimentally validated and predicted miRNA-lncRNA associations, which we defined as basic interactions. We next constructed an individual-specific miRNA-lncRNA network (ISMLN) for each cancer sample and a basic miRNA-lncRNA network (BMLN) for each type of cancer by examining the expression profiles of miRNAs and lncRNAs in the TCGA (The Cancer Genome Atlas) database. We then selected potential miRNA-lncRNA biomarkers based on the BLMN. Using this method, we identified cancer-related miRNA-lncRNA biomarkers and modules specific to a certain cancer. This method of profiling will contribute to the diagnosis and treatment of cancers at the level of gene regulatory networks.
[Objective] We identified 16S rRNA genes in genomes of prokaryotes.[Methods] We constructed a 3-layer filtering model based on the three features of GC bases content of the gene sequences,3-base periodicity and Markov chain to recognize the 16S rRNA genes from prokaryotic genomes.[Results] The specificity,sensitivity and Matthews correlation coefficients of the model were 99.58%,91.60% and 91.49%,respectively.[Conclusion] The results showed that the 16S rRNA genes can be identified efficiently and accurately by using our model.
Ditch-buried straw return(DB-SR)is a novel soil tillage practice which forms a special"straw layer"structure. In order to elabo-rate the role of straw layer on soil nitrogen distribution and microbial community, a field experiment was conducted under DB-SR with three burial depths(20 cm:DB-SR-20;40 cm:DB-SR-40 and CK). NH+4-N, NO-3-N, microbial biomass carbon(MBC)and community level physiological profile(CLPP)were determined in the straw layer and its interface soil layers under different treatments. Results showed that the structure of"straw layer"had positive effect on nitrogen retentions. In DB-SR-20, the straw layer increased MBC at the interface of straw layer, but no significant effect was found for the functional diversity. In DB-SR-40, MBC decreased at first and then increased at the interface of rice straw layer, but the pattern was reversed for wheat straw layer. Microbial diversity index(H)increased for rice straw but decreased at first and then increased for wheat straw layer over time. CLPP suggested that the microorganisms of straw layer could utilize various carbon sources, and their metabolic activity was higher than CK. In DB-SR-20 for wheat straws, NH+4-N and NO-3-N were significantly related to variation of microbial community in the straw-soil interface, but MBC was correlated to the microbial community in the straw layers. In DB-SR-40, NH+4-N、NO-3-N and MBC were significantly correlated to the variation of microbial communities in the straw layers. This study sug-gested that the"straw layer"could effectively increase soil N retention, and increase the functional diversity of soil microbial community.
As a novel soil tillage practice, ditch-buried straw return (DB-SR) has exhibited positive effects on soil carbon sequestration, nitrogen retention and rice yield in previous studies. However, little is known about how long-term DB-SR affects soil hydrothermal and microbial processes. Our objective is to test whether DB-SR will alter the soil water potential, temperature and microbial community in a wheat field following rice cultivation. In this study, we found significant alterations in soil water potential, temperature, and microbial communities driven by DB-SR. On average, soil water potential was significantly reduced by 37.33% and 17.56% under DB-SR to a depth of 20 cm (DB-SR-20) and 40 cm (DB-SR-40), respectively. DB-SR-20 increased soil mean daily temperature and daily range of temperature more than DB-SR-40, possibly caused by decreased water content, especially at soil depths of 10 and 15 cm. Both DB-SR-20 and DB-SR-40 led to distinct shifts in soil bacterial and fungal community composition. DB-SR-20 significantly increased the activities of peroxidase, cellobiohydrolase, urease, and acid phosphatase by 3.5%, 75.0%, 81.4% and 41.7%, respectively, but had no effects on beta-D-glucosidase activity. DB-SR-40, in contrast, significantly increased the activities of peroxidase and cellobiohydrolase by 2.4% and 36.0%, respectively, but showed no effects on urease and acid phosphatase. It did, however, reduce p-o-glucosidase activity by 15.0%. Overall functional diversity was increased by 29.9% under DBSR-20 but was not affected by DB-SR-40. Our results suggest that these improvements in soil ecological processes driven by DB-SR will promote wheat yield in a rice-wheat rotation system. (C) 2016 Elsevier B.V. All rights reserved.
Arbuscular mycorrhizal fungi ( AMF) are a group of ecologically important soil microbes and show wide geographic distribution across the globe. AMF form obligate symbiosis with roots of -80% land plants. In the symbiosis, host plants provide carbon for AMF in return for several benefits, i.e., promoting nutrient uptake, tolerating drought and salt stress, resisting pathogens and herbivores, etc. AMF also can redistribute resources ( i.e., C, N and P) between plants and alter their competitive interactions, and thus drive plant population dynamics and community processes. AMF diversity is one of the most important components in soil diversity. In the past decades, AMF are found in almost all terrestrial habitats, including grassland, forest, desert, wetland, alpine meadow, polar region and mangrove, etc. This suggests that AMF have high species diversity. Although AMF diversity has a relatively long research history, most studies only tried to investigate species composition in AMF communities, little is known about the functioning of AMF diversity. In this mini⁃ review, we summarized the new advances in the AMF diversity field, including ecological functioning, determinants and assembling rules. AMF diversity has important ecological functioning. Here, we discussed three aspects: the effects on plant system diversity, stability and productivity. First, several studies reported that AMF diversity is an important determinant for plant diversity. This might be caused by mycorrhizal dependence of subordinate plants. Some studies found that host plants have some preferentially selection towards AMF. Thus, with increasing AMF diversity, subordinate plants will have a higher probability to meet their best AMF partner. Another possibility is that negative plant⁃mycorrhiza feedbacks might generate positive AMF diversity⁃plant diversity patterns. This might be caused by host selection towards specific AMF communities. Distinctive AMF communities will make host plants occupy different niche for soil resources. Secondly, AMF diversity could stabilize plant community. Two possibilities can be used to explain this pattern. One is functioning redundancy for several
The arbuscular mycorrhiza (AM) is among the most ubiquitous symbiosis in the world. A meta-analysis of 759 articles (1978–2012) was conducted to test whether ecologically important host plant traits (N-fixation and C-fixation pathway) affect the response of the plant to mycorrhizal colonization. We found that the effect of N-fixation on mycorrhizal growth response (MGR) depended on whether the plant was woody or a forb. N-fixing forbs had a higher MGR than non-N-fixing forbs, but the reverse was true for woody plants. Moreover, C4-grasses had significantly higher MGR than C3-grasses, but no significant difference was found between C3 and C4 forbs, or between C3 and C4 woody species. Overall, woody species had higher MGR than any other functional group. These results demonstrate that MGR does depend on host functional characteristics, but neither N-fixation capacity nor C-fixation pathway are apparently fundamental controllers of MGR. Instead, it would appear possible that these traits influence MGR only insofar as they influence more fundamental functions such as P demand and P supply.
通过5.5a的大田定位试验,将上季秸秆全量沟埋还田,设置秸秆沟埋还田深度为20、40 cm以及免耕秸秆不还田(对照)3个处理.研究秸秆沟埋还田对麦田土壤水势、温度的影响以及长期秸秆沟埋还田方式下,沟埋还田20 cm处理各埋草沟土壤容重、总孔隙度的变化.结果表明:秸秆沟埋还田具有降低土壤容重,增加土壤总孔隙度的作用,随着还田时间的增加,这种作用逐渐降低.当降雨量较大(26.6 mm)时,沟埋还田各处理水势值在短时间内上升的较快,而对照则相对较慢;当降雨量较小(10 mm)时,沟埋还田40 cm处理水势值上升速度大于沟埋还田20 cm,对照处理最慢;降雨过后的12d内,沟埋还田各处理水势值下降速度较对照更快;连续40d各处理土壤水势日均值大小为对照>沟埋还田40 cm>沟埋还田20 cm.土壤0-15 cm温度日较差大小为沟埋还田20 cm>对照>沟埋还田40 cm,土壤20 cm处日较差对照最大;沟埋还田20 cm处理0-15 cm以及沟埋还田40 cm处理0-20 cm土壤日均温高于对照,沟埋还田20 cm处理20 cm处土壤日均温与对照较为接近.在沿江稻麦轮作地区,秸秆集中沟埋还田具有较好的改善土壤物理性质的作用.
In microbial ecology, the "everything is everywhere" hypothesis has long been controversial. In the present study, we performed data-mining for 18S rDNA sequences of glomeromycotan fungi in order to test this hypothesis. 18S rDNA sequences targeted using AM1–NS31 fragments were retrieved from GenBank, with a total of 1768 sequences collected from 34 sites worldwide. In total, 229, 330 and 518 operational taxonomic units (OTUs) were defined based on 97, 98 and 99 % similarity, respectively. The 97 % OTUs showed a limited geographical range of glomeromycotan fungi. Among the OTUs, 58.1 % were endemic, and 17.9 % and 9.2 % were found in two and three sites, respectively. The most widespread OTU was shared by 17 sites. Phylogenetic structure analysis demonstrated that most local communities (26 of 34) were clustered. OTUs with larger host breadth had wider geographic ranges. A significant distance–decay relationship was revealed that was independent of habitat. Cluster analysis showed that fungal composition was not related to habitat, while Fast UniFrac analysis indicated that the distribution of Glomeromycota was affected by temperature. Taken together, these results suggest that glomeromycotan fungi were not randomly distributed under natural conditions; rather, they were affected by host plants, dispersal ability and temperature. Thus, the distribution of glomeromycotan fungi argues against the hypothesis that "everything is everywhere."
Ditch-buried straw return (DBSR) is a novel farming system that not only efficiently eliminates the need to burn straw, but also shows positive effects on soil carbon sequestration and crop yields. Implementation of DBSR, however, may penetrate the tillage pan, increasing the risk of N leaching losses. We therefore determined whether N retention could be increased by DBSR in order to reduce the risk of N loss to the environment. A four-year field experiment and a complementary greenhouse experiment were conducted to test the effects of DBSR on N retention in a rice-wheat rotation system. We found that DBSR altered the spatial distribution of fertilizer N. N content was significantly increased above but reduced below the straw layer in the field experiment. The greenhouse experiment further confirmed the N retention effects by the straw layer. In theory, a maximum of 9.09mg urea-N could be adsorbed by one gram dry wheat straw. Our results suggest that DBSR has the potential to increase N retention in the soil, thus increasing crop uptake and minimizing leaching N loss in the rice-wheat rotation system.
为弄清丛枝菌根(arbuscular mycorrhiza,AM)真菌群落随宿主植物演化的变异规律,通过对MaarjAM数据库进行数据挖掘,根据每个分子虚拟种(virtual taxa,VT)包含的DNA序列不少于5条的标准,筛选出188种菌根植物.通过分析植物与其根内AM真菌的关系发现:AM真菌的物种丰富度随着寄主植物的分化而增加;在不同的植物系统类群中,AM真菌的物种丰富度显著不同;在起源时间较晚的被子植物和裸子植物中,AM真菌的物种丰富度显著高于起源较早的苔类、角苔类和蕨类植物类群,而与寄生植物共生的AM真菌物种丰富度与早期植物无显著差异;不同寄主植物进化类群间AM真菌组成差异显著.以上结果表明:AM真菌群落随着寄主植物进化而发生变化.在进化过程中,寄主植物倾向于选择保留共生效率较高的AM真菌.