Autism spectrum disorder (ASD) and Alzheimer’s disease (AD) are neurodevelopmental and neurodegenerative disorders, respectively. While exome sequencing is routinely employed during the early stages of ASD diagnosis, it rarely influences therapeutic strategies. To address this gap, we have reconstructed and analyzed the gene networks linking autism spectrum disorders, Alzheimer’s disease, and mTOR signaling. In addition, we have performed a phylostratigraphic analysis that reveals similarities and differences in the evolution of both ASD and Alzheimer’s disease predisposition genes. We have shown that almost half of the genes predisposing to autism and two-fifths of the genes predisposing to Alzheimer’s disease are directly related to the mTOR signaling pathway. Analysis of Phylostratigraphic Age Index (PAI) value distributions revealed a significant enrichment of evolutionarily ancient genes in both ASD- and AD-related gene sets. When studying the distribution of ASD predisposition genes by Divergence Index (DI) values, a significant enrichment with genes having extremely low DI = 0 has been found. Such low DI values indicate that most likely these genes are under stabilizing selection. Using the ANDVisio tool, both pharmacological and natural mTOR regulators with potential for ASD treatment were selected, such as propofol, dexamethasone, celecoxib, statins, berberine, resveratrol, quercetin, myricetin, mio-inositol, and several amino acids.
We report the near-complete genome of Cydia pomonella granulovirus BZR GV 10. The genome sequence is 123,362 bp long and shows a degree of similarity of more than 99% to the reference isolate Mexico/1963.
The use of CpGV strains as the basis for bioinsecticides is an effective and safe way to control Cydia pomonella. The research is aimed at the identification and study of new CpGV strains. Objects of identification and bioinformatic analysis: 18 CpGV strains. Sequencing was carried out on a NextSeq550. Genome assembly and annotation were carried out using Spades, Samtools 1.9, MinYS, Pilon, Gfinisher, Quast, and Prokka. Comparative genomic analysis was carried out in relation to the reference genome present in the «Madex Tween» strain-producer (biological standard) according to the average nucleotide identity (ANI) criterion. The presence/absence of IAP, cathepsin, MMP, and chitinase in the genetic sequences of the strains was determined using simply phylogeny. Entomopathogenic activity was assessed against C. pomonella according to the criterion of biological efficacy. Thus, molecular genetic identification revealed that 18 CpGV strains belong to a genus of Betabaculovirus. For all the strains under study ANI values of 99% or more were obtained, and the presence of the cathepsin, chitinase, IAP, and MMP genes was noted. The strains BZR GV 1, BZR GV 3, BZR GV 7, BZR GV 10, and BZR GV L-8 showed the maximum biological efficacy: 100% on the 15th day of observation. Strains BZR GV 4, BZR GV 8, and BZR GV 13 showed efficacy at the level of the «Madex Tween» preparation: 89.5% on the 15th day of observation. The strains with the highest mortality rate of the host insect were identified: BZR GV 9, BZR GV 10, BZR GV L-6, and BZR GV L-8.
Technologies for the production of a range of compounds using microorganisms are becoming increasingly popular in industry. The creation of highly productive strains whose metabolism is aimed to the synthesis of a specific desired product is impossible without complex directed modifications of the genome using mathematical and computer modeling methods. One of the bacterial species actively used in biotechnological production is Corynebacterium glutamicum. There are already 5 whole-genome flux balance models for it, which can be used for metabolism research and optimization tasks. The paper presents fluxMicrobiotech, a software module developed at the Institute of Cytology and Genetics of the Siberian Branch of the Russian Academy of Sciences, which implements a series of computational protocols designed for high-performance computer analysis of C. glutamicum whole-genome flux balance models. The tool is based on libraries from the opencobra community (https://opencobra.github.io) within the Python programming language (https://www.python.org), using the Pandas (https://pandas.pydata.org) and Escher (https://escher.readthedocs.io) libraries . It is configured to operate on a ‘file-in/file-out’ basis. The model, environmental conditions, and model constraints are specified as separate text table files, which allows one to prepare a series of files for each section, creating databases of available test scenarios for variations of the model. Or vice versa, allowing a single model to be tested under a series of different cultivation conditions. Post-processing tools for modeling data are set up, providing visualization of summary charts and metabolic maps.
In this work, we for the first time performed a comprehensive bioinformatics analysis of 568 human genes that, according to the NCBI Gene database as on September 15, 2024, were associated with pain generation, perception and anesthesia. The SCN9A gene encoding the sodium voltage-gated channel α subunit 9 and expressed in sensory neurons for transferring signals to the central nervous system about tissue damage was the only one involved in all the processes of interest at once as a hub gene. First, with our tool called OrthoWeb, we estimated the phylostratigraphic age indices (PAIs) for each of the genes, that is, identified the taxon of the most recent common ancestor of the organisms for which that gene has been sequenced. The mean PAI for all genes under study, including SCN9A as a hub gene for pain generation, perception, response and anesthesia, was ‘4’. On the evolutionary scale by the Kyoto Encyclopedia of Genes and Genomes (KEGG), the ancestor is the phylum Chordata, some of the most ancient of which evolved the central and the peripheral nervous system. Next, with our tool called ANDSystem, we found that phosphorylation of ion channels is a centerpiece in pain generation, perception, response and anesthesia, on which the efficiency of signal transduction from the peripheral to the central system depends. This conclusion was consistent with literature data on a key role an efficient signal transduction from the peripheral to the central system from the peripheral to the central system for adjusting the human circadian rhythm through detection of a change from the dark of night to the light of day and for identification of the direction of the source of sound by auditory brainstem nuclei, for generating the response to cold stress and for physical coordination. 21 candidate SNP marker of significant SCN9A over- and underexpression. Finally, the ratio of SCN9A upregulating to downregulating SNPs was compared to that for all known human genes estimated by the 1000 Genomes Project Consortium. It was found that SCN9A as a hub gene for pain generation, perception, pain response and anesthesia is acted on by natural selection against its downregulation, to keep the nervous system highly informed on the status of the organism and the environment.
This article introduces Orthoweb (https://orthoweb.sysbio.cytogen.ru/), a software package developed for the calculation of evolutionary indices, including phylostratigraphic indices and divergence indices (Ka/Ks) for individual genes as well as for gene networks. The phylostratigraphic age index (PAI) allows the evolutionary stage of a gene’s emergence (and thus indirectly the approximate time of its origin, known as “evolutionary age”) to be assessed based on the analysis of orthologous genes across closely and distantly related taxa. Additionally, Orthoweb supports the calculation of the transcriptome age index (TAI) and the transcriptome divergence index (TDI). These indices are important for understanding the dynamics of gene expression and its impact on the development and adaptation of organisms. Orthoweb also includes optional analytical features, such as the ability to explore Gene Ontology (GO) terms associated with genes, facilitating functional enrichment analyses that link evolutionary origins of genes to biological processes. Furthermore, it offers tools for SNP enrichment analysis, enabling the users to assess the evolutionary significance of genetic variants within specific genomic regions. A key feature of Orthoweb is its ability to integrate these indices with gene network analysis. The software offers advanced visualization tools, such as gene network mapping and graphical representations of phylostratigraphic index distributions of network elements, ensuring intuitive interpretation of complex evolutionary relationships. To further streamline research workflows, Orthoweb includes a database of pre-calculated indices for numerous taxa, accessible via an application programming interface (API). This feature allows the users to retrieve pre-computed phylostratigraphic and divergence data efficiently, significantly reducing computational time and effort.
Wheat heading time is primarily governed by two loci: VRN-1 (response to vernalization) and PPD-1 (response to photoperiod). Five sets of near-isogenic lines (NILs) were studied with the aim of investigating the effect of the aforementioned genes on wheat vegetative period duration and 14 yield-related traits. Every NIL was sown in the hydroponic greenhouse of the Institute of Cytology and Genetics, SB RAS. To assess their allelic composition at the VRN-1 and PPD-1 loci, molecular markers were used. It was shown that HT in plants with the Vrn-A1vrn-B1vrn-D1 genotype was reduced by 29 and 21 days (p < 0.001) in comparison to HT in plants with the vrn-A1Vrn-B1vrn-D1 and the vrn-A1vrn-B1Vrn-D1 genotypes, respectively. In our study, we noticed a decrease in spike length as well as spikelet number per spike parameter for some NIL carriers of the Vrn-A1a allele in comparison to carriers of the Vrn-B1 allele. PCA revealed three first principal components (PC), together explaining more than 70% of the data variance. Among the studied genetic traits, the Vrn-A1a and Ppd-D1a alleles showed significant correlations with PCs. Regarding genetic components, significant correlations were calculated between PC3 and Ppd-B1a (−0.26, p < 0.05) and Vrn-B1 (0.57, p < 0.05) alleles. Thus, the presence of the Vrn-A1a allele affects heading time, while Ppd-D1a is associated with plant height reduction.
Cholesterol is an essential structural component of cell membranes and a precursor of vitamin D, as well as steroid hormones. Humans and other animal species can absorb cholesterol from food. Cholesterol is also synthesized de novo in the cells of many tissues. We have previously reconstructed the gene network regulating intracellular cholesterol levels, which included regulatory circuits involving transcription factors from the SREBP (Sterol Regulatory Element-Binding Proteins) subfamily. The activity of SREBP transcription factors is regulated inversely depending on the intracellular cholesterol level. This mechanism is implemented with the participation of proteins SCAP, INSIG1, INSIG2, MBTPS1/S1P and MBTPS2/S2P. This group of proteins, together with the SREBP factors, is designated as “cholesterol sensor”. An elevated cholesterol level is a risk factor for the development of cardiovascular diseases and may also be observed in obesity, diabetes and other pathological conditions. Systematization of information about the molecular mechanisms controlling the activity of SREBP factors and cholesterol biosynthesis in the form of a gene network and building new knowledge about the gene network as a single object is extremely important for understanding the molecular mechanisms underlying the predisposition to diseases. With a computer tool, ANDSystem, we have built a gene network regulating cholesterol biosynthesis. The gene network included data on: (1) the complete set of enzymes involved in cholesterol biosynthesis; (2) proteins that function as part of the “cholesterol sensor”; (3) proteins that regulate the activity of the “cholesterol sensor”; (4) genes encoding proteins of these groups; (5) genes whose transcription is regulated by SREBP factors (SREBP target genes). The gene network was analyzed and feedback loops that control the activity of SREBP factors were identified. These feedback loops involved the PPARG, NR0B2/SHP1, LPIN1, and AR genes and the proteins they encode. Analysis of the phylostratigraphic age of the genes showed that the ancestral forms of most human genes encoding the enzymes of cholesterol biosynthesis and the proteins of the “cholesterol sensor” may have arisen at early evolutionary stages (Cellular organisms (the root of the phylostratigraphic tree) and the stages of Eukaryota and Metazoa divergence). However, the mechanism of gene transcription regulation in response to changes in cholesterol levels may only have formed at later evolutionary stages, since the phylostratigraphic age of the genes encoding the transcription factors SREBP1 and SREBP2 corresponds to the stage of Vertebrata divergence.
We report a genome of CpGV from the bioresource collection of the Federal Research Center of Biological Plant Protection "State Collection of Entomoacariphages and Microorganisms." Its sequence is 123,862 bp. The genome under study demonstrates a degree of similarity of more than 99% with reference NC_002816 from the NCBI RefSeq database.
Cydia pomonella granulovirus is a natural pathogen for Cydia pomonella that is used as a biocontrol agent of insect populations. The study of granulovirus virulence is of particular interest since the development of resistance in natural populations of C. pomonella has been observed during the long-term use of the Mexican isolate CpGV. In our study, we present the genomes of 18 CpGV strains endemic to southern Russia and from Kazakhstan, as well as a strain included in the commercial preparation “Madex Twin”, which were sequenced and analyzed. We performed comparative genomic analysis using several tools. From comparisons at the level of genes and protein products that are involved in the infection process of virosis, synonymous and missense substitution variants have been identified. The average nucleotide identity has demonstrated a high similarity with other granulovirus genomes of different geographic origins. Whole-genome alignment of the 18 genomes relative to the reference revealed regions of low similarity. Analysis of gene repertoire variation has shown that BZR GV 4, BZR GV 6, and BZR GV L-7 strains have been the closest in gene content to the commercial “Madex Twin” strain. We have confirmed two deletions using read depth coverage data in regions lacking genes shown by homology analysis for granuloviruses BZR GV L-4 and BZR GV L-6; however, they are not related to the known genes causing viral pathogenicity. Thus, we have isolated novel CpGV strains and analyzed their potential as strains producing highly effective bioinsecticides against C. pomonella.
Data on the genetics and molecular biology of diabetes are accumulating rapidly. This poses the challenge of creating research tools for a rapid search for, structuring and analysis of information in this field. We have developed a web resource, GlucoGenes®, which includes a database and an Internet portal of genes and proteins associated with high glucose (hyperglycemia), low glucose (hypoglycemia), and both metabolic disorders. The data were collected using text mining of the publications indexed in PubMed and PubMed Central and analysis of gene networks associated with hyperglycemia, hypoglycemia and glucose variability performed with ANDSystems, a bioinformatics tool. GlucoGenes® is freely available at: https://glucogenes.sysbio.ru/genes/main. GlucoGenes® enables users to access and download information about genes and proteins associated with the risk of hyperglycemia and hypoglycemia, molecular regulators with hyperglycemic and antihyperglycemic activity, genes up-regulated by high glucose and/or low glucose, genes down-regulated by high glucose and/or low glucose, and molecules otherwise associated with the glucose metabolism disorders. With GlucoGenes®, an evolutionary analysis of genes associated with glucose metabolism disorders was performed. The results of the analysis revealed a significant increase (up to 40 %) in the proportion of genes with phylostratigraphic age index (PAI) values corresponding to the time of origin of multicellular organisms. Analysis of sequence conservation using the divergence index (DI) showed that most of the corresponding genes are highly conserved (DI < 0.6) or conservative (DI < 1). When analyzing single nucleotide polymorphism (SNP) in the proximal regions of promoters affecting the affinity of the TATA-binding protein, 181 SNP markers were found in the GlucoGenes® database, which can reduce (45 SNP markers) or increase (136 SNP markers) the expression of 52 genes. We believe that this resource will be a useful tool for further research in the field of molecular biology of diabetes.
Расстройства аутистического спектра (РАС) — это сложное нарушение нейропсихического развития, диагностируемое в настоящее время более, чем у 2 % детей. Основные симптомы РАС: снижение коммуникативных и социальных функций, повышение стереотипий во всех формах поведения. Для РАС характерна как симптоматическая, так и генетическая гетерогенность, что является препятствием для разработки эффективной терапии. Разделение аутизма на несколько подтипов, основанных на общих патогенетических механизмах, становится все более актуальным. Одним из таких подтипов стал аутизм, связанный с материнской иммунной активацией в процессе беременности, в результате которого организмом матери нарабатываются аутоантитела к нейрональным белкам плода и тем самым нарушается нормальное нейроразвитие. Другими сложными для дифференциальной диагностики РАС считаются синдромы PANS/PANDAS — постинфекционные аутоиммунные осложнения, имеющие ярко выраженную нейропсихическую симптоматику. Также обсуждается связь генетических и иммунных нарушений при РАС с сигнальным путем mTOR, гиперактивация которого часто наблюдается при аутизме.
We propose the trait-based method for quantifying the activity of functional groups in the human gut microbiome based on metatranscriptomic data. It allows one to assess structural changes in the microbial community comprised of the following functional groups: butyrate-producers, acetogens, sulfate-reducers, and mucin-decomposing bacteria. It is another way to perform a functional analysis of metatranscriptomic data by focusing on the ecological level of the community under study. To develop the method, we used published data obtained in a carefully controlled environment and from a synthetic microbial community, where the problem of ambiguity between functionality and taxonomy is absent. The developed method was validated using RNA-seq data and sequencing data of the 16S rRNA amplicon on a simplified community. Consequently, the successful verification provides prospects for the application of this method for analyzing natural communities of the human intestinal microbiota.
Modern computational biology makes widespread use of mathematical models of biological systems, in particular systems of ordinary differential equations, as well as models of dynamic systems described in other formalisms, such as agent-based models. Parameters are numerical values of quantities reflecting certain properties of a modeled system and affecting model solutions. At the same time, depending on parameter values, different dynamic regimes—stationary or oscillatory, established as a result of transient modes of various types—can be observed in the modeled system. Predicting changes in the solution dynamics type depending on changes in model parameters is an important scientific task. Nevertheless, this problem does not have an analytical solution for all formalisms in a general case. The routinely used method of performing a series of computational experiments, i.e., solving a series of direct problems with various sets of parameters followed by expert analysis of solution plots is labor-intensive with a large number of parameters and a decreasing step of the parametric grid. In this regard, the development of methods allowing the obtainment and analysis of information on a set of computational experiments in an aggregate form is relevant. This work is devoted to developing a method for the visualization and classification of various dynamic regimes of a model using a composition of the dynamic time warping (DTW-algorithm) and principal coordinates analysis (PCoA) methods. This method enables qualitative visualization of the results of the set of solutions of a mathematical model and the performance of the correspondence between the values of the model parameters and the type of dynamic regimes of its solutions. This method has been tested on the Lotka–Volterra model and artificial sets of various dynamics.
Modern investigations in biology often require the efforts of one or more groups of researchers. Often these are groups of specialists from various scientific fields who generate and share data of different formats and sizes. Without modern approaches to work automation and data versioning (where data from different collaborators are stored at different points in time), teamwork quickly devolves into unmanageable confusion. In this review, we present a number of information systems designed to solve these problems. Their application to the organization of scientific activity helps to manage the flow of actions and data, allowing all participants to work with relevant information and solving the issue of reproducibility of both experimental and computational results. The article describes methods for organizing data flows within a team, principles for organizing metadata and ontologies. The information systems Trello, Git, Redmine, SEEK, OpenBIS and Galaxy are considered. Their functionality and scope of use are described. Before using any tools, it is important to understand the purpose of implementation, to define the set of tasks they should solve, and, based on this, to formulate requirements and finally to monitor the application of recommendations in the field. The tasks of creating a framework of ontologies, metadata, data warehousing schemas and software systems are key for a team that has decided to undertake work to automate data circulation. It is not always possible to implement such systems in their entirety, but one should still strive to do so through a stepbystep introduction of principles for organizing data and tasks with the mastery of individual software tools. It is worth noting that Trello, Git, and Redmine are easier to use, customize, and support for small research groups. At the same time, SEEK, OpenBIS, and Galaxy are more specific and their use is advisable if the capabilities of simple systems are no longer sufficient.
Nonribosomal peptides play an important role in the vital activity of bacteria and have an extremely broad field of biological activity. In particular, they act as antibiotics, toxins, surfactants, siderophores, and also perform a number of other specific functions. Biosynthesis of these molecules does not occur on ribosomes but by special enzymes that form gene clusters in bacterial genomes. We hypothesized that the presence of nonribosomal peptide synthesis pathways is a specific feature of bacterial metabolism, which may affect other vital processes of the cell, including translational ones. This work was the first to show the relationship between the translation regulation mechanism of protein-coding genes in bacteria, which is largely determined by the efficiency of translation elongation, and the presence of gene clusters in the genomes for the biosynthesis of nonribosomal peptides. Bioinformatic analysis of the translation elongation efficiency of protein-coding genes was performed in 11 679 bacterial genomes, some of which contained gene clusters of nonribosomal peptide biosynthesis and some of which did not. The analysis showed that bacteria whose genomes contained clusters of nonribosomal peptide biosynthetic genes and those without such gene clusters differ significantly in the molecular mechanisms that ensure translation efficiency. Thus, among microorganisms whose genomes contain gene clusters of nonribosomal peptide synthetases, a significantly smaller part of them is characterized by optimized regulation of the number of local inverted repeats, while most of them have genomes optimized by the averaged energy of inverted repeats studs in mRNA and additionally by codon composition. Our results suggest that the presence of nonribosomal peptide biosynthetic pathways in bacteria may influence the structure of the overall bacterial metabolism, which is also expressed in the specific mechanisms of ribosomal protein biosynthesis.
Translation efficiency modulates gene expression in prokaryotes. The comparative analysis of translation elongation efficiency characteristics of Ralstonia genus bacteria genomes revealed that these characteristics diverge in accordance with the phylogeny of Ralstonia. The first branch of this genus is a group of bacteria commonly found in moist environments such as soil and water that includes the species R. mannitolilytica, R. insidiosa, and R. pickettii, which are also described as nosocomial infection pathogens. In contrast, the second branch is plant pathogenic bacteria consisting of R. solanacearum, R. pseudosolanacearum, and R. syzygii. We found that the soil Ralstonia have a significantly lower number and energy of potential secondary structures in mRNA and an increased role of codon usage bias in the optimization of highly expressed genes' translation elongation efficiency, not only compared to phytopathogenic Ralstonia but also to Cupriavidus necator, which is closely related to the Ralstonia genus. The observed alterations in translation elongation efficiency of orthologous genes are also reflected in the difference of potentially highly expressed gene' sets' content among Ralstonia branches with different lifestyles. Analysis of translation elongation efficiency characteristics can be considered a promising approach for studying complex mechanisms that determine the evolution and adaptation of bacteria in various environments.
В статье описан вклад Сергея Ивановича Бажана в развитие методов компьютерного моделирования сложных биологических систем. Химико-кинетический метод моделирования, предложенный С.И. Бажаном и его коллегой В.А. Лихош- ваем во время работы в ФБУН ГНЦ ВБ «Вектор» в 1970-х гг., оказался исключительно удачным и эффективным инструментом исследования динамики сложных, иерархически организованных биологических систем. Данный способ представляет собой одно из важнейших достижений сибирской школы математической/системной биологии и биоинформатики. Концепции, почти полвека назад ставшие основой этого подхода, до сих пор соответствуют тенденциям современной системной биологии.
Cancer is a complex and heterogeneous disease characterized by the accumulation of genetic alterations that drive uncontrolled cell growth and proliferation. Evolutionary dynamics plays a crucial role in the emergence and development of tumors, shaping the heterogeneity and adaptability of cancer cells. From the perspective of evolutionary theory, tumors are complex ecosystems that evolve through a process of microevolution influenced by genetic mutations, epigenetic changes, tumor microenvironment factors, and therapyinduced changes. This dynamic nature of tumors poses significant challenges for effective cancer treatment, and understanding it is essential for developing effective and personalized therapies. By uncovering the mechanisms that determine tumor heterogeneity, researchers can identify key genetic and epigenetic changes that contribute to tumor progression and resistance to treatment. This knowledge enables the development of innovative strategies for targeting specific tumor clones, minimizing the risk of recurrence and improving patient outcomes. To investigate the evolutionary dynamics of cancer, researchers employ a wide range of experimental and computational approaches. Traditional experimental methods involve genomic profiling techniques such as nextgeneration sequencing and fluorescence in situ hybridization. These techniques enable the identification of somatic mutations, copy number alterations, and structural rearrangements within cancer genomes. Furthermore, singlecell sequencing methods have emerged as powerful tools for dissecting intratumoral heterogeneity and tracing clonal evolution. In parallel, computational models and algorithms have been developed to simulate and analyze cancer evolution. These models integrate data from multiple sources to predict tumor growth patterns, identify driver mutations, and infer evolutionary trajectories. In this paper, we set out to describe the current approaches to address this evolutionary complexity and theories of its occurrence.
To overcome immune tolerance to cancer, the immune system needs to be exposed to a multi-target action intervention. Here, we investigated the activating effect of CpG oligodeoxynucleotides (ODNs), mesyl phosphoramidate CpG ODNs, anti-OX40 antibodies, and OX40 RNA aptamers on major populations of immunocompetent cells ex vivo. Comparative analysis of the antitumor effects of in situ vaccination with CpG ODNs and anti-OX40 antibodies, as well as several other combinations, such as mesyl phosphoramidate CpG ODNs and OX40 RNA aptamers, was conducted. Antibodies against programmed death 1 (PD1) checkpoint inhibitors or their corresponding PD1 DNA aptamers were also added to vaccination regimens for analytical purposes. Four scenarios were considered: a weakly immunogenic Krebs-2 carcinoma grafted in CBA mice; a moderately immunogenic Lewis carcinoma grafted in C57Black/6 mice; and an immunogenic A20 B cell lymphoma or an Ehrlich carcinoma grafted in BALB/c mice. Adding anti-PD1 antibodies (CpG+αOX40+αPD1) to in situ vaccinations boosts the antitumor effect. When to be used instead of antibodies, aptamers also possess antitumor activity, although this effect was less pronounced. The strongest effect across all the tumors was observed in highly immunogenic A20 B cell lymphoma and Ehrlich carcinoma.