The effect of pH changes in external ionic solutions on liposomes entrapping non-permeable weak acid/base pairs was simulated using a simple "toy" model. Membrane semi-permeability and selective passive diffusion lead to internal stationary states that inevitably exhibit electroneutrality breaking. The resulting electrostatic effects on species' potentials and ion flow rates depend strongly on compartment volume, in agreement with the biological surface-to-volume rule. Although electrostatic effects on membrane stability, species' chemical potentials, and transmembrane ion migration are negligible at very small volumes (pL range), the electrical potential differences between external and internal phases fall within the range observed in cellular compartments. Furthermore, proton gradients generate electrical potential differences across the membrane that drive electrons in the same direction as protons, offering a novel perspective on the proton's universal role as a driving force in living compartments via its coupling to redox reactions. Since all these effects are inherently tied to very small volumes, where electroneutrality deviations are possible, in contrast to macroscopic electrolyte solutions, this supports the emergence of life and cellular chemistry in small semi-permeable compartments.
RNA virus populations consist of complex and dynamic mutant spectra in which most individual genomes differ in one or more positions from the other genomes of the same population. This behavior, known as quasispecies dynamics, applies to SARS-CoV-2 which exhibits intrahost genetic and functional heterogeneity while evolving at a high rate in the human population. In the present study, we describe a remarkable reduction in mutant spectrum complexity (intrahost viral genome heterogeneity) in SARS-CoV-2 isolates of late relative to early COVID-19 waves, as they reached Madrid (Spain) from 2020 until 2022. In contrast, the consensus (average) sequence of the corresponding isolates displayed a continuing divergence from the initial Wuhan-Hu-1 virus as the pandemic advanced. The mutant spectrum complexity developed upon replication in Vero E6 cells of the isolates from the first and sixth COVID-19 waves, as well as of biological clones retrieved from them, was similar. Therefore, the mutant spectrum complexity reduction observed in vivo was not due to an increased accuracy of the viral replicative machinery, but rather to other factors related to viral epidemiology or pathogenesis. Such possible factors and their implications for viral trait modifications in the course of a viral pandemic are discussed. The results establish that mutant spectrum complexity of genetically variable viruses can be an epidemiologically evolvable trait.
On May 23-24, 2024, the 1st Spanish Conference on Genomic Medicine convened in Madrid, Spain. An international and multidisciplinary group of experts gathered to discuss the current state and prospects of genomic medicine in the Spanish-speaking world. There were 278 attendees from Latin America, US, UK, Germany, and Spain, and the topics covered included rare diseases, genome medicine in national health systems (NHSs), artificial intelligence, and commercial development ventures. One particular area of attention was our still sketchy understanding of genome variants. This is evidenced by the fact that many diagnoses in rare diseases continue to yield odysseys that take years, with up to 50% of cases that may go undiagnosed. Since a lot of the genome remains to poorly understood, as new technologies such as long read sequencing become more ubiquitous and cheaper, it is expected that current gaps in genome references will improve. However, disparities within the NHSs suggest that advancements do not necessarily rely on resources but the appropriate regulation and pathways for education of professionals being properly implemented. This is where Genomics England can be a clinical genomic implementation example for routine health care. Ethical challenges, including privacy, informed consent, equity, representation, and genetic discrimination, also require the need for robust legal frameworks and culturally sensitive practices. The future of genomics in Spanish-speaking countries depends on addressing all of these issues. By navigating these challenges responsibly, Spanish-speaking countries can harness the power of genomics to improve health outcomes and advance scientific knowledge, ensuring that the benefits of personalized medicine are realized in an inclusive and equitable manner.
Since its introduction in the human population, SARS-CoV-2 has evolved into multiple clades, but the events in its intrahost diversification are not well understood. Here, we compare three-dimensional (3D) self-organized neural haplotype maps (SOMs) of SARS-CoV-2 from thirty individual nasopharyngeal diagnostic samples obtained within a 19-day interval in Madrid (Spain), at the time of transition between clades 19 and 20. SOMs have been trained with the haplotype repertoire present in the mutant spectra of the nsp12- and spike (S)-coding regions. Each SOM consisted of a dominant neuron (displaying the maximum frequency), surrounded by a low-frequency neuron cloud. The sequence of the master (dominant) neuron was either identical to that of the reference Wuhan-Hu-1 genome or differed from it at one nucleotide position. Six different deviant haplotype sequences were identified among the master neurons. Some of the substitutions in the neural clouds affected critical sites of the nsp12-nsp8-nsp7 polymerase complex and resulted in altered kinetics of RNA synthesis in an in vitro primer extension assay. Thus, the analysis has identified mutations that are relevant to modification of viral RNA synthesis, present in the mutant clouds of SARS-CoV-2 quasispecies. These mutations most likely occurred during intrahost diversification in several COVID-19 patients, during an initial stage of the pandemic, and within a brief time period.
The creation of fitness maps from viral populations especially in the case of RNA viruses, with high mutation rates producing quasispecies, is complex since the mutant spectrum is in a very high-dimensional space. In this work, a new approach is presented using a class of neural networks, Self-Organized Maps (SOM), to represent realistic fitness landscapes in two RNA viruses: Human Immunodeficiency Virus type 1 (HIV-1) and Hepatitis C Virus (HCV). This methodology has proven to be very effective in the classification of viral quasispecies, using as criterium the mutant sequences in the population. With HIV-1, the fitness landscapes are constructed by representing the experimentally determined fitness on the sequence map. This approach permitted the depiction of the evolutionary paths of the variants subjected to processes of fitness loss and gain in cell culture. In the case of HCV, the efficiency was measured as a function of the frequency of each haplotype in the population by ultra-deep sequencing. The fitness landscapes obtained provided information on the efficiency of each variant in the quasispecies environment, that is, in relation to the entire spectrum of mutants. With the SOM maps, it is possible to determine the evolutionary dynamics of the different haplotypes.
Populations of RNA viruses are composed of complex and dynamic mixtures of variant genomes that are termed mutant spectra or mutant clouds. This applies also to SARS-CoV-2, and mutations that are detected at low frequency in an infected individual can be dominant (represented in the consensus sequence) in subsequent variants of interest or variants of concern. Here we briefly review the main conclusions of our work on mutant spectrum characterization of hepatitis C virus (HCV) and SARS-CoV-2 at the nucleotide and amino acid levels and address the following two new questions derived from previous results: (i) how is the SARS-CoV-2 mutant and deletion spectrum composition in diagnostic samples, when examined at progressively lower cut-off mutant frequency values in ultra-deep sequencing; (ii) how the frequency distribution of minority amino acid substitutions in SARS-CoV-2 compares with that of HCV sampled also from infected patients. The main conclusions are the following: (i) the number of different mutations found at low frequency in SARS-CoV-2 mutant spectra increases dramatically (50- to 100-fold) as the cut-off frequency for mutation detection is lowered from 0.5% to 0.1%, and (ii) that, contrary to HCV, SARS-CoV-2 mutant spectra exhibit a deficit of intermediate frequency amino acid substitutions. The possible origin and implications of mutant spectrum differences among RNA viruses are discussed.
Fitness landscapes reflect the adaptive potential of viruses. There is no information on how fitness peaks evolve when a virus replicates extensively in a controlled cell culture environment. Here we report the construction of Self-Organized Maps (SOMs), based on deep sequencing reads of three amplicons of the NS5A-NS5B-coding region of hepatitis C virus (HCV). A two-dimensional neural network was constructed and organized according to sequence relatedness. The third dimension of the fitness profile was given by the haplotype frequencies at each neuron. Fitness maps were derived for 44 HCV populations that share a common ancestor that was passaged up to 210 times in human hepatoma Huh-7.5 cells. As the virus increased its adaptation to the cells, the number of fitness peaks expanded, and their distribution shifted in sequence space. The landscape consisted of an extended basal platform, and a lower number of protruding higher fitness peaks. The function that relates fitness level and peak abundance corresponds a power law, a relationship observed with other complex natural phenomena. The dense basal platform may serve as spring-board to attain high fitness peaks. The study documents a highly dynamic, double-layer fitness landscape of HCV when evolving in a monotonous cell culture environment. This information may help interpreting HCV fitness landscapes in complex in vivo environments.IMPORTANCE The study provides for the first time the fitness landscape of a virus in the course of its adaptation to a cell culture environment, in absence of external selective constraints. The deep sequencing-based self-organized maps document a two-layer fitness distribution with an ample basal platform, and a lower number of protruding, high fitness peaks. This landscape structure offers potential benefits for virus resilience to mutational inputs.
RNA viruses replicate as complex mutant spectra termed viral quasi-species. The frequency of each individual genome in a mutant spectrum depends on its rate of generation and its relative fitness in the replicating population ensemble. The advent of deep sequencing methodologies allows for the first-time quantification of haplotype abundances within mutant spectra. There is no information on the haplotype profile of the resident genomes and how the landscape evolves when a virus replicates in a controlled cell culture environment. Here, we report the construction of intramutant spectrum haplotype landscapes of three amplicons of the NS5A-NS5B coding region of hepatitis C virus (HCV). Two-dimensional (2D) neural networks were constructed for 44 related HCV populations derived from a common clonal ancestor that was passaged up to 210 times in human hepatoma Huh-7.5 cells in the absence of external selective pressures. The haplotype profiles consisted of an extended dense basal platform, from which a lower number of protruding higher peaks emerged. As HCV increased its adaptation to the cells, the number of haplotype peaks within each mutant spectrum expanded, and their distribution shifted in the 2D network. The results show that extensive HCV replication in a monotonous cell culture environment does not limit HCV exploration of sequence space through haplotype peak movements. The landscapes reflect dynamic variation in the intramutant spectrum haplotype profile and may serve as a reference to interpret the modifications produced by external selective pressures or to compare with the landscapes of mutant spectra in complex in vivo environments. IMPORTANCE The study provides for the first time the haplotype profile and its variation in the course of virus adaptation to a cell culture environment in the absence of external selective constraints. The deep sequencing-based self-organized maps document a two- layer haplotype distribution with an ample basal platform and a lower number of protruding peaks. The results suggest an inferred intramutant spectrum fitness landscape structure that offers potential benefits for virus resilience to mutational inputs.
An accurate analysis of user behaviour in online learning environments is a useful means of early follow up of students, so that they can be better supported to improve their performance and achieve the expected competences. However, that task becomes challenging due to the massive data that learning management systems store and categories. With the COVID-19 pandemic still on-going, face-to-face learning settings have migrate into online and blended ones, meaning an increase of online students and teachers in need for a tailored and effective support to their needs. A novel unsupervised clustering technique based on the Self-Organizing Map (SOM) artificial neural network model is used in this research to analyse 1,709,189 records of online students enrolled from 2015 to 2019 at Universidad Internacional de La Rioja (UNIR), a fully online Higher Education institution. SOM performs a precise and diverse user clustering based on those records. Results highlight that specific clusters are linked to the intake average profile at the university, with a clear relation between user interaction and a higher performance. Further, results show that, out of a targeted desk research compared to the analysis in this paper, face-to-face and online settings are connected through the methodological approach beyond the technology-based environment, which presents a similar behaviour in both contexts.
Catalytic reaction networks consist of molecular arrays interconnected by autocatalysis and cross-catalytic pathways among the reactants, and serve as bottom-up models for the design and understanding of molecular evolution and emergent phenomena. An important example of the latter is the emergence of homochirality in biomolecules during chemical evolution. This chiral symmetry breaking is triggered by bistability and bifurcation in networks of chiral replicators. Spontaneous mirror symmetry breaking (SMSB) results from hypercyclic connectivity when the chirality and enantioselectivity of the replicators are taken into account. Heretofore, SMSB has been generally understood as involving chemical transformations yielding scalemic outcomes as non-equilibrium steady states (NESS). Here, in marked contrast, we consider the chaotic regime, in which steady states do not exist. The dissipation, or entropy production, is chaotic as is the exchange entropy. The rate of change of the total system entropy, governed by the entropy balance equation, is also chaotic. Subsequent to the mirror symmetry breaking transition, the time averaged entropy production is minimized in the final chaotic chiral state with respect to the former chaotic racemic state. The chemical forces (i.e., the affinities) evolve in time so as to lower the sum of the entropy production and the exchange entropy, in compliance with the general evolution criterion extended to reaction networks subject to volumetric open flow.
Data clustering is aimed at finding groups of data that share common hidden properties. These kinds of techniques are especially critical at early stages of data analysis where no information about the dataset is available. One of the mayor shortcomings of the clustering algorithms is the difficulty for non-experts users to configure them and, in some cases, interpret the results. In this work a computational approach with a two-layer structure based on Self-Organizing Map (SOM) is presented for cluster analysis. In the first level, a quantization of the data samples using topology-preserving metrics to automatically determine the number of units in the SOM is proposed. In the second level the obtained SOM prototypes are clustered by means of a connectivity analysis to explore the quality of the partitioning with different number of clusters. The most important benefit of this two-layer procedure is that computational load decreases considerably in comparison with data based clustering methods, making it possible to cluster large data sets and to consider several different clustering alternatives in a limited time. This methodology produces a two-dimensional map representation of the, usually, high dimensional input space, along with quantitative information on viable clustering alternatives, which facilitates the exploration of the possible partitions in a dataset. The efficiency and interpretation of the methodology is illustrated by its application to artificial, benchmark and real complex biological datasets. The experimental results demonstrate the ability of the method to identify possible segmentations in a dataset, compared to algorithms that only yield a single clustering solution. The proposed algorithm tackles the intrinsic limitations of SOM and the parameter settings associated with the clustering methodology, without requiring the number of clusters or the SOM architecture as a prerequisite, among others. This way, it makes possible its application even by researchers with a limited expertise in machine learning. (C) 2017 Elsevier Ltd. All rights reserved.
Background We describe the pioneering experience of a Spanish family pursuing the goal of understanding their own personal genetic data to the fullest possible extent using Direct to Consumer (DTC) tests. With full informed consent from the Corpas family, all genotype, exome and metagenome data from members of this family, are publicly available under a public domain Creative Commons 0 (CC0) license waiver. All scientists or companies analysing these data (“the Corpasome”) were invited to return results to the family. Methods We released 5 genotypes, 4 exomes, 1 metagenome from the Corpas family via a blog and figshare under a public domain license, inviting scientists to join the crowdsourcing efforts to analyse the genomes in return for coauthorship or acknowldgement in derived papers. Resulting analysis data were compiled via social media and direct email. Results Here we present the results of our investigations, combining the crowdsourced contributions and our own efforts. Four companies offering annotations for genomic variants were applied to four family exomes: BIOBASE, Ingenuity, Diploid, and GeneTalk. Starting from a common VCF file and after selecting for significant results from company reports, we find no overlap among described annotations. We additionally report on a gut microbiome analysis of a member of the Corpas family. Conclusions This study presents an analysis of a diverse set of tools and methods offered by four DTC companies. The striking discordance of the results mirrors previous findings with respect to DTC analysis of SNP chip data, and highlights the difficulties of using DTC data for preventive medical care. To our knowledge, the data and analysis results from our crowdsourced study represent the most comprehensive exome and analysis for a family quartet using solely DTC data generation to date.
MOTIVATION:Self-organizing maps (SOMs) are readily available bioinformatics methods for clustering and visualizing high-dimensional data, provided that such biological information is previously transformed to fixed-size, metric-based vectors. To increase the usefulness of SOM-based approaches for the analysis of genomic sequence data, novel representation methods are required that automatically and objectively transform aligned nucleotide sequences into numeric vectors, dealing with both nucleotide ambiguity and gaps derived from sequence alignment.RESULTS:Six different codification variants based on Euclidean space, just like SOM processing, have been tested using two SOM models: the classical Kohonen's SOM and growing cell structures. They have been applied to two different sets of sequences: 32 sequences of small sub-unit ribosomal RNA from organisms belonging to the three domains of life, and 44 sequences of the reverse transcriptase region of the pol gene of human immunodeficiency virus type 1 belonging to different groups and sub-types. Our results show that the most important factor affecting the accuracy of sequence clustering is the assignment of an extra weight to the presence of alignment-derived gaps. Although each of the codification variants shows a different level of taxonomic consistency, the results are in agreement with sequence-based phylogenetic reconstructions and anticipate a broad applicability of this codification method.
The prediction of links among variables from a given dataset is a task referred to as network inference or reverse engineering. It is an open problem in bioinformatics and systems biology, as well as in other areas of science. Information theory, which uses concepts such as mutual information, provides a rigorous framework for addressing it. While a number of information-theoretic methods are already available, most of them focus on a particular type of problem, introducing assumptions that limit their generality. Furthermore, many of these methods lack a publicly available implementation. Here we present MIDER, a method for inferring network structures with information theoretic concepts. It consists of two steps: first, it provides a representation of the network in which the distance among nodes indicates their statistical closeness. Second, it refines the prediction of the existing links to distinguish between direct and indirect interactions and to assign directionality. The method accepts as input time-series data related to some quantitative features of the network nodes (such as e.g. concentrations, if the nodes are chemical species). It takes into account time delays between variables, and allows choosing among several definitions and normalizations of mutual information. It is general purpose: it may be applied to any type of network, cellular or otherwise. A Matlab implementation including source code and data is freely available (http://www.iim.csic.es/~gingproc/mider.html). The performance of MIDER has been evaluated on seven different benchmark problems that cover the main types of cellular networks, including metabolic, gene regulatory, and signaling. Comparisons with state of the art information-theoretic methods have demonstrated the competitive performance of MIDER, as well as its versatility. Its use does not demand any a priori knowledge from the user; the default settings and the adaptive nature of the method provide good results for a wide range of problems without requiring tuning.
Human Immunodeficiency Virus type 1 (HIV-1) because of high mutation rates, large population sizes, and rapid replication, exhibits complex evolutionary strategies. For the analysis of evolutionary processes, the graphical representation of fitness landscapes provides a significant advantage. The experimental determination of viral fitness remains, in general, difficult and consequently most published fitness landscapes have been artificial, theoretical or estimated. Self-Organizing Maps (SOM) are a class of Artificial Neural Network (ANN) for the generation of topological ordered maps. Here, three-dimensional (3D) data driven fitness landscapes, derived from a collection of sequences from HIV-1 viruses after “in vitro” passages and labelled with the corresponding experimental fitness values, were created by SOM. These maps were used for the visualization and study of the evolutionary process of HIV-1 “in vitro” fitness recovery, by directly relating fitness values with viral sequences. In addition to the representation of the sequence space search carried out by the viruses, these landscapes could also be applied for the analysis of related variants like members of viral quasiespecies. SOM maps permit the visualization of the complex evolutionary pathways in HIV-1 fitness recovery. SOM fitness landscapes have an enormous potential for the study of evolution in related viruses of “in vitro” works or from “in vivo” clinical studies with human, animal or plant viral infections.
1 poster presentado al European Conference on Computational Biology 2014, 7 - 10 Sep 2014.-- This poster is open access subject to the CC BY-NC Creative Commons 4.0 License
Background The availability of open access genomic data is essential for the personal genomics field. Public genomic data allow comparative analyses, testing of new tools and genotype-phenotype association studies. Personal genomics data of unrelated individuals are available in the public domain, notably the Personal Genome Project; however, to date genomics family data and metadata are severely lacking, mainly due to cost, privacy concerns or restricted access to Next Generation Sequencing (NGS) technology. Family data have a lot to offer as they allow the study of heritability, something which is impossible to do just by using unrelated individuals. Findings A whole family from Southern Spain decided to genotype, sequence and analyse their personal genomes making them publicly available under a Creative Commons 0 license (CC0; commonly denominated as public domain). These data include a) five 23andMe SNP chip genotype bed files, b) four raw exomes with their assorted bam files and VCF files, c) a metagenomic raw sequencing data file and d) derived data of likely phenotypes using SNPedia-derived tools. Conclusions To our knowledge this is the first CC0 released set of genomic, phenotypic and metagenomic data for a whole family. This dataset is also unique in that it was obtained through direct-to-consumer genetic tests. Hence any ordinary citizen with enough budget and samples should be able to reproduce this experiment. We envisage this dataset to be a useful resource for a variety of applications in the personal genomics field as a) negative control data for trait association discovery, b) testing data for development of new software and c) sample data for heritability studies. We encourage prospective users to share with us derived results so that they can be added to our existing collection.
The mountains of data thrusting from the new landscape of modern high-throughput biology are irrevocably changing biomedical research and creating a near-insatiable demand for training in data management and manipulation and data mining and analysis. Among life scientists, from clinicians to environmental researchers, a common theme is the need not just to use, and gain familiarity with, bioinformatics tools and resources but also to understand their underlying fundamental theoretical and practical concepts. Providing bioinformatics training to empower life scientists to handle and analyse their data efficiently, and progress their research, is a challenge across the globe. Delivering good training goes beyond traditional lectures and resource-centric demos, using interactivity, problem-solving exercises and cooperative learning to substantially enhance training quality and learning outcomes. In this context, this article discusses various pragmatic criteria for identifying training needs and learning objectives, for selecting suitable trainees and trainers, for developing and maintaining training skills and evaluating training quality. Adherence to these criteria may help not only to guide course organizers and trainers on the path towards bioinformatics training excellence but, importantly, also to improve the training experience for life scientists.
The goal of synthetic biology is to create artificial organisms. To achieve this it is essential to understand what life is. Metabolism-replacement systems, or (M, R)-systems, constitute a theory of life developed by Robert Rosen, characterized in the statement that organisms are closed to efficient causation, which means that they must themselves produce all the catalysts they need. This theory overlaps in part with other current theories, including autopoiesis, the chemoton, and autocatalytic sets, all of them invoking some idea of closure. A simple model of an (M, R)-system has been implemented in the computer, and behaves in ways that may shed light on the requirements for a prebiotic self-organizing system. In addition to a trivial steady state in which nothing happens, it can establish a non-trivial steady state in which all intermediates have finite concentrations, with their rates of degradation balanced by their rates of synthesis. The system can be regenerated from the set of food components plus a single intermediate, and maintain itself in that state indefinitely, despite continuous degradation. At the very low compartment volumes that may have existed in prebiotic conditions, for example in cavities in minerals, or in micelles formed by simple amphiphiles, statistical fluctuations in the numbers of molecules need to be taken into account. With the stochastic approach there is no non-trivial steady state in strict mathematical terms, because the system will always collapse to the trivial state after sufficient time. However, the average time before collapse is so long for volumes greater than \(10^{-19}\;{\textsc{l}}\) (much smaller than the volume of the order of \(10^{-15}\;{\textsc{l}}\) for a typical bacterial cell) that for practical purposes the self-maintaining state of non-null concentrations becomes significant, recalling the situation of bistability that is observed in deterministic analysis. In turn, there exists a minimum size below which the self-organizing system cannot maintain itself on chemically relevant time scales. The value of the critical volume depends on the particular concentrations and rate constants assumed, but the principle could apply generally.
J. Merelo合作论文数Dept. of Computer Technology and Architecture;Universidad de Granada3