Small peptides are an important component of the vertebrate immune system. They are important molecules for distinguishing proteins that originate in the host from proteins derived from a pathogenic organism, such as a virus or bacterium. Consequently, these peptides are central for the vertebrate host response to intracellular and extracellular pathogens. Computational models for prediction of these peptides have been based on a narrow sample of data with an emphasis on the position and chemical properties of the amino acids. In past literature, this approach has resulted in higher predictability than models that rely on the geometrical arrangement of atoms. However, protein structure data from experiment and theory are a source for building models at scale, and, therefore, knowledge on the role of small peptides and their immunogenicity in the vertebrate immune system. The following sections introduce procedures that contribute to theoretical prediction of peptides and their role in immunogenicity. Lastly, deep learning is discussed as it applies to immunogenetics and the acceleration of knowledge by a capability for modeling the complexity of natural phenomena.
Tokenization is a procedure for recovering the elements of interest in a sequence of data. This term is commonly used to describe an initial step in the processing of programming languages, and also for the preparation of input data in the case of artificial neural networks; however, it is a generalizable concept that applies to reducing a complex form to its basic elements, whether in the context of computer science or in natural processes. In this entry, the general concept of a token and its attributes are defined, along with its role in different contexts, such as deep learning methods. Included here are suggestions for further theoretical and empirical analysis of tokenization, particularly regarding its use in deep learning, as it is a rate-limiting step and a possible bottleneck when the results do not meet expectations.
In deep learning, large language models are typically trained on data from a corpus as representative of current knowledge. However, natural language is not an ideal form for the reliable communication of concepts. Instead, formal logical statements are preferable since they are subject to verifiability, reliability, and applicability. Another reason for this preference is that natural language is not designed for an efficient and reliable flow of information and knowledge, but is instead designed as an evolutionary adaptation as formed from a prior set of natural constraints. As a formally structured language, logical statements are also more interpretable. They may be informally constructed in the form of a natural language statement, but a formalized logical statement is expected to follow a stricter set of rules, such as with the use of symbols for representing the logic-based operators that connect multiple simple statements and form verifiable propositions.
EDITORIAL article Front. Genet., 02 June 2023Sec. Neurogenomics Volume 14 - 2023 | https://doi.org/10.3389/fgene.2023.1220750
Nature is composed of elements at various spatial scales, ranging from the atomic to the astronomical level. In general, human sensory experience is limited to the mid-range of these spatial scales, in that the scales which represent the world of the very small or very large are generally apart from our sensory experiences. Furthermore, the complexities of Nature and its underlying elements are not tractable nor easily recognized by the traditional forms of human reasoning. Instead, the natural and mathematical sciences have emerged to model the complexities of Nature, leading to knowledge of the physical world. This level of predictiveness far exceeds any mere visual representations as naively formed in the Mind. In particular, geometry has served an outsized role in the mathematical representations of Nature, such as in the explanation of the movement of planets across the night sky. Geometry not only provides a framework for knowledge of the myriad of natural processes, but also as a mechanism for the theoretical understanding of those natural processes not yet observed, leading to visualization, abstraction, and models with insight and explanatory power. Without these tools, human experience would be limited to sensory feedback, which reflects a very small fraction of the properties of objects that exist in the natural world. As a consequence, as taught during the times of antiquity, geometry is essential for forming knowledge and differentiating opinion from true belief. It not only provides a framework for understanding astronomy, classical mechanics, and relativistic physics, but also the morphological evolution of living organisms, along with the complexities of the cognitive systems. Geometry also has a role in the information sciences, where it has explanatory power in visualizing the flow, structure, and organization of information in a system. This role further impacts the explanations of the internals of deep learning systems as developed in the fields of computer science and engineering.
Supplementary Figures S1-S2 from Evaluation of Lapatinib and Topotecan Combination Therapy: Tissue Culture, Murine Xenograft, and Phase I Clinical Trial Data
The nematode worm Caenorhabditis elegans has a relatively simple neural system for analysis of information transmission from sensory organ to muscle fiber. Consequently, this study includes an example of a neural circuit from the nematode worm, and a procedure is shown for measuring its information optimality by use of a logic gate model. This approach is useful where the assumptions are applicable for a neural circuit, and also for choosing between competing mathematical hypotheses that explain the function of a neural circuit. In this latter case, the logic gate model can estimate computational complexity and distinguish which of the mathematical models require fewer computations. In addition, the concept of information optimality is generalized to other biological systems, along with an extended discussion of its role in genetic-based pathways of organisms.
This review is of basic models of the interactions between a pathogenic virus and vertebrate animal host. The interactions at the population level are described by a predatory-prey model, a common approach in the ecological sciences, and depend on births and deaths within each population. This ecological perspective is complemented by models at the genetical level, which includes the dynamics of gene frequencies and the mechanisms of evolution. These perspectives are symmetrical in their relatedness and reflect the idealized forms of processes in natural systems. In the latter sections, the general use of deep learning methods is discussed within the above context, and proposed for effective modeling of the response of a pathogenic virus in a pathogen–host system, which can lead to predictions about mutation and recombination in the virus population.
The nematode worm Caenorhabditis elegans has a relatively simple neural system for analysis of information transmission from sensory organ to muscle fiber. Therefore, an example of a neural circuit is analyzed that originates in the nematode worm, and a method is applied for measuring its information flow efficiency by use of a model of logic gates. This model-based approach is useful where the assumptions of a logic gate design are applicable. It is also an useful approach where there are competing mathematical models for explaining the role of a neural circuit since the logic gate model can estimate the computational complexity of a network, and distinguish which of the mathematical models require fewer computations. In addition, for generalization of the concept of information optimality in biological systems, there is an extensive discussion of its role in the genetic-based pathways of organisms.
This editorial addresses the universality and importance of the science of perception. In particular, recently published studies in this journal illustrate the natural variations in perception. These articles are a reminder of perception as a natural process with inherent variations and that any two individuals are not guaranteed to form the same representation of an object, regardless of whether it originates from the senses or not. Since perception is a foundation for higher cognition, it also has an immense influence on studies of humanity and interpretations of natural processes.
An animal neural system ranges from a cluster of a few neurons to a brain of billions. At the lower range, it is possible to test each neuron for its role across a set of environmental conditions. However, the higher range requires another approach. One method is to disentangle the organization of the neuronal network. In the case of the entorhinal cortex in a rodent, a set of neuronal cells involved in spatial location activate in a regular grid-like arrangement. Therefore, it is of interest to develop methods to find these kinds of patterns in a neural network. For this study, a square grid arrangement of neurons is quantified by network metrics and then applied for identification of square grid structure in areas of the fruit fly brain. The results show several regions with contiguous clusters of square grid arrangements in the neural network, supportive of specialization in the information processing of the system.
Cognition is the acquisition of knowledge by the mechanical process of information flow in a system. In cognition, input is received by the sensory modalities and the output may occur as a motor or other response. The sensory information is internally transformed to a set of representations, which is the basis for downstream cognitive processing. This is in contrast to the traditional definition based on mental processes, a phenomenon of the mind that originates in past ideas of philosophy.
The nematode worm Caenorhabditis elegans is a model for deciphering the neural circuitry that transmits information from sensory organ to muscle tissue. It is also studied for disentangling the characteristics of the network, the efficiency of its design, and for testing theoretical models on how information is encoded. For this study, the efficiency of the synaptic connections was studied by testing the robustness of the neural network. A randomization test of robustness was applied to previously computed neural modules of the pharynx of C. elegans. The results support robustness as a reason for the observed over connectiveness across the pharyngeal system. In addition, rare events of single-neuron loss may expectedly lead to loss of function in a neural system.
Cognition is often defined as a dual process of physical and non-physical mechanisms. This duality originated from past theory on the constituent parts of the natural world. Even though material causation is not an explanation for all natural processes, phenomena at the cellular level of life are modeled by physical causes. These phenomena include explanations for the function of organ systems, including the nervous system and information processing in the cerebrum. This review restricts the definition of cognition to a mechanistic process and enlists studies that support an abstract set of proximate mechanisms. Specifically, this process is approached from a large-scale perspective, the flow of information in a neural system. Study at this scale further constrains the possible explanations for cognition since the information flow is amenable to theory, unlike a lower-level approach where the problem becomes intractable. These possible hypotheses include stochastic processes for explaining the processes of cognition along with principles that support an abstract format for the encoded information.
Here is a review of several empirical examples of information processing that occur in the primate cerebral cortex. These include visual processing, object identification and perception, information encoding, and memory. Also, there is a discussion of the higher scale neural organization, mainly theoretical, which suggests hypotheses on how the brain internally represents objects. Altogether they support the general attributes of the mechanisms of brain computation, such as efficiency, resiliency, data compression, and a modularization of neural function and their pathways. Moreover, the specific neural encoding schemes are expectedly stochastic, abstract and not easily decoded by theoretical or empirical approaches.
A computational toolset was developed to automate use of the web database at NeuroMorpho.Org. It is platform independent and includes code to retrieve curated neuronal data in the JSON file format. The source of the data is from past studies which is accessible by the client–server HTTP protocol. Tools include methods to convert the database format to plain text, combine tables, and validate files.
Neuron morphology is highly variable across the mammalian brain. It is thought that these attributes of neuronal cell shape, such as soma surface area and branching frequency, are determined by biological function and information processing. In this study, a large data set of neurons across the rat neocortex were clustered by their anatomical characters for evidence of distinctiveness among neocortical regions and the somatosensory layers. This data set of neuronal morphologies was compiled from 31 different lab sources with a validation procedure so that data records are potentially comparable across research studies. With this large set of heterogeneous data and by clustering analysis, this study shows that neuronal morphological traits overlap among neocortical and somatosensory regions. In the context of past neuroanatomical studies, this result is not congruent with tissue level analysis and strongly suggests further sampling of neuronal data to lessen the effect of confounding factors, such as the influence of different methodologies from use of heterogeneous samples of neuronal data.
The somatic nervous system of the nematode worm Caenorhabditis elegans is a model for understanding the physical characteristics of the neurons and their interconnections. Its neurons show high variation in morphological attributes. This study investigates the relationship of neuronal morphology to the number of synapses per neuron. Morphology is also examined for any detectable association with neuron cell type or ganglion membership.
Phylogenetic studies aim to discover evolutionary relationships and histories. These studies are based on similarities of morphological characters and molecular sequences. Currently, widely accepted phylogenetic approaches are based on multiple sequence alignments, which analyze shared gene datasets and concatenate/coalesce these results to a final phylogeny with maximum support. However, these approaches still have limitations, and often have conflicting results with each other. Reconstructing ancestral genomes helps us understand mechanisms and corresponding consequences of evolution. Most existing genome level phylogeny and ancestor reconstruction methods can only process simplified real genome datasets or simulated datasets with identical genome content, unique genome markers, and limited types of evolutionary events. Here, we provide an alternative way to resolve phylogenetic problems based on analyses of real genome data. We use phylogenetic signals from all types of genome level evolutionary events, and overcome the conflicting issues existing in traditional phylogenetic approaches. Further, we build an automated computational pipeline to reconstruct phylogenies and ancestral genomes for two high-resolution real yeast genome datasets. Comparison results with recent studies and publications show that we reconstruct very accurate and robust phylogenies and ancestors. Finally, we identify and analyze the conserved syntenic blocks among reconstructed ancestral genomes and present yeast species.