
The evolution of the human brain is characterized by profound changes in structure and function, despite relatively limited divergence in protein-coding genes compared to other primates. This paradox has led to increasing recognition of gene regulatory elements (GREs) as primary drivers of evolutionary innovation. In this review, we synthesize current knowledge on the role of conserved noncoding elements (CNEs), human accelerated regions (HARs), and transposable element (TE)-derived sequences in shaping gene regulatory networks (GRNs) underlying brain development. Comparative analyses across humans and closely related primates, including the chimpanzee, gorilla, and orangutan, reveal that while core regulatory architectures are highly conserved, subtle changes in regulatory elements drive species-specific gene expression patterns. We highlight how CNEs provide a stable regulatory framework, whereas HARs and TE-derived elements introduce lineage-specific modifications that fine-tune neurodevelopmental processes. Advances in functional genomics, including CRISPR-based perturbations, massively parallel reporter assays, and single-cell multi-omics, have enabled direct interrogation of regulatory function, linking sequence variation to cellular phenotypes. Furthermore, we discuss how regulatory evolution contributes to both cognitive innovation and susceptibility to neurological disorders. Despite significant progress, challenges remain in establishing causal relationships between regulatory variation and phenotypic outcomes. Future integration of multi-omics data and comparative models will be essential for resolving these complexities. Together, this review provides a comprehensive framework for understanding the molecular basis of primate brain evolution through the lens of gene regulation.
Reduced oxidation state phosphorus species, such as phosphite, have been proposed as more soluble phosphorus sources, yet their role in phosphorylation chemistry remains incompletely understood. Here, we investigate the concurrent oxidation of phosphite and phosphorylation/phosphonylation of adenosine under prebiotically plausible conditions. Using aqueous reaction systems containing phosphite, urea, hydrogen peroxide (H₂O₂), and adenosine, we show synthesis of adenosine phosphates, adenosine phosphites, adenosyl adenosine phosphate, and trace amounts of ADP (adenosine diphosphate). Reactions conducted under warm evaporating pool conditions (75–78 °C) yielded up to 73
The emergence of the first nucleic acids required the prebiotic formation of nucleosides and nucleotides under chemically challenging conditions. Because assembly of the canonical ribose–phosphate framework is disfavored in water, simpler ancestral backbones may have preceded RNA and DNA. Semiempirical prescreening and density functional theory (DFT) calculations are used to evaluate the thermodynamics of glycerol-based nucleosides and nucleotides as possible proto-nucleic-acid building blocks. Two assembly routes were examined: a classic pathway, in which glycerol first condenses with a recognition unit and then with an ionized linker, and an alternative pathway, in which glycerol first condenses with the ionized linker and only then with the recognition unit. Although Gibbs free energy is a state function, the two pathways access different regions of the potential energy surface and converge to distinct local minima, leading to pathway-dependent differences in the computed thermodynamic quantities. Glycerol-derived nucleosides are found to form favorably in both vacuum and implicit aqueous solution when combined with either canonical or the studied non-canonical bases. Glycerol-based nucleotides are, likewise, thermodynamically accessible by both pathways, although the classic route is consistently more favorable than the alternative route. Among the most stabilized products are derivatives containing adenine, uracil, and C-glycosylated barbituric acid. Replacement of phosphate by arsenate produces only modest energetic and structural changes, indicating that arsenate-containing analogues cannot be excluded on thermodynamic grounds alone. However, the longer As–OC3 bond relative to P–OC3 suggests a weaker arsenate-based backbone compared to the predominant present-day phosphate-based backbone and therefore a plausible hydrolytic disadvantage for arsenate. These results support the view that glycerol-based backbones could have participated in early proto-nucleic-acid chemistry and suggest that phosphate may have been selected not because arsenate analogues were thermodynamically unstable, but because phosphate-based backbones were more kinetically persistent and more resistant to hydrolysis than arsenate-based ones.
India’s indigenous pig genetic resources are central to sustainable meat production and rural livelihoods, yet their genomic potential remains underexplored. The Ghurrah pig, a native breed from northern India, is among the most genetically distinct and adaptively resilient populations, shaped by ancestral admixture from Asian, European, and Indian wild boar lineages. This study elucidates the genetic uniqueness, population structure, and adaptive introgression of the Ghurrah pig by jointly analysing 35,777 high-quality SNPs across 270 global pig genotypes. Higher observed heterozygosity (Ho = 0.287) and lower inbreeding coefficient (FIS = 0.206) in Ghurrah indicated substantial genetic variability and minimal inbreeding, while its effective population size (Ne = 26 at 13 generations ago) revealed demographic constraint. Linkage disequilibrium (r² = 0.325) and haplotype block structure (average block size = 415.56 kb) reflected reduced recombination and hybrid ancestry. Admixture analysis at K = 5 positioned Ghurrah intermediate between Asian and European clusters, with 39.07
The phenotypic diversity between castes in social species has been associated with increased regulation of gene expression. while the regulatory role of transcription factors (TFs) in the social transitions of Hymenoptera (bees, ants and wasps) is well studied, it remains poorly explored in Blattodea (cockroaches and termites). By analysing three cockroach and five termite genomes, we show that while Blattodea share some convergent regulatory signatures with Hymenoptera, their transition is mostly characterised by distinct regulatory changes. We identified convergent patterns such as TF families with more relaxed selection compared to intensified selection and lineage-specific gene family expansions in termites, which has also been reported in Hymenoptera. There are also key differences in Blattodea TF regulation in comparison with Hymenoptera with contractions in TF gene families and no compensatory changes in TF DNA binding motifs either in frequency or diversity in TF promoter regions. Furthermore, we show that an increase in social complexity associates with greater diversity in TF activating domains, one of the evolutionary and structural building blocks of TFs; meanwhile, DNA-binding domains, undergo very little change. This study highlights similarities in social transitions between Hymenoptera and Blattodea, with evidence of large changes in transcriptional regulation followed by lineage specific adaptations. Our results indicate that at least in the case of Blattodea, the transcriptional diversity linked to social complexity is not attributable to transcription factors alone, but is instead likely driven in combination with alternative mechanisms.
Charged Clusters (CC), subclasses of Intrinsically Disordered Protein Regions (IDPR), are vital for protein function by mediating electrostatic interactions, metal binding, and serving as hotspots for post-translational modifications. The evolutionary and functional constraints driving the Conserved CC (CCC) and their role as potentially important regions within IDPR are poorly understood. This study aimed to identify and analyze Positive and Negative CCC (PCCC and NCCC) across green plant species. Conservation rates for non-redundant CCC ranged between 9.8
Cetaceans and sirenians independently transitioned from land to water, evolving unique and convergent sensory adaptations shaped by aquatic environments. Among sensory receptors, the Transient Receptor Potential (TRP) channel superfamily is central to thermo-, chemo-, and mechanosensation, but its evolutionary history in fully aquatic mammals remains poorly characterized. Here, we investigated the molecular evolution of TRP channels in these lineages. Orthology and phylogenetic relationships were inferred using Maximum Likelihood and Bayesian approaches. Signals of positive selection and molecular convergence were evaluated with codon and amino acid models. Amino acid substitutions, protein structure, and stability were assessed using 3D protein modeling. Our analyses reveal accelerated evolutionary rates in aquatic mammals, including multiple positively selected sites, lineage-specific amino acid substitutions, and convergent evolution across cetaceans and sirenians. Protein-level assessments identified substitutions with potential functional consequences, and evidence of pseudogenization was detected in cetacean PKD1L3, PKD2L1, TRPA1, and TRPM5, in contrast to intact copies in sirenians. These patterns suggest lineage-specific sensory trajectories, including reduced chemosensory repertoires in cetaceans, conservation of taste-related genes in sirenians, and adaptations in somatosensory associated genes that reflect both convergent requirements of fully underwater living and distinct aquatic environments. Overall, our findings advance understanding of the molecular mechanisms underlying sensory evolution during the land-to-water transition in mammals.
Synonymous codon sites were historically considered largely neutral; however, accumulating evidence indicates that codon usage may influence gene regulation, translation dynamics, and evolutionary fitness. In plants, gene body methylation (gbM) and chromatin-based stress memory represent stable epigenetic features whose underlying sequence determinants remain poorly understood. Here, we propose that synonymous codon architecture may contribute to chromatin stability in plant genomes through two complementary mechanisms. First, codon choice may influence the density and distribution of CG dinucleotides within coding regions, thereby shaping the substrate landscape for CG gene body methylation maintained by DNA methyltransferases. Second, synonymous codon usage may influence translation elongation kinetics, which may affect co-translational folding and post-translational modification dynamics of chromatin-associated proteins, particularly histones. Because histone proteins are highly conserved at the amino acid level, synonymous variation in histone genes may provide a powerful model system to detect regulatory constraints acting on coding sequences. We propose that selection acting on synonymous codons may therefore contribute not only to translational efficiency but also to the stability of epigenetic states and chromatin regulation. This perspective integrates codon usage bias, plant epigenomics, and chromatin biology and generates testable predictions regarding how coding sequence composition may influence gene body methylation patterns, chromatin stability, and long-term epigenetic inheritance.
The transition from marine to freshwater environments represents a remarkable evolutionary shift in cetaceans, yet the genomic underpinnings of this adaptation remain poorly understood. This study investigates the genomic signatures of freshwater adaptation in riverine cetaceans through comparative evolutionary analyses of six species: Inia geoffrensis, Lipotes vexillifer, Platanista gangetica, Platanista minor, Neophocaena asiaeorientalis asiaeorientalis, and Sotalia fluviatilis. By integrating positive selection analysis, gene family dynamics, and evolutionary rate convergence, we identified key molecular adaptations associated with freshwater colonization. The analysis revealed that five of the six species exhibited positive selection in the NSMAF and CTRL genes, suggesting widespread selective pressures related to inflammatory responses and digestive adaptations, respectively. Functional enrichment analyses revealed adaptive signatures in hematopoiesis, osmoregulation, skeletal development, and immune responses, reflecting the physiological challenges of freshwater environments. Gene family evolution analyses using CAFE identified dynamic patterns of expansions and contractions in immune-related genes, transcriptional regulation, and cell adhesion pathways across riverine lineages. Relative evolutionary rate (RER) analysis using RERconverge identified 95 genes showing convergent rate shifts associated with freshwater adaptation, including genes involved in cellular nitrogen compound responses and transcriptional regulation. Despite positively selected genes overlap was limited and did not follow simple clade-wide patterns, our results demonstrate that freshwater adaptation in cetaceans involves putative convergent evolution of fundamental biological systems, including immune responses, metabolic regulation, and morphological development. These findings provide new insights into the molecular mechanisms underlying speciation in aquatic mammals and highlight critical biological pathways that have enabled the successful colonization of freshwater ecosystems by multiple independent cetacean lineages.
In origins-of-life research, a key challenge is to explain the emergence of polymers of sufficient length to confer the complex functions needed for genetic inheritance. Previous studies have demonstrated that prebiotic environments enabling multilevel selection can facilitate the survival of cooperating polymers such as ribozymes, which allows that complex functions undertaken today by long polymers might have originated from multiple simpler functions undertaken by shorter polymers. To further investigate this possibility, we developed a new computational model of cooperative catalytic and replicating polymer systems that avoids tracking all possible polymer sequences. This approach scales well computationally, avoiding the need to set an a priori cap on polymer length. We first validated this model by replicating key conclusions of previous studies, for example that the persistence of cooperative synthetase-ligase systems is facilitated by both intrinsic factors (shorter length, higher catalytic efficiency) and by factors that promote multilevel-selection (compartmentalization, slower diffusion). We then explored the effects of introducing a mutation inhibitor into a cooperative synthetase-ligase system. The results support the possibility that mutation inhibition could have arisen, not through the appearance of a single proofreading polymerase, but through the emergence of distinct, mutation inhibiting catalysts.
Nup98 is a key component of the nuclear pore complex that plays essential roles in nucleocytoplasmic transport and gene regulation. To investigate the evolutionary diversification of Nup98 in ciliates, homologous sequences were collected from genomic and transcriptomic databases representing major ciliate lineages. Most species possessed multiple Nup98 paralogs, indicating that gene duplication events have occurred repeatedly during ciliate evolution. Phylogenetic analyses based on the conserved Nucleoporin2 domain revealed that Nup98 proteins from each ciliate class tend to form distinct monophyletic groups, suggesting lineage-specific diversification patterns. These results are consistent with a scenario in which early ciliates possessed a limited number of Nup98 genes, followed by independent duplication and diversification events in multiple lineages. However, alternative evolutionary scenarios remain possible. Differences in repeat motif composition were observed among paralogs across lineages, although these patterns were not conserved across all ciliates. While such variation may suggest potential functional diversification, phylogenetic evidence alone does not demonstrate functional differences among paralogs. Overall, the present diversity of Nup98 genes in ciliates is best interpreted as the result of lineage-specific evolutionary processes. These findings provide a framework for understanding Nup98 evolution and generate testable hypotheses regarding its potential roles in the diversification of nuclear dualism.
Substitution rate estimates are central to evolutionary biology, underpinning divergence-time inference and a wide range of macroevolutionary analyses. Mitochondrial DNA (mtDNA) rates are widely used for this purpose, yet they are often derived from a limited set of genes, closely related taxa, or a small number of model organisms. Here, we use nearly complete mitogenomes from 27 pleurodontan species (Squamata: Pleurodonta) to estimate substitution rates across the mitochondrial genome, explicitly evaluating the effects of data partitioning, calibration strategies, and model specification. Bayesian analyses revealed pronounced heterogeneity in substitution rates among codon positions and between coding and non-coding regions. Estimated rates ranged from approximately 0.004 to 0.02 substitutions per site per million years, consistent with previous lineage-specific estimates. Commonly used rates closely matched those estimated for third codon positions and for analyses based on combined partitions, suggesting that widely adopted values may primarily reflect signals from faster-evolving sites or aggregated partitioning schemes. Calibrated analyses generally yielded lower substitution rate estimates with reduced variance relative to non-calibrated analyses. However, substantial overlap in 95% highest posterior density intervals indicates limited evidence for systematic differences between these approaches, suggesting that much of the relative rate structure is already captured by the molecular data under a partitioned relaxed-clock framework. Comparisons between alternative partitioning strategies further showed that, although data-driven schemes recover broad patterns of rate variation, they may do so at the cost of reduced parameter resolution, particularly in the absence of calibration. Together, these results highlight that substitution rate estimates are sensitive to partitioning and modeling choices, and that model complexity should be evaluated in terms of parameter resolution and data informativeness rather than parameter count alone. By providing partition-specific rate estimates across the mitogenome, this study offers a robust empirical framework for improving molecular dating and evolutionary inference in squamates and other non-model systems.
Non‑pharmaceutical interventions such as contact tracing, quarantine, and targeted restrictions remain central to outbreak control. Yet, their success depends on understanding how pathogens spread through heterogeneous populations. Here, I highlight an innovative network‑informed Bayesian framework that integrates genomic and contact data across time to reconstruct transmission pathways more accurately (Xu et al. 2026). By modeling network structure as a prior, the approach captures individual‑level heterogeneity and resolves ambiguities that arise when genetic data are sparse . Notably, this framework captures genomic variation using mutation loci rather than complete genomes preserves evolutionary signal while greatly reducing computational demands, and also enables more efficient integration of key metadata. Together, these developments provide a critical step toward designing interventions that more precisely disrupt transmission while minimizing social and economic costs.
The alignment of multiple protein or DNA sequences is a common task in many biological fields. These alignments may then be used to identify conserved regions or positions, which could in turn be used to score the likely pathogenicity of a sequence change in a patient, detect functional domains that can determine the activity of an unknown gene, or map evolutionary relationships by tracking changes in conserved regions. The development of high-throughput sequencing technologies has trivialised the generation of sequencing data, which is made available through public repositories, such as GenBank, hosted by the NCBI. In October 2024, GenBank held approximately 4.7 billion sequences from over 580,000 species. Consequently, the constraints of multiple sequence alignment estimation have moved from data generation to data retrieval, extraction and filtering. While human-curated datasets for widely used sequences have been released, retrieving datasets for less commonly used sequences can be onerous. Therefore, we have developed GeneMatrix, an application that aids the extraction, filtering, aggregation, and alignment of DNA and protein sequences from GenBank-formatted files of single gene sequences or genomes such as viruses and mitochondria.
Substitution model selection is central to phylogenetic inference and is commonly treated as a problem of identifying the substitutional complexity required to describe sequence evolution along a single tree. This framework implicitly assumes a shared genealogy across all sites, an assumption that is routinely violated in phylogenomic data by incomplete lineage sorting and other sources of gene-tree discordance. Recent work by Lozano et al. (2026) demonstrates that unmodeled genealogical heterogeneity can systematically distort substitution model selection, creating spurious support for parameter-rich models even when the underlying substitution process is simple. In this perspective, we examine the consequences of this confounding for phylogenetic estimation, divergence-time inference, and the biological interpretation of substitution model parameters, highlighting the need to more explicitly integrate genealogical heterogeneity into model selection and adequacy assessment in phylogenomics. This perspective reframes substitution model choice not only as a tool for describing molecular evolution, but also as a potential diagnostic for violations of shared-genealogy assumptions that are central to modern systematics.
Transposable elements (TEs) are dynamic DNA sequences that play a significant role in shaping genome structure and function in eukaryotic species. Advances in next-generation sequencing technologies have enhanced our understanding of the abundance and diversity of transposable element families. Transcriptionally active TEs contribute to intra-species genetic variability and facilitate adaptation to environmental stressors, such as heat, drought, and salinity, by inducing mutations, modulating gene expression, and promoting genome rearrangements. Recent studies highlight the important role of horizontal transfer and vertical transmission mechanisms in the evolution of Class I and Class II TE families. The Opie and Ji families of LTR elements serve as examples of conserved TEs that contribute to the expansion of the maize genome. In contrast to RIRE1, which remains relatively stable, Tos17 is largely inactive under normal conditions but can be activated under stress, such as tissue culture, thereby contributing to genome dynamics. This review explores key examples of horizontal transfer and vertical transmission of TEs in plant species, along with their structural features, evolutionary trajectories, and divergence patterns.
Transcriptional adaptation (TA) works through a mechanism in which ILF3 binds degradation fragments of aberrant mRNAs and upregulates a sequence-related “adapting gene” for expressional compensation. It is worth thinking whether the entire TA mechanism can be regarded as adaptive given its potential off-targeting, sensitivity to mutations, reliance on paralogous genes, and the counterproductive self-TA. Here, we ask how exactly adaptation should be defined. We need to distinguish between adaptation at the case-study level versus the global-mechanism level, possibly through the lens of drift-barrier hypothesis. A few beneficial cases cannot justify the adaptation of the entire mechanism. Using RNA editing as an example, we illustrate that many functional molecular processes may actually arise through constructive neutral evolution (CNE) and be bufferred by genetic capacitors, a scenario echoing the spandrels of San Marco and the molecular error hypothesis. The existence of such mechanisms permitted the accumulation of otherwise deleterious mutations. We propose that functionality is a necessary but not sufficient condition for adaptation. Functional cases should not be automatically extended to adaptation unless its origin and evolution have been investigated from an omics angle. Our perspectives provide insights into the origins of biological mechanisms and the nature of evolution.
Phylogenetic conflict—particularly hidden gene-tree discordance generated by incomplete lineage sorting (ILS)—is pervasive in multilocus and phylogenomic datasets, yet its consequences for nucleotide substitution model selection remain poorly understood. Modern molecular studies increasingly collect and concatenate large sets of independent loci sampled across distant and often poorly characterized regions of the genome, creating significant potential for heterogeneity when analyzed in combination. Here, we examine whether intra-alignment genealogical conflict can influence standard model selection procedures to favor parameter-rich substitution models even when sequences evolve under a simple substitution process. Through a series of in silico case studies, we simulated sequence evolution under the simplest rate-homogeneous Jukes–Cantor (JC69) model and generated concatenated alignments as mosaics of multiple loci, each evolving on its own gene tree drawn under the multispecies coalescent. Conflict was increased by manipulating conditions expected to elevate ILS and gene-tree heterogeneity and embedding progressively more hidden genealogies within alignments while holding total alignment length constant. Despite all data being generated under JC69, model selection frequently favored more complex models, with varying sensitivity depending on the number of taxa, the expected amount of conflict, and the specific selection criterion applied. A dominant pattern was frequent inclusion of among-site rate variation parameters (+ G4 and/or + I), and under extreme conflict, model selection increasingly favored richer substitution models (e.g., SYM, GTR). Broadly, our results showed that hidden conflict can manifest as substitutional and rate heterogeneity, driving selection procedures to compensate with additional parameters in concatenated analyses under high conflict. Broadly, our study contributes to a greater understanding and appreciation of the challenges in modeling molecular evolution in the era of multilocus phylogenetics.
Synaptic vesicle proteins, including the synaptophysin, synaptogyrins, and synaptic vesicle glycoprotein 2 family are fundamental for neurotransmitter release and synaptic function, influencing numerous physiological processes. Although these proteins hold promise as therapeutic targets, their study has remained complicated due to their location within the cell membrane. To tackle this, we performed comparative analyses on these proteins and their water-soluble variants, which were designed using the QTY code. This approach involves systematically replacing hydrophobic amino acids L (leucine), V/I (valine/isoleucine), and F (phenylalanine) with hydrophilic amino acids Q (glutamine), T (threonine), and Y (tyrosine). The water-soluble QTY variants generated in our study, despite having significant differences in their transmembrane sequences up to 55
This work is presented as a hypothesis-driven review article, offering a conceptual framework that integrates prebiotic chemistry, simulation data, and modern peptide analogues to propose β-strand oligomers as a plausible solution to the protocell permeability problem. A persistent challenge in origins-of-life research is explaining how primitive fatty-acid vesicles could exchange molecules with their surroundings. While such compartments assemble readily under prebiotic conditions, their permeability to nucleotides, peptides, and divalent ions is generally limited and condition-dependent, posing a challenge for sustained chemical exchange and protocell growth. I evaluate the hypothesis that short, abiotically produced β-strand peptides could have modulated protocell permeability. Unlike α-helical pore formers, which require helix-stabilizing residues rare in prebiotic syntheses, β-strand oligomers are strongly favored by the glycine-, alanine-, and valine-rich inventories documented in meteorites and laboratory experiments. Taken together, these results suggest that irregular β-sheet oligomers may have provided the first, non-selective permeability in fatty-acid protocells. This hypothesis is testable; Gly/Ala/Val-rich peptides synthesized under prebiotic cycling conditions should measurably increase leakage across fatty-acid vesicles, providing a tractable experimental framework for probing early membrane function.