
This review discusses recent computational approaches aimed at enhancing the precision and efficiency of the CRISPR-Cas9 gene editing system. These approaches leverage data-driven biophysical modeling, integrating insights from high-throughput experimental datasets and atomistic molecular dynamics simulations to elucidate the underlying molecular mechanisms. We evaluate these computational frameworks in terms of their ability to accurately predict CRISPR-Cas9 cleavage efficiency. Importantly, we highlight how the synergy between computational modeling and experimental validation accelerates the development of more robust predictive tools, while also minimizing the reliance on costly trial-and-error strategies in single-guide RNA (sgRNA) design. Looking ahead, the implementation of closed-loop feedback systems, where computational predictions guide experiments and experimental outcomes refine models, will be essential for realizing the full potential of CRISPR-based therapeutics and applications.
Antibodies are key immune system proteins that recognize and neutralize antigens through their variable domains, which contain complementarity-determining regions (CDRs). Knowledge of antibody-antigen structures is essential for understanding immune recognition and guiding therapeutic antibody development. High-resolution experimental methods such as X-ray crystallography, NMR, and cryo-EM are resource-intensive and impractical for all antibodies due to the vast diversity of antigen-specific antibodies. As a result, computational structure prediction methods have become increasingly valuable for modeling antibody-antigen complexes. This study provides an overview of commonly used computational approaches for predicting antibody-antigen structures, including molecular docking and recent machine learning-driven approaches.
With the rapid development of genome sequencing technology, genomic sequence analysis has become an important field in modern biological research. However, sequencing errors, repetitive regions, and complex biological processes often lead to missing or ambiguous bases in genomic sequences, which are typically represented by non-standard symbols (such as R, Y, S, W, K, etc.). These issues severely affect the accuracy of genomic data, especially in tasks such as gene assembly and variant detection. To address this issue, this study proposes an encoding method based on asymmetric covariance natural vectors to characterize genomic sequences and predict ambiguous bases using Gated Recurrent Unit (GRU). Experimental results demonstrate that, compared with traditional encoding methods (such as one-hot encoding), the asymmetric covariance natural vector can more effectively utilize the information surrounding missing nucleotides for prediction, showing significant advantages in recovering nucleotides missing from intermediate positions. Additionally, this method also performs well on the SARS-CoV-2 Alpha variant dataset, with an error rate of only 1.09% in predicting non-standard bases during the encoding recovery process, further validating its effectiveness and potential for practical genomic data analysis.
Estimation of stochastic systems on very large networks is intractable computationally. Graphon theory provides limit objects for infinite sequences of graphs by mapping adjacency matrices to the unit square, enabling the modelling of dynamical systems on arbitrarily large graphs via functional analytic methods. In previous work (Dunyak and Caines, 2022, 2023, 2024), Q-noise was used to extend stochastic systems on large graphs to stochastic systems in Hilbert spaces on graphons. In this paper, the linear system state estimation problem on large networks and their graphon limits are analyzed, and the Separation Principle of control and estimation is introduced. Convergence of finite network linear system state estimates and the corresponding Kalman filter systems to their graph limit counterparts is established. Subject to the new hypothesis of Q-observability, the Separation Principle is extended to the case of control via finite-dimensional approximating subsystems. This control method is demonstrated on an Erdos-Renyi graph system.
Networked industrial systems involve many subsystems whose local controllers must collaborate and collectively achieve stability and optimal control objectives. These systems and control problems are broadly exemplified by frequency/voltage regulations in modern power grids, load balancing in networked computing, team formation and movements in autonomous teams, temperature and energy management in smart buildings, distribution in water and energy networks, among many others. In many of these applications, the maximum control values are limited by random environmental and operational conditions, such as wind speed, solar radiation, vehicle battery states of charge, control ranges of appliances, etc. Such stochastic bounds create random saturations and have a substantial detrimental impact on system stability, performance, and optimality. Focusing on coordinated linear quadratic (LQ) optimal control problems with multiple local controllers, this paper introduces new adaptation methods to deal with stochastic saturations for enhanced stability and performance. New algorithms are introduced and their main properties are established. Examples and case studies are used to illustrate the advantages of the new methods over fixed (non-adaptive) control strategies.
Filtering is a subject of providing sequential estimations of a given stochastic dynamical system based on noisy observations. At the beginning of this century, a two-stage algorithm framework was proposed and analyzed for general nonlinear filtering problems, which is now referred to as the Yau-Yau algorithm. The two-stage structure of this framework theoretically guarantees the potential of solving nonlinear filtering problems in a real-time manner, which is crucial for practical applications. With the introduction of spectral method and neural networks, numerous model-driven and data-driven implementations of Yau-Yau algorithms have been proposed in the last decade. In this paper, we will present a thorough review of the development of nonlinear filters under the Yau-Yau algorithm framework, from model-driven approaches to data-driven approaches, which serves as a guidance for practitioners. Current status and promising future directions in the research of Yau-Yau nonlinear filtering algorithm framework are also summarized and discussed in this paper.
This paper presents a systematic review of recent advances in nonlinear filtering algorithms, structured into three principal categories: Kalman-type methods, Monte Carlo methods, and the Yau-Yau algorithm. For each category, we provide a comprehensive synthesis of theoretical developments, algorithmic variants, and practical applications that have emerged in recent years. Importantly, this review addresses both continuous-time and discrete-time system formulations, offering a unified review of filtering methodologies across different frameworks. Furthermore, our analysis reveals the transformative influence of artificial intelligence breakthroughs on the entire nonlinear filtering field, particularly in areas such as learning-based filters, neural network-augmented algorithms, and data-driven approaches.
Sequential Monte Carlo (SMC) methods have recently shown successful results for conditional sampling of generative diffusion models. In this paper we propose a new diffusion posterior SMC sampler achieving improved statistical efficiencies, particularly under outlier conditions or highly informative likelihoods. The key idea is to construct an observation path that correlates with the diffusion model and to design the sampler to leverage this correlation for more efficient sampling. Empirical results conclude the efficiency.
Viral mutations pose significant threats to public health by increasing infectivity, strengthening vaccine resistance, and altering disease severity. To track these evolving patterns, agencies like the CDC annually evaluate thousands of virus strains, underscoring the urgent need to understand viral mutagenesis and evolution in depth. In this study, we integrate genomic analysis, clustering, and three leading dimensionality reduction approaches, namely, principal component analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE), and uniform manifold approximation and projection (UMAP)-to investigate the effects of COVID-19 on influenza virus propagation. By applying these methods to extensive pre- and post-pandemic influenza datasets, we reveal how selective pressures during the pandemic have influenced the diversity of influenza genetics. Our findings indicate that combining robust dimension reduction with clustering yields critical insights into the complex dynamics of viral mutation, informing both future research directions and strategies for public health intervention.
The Lotka-Volterra model reflects real ecological interactions where species compete for limited resources, potentially leading to coexistence, dominance of one species, or extinction of another. Comprehending the mechanisms governing these systems can yield critical insights for developing strategies in ecological management and biodiversity conservation. In this work, we investigate the controllability of a Lotka-Volterra system modeling weak competition between two species. Through constrained controls acting on the boundary of the domain, we establish conditions under which the system can be steered towards various target states. More precisely, we show that the system is controllable in finite time towards a state of coexistence whenever it exists, and asymptotically controllable towards single-species states or total extinction, depending on domain size, diffusion, and competition rates. Additionally, we determine scenarios where controllability is not possible and, in these cases, we construct barrier solutions that prevent the system from reaching specific targets. Our results offer critical insights into how competition, diffusion, and spatial domain influence species dynamics under constrained controls. Several numerical experiments complement our analysis, confirming the theoretical findings.
Over the past four decades, molecular studies have yielded two primary models for understanding the uniparental DNA phylogenetic trees of modern humans: the Out of Africa (OOA) and the Out of East Asia (OOEA) models. These models differ in their underlying assumptions, particularly in relation to early stem haplotypes, even though they share many haplotype relationships. Leveraging the wealth of new genetic variants unveiled through the comprehensive sequencing of 43 diverse human Y chromosomes, we here investigated the presence of shared variants among different haplotypes to determine which model better aligns with the genetic data. We validated our approach by confirming numerous well-established haplotype relationships that are consistent with both the OOA and OOEA models. Remarkably, our analysis revealed a compelling pattern: we were able to corroborate the existence of stem haplotypes specific to the OOEA model, but not those exclusive to the OOA model. For instance, we found that A0b and A1a shared the most variants with each other, aligning with the notion that both fall under the A00A1a stem haplotype of the OOEA model. So, it becomes evident that the genetic data obtained from the complete sequencing of the 43 newly analyzed human Y chromosomes lends robust support to the OOEA model as the more accurate representation of modern human origins.
In generative AI, the approach called diffusion-based generative modeling has introduced the problem of modifying a probability through a diffusion process on a given interval of time, in order to obtain a target probability. If, instead of speaking of probability distribution one speaks of random variables, this problem looks very similar to a controllability problem for a stochastic dynamic system. It is thus meaningful to explore the connection between the two problems, which has not been considered in the literature so far. The approach of controllability can open new possibilities for generative modeling. At the same time, many differences occur. It is interesting to notice that the concept of probability distribution and that of random variable, although representing an equivalent uncertainty, are complementary and not identical. We survey in this work the two domains and how to make use of the two methodologies.
In this paper, we propose a robust estimation method for binary classification based on convolutional neural networks (CNNs). Unlike classical approaches such as logistic regression, which are highly sensitive to model misspecification and data contamination, our CNN-based method offers improved resilience to deviations from ideal data distributional assumptions. An oracle inequality is derived for the resulting estimator, showing that its risk is bounded by the sum of a CNN approximation error and a stochastic error. When the model is well-specified and the true target function is alpha-Holder smooth, the estimator achieves the standard convergence rate of n(-2 alpha/(2 alpha+d)).
Finite-dimensional filters (FDFs), which are completely characterized by finite-dimensional statistics, have emerged as a crucial area of study in nonlinear filtering theory, with the algebraic classification of such filters representing a fundamental challenge in the field; while the complete classification of maximal rank estimation algebras has been well-established, recent research efforts have increasingly focused on the more complex non-maximal rank cases, leading to three key advances reviewed in this paper: (1) the verification of the linear structure of the Omega-matrix-an antisymmetric matrix derived from drift terms-for arbitrary state dimensions, resolving a longstanding open question; (2) the derivation of a sufficient condition for proving the weak form of Mitter's conjecture through extensions of quadratic form theory, with explicit verification for the case of dimension n = 4 and rank r = 3; and (3) the construction of novel FDFs that significantly expand the known classes beyond classical paradigms such as the Kalman-Bucy, Benes, and Yau filters, thereby opening new possibilities for solving previously intractable nonlinear filtering problems and advancing both theoretical understanding and practical applications in the field.
In this paper, we consider the Cauchy problem of the 3D isentropic compressible Navier-Stokes equations with degenerate viscosities and vacuum. It is proved that if the viscosities are the power law of density with different powers, i.e., mu(rho) = alpha(1)rho(delta 1) + alpha(2)rho(delta 2), lambda(rho) = beta rho(delta 2)(delta(1) not equal delta(2), delta(1), delta(2) > 1), then the Cauchy problem to 3D compressible Navier-Stokes equations admit a unique local classical solution (rho, u). Note that the initial data can contain vacuum in an open set, and there are no initial compatibility conditions.
In this paper we construct traveling gravity-capillary water waves with vortex sheets of both finite and infinite depth. The water waves obtained in this paper are flows with vorticity supported on a closed curve perturbed from a small circle. The construction is based on reformulation of boundary conditions and the use of Birkhoff-Rott operator, and accomplished by applying implicit function theorem at a point vortex with appropriately chosen function spaces.