Understanding the interactions and regulatory relationships among biomolecules is essential for deciphering complex biological systems and elucidating the mechanisms behind diverse biological functions. Traditionally, the collection of such molecular interaction data has relied on expert curation, a process that is both time-consuming and labor-intensive. To address these limitations, this study explores the use of large language models (LLMs) to automate the genome-scale extraction of molecular interaction knowledge. We evaluate the performance of various LLMs on key biological tasks, including the identification of protein-protein interactions, detection of genes associated with pathways influenced by low-dose radiation, and inference of gene regulatory relationships. Our findings demonstrate that larger LLMs tend to perform better, particularly in extracting intricate gene and protein interactions. Despite their strengths, these models face challenges in recognizing functionally diverse gene groups and highly correlated regulatory relationships. Through a comprehensive analysis using established molecular interaction and pathway databases, we show that LLMs possess the potential to identify relevant biomolecules and predict their interactions, offering valuable insights and marking a significant step toward AI-driven biological knowledge discovery.
High-fidelity direct numerical simulation of turbulent flows for most real-world applications remains an outstanding computational challenge. Several machine learning approaches have recently been proposed to alleviate the computational cost even though they become unstable or unphysical for long time predictions. We identify that the Fourier neural operator (FNO) based models combined with a partial differential equation (PDE) solver can accelerate fluid dynamic simulations and thus address computational expense of large-scale turbulence simulations. We treat the FNO model on the same footing as a PDE solver and answer important questions about the volume and temporal resolution of data required to build pre-trained models for turbulence. We also discuss the pitfalls of purely data-driven approaches that need to be avoided by the machine learning models to become viable and competitive tools for long time simulations of turbulence.
Particle-resolved direct numerical simulations (PR-DNS) play an increasing role in investigating aerosol-cloud-turbulence interactions at the most fundamental level of processes. However, the high computational cost associated with high resolution simulations poses considerable challenges for large domain or long duration simulation using PR-DNS. To address these issues, here we present an emulator of the complex physics-based PR-DNS developed by use of the data-driven Fourier Neural Operator (FNO) method. The effectiveness of the method is showcased by presenting turbulence and temperature fields in a two-dimensional space. The results demonstrate high accuracy at various resolutions and the emulator is two orders of magnitude cheaper in terms of computational demand compared to the physics-based PR-DNS model. Furthermore, the FNO emulator exhibits strong generalization capabilities for different initial conditions and ultra-high-resolution without the need to retrain models. These findings highlight the potential of the FNO method as a promising tool to simulate complex fluid dynamics problems with high accuracy, computational efficiency, and generalization capabilities, enhancing our understanding of the aerosol-cloud-precipitation system. Particle-resolved direct numerical simulations (PR-DNS) are an important model to enhance our understanding of the fundamental processes involved in aerosol-cloud-turbulence interactions. However, achieving the extra-high-resolution simulations comes at very expensive computational cost. Although high-performance computing can accelerate PR-DNS simulations, it requires considering various factors, such as efficient message passing interface communications and graphics processing unit memory utilization. The machine learning (ML) emulators require much less computation cost compared to traditional numerical methods. Nevertheless, conventional ML models can only learn mappings between specific finite-dimensional spaces. The Fourier Neural Operator (FNO) method has recently been proposed to learn in the mesh-free and infinite dimensional space. In this study, we first present an emulator of the complex PR-DNS developed by using of the FNO method. The results show that the FNO model can achieve high accurate prediction, require low computational cost, and perform well with different initial conditions and resolutions, without re-training the ML models. The Fourier Neural Operator (FNO) model can accurately emulate particle-resolved direct numerical simulations (PR-DNS) at various resolutions The computational time of the physics-based PR-DNS model is reduced by two orders of magnitude with the FNO model The FNO model demonstrates robust and zero-shot generalization for various initial conditions and ultra-high resolutions
One among several advantages of measure transport methods is that they allow or a unified framework for processing and analysis of data distributed according to a wide class of probability measures. Within this context, we present results from computational studies aimed at assessing the potential of measure transport techniques, specifically, the use of triangular transport maps, as part of a workflow intended to support research in the biological sciences. Scenarios characterized by the availability of limited amount of sample data, which are common in domains such as radiation biology, are of particular interest. We find that when estimating a distribution density function given limited amount of sample data, adaptive transport maps are advantageous. In particular, statistics gathered from computing series of adaptive transport maps, trained on a series of randomly chosen subsets of the set of available data samples, leads to uncovering information hidden in the data. As a result, in the radiation biology application considered here, this approach provides a tool for generating hypotheses about gene relationships and their dynamics under radiation exposure.
Radiation exposure poses a significant threat to human health. Emerging research indicates that even low-dose radiation once believed to be safe, may have harmful effects. This perception has spurred a growing interest in investigating the potential risks associated with low-dose radiation exposure across various scenarios. To comprehensively explore the health consequences of low-dose radiation, our study employs a robust statistical framework that examines whether specific groups of genes, belonging to known pathways, exhibit coordinated expression patterns that align with the radiation levels. Notably, our findings reveal the existence of intricate yet consistent signatures that reflect the molecular response to radiation exposure, distinguishing between low-dose and high-dose radiation. Moreover, we leverage a pathway-constrained variational autoencoder to capture the nonlinear interactions within gene expression data. By comparing these two analytical approaches, our study aims to gain valuable insights into the impact of low-dose radiation on gene expression patterns, identify pathways that are differentially affected, and harness the potential of machine learning to uncover hidden activity within biological networks. This comparative analysis contributes to a deeper understanding of the molecular consequences of low-dose radiation exposure.
We present our ongoing work aimed at accelerating a particle-resolved direct numerical simulation model designed to study aerosol-cloud-turbulence interactions. The dynamical model consists of two main components-a set of fluid dynamics equations for air velocity, temperature, and humidity, coupled with a set of equations for particle (i.e., cloud droplet) tracing. Rather than attempting to replace the original numerical solution method in its entirety with a machine learning (ML) method, we consider developing a hybrid approach. We exploit the potential of neural operator learning to yield fast and accurate surrogate models and, in this study, develop such surrogates for the velocity and vorticity fields. We discuss results from numerical experiments designed to assess the performance of ML architectures under consideration as well as their suitability for capturing the behavior of relevant dynamical systems.
Gilchan Park, Byung-Jun Yoon, Xihaier Luo, Vanessa Lpez-Marrero, Patrick Johnstone, Shinjae Yoo, Francis Alexander. The 22nd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks. 2023.
There are various sources of ionizing radiation exposure, where medical exposure for radiation therapy or diagnosis is the most common human-made source. Understanding how gene expression is modulated after ionizing radiation exposure and investigating the presence of any dose-dependent gene expression patterns have broad implications for health risks from radiotherapy, medical radiation diagnostic procedures, as well as other environmental exposure. In this paper, we perform a comprehensive pathway-based analysis of gene expression profiles in response to low-dose radiation exposure, in order to examine the potential mechanism of gene regulation underlying such responses. To accomplish this goal, we employ a statistical framework to determine whether a specific group of genes belonging to a known pathway display coordinated expression patterns that are modulated in a manner consistent with the radiation level. Findings in our study suggest that there exist complex yet consistent signatures that reflect the molecular response to radiation exposure, which differ between low-dose and high-dose radiation.
The 4M-2N complexities (Multiscale, Multiphysics, Multibody, Multidimension, Non-linearitiy, and Non-Gaussianality) of atmospheric aerosol-cloud-precipitation-turbulence-radiation system poses physical and computational challenges to further advance predictive models; We plan to address the challenges by developing an AI-enhanced modeling framework that facilitates automated calibration and improvement of subgrid parameterizations, enhance data assimilation of measurements to improve initial and boundary conditions used to drive the physical model, and optimally blends data-driven and physics-based forecasting models.