The nsp16 2'-O-Methyltransferase is an essential non-structural protein of SARS-CoV-2, which methylates the viral mRNA cap structure, enabling it to evade the host immune response for higher translation efficiency. However, nsp16 is only active when it is bound to its cofactor, namely the non-structural protein-10 (nsp10). Understanding how nsp10 binds to and activates nsp16 function can help to develop targeted inhibitors; however, given the varying degree of disorder in both nsp10 and nsp16, characterizing this interaction has been challenging. Using long-timescale molecular dynamics simulations and AI/ML methods, we posit that the nsp16/nsp10 binding process is mediated by a hydrophobic latch formed with Leu4298 from nsp10 and a hydrophobic concave on the nsp16 protein surface. Our study highlights how the nsp16 S-adenosyl-L-methionine (SAM) pocket closes in its monomer state, which in turn deactivates the MTase function. We also observe that the nsp16/nsp10 complex allows for the RNA binding site to open with the empty SAM pocket. The results reveal how the SAM pocket loops facilitate SAM binding while allowing for the by-product S-adenosyl-L-homocysteine (SAH) to exit. Our study thus provides valuable atomistic-level mechanistic insights into understanding the activation of nsp16 MTase function while highlighting the challenges of studying protein-protein interactions mediated by largely fiexible/disordered regions.
Enzymes are powerful and sustainable catalysts, but their widespread application is limited by the difficulty of identifying functional starting points for optimization, creating a major bottleneck in early- stage biocatalyst discovery. Designing libraries of such starting enzymes remains particularly challenging. Here, we use the GenSLM protein language model to generate novel β -subunit of tryptophan synthase (TrpB) enzymes that express in Escherichia coli and are both stable and catalytically active. Many generated TrpBs also display significant substrate promiscuity, outperforming their natural counterparts on non-native substrates. Some even surpass laboratory-evolved TrpBs. Comparison of the most-active and most-promiscuous generated TrpB to its closest natural homolog confirms that the enhanced versatility is absent from the natural enzyme, highlighting the creative potential of generative models. These results demonstrate that the generated TrpBs not only preserve natural structure and function but also acquire non-natural properties, establishing generative models as powerful tools for biocatalyst discovery and engineering.
Automating experimental protocol design and execution remains as a fundamental bottleneck in realizing self-driving laboratories. We introduce PRISM (Protocol Refinement through Intelligent Simulation Modeling), a framework that automates the design, validation, and execution of experimental protocols on a laboratory platform composed of off-the-shelf robotic instruments. PRISM uses a set of language-model-based agents that work together to generate and refine experimental steps. The process begins with automatically gathering relevant procedures from web-based sources describing experimental workflows. These are converted into structured experimental steps (e.g., liquid handling steps, deck layout and other related operations) through a planning, critique, and validation loop. The finalized steps are translated into the Argonne MADSci protocol format, which provides a unified interface for coordinating multiple robotic instruments (Opentrons OT-2 liquid handler, PF400 arm, Azenta plate sealer and peeler) without requiring human intervention between steps. To evaluate protocol-generation performance, we benchmarked both single reasoning models and multi-agent workflow across constrained and open-ended prompting paradigms. The resulting protocols were validated in a digital-twin environment built in NVIDIA Omniverse to detect physical or sequencing errors before execution. Using Luna qPCR amplification and Cell Painting as case studies, we demonstrate PRISM as a practical end-to-end workflow that bridges language-based protocol generation, simulation-based validation, and automated robotic execution.
Molecular dynamics simulations allow the investigation of the time-resolved mechanics of large, complex biological systems such as the Gram-negative bacterial cell envelope. Such simulations pose challenges due to their chemical diversity and crowded environments. We review artificial intelligence-based approaches that can support simulations of large biological systems, focussing on trajectory analysis and propagation whilst highlighting the difficulties of feature representation. Integrating trajectory analysis and propagation is powerful, but we also consider the serious data and resource requirements involved. In this context, we summarise the current state of cell envelope simulations and then ask a practical question: where can such approaches be applied to understand these crowded, chemically diverse environments?
Despite the large corpus of biology training text, the impact of reasoning models on biological research generally lags behind math and coding. In this work, we show that biology questions from current large-scale reasoning datasets do not align well with modern research topic distributions in biology, and that this topic imbalance may negatively affect performance. In addition, we find that methods for extracting challenging and verifiable research problems from biology research text are a critical yet underdeveloped ingredient in applying reinforcement learning for better performance on biology research tasks. We introduce BioAlchemy, a pipeline for sourcing a diverse set of verifiable question-and-answer pairs from a scientific corpus of biology research text. We curate BioAlchemy-345K, a training dataset containing over 345K scientific reasoning problems in biology. Then, we demonstrate how aligning our dataset to the topic distribution of modern scientific biology can be used with reinforcement learning to improve reasoning performance. Finally, we present BioAlchemist-8B, which improves over its base reasoning model by 9.12
Self-driving laboratories increasingly rely on low-cost liquid handlers such as the Opentrons OT-2, which ship without the pressure-based aspiration monitoring of Hamilton or Tecan systems and are typically run open-loop. Two failure modes go undetected: protocols that are syntactically valid but violate assay-specific invariants (e.g., tip reuse between a PCR template and a no-template control), and physical execution failures (partial dispense, air bubbles, missing tips) at runtime. We present AEGIS, a two-layer guardian for both. Layer 1 pairs a curated machine-readable assay rule database with an LLM that reasons over OT-2 Python code, reaching an adjusted F1 of 0.97 on a 24-protocol benchmark across five assay families and beating rules-only and LLM-only ablations across five backends; a free open-weight model ties the best proprietary one, so no paid API is required. Layer 2 fits a PCA world model to YOLO-cropped four-frame pipette trajectories; under a leakage-free leave-one-plate-out evaluation it reaches average precision 0.89 and operating-point F1 0.71 (AUROC 0.80), a deployment-faithful number that matches the live demonstration, and we characterize the small-pipette (p20) resolution limit (F1 0.47). A live demonstration on a physical OT-2 (five replicates per condition) catches planted no-tip failures deterministically and partial dispense on coloured dyes, with an always-VLM self-vote gate lifting partial-dispense recall to 5/5; transparent water is a principled limit of any front-view-only monitor, which AEGIS surfaces as low-confidence VLM reasoning rather than a wrong verdict. Cascade triage holds VLM cost near 1.63 per plate versus10.33 for an always-VLM baseline. AEGIS is open source and, to our knowledge, the first system to unify pre-flight assay-aware validation with runtime visual monitoring for an open-source liquid handler.
AI agents have proliferated in scientific applications as a means of accelerating discovery. Research infrastructure presents unique challenges in terms of building and deploying these applications; for instance diverse computing resources, federated authentication and access, restrictive networking policies, and variable up-times. These challenges are shaping how agentic applications for science are being constructed and deployed in practice. Specifically, in this work, we examine applications built using the Academyframework, a middleware for deploying agents on federated research infrastructure. We survey the architecture and methods employed in agentic applications across eleven different scientific applications, and include insights from interviewing the scientists who designed and built them. From these experiences we distill patterns for constructing future agentic science applications, and implications for the systems research community.
Cellular senescence, a stable growth-arrested state induced by stress or chemotherapeutic agents, is accompanied by metabolic remodeling that supports the senescence-associated secretory phenotype (SASP). Among these pathways, lipid and arachidonic acid (AA) metabolism play central roles in maintaining and propagating the senescent state. Here, we used hyperspectral confocal Raman microscopy to visualize biochemical remodeling in MCF7 human breast adenocarcinoma cells undergoing doxorubicin-induced senescence. Raman spectral analysis and principal component decomposition revealed time-dependent alterations in lipid-associated vibrational modes, particularly CH2 and CC stretching, consistent with enhanced lipid accumulation and remodeling between Days 10 and 15 after DNA damage induction. PCA of lipid-rich compartments isolated by using true component analysis also confirms progressive increases in triacylglycerol and unsaturated lipid signatures. Using deuterated arachidonic acid (AA-d 8) and COX2 inhibition, we further demonstrated real-time intracellular AA metabolism by tracking (CC)D stretching peaks (2220-2254 cm-1) in the Raman-silent window. The ratio of these deuterium bands to CH2 stretching provided a label-free quantitative metric for COX2-dependent AA turnover in senescent cells. Together, these findings establish hyperspectral Raman imaging as a powerful, nonperturbative tool to map lipid and oxylipin metabolism during cellular senescence, offering new avenues to identify metabolic vulnerabilities in senescent tumor cells.
Intrinsically disordered proteins (IDPs) represent crucial therapeutic targets due to their significant role in disease – approximately 80% of cancer-related proteins contain long disordered regions – but their lack of stable secondary/tertiary structures makes them "undruggable". While recent computational advances, such as diffusion models, can design high-affinity IDP binders, translating these to practical drug discovery requires autonomous systems capable of reasoning across complex conformational ensembles and orchestrating diverse computational tools at scale.To address this challenge, we designed and implemented StructBioReasoner, a scalable multi-agent system for designing biologics that can be used to target IDPs. StructBioReasoner employs a novel tournament-based reasoning framework where specialized agents compete to generate and refine therapeutic hypotheses, naturally distributing computational load for efficient exploration of the vast design space. Agents integrate domain knowledge with access to literature synthesis, AI-structure prediction, molecular simulations, and stability analysis, coordinating their execution on HPC infrastructure via an extensible federated agentic middleware, Academy. We benchmark StructBioReasoner across Der f 21 and NMNAT-2 and demonstrate that over 50% of 787 designed and validated candidates for Der f 21 outperformed the human-designed reference binders from literature, in terms of improved binding free energy. For the more challenging NMNAT-2 protein, we identified three binding modes from 97,066 binders, including the well-studied NMNAT2:p53 interface. Thus, StructBioReasoner lays the groundwork for agentic reasoning systems for IDP therapeutic discovery on Exascale platforms.
Physics-based models of biomolecular systems that explicitly represent biomolecular structure and mechanics, such as atomistic molecular dynamics simulations, are well established because experimental data have been available to iteratively improve and validate models. Now, simulations of the biological mesoscale are growing in importance because of the improvements in experimental tools to visualize this regime. This includes techniques such as cryo-electron microscopy and tomography, microscopies that follow individual proteins in their cellular contexts, in situ scattering to follow the dynamic evolution of biomolecular assembly, and omics tools. Together, these approaches, alone and in combination, have revealed the importance of interactomes that bridge multiple scales. Here, we describe the theoretical, computational, and cultural challenges that need to be overcome to gain an understanding of the biological mesoscale and offer potential solutions. This commentary is the result of a joint CECAM/CCPBioSim discussion workshop on how the community should address the challenges of biomolecular simulations at the mesoscale, held in Trento, Italy, in the summer of 2024. The aim is to provide a broad overview of the tools and techniques relevant to the biological mesoscale and to signpost readers to more detailed discussions in the cited literature.
The Bacterial and Viral Bioinformatics Resource Center (BV-BRC; https://www.bv-brc.org) is a comprehensive resource supporting research on bacterial and viral pathogens. It currently hosts over 14 million publicly available genomes and 33 high-throughput bioinformatic analysis services with numerous visual analytic tools allowing researchers to analyze their private data, generate comparisons with public data, and share data and results with colleagues. In recent years, the BV-BRC has added several new analysis services to support rapid comparative genomics and epidemiological analysis, viral genome assembly and annotation, viral subspecies classification, wastewater analysis, and molecular docking. In addition, several existing services have been updated to incorporate state-of-the-art tools, including assembly, annotation, taxonomic classification, metagenomic read mapping, and RNA-seq analysis. A new tool, called BV-BRC Copilot, provides an AI-powered natural-language interface that combines large language models with retrieval-augmented generation to guide users through data exploration, analysis workflows, and knowledge integration. With expanded outbreak tracking pages, training and educator engagement, and continued development of novel AI-driven analytics, BV-BRC continues to provide a unified resource to meet the evolving needs of the global research community.
TFs combine DBDs, which anchor them to DNA, with EDs that regulate transcription through activation or repression, yet the sequence logic linking ED composition to function remains unclear. Here, we systematically define proxy regions —disordered segments adjacent to DBDs—to enable quantitative analysis of ED-like sequences across the human TF repertoire. Using a biophysically interpretable 22-feature classifier (FALK22) together with an embedding-based model (ESM), we map ED diversity and identify composition and charge-pattern signatures that correspond to regulatory activity along a disorder continuum, separating activation-from repression-associated regions. FALK22 identified classes align well with those identified from ESM while providing transparent, sequence-level features. Proxy regions near C-termini exhibit gradients that track DBD families, suggesting that EDs and DBDs might have co-evolved rather than evolved independently. These results establish proxy regions and FALK22 as a framework to connect sequence features with transcriptional activity and to generate testable hypotheses about effector-domain function and co-evolution with DNA-binding domains. HIGHLIGHTS ### Competing Interest Statement The authors have declared no competing interest. Cancer Prevention and Research Institute of Texas, RR220008 Welch Foundation, https://ror.org/00np6vq88, E-2221, V-E-001 U.S. National Science Foundation, https://ror.org/021nxhr62, CBET- 2442006
Automated Program Repair tools are developed for generating feedback and suggesting a repair method for erroneous code. State of the art (SOTA) code repair methods rely on data-driven approaches and often fail to deliver solution for complicated programming questions. To interpret the natural language of unprecedented programming problems, using Large Language Models (LLMs) for code-feedback generation is crucial. LLMs generate more comprehensible feedback than compiler-generated error messages, and Reinforcement Learning with Human Feedback (RLHF) further enhances quality by integrating human-in-the-loop which helps novice students to lean programming from scratch interactively. We are applying RLHF fine-tuning technique for an expected Socratic response such as a question with hint to solve the programming issue. We are proposing code feedback generation tool by fine-tuning LLM with RLHF, Automated Code Evaluation with RLHF (ACE-RLHF), combining two open-source LLM models with two different SOTA optimization techniques. The quality of feedback is evaluated on two benchmark datasets containing basic and competition-level programming questions where the later is proposed by us. We achieved 2-5 using Llama-3-7B-Proximal-policy optimization in automated evaluation and similar or slightly higher accuracy compared to reward model-free RL with AI Feedback (RLAIF). We achieved almost 40 optimization while performing manual evaluation.
Language models for scientific tasks are trained on text from scientific publications, most distributed as PDFs that require parsing. PDF parsing approaches range from inexpensive heuristics (for simple documents) to computationally intensive ML-driven systems (for complex or degraded ones). The choice of the "best" parser for a particular document depends on its computational cost and the accuracy of its output. To address these issues, we introduce an Adaptive Parallel PDF Parsing and Resource Scaling Engine (AdaParse), a data-driven strategy for assigning an appropriate parser to each document. We enlist scientists to select preferred parser outputs and incorporate this information through direct preference optimization (DPO) into AdaParse, thereby aligning its selection process with human judgment. AdaParse then incorporates hardware requirements and predicted accuracy of each parser to orchestrate computational resources efficiently for large-scale parsing campaigns. We demonstrate that AdaParse, when compared to state-of-the-art parsers, improves throughput by 17× while still achieving comparable accuracy (0.2 percent better) on a benchmark set of 1000 scientific documents. AdaParse's combination of high accuracy and parallel scalability makes it feasible to parse large-scale scientific document corpora to support the development of high-quality, trillion-token-scale text datasets. The implementation is available at https://github.com/7shoe/AdaParse/
Enzymes offer unparalleled selectivity and sustainability for chemical synthesis, yet their widespread industrial application is often hindered by the slow and uncertain process of discovering and optimizing suitable biocatalysts. While directed evolution remains the gold standard for enzyme optimization, its success hinges on the availability of a starting enzyme with measurable activity, a persistent bottleneck for many desired functions. Designing libraries likely to contain such functional starting points remains a major challenge. In this work, we use the GenSLM protein language model (PLM) along with a series of filters to generate novel sequences of the β -subunit of tryptophan synthase (TrpB) that express in Escherichia coli , are stable, and are catalytically active in the absence of a TrpA partner. Many generated TrpBs also demonstrated significant substrate promiscuity, accepting non-canonical substrates typically inaccessible to natural TrpBs. Remarkably, several outperformed both natural and laboratory-optimized TrpBs on native and non-canonical substrates. Comparative analysis of the most active and promiscuous generated TrpB and its closest natural homolog confirmed that this enhanced functional versatility does not stem from the natural enzyme, highlighting the creative potential of generative models. Our results demonstrate that the model can generate enzymes which not only preserve natural structure and function but also acquire non-natural properties, establishing PLMs as powerful tools for biocatalyst discovery and engineering, with the potential in some cases to bypass further optimization. ### Competing Interest Statement The authors have declared no competing interest. U.S. Department of Energy, Office of Science, Office of Basic Energy Sciences, DE-SC0022218 Fulbright France Schmidt AI2050
The evaluation of NAD+-boosting compounds in human skeletal muscle is hindered by limitations of traditional 2D cultures and animal models. Human-relevant, three-dimensional (3D) engineered skeletal muscle organoids offer a promising platform to assess the biological effects of metabolic modulation. Here we engineered 3D human skeletal muscle organoids to investigate the impact of dihydronicotinamide riboside (NRH), a potent NAD+ precursor. Sustained exposure to high NRH concentrations (500 µM) enhanced early differentiation markers, including increased myotube fusion and fast-twitch fiber area, but concurrently induced structural defects such as disrupted sarcomeric organization, enlarged acetylcholine receptor clusters, and impaired acetylcholine-stimulated calcium signaling. These results reveal that excessive and sustained NAD+ elevation can uncouple rapid differentiation from proper maturation in muscle tissue. Our results highlight the importance of dose and duration optimization for NAD+-boosting compounds and establish 3D engineered muscle organoids as a valuable non-animal platform for mechanistic toxicology and preclinical safety assessment.
The volume of scientific literature is growing exponentially, leading to underutilized discoveries, duplicated efforts, and limited cross-disciplinary collaboration. Retrieval-Augmented Generation (RAG) offers a way to assist scientists by improving the factuality of Large Language Models (LLMs) in processing this influx of information. However, scaling RAG to handle millions of articles introduces significant challenges, including the high computational costs associated with parsing documents and embedding scientific knowledge, as well as the algorithmic complexity of aligning these representations with the nuanced semantics of scientific content. To address these issues, we introduce HiPerRAG, a RAG workflow powered by high performance computing (HPC) to index and retrieve knowledge from more than 3.6 million scientific articles. At its core are Oreo, a high-throughput model for multimodal document parsing, and ColTrast, a query-aware encoder fine-tuning algorithm that enhances retrieval accuracy by using contrastive learning and late-interaction techniques. HiPerRAG delivers robust performance on existing scientific question answering (Q/A) benchmarks and two new benchmarks introduced in this work, achieving 90% accuracy on SciQ and 76% on PubMedQA-outperforming both domain-specific models like PubMedGPT and commercial LLMs such as GPT-4. Scaling to thousands of GPUs on the Polaris, Sunspot, and Frontier supercomputers, HiPerRAG delivers million document-scale RAG workflows for unifying scientific knowledge and fostering interdisciplinary innovation.