Automated cell identification enables systematic phenotyping in developing embryos. Existing methods typically require nearly complete cellular configurations, but experimental constraints including phototoxicity and tissue-specific labeling often limit observation to cell subsets. We present a method that identifies Caenorhabditis elegans embryonic cells from partial observations containing 5–20 nuclei. Each cell is represented by geometric features computed from its local neighborhood, including pairwise relational statistics derived from the distance matrix. A joint attention encoder learns context-dependent embeddings, and cells are identified by nearest-neighbor lookup in a manifold built from training data. The model is trained on simulated embryos and fine-tuned on real data following the Twin Attention transfer learning approach. The method achieves 88.4
Cell migration is a fundamental phenomenon in biology that underlies normal development as well as cancer. Recently, a data-driven approach was introduced that uses deep reinforcement learning(DRL) and 3-D live images to study cell migration. This approach formulates the cell migration process as a sequential Markov decision process (MDP), so that hypotheses of the underlying mechanism of the observed migration can easily be incorporated as high-level regulatory rules and constraints for DRL. The application of the approach successfully uncovered a novel mechanism of cell migration in C. elegans embryogenesis that involves a modular organization of cells by using ubiquitous labels of cell nuclei and simple rules based on empirical statistics of the images. This success demonstrates new opportunities to use DRL to infer the biology of cell migration without prior knowledge. This paper presents an open framework, CellMigrationGym, to standardize the DRL approach to study cell migration. Built upon common packages (OpenAI Gym, PyBullet, and DRL libraries), CellMigrationGym provides powerful and flexible functions to investigate cell migration behavior. Through a case study, we demonstrate the critical functions of CellMigrationGym with technical details, such as 1) preparation and standardization of multiple observational data, 2) reward formulation and DRL model configuration appertaining to the hypotheses of migration mechanism (such as gradient-driven and collective cell behavior-driven mechanisms), 3) exploration of migration scenarios under hypothesized mechanisms, and 4) evaluation of neighboring cell’s influence on the cell migration.
This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km & times; 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component's efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.
We present Python and Julia bindings (PyPaRSEC and PaRSEC.jl) for the PaRSEC task-based runtime system, providing high-level access to its efficient and portable execution model on distributed heterogeneous architectures. The bindings expose PaRSEC ’s core abstractions—including data management, task creation, and runtime execution—allowing domain scientists to construct parallel workflows from high-level code while preserving PaRSEC ’s execution semantics. We evaluate these bindings using a 1D stencil workload (through PaRSEC ’s Parameterized Task Graph (PTG)) and a general matrix–matrix multiplication (GEMM) workload expressed via Dynamic Task Discovery (DTD). The results show that the bindings reproduce the official workflows with modest overhead, and that performance remains close to native implementations when leveraging optimized numerical backends on CPUs and vendor-provided accelerator libraries on GPUs, while callback-based user kernels provide flexibility at higher overhead. This work demonstrates that PaRSEC ’s high-performance runtime can be effectively leveraged from productive, high-level languages, bridging the gap between scientific prototyping and exascale execution.
The Software Package for E3SM Land Model (SPEL) is an innovative tool aimed at automating the development and validation of unit tests for subroutines within the ELM codebase. It systematically analyzes variable usage, control flow influenced by macros and namelist configurations, and the Fortran code structure of ELM. By capturing the context, dependencies, and call tree information, SPEL addresses the complexities associated with understanding intricate scientific codes. Additionally, it automates the creation of independent unit test programs, facilitating dependency analysis, code generation, and bit-for-bit verification. By enhancing testing, debugging, and code analysis processes, SPEL proves to be an essential resource for the development and maintenance of ELM. Beyond automated unit-test generation, SPELs metadata and dependency structure also support broader workflows such as constructing training datasets and verification pathways for AI-based surrogate models of ELM components.
General Matrix Multiplication (GEMM) is a critical operation underpinning a wide range of applications in high-performance computing (HPC) and artificial intelligence (AI). The emergence of hardware optimized for low-precision arithmetic necessitates a reevaluation of numerical algorithms to leverage mixed-precision computations, achieving improved performance and energy efficiency. This research presents an adaptive mixed-precision GEMM framework that enables support for various precision formats at fine-grained tile and block levels, offering a reliable foundation for trustworthy mixed-precision computations. Furthermore, we leverage the PaRSEC runtime system to effectively balance workloads across diverse architectures. The performance exhibits strong scalability across both homogeneous platforms (Intel CPU-based systems and the ARM CPU-based Fugaku supercomputer) and heterogeneous systems (Nvidia V100, A100, and H100 GPU-based platforms, as well as the AMD GPU-based Frontier supercomputer). This work aims to improve computational efficiency and accuracy by bridging algorithmic innovations with hardware capabilities, fostering transformative advancements across a wide range of applications.
Complex tissues are now characterizable at single-cell resolution, but the spatial logic underlying tissue organization remains challenging to access without effective single-cell spatial descriptors. We present a learned single-cell manifold of cell positions over time, with which one can measure tissue morphology and trace dynamics. Learning co-processes pairs of cell point clouds sampled along time using a Transformer encoder with inter-sample attention, a strategy that promotes efficient joint spatiotemporal learning. The manifold shows desirable properties of a general descriptor, e.g. interpretable cell type clusters, preserved local distances, and a pseudo-time axis, and enables common but challenging spatial reasoning tasks such as annotation of anatomical landmarks at cellular resolution and detection of subtle, transient phenotypes in large screens. Our study demonstrates a widely applicable cell-based learning strategy and representation for studying tissue biology.
Large-scale numerical simulations underpin modern scientific discovery but remain constrained by prohibitive computational costs. AI surrogates offer acceleration, yet adoption in mission-critical settings is limited by concerns over physical plausibility, trustworthiness, and the fusion of heterogeneous data. We introduce PHASE, a modular deep-learning framework for physics-integrated, heterogeneity-aware surrogates in scientific simulations. PHASE combines data-type-aware encoders for heterogeneous inputs with multi-level physics-based constraints that promote consistency from local dynamics to global system behavior. We validate PHASE on the biogeochemical (BGC) spin-up workflow of the U.S. Department of Energy's Energy Exascale Earth System Model (E3SM) Land Model (ELM), presenting-to our knowledge-the first scientifically validated AI-accelerated solution for this task. Using only the first 20 simulation years, PHASE infers a near-equilibrium state that otherwise requires more than 1,200 years of integration, yielding an effective reduction in required integration length by at least 60x. The framework is enabled by a pipeline for fusing heterogeneous scientific data and demonstrates strong generalization to higher spatial resolutions with minimal fine-tuning. These results indicate that PHASE captures governing physical regularities rather than surface correlations, enabling practical, physically consistent acceleration of land-surface modeling and other complex scientific workflows.
Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O3) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O3 pollution. The results indicate that O3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R2 = 0.9043). Feature importance identifies NO2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O3 transport with local emissions and meteorological conditions further increase the O3 pollution level. This study highlights the need to strengthen coordinated NOX and VOCs emission reductions and enhance regional joint prevention and control strategies in China.
This paper outlines a distributed data infrastructure designed to enhance open bioinformatics research, particularly in enzyme structure investigation. Utilizing high-throughput computing and cloud resources, such as AWS, the architecture facilitates collaborative workflows while ensuring robust user management, security, data integrity, and scalability. Powered by a Django backend and an Angular frontend, the infrastructure promotes seamless communication and data handling. A case study demonstrates its application in enzyme structure analysis using computational tools like AlphaFold and AlphaFill, revealing improvements in understanding enzymatic functions and accelerating drug discovery. Ultimately, this infrastructure aims to improve the efficiency and accessibility of bioinformatics research.
Sparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74-98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R-2 scores in range of 0.98-0.99 against observation data.
This paper presents advancements in scaling the ultrahigh-resolution E3SM Land Model (uELM) for deployment on leadership-class supercomputers, addressing the increased demand for km-scale Earth system modeling. By focusing on km-scale ELM simulations, we enhance predictive capabilities for climate interactions, facilitating improved responses to climate change impacts on energy systems, agriculture, and water resources. Our approach leverages innovative software architecture optimizations, sophisticated data handling techniques, and advanced parallel processing, achieving strong scalability on two leadership supercomputers (2400 nodes (105,600 cores) on Summit, and 1200 nodes (76,800 cores) on Frontier). Results from extensive scalability assessments on the Summit and Frontier also demonstrate outstanding I/O performance (close to 400 GB/s write throughput) and the model's ability to efficiently handle increasing computational demands. This study not only establishes uELM's capability for high-resolution simulations over vast geographical domains, but also sets a foundation for future Earth system modeling breakthroughs.
General Matrix Multiplication (GEMM) is a critical operation underpinning a wide range of applications in high-performance computing (HPC) and artificial intelligence (AI). The emergence of hardware optimized for low-precision arithmetic necessitates a reevaluation of numerical algorithms to leverage mixed-precision computations, achieving improved performance and energy efficiency. This research introduces an adaptive mixed-precision GEMM framework that supports different precision formats at fine-grained tile/block levels. We utilize the PaRSEC runtime system to balance workloads across various architectures. The performance scales well on ARM CPU-based Fugaku supercomputer, Nvidia GPU-based A100 DGX, and AMD GPU-based Frontier supercomputer. This research aims to enhance computational efficiency and accuracy by bridging algorithmic advancements and hardware innovations, driving transformative progress in various applications.
The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km × 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.
The development of a kilometer-scale E3SM Land Model (km-scale ELM) is an integral part of the E3SM project, which seeks to advance energy-related Earth system science research with state-of-the-art modeling and simulation capabilities on exascale computing systems. Through the utilization of high-fidelity data products, such as atmospheric forcing and soil properties, the km-scale ELM plays a critical role in accurately modeling geographical characteristics and extreme weather occurrences. The model is vital for enhancing our comprehension and prediction of climate patterns, as well as their effects on ecosystems and human activities. This study showcases the first set of full-capability, km-scale ELM simulations over various computational domains, including simulations encompassing 21.6 million land gridcells, reflecting approximately 21.5 million square kilometers of North America at a 1 km x 1 km resolution. We present the largest km-scale ELM simulation using up to 100,800 CPU cores across 2,400 nodes. This continental-scale simulation is 300 times larger than any previous studies, and the computational resources used are about 400 times larger than those used in prior efforts. Both strong and weak scaling tests have been conducted, revealing exceptional performance efficiency and resource utilization. The km-scale ELM uses the common E3SM modeling infrastructure and a general data toolkit known as KiloCraft. Consequently, it can be readily adapted for both fully-coupled E3SM simulations and data-driven simulations over specific areas, ranging from a single gridcell to the entire North America.
Designing and optimizing complex scientific code for new computing architectures is a challenging task. To address this issue in the E3SM land model (ELM) development, we developed a software tool called SPEL, which facilitates code generation, verification, and performance tuning using compiler directives within a Function Unit Test framework. In this paper, we present a SPEL extension that leverages the version control system (e.g., Git) to autonomous code generation and demonstrate its application to continuous code integration and development of the ELM software system. The study can benefit the scientific software development community.
Despite recent enhancements in China's anthropogenic emission controls, ozone (O3) concentrations have continuously increased owing to its nature as a secondary pollutant and the complexities of its production and consumption processes. This study quantified the contributions of urban and sectoral cross-emission sources to O3 levels and identified the anthropogenic emission sources requiring targeted control. Moreover, O3 sensitivity tests were conducted to determine optimal reduction ratios for nitrogen oxides (NOx) and volatile organic compounds (VOCs) emissions. The results were used to recommend effective measures for controlling O3 pollution in the Central Plains urban agglomeration (CPUA). The top 35 cities and sectoral cross-emission sources accounted for 80% of the O3 concentrations in the region, indicating the need for prioritized management of these sources. To achieve reductions in O3 concentrations across all cities, it was found that a 10% reduction in total NOx emissions would require a minimum of 18% reduction in VOCs emissions. Our results indicated that the appropriate coordination of reductions in VOCs and NOx emissions reduced the maximum daily 8-h average O3 (MDA8) concentrations in CPUA by 0.14%-4.78%. Enhancing control measures for prioritized emission sources reduced MDA8 concentrations by 0.78%-7.09%. Furthermore, adjusting the production and emission hours of the industrial sector resulted in a decrease in MDA8 concentrations by 1.10%-12.62%. Overall, our findings indicate that appropriately coordinated reduction of precursor emissions can reduce O3 levels. Further efforts to mitigate O3 pollution should include optimizing the timing of emissions from the industrial sector and other major sources of VOCs emissions.
Fengguang Song合作论文数University of Tennessee9