
Abstract Constructing multi-step bioinformatics workflows, from read quality control through genome assembly to functional annotation, requires expertise in both biology and computational tool selection, creating a bottleneck for scalable and reproducible analysis. We present the KBase Research Agent, a multi-agent system for automating such workflows within the DOE Systems Biology Knowledgebase (KBase). Given a set of sequencing reads and a research objective, the agent constructs an analysis plan grounded in KBase documentation and a Knowledge Graph (KG) of the KBase application catalog, then selects, parameterizes, validates and executes appropriate KBase applications to carry out the workflow. The resulting analysis is preserved as a reproducible KBase Narrative. We evaluate the system’s planning and execution quality against ground truth constructed from reference workflows derived from peer-reviewed Microbiology Resource Announcements. We further apply the agent to 100 previously unanalyzed bacterial isolate genomes from the JGI IMG/M database, where it autonomously performed read quality control, genome assembly, taxonomic classification with GTDB-Tk, and downstream analysis producing annotated genomes, reproducible Narratives, and draft manuscripts without human intervention. Across these experiments, the KBase Research Agent demonstrates the feasibility of domain-grounded, end-to-end scientific workflow automation in a production bioinformatics platform.
Modern High Performance Computing (HPC) centers face growing challenges in ingesting large and diverse data streams. These issues often create bottlenecks that limit bandwidth utilization and delay scientific progress. Traditional static allocation and simple queuing methods are often insufficient. This paper presents a dynamic, value-based approach to bandwidth allocation. We formalize the problem by incorporating both network and processing constraints. To address it, we introduce two auction-based mechanisms: the Greedy Value Density Auction, which is computationally efficient, and the Vickrey–Clarke–Groves (VCG) Knapsack Auction, which provides strong theoretical guarantees. Both mechanisms rely on user bids that specify data requirements and scientific value. The objective is to maximize the total value of successful transfers, commonly referred to as social welfare. Simulation results demonstrate that the proposed mechanisms significantly outperform First Come First Served (FCFS) baselines. Under high-load conditions, they reduce average and tail completion delays by more than 80
The recent development of differentiable simulation codes for particle accelerators has enabled gradient-based workflows that promise finer control and more realistic modeling of accelerator facilities. However, when using reverse-mode automatic differentiation, the memory usage continuously increases during the simulation, and can potentially exceed the available hardware memory – especially when costly space charge computation is included. To study the memory requirements for differentiable simulations, we have implemented space charge in Cheetah, a PyTorch-based beam tracking code that supports reverse-mode differentiation. We find that the memory usage for reverse-mode differentiation grows linearly with the number of macroparticles and cells, and that it is proportional to the number of space charge kicks involved in the simulation. This general scaling can be used to evaluate whether a given differentiable simulation is feasible given hardware memory constraints.