Codon-based model testing is fundamental to molecular evolution studies. The growing complexity of these analyses – driven by computationally intensive likelihood estimations, memory-demanding datasets, and the need to scale across hundreds or thousands of genes – necessitates efficient use of high-performance computing resources. To address this need, we present HighSPA, a scalable and reproducible framework that integrates two widely used tools for evolutionary analysis – CodeML and HyPhy – into parallel workflows using the Parsl library. We validated HighSPA using DENV genomes from Brazil (serotypes 1–4), applying six codon substitution models. When using the CodeML workflow, HighSPA identified high-confidence positively selected sites (PSS) with serotype-specific patterns. In contrast, the HyPhy workflow detected fewer PSS, likely due to its more conservative inference approach. In terms of performance, HighSPA-HyPhy significantly reduced makespan and increased throughput – by an average of 87 × and 89 × , respectively – compared to the sequential execution of the analyses. These results support the presence of adaptive evolution in key genes such as E, NS3, and NS5, and demonstrate HighSPA’s effectiveness for large-scale evolutionary analysis.
As High-Performance Computing (HPC) systems become faster and more complex, their energy consumption continues to increase, raising both economic and environmental concerns. Therefore, it is essential to manage energy usage efficiently, balancing energy consumption with application performance. In this work, we extend our prior research on predictive modeling for energy-efficient resource allocation in HPC systems, which uses the Energy-Delay Product (EDP) as a measure of energy efficiency. Furthermore, instead of focusing solely on predictive accuracy, we aim to provide a deeper understanding of model behavior and its impact on configuration decisions. While previous results demonstrated that an Extra Trees Regressor (ETR) can approximate Oracle configurations with high accuracy, little is known about when and why prediction errors occur. In this work, we present an analysis of prediction errors, an initial uncertainty-aware perspective, and model interpretability in EDP prediction. Our results show that although exact-match accuracy is 81.25
In High-Performance Computing (HPC) systems, multiple processes simultaneously consume resources such as CPU time, memory, and electrical power, among others. Accurately predicting the resource consumption of a process based on its execution parameters enables more efficient resource allocation, ultimately improving the overall performance of the HPC system. While many studies have explored this topic, fewer explicitly examine the underlying assumptions of their approaches. This work contributes to filling that gap by proposing, experimenting with, and discussing a protocol to approach this problem, covering from the collection of processes footprint data to the experimental evaluation of Machine Learning models based on such data. The reported results of the assessment of this protocol in a case study of the RAxML bioinformatics application on a real supercomputer highlight not only its effectiveness (R2 values greater than 0.9 were achieved in most tests) but also the reasonableness of the assumptions considered.
O uso de Inteligência Artificial (IA) em larga escala tem crescido rapidamente, aumentando a pressão sobre infraestruturas de computação de alto desempenho (HPC). Em particular, com o surgimento de AI Factories, é necessária a adoção de estratégias eficientes de alocação de recursos para maximizar o throughput e reduzir o consumo energético. Nesse cenário, o compartilhamento de GPUs surge como uma alternativa promissora para melhorar a utilização dos aceleradores e equilibrar desempenho e eficiência energética. Este artigo investiga o impacto do compartilhamento da GPU na eficiência energética durante a inferência de modelos de IA usando o acelerador Intel Data Center GPU Max 1550 no supercomputador de classe exascale Aurora. Avaliamos quatro modelos distintos de IA de diferentes domínios de aplicação, como visão computacional, processamento de linguagem natural e geração de texto. Os resultados mostram que a co-localização de aplicações com perfis de uso de recursos complementares pode reduzir o tempo de execução em até 50% e o consumo de energia em até 43%. Por fim, demonstramos que a alocação de tarefas concorrentes em GPU, baseada na caracterização prévia das aplicações, é uma estratégia promissora para maximizar a eficiência em ambientes HPC.
O presente trabalho propõe a avaliação de uma metodologia baseada em aprendizado de máquina para apoiar a seleção de configurações de execução em ambientes de Computação de Alto Desempenho (HPC). Com o uso do modelo Extra Trees Regressor (ETR) e dados de execução da aplicação RAxML, foi possível estimar combinações de número de nós e de threads que minimizam o Energy-Delay Product (EDP). Os resultados demonstraram baixo erro preditivo (MAE = 0,05), evidenciando a viabilidade da abordagem para reduzir o consumo energético e otimizar a utilização de supercomputadores.
High-performance computing is pivotal for processing large datasets and executing complex simulations, ensuring faster and more accurate results. Improving the performance of software and scientific workflows in such environments requires careful analysis of their computational behavior and energy consumption. Therefore, maximizing computational throughput in these environments, through adequate software configuration and resource allocation, is essential for improving performance. The work presented in this paper focuses on leveraging regression-based machine learning and decision trees to analyze and optimize resource allocation in high-performance computing environments based on application's performance and energy metrics. Applied to a bioinformatics case study, these models enable informed decision-making by selecting the appropriate computing resources to enhance the performance of a phylogenomics software. Our contribution is to better explore and understand the efficient resource management of supercomputers, namely Santos Dumont. We show that the predictions for application's execution time using the proposed method are accurate for various amounts of computing nodes, while energy consumption predictions are less precise. The application parameters most relevant for this work are identified and the relative importance of each application parameter to the accuracy of the prediction is analysed.
Este trabalho apresenta um estudo comparativo entre os Sistemas de Gerenciamento de Workflows Científicos (SGWfCs) Parsl e PyCOMPSs, avaliando tanto a perspectiva do usuário na implementação de um workflow científico de bioinformática quanto em termos de desempenho. Os resultados obtidos, aliados à experiência prática de uso, indicam que, embora ambos apresentem desempenho semelhante – com o Parsl sendo, em média, apenas 6 segundos mais rápido –, o Parsl se destaca em termos de usabilidade e facilidade de integração para o usuário final.
The increasing gap between compute and I/O speeds in high-performance computing (HPC) systems imposes the need for techniques to improve applications' I/O performance. Such techniques must rely on assumptions about I/O behavior in order to efficiently allocate I/O resources such as burst buffers, to schedule accesses to the shared parallel file system or to delay certain applications at the batch scheduler level to prevent contention, for instance. In this paper, we verify these common assumptions about I/O behavior, specifically about temporal behavior, using over 440,000 traces from real HPC systems. By combining traces from diverse systems, we characterize the behaviors observed in real HPC workloads. Among other findings, we show that I/O activity tends to last for a few seconds, and that periodic jobs are the minority, but responsible for a large portion of the I/O time. Furthermore, we make projections for the expected improvement yielded by popular approaches for I/O performance improvement. Our work provides valuable insights to everyone working to alleviate the I/O bottleneck in HPC.
O Alinhamento Múltiplo de Sequências (AMS) é uma etapa fundamental na biologia evolutiva molecular, com impacto direto na identificação de marcadores associados a doenças genéticas e infecciosas. A qualidade dos alinhamentos é determinante para a confiabilidade das interpretações biológicas. No entanto, algoritmos exatos para AMS enfrentam limitações quanto ao número de sequências que podem ser processadas, mesmo em ambientes de Processamento de Alto Desempenho (PAD), devido à natureza NP-difícil do problema. Assim, a prática usual envolve a seleção manual de subconjuntos reduzidos de sequências. Neste trabalho, propomos e avaliamos um workflow científico em PAD para AMS com algoritmos exatos, incorporando a seleção automática do subconjunto representativo de sequências, de modo a contornar as restrições das ferramentas disponíveis. Os experimentos, considerando implementações do workflow com PyCOMPSs e com scripts Shell, apresentaram ganhos de 2, 08× e 3, 95×, respectivamente, em relação à execução sequencial das tarefas. Em particular, a implementação com PyCOMPSs mostrou melhor escalabilidade, alcançando 80,8% de ganho no alinhamento de 38 sequências.
Este estudo avaliou o desempenho do programa de bioinformática PA-Star, utilizado para alinhamento múltiplo de sequências, em duas arquiteturas do supercomputador Santos Dumont: o Mesca2, com maior capacidade computacional e de memória, e o nó Intel Xeon Ivy Bridge. As sequências biológicas processadas variaram em tamanho, quantidade e nível de similaridade. Observou-se que o Mesca2 apresentou desempenho superior para arquivos com alta demanda de RAM, enquanto o Intel Xeon Ivy Bridge foi mais eficiente para arquivos com menor demanda. Esses resultados sugerem que o Mesca2 é essencial para o processamento de alinhamentos com elevado consumo de memória RAM no Santos Dumont.
Este trabalho apresenta um estudo de desempenho da aplicação de inferência bayesiana BEAST 1.10, acoplada à biblioteca de alto desempenho BEAGLE 3, em execuções realizadas nos nós do supercomputador Santos Dumont. Nos experimentos de filogenia, utilizamos dados genômicos do vírus da Dengue, sorotipo DENV-1, em formato XML. Analisamos a variabilidade do tamanho dos genomas, o chainLength e modelos evolutivos do BEAST 1.10, o número de threads e o ambiente computacional (CPU e GPU) do SDumont. Os resultados do estudo do desempenho do BEAST no BioInfo-Portal, possibilitam uma utilização mais eficiente dos recursos computacionais do SDumont, segundo os parâmetros alocados na submissão dos jobs.
O artigo apresenta uma análise do desempenho computacional da etapa de alinhamento mais custosa do workflow Parsl-RNA Seq, utilizando o perfilador Intel VTune no supercomputador Santos Dumont. O objetivo do trabalho é analisar o desempenho computacional do software Bowtie2, responsável pelo alinhamento de sequências que é a que mais demanda recursos computacionais. O Bowtie2 foi executado com suporte da biblioteca Parsl e monitorado pelo VTune, com o objetivo de avaliar sua eficácia e escalabilidade em ambientes de computação de alto desempenho.
Experimentos científicos em larga escala são complexos devido à modelagem, execução e análise de grandes volumes de dados. Na bioinformática, esses experimentos são estruturados como workflows científicos, usando sistemas de gerência de workflows e computação de alto desempenho. Este artigo apresenta uma análise computacional do desempenho do workflow de transcriptômica ParslRNA-Seq entre arquiteturas com memória distribuída e memória compartilhada mais recentes do supercomputador Santos Dumont, demonstrando como a escolha das máquinas pode influenciar o desempenho.
Este estudo apresenta uma análise de desempenho e eficiência energética do Simulador Multifásico baseado no método Lattice Boltzmann (LBM) utilizando as arquiteturas NVIDIA V100 e NVIDIA GH200. O método LBM é amplamente utilizado para simulações de dinâmica de fluidos computacional devido à sua capacidade de modelar fluxos de permeabilidade relativa e complexos. Neste trabalho, foram realizadas simulações com malhas sintéticas variadas para avaliar o desempenho e o consumo energético em ambas as arquiteturas. Os resultados indicam que a arquitetura GH200 supera consistentemente a V100 em termos de desempenho, especialmente em malhas maiores, com ganhos de até 2.82 vezes. Além disso, a GH200 demonstrou uma eficiência energética superior, consumindo até 53% menos energia em comparação à V100. A análise de perfilagem com o NVIDIA NSight Compute identificou os principais fatores que contribuem para a perda de desempenho, destacando a necessidade de otimização no acesso à memória.
SummaryThe breadth‐first search procedure is an algorithm that traverses the vertices of a graph, determining the distance from each vertex to the initial vertex. The distance is infinite for a non‐reachable vertex from the starting vertex. Despite having an efficient serial version, this important algorithm is irregular, making its effective parallel implementation a daunting task. This paper shows the results of an OpenMP‐based implementation of the breadth‐first search procedure using the bag data structure. Furthermore, the code relied on the C++ programming language. This paper reimplements an existing proposal coded using the Cilk++ programming language. The experiments relied on 32 strongly connected graphs and 31 disconnected graphs in executions performed on two machines. The first machine contained 28 cores and two threads per core. The second machine comprised 48 processing cores, with hyperthreading disabled. Regarding the serial version, the parallel implementation yielded a speedup of up to 20× when using 28 processing cores and up to 25× when using 56 threads in tests performed on a machine with the first generation of Intel® Xeon® Scalable processors. Furthermore, the new parallel implementation yielded speedups of up to 45× when using 48 cores in experiments performed on a machine with the second generation of Intel® Xeon® Scalable processors.
O estudo avalia o desempenho do software Bowtie2, amplamente utilizado em tarefas de alinhamento genético e é uma das etapas mais custosas no workflow ParslRNA-Seq. A análise foi realizada nos nós Ivy Bridge e Ivy Bridge (MESCA2), ambos com arquitetura de memória compartilhada. As execuções foram gerenciadas pela biblioteca Parsl e monitoradas pelo perfilador Intel VTune para observar a distribuição e eficiência da tarefa em ambos os nós. Os resultados mostraram que o Ivy Bridge proporcionou uma melhor distribuição de tarefas entre os núcleos, enquanto o Ivy Bridge (MESCA2) apresentou limitações no uso eficiente dos núcleos, resultando em desempenho inferior devido à subutilização dos recursos disponíveis.
In petroleum reservoir simulations, the level of detail incorporated into the geologic model typically exceeds the capabilities of traditional flow simulators. In this sense, such simulations demand new high-performance computing techniques to deal with a large amount of data allocation and the high computational cost of computing the behavior of the fluids in the porous media. This paper presents optimizations performed on a code that implements an explicit numerical scheme to provide an approximate solution to the governing differential equation for water saturation in a two-phase flow problem with heterogeneous permeability and porosity fields. The experiments were performed on the SDumont Supercomputer using 2nd Generation Intel®Xeon®Scalable Processors (formerly Cascade Lake architecture). The paper employs a direct memory data access scheme to reduce the execution times of the numerical method. The article analyzes the performance gain using direct memory access related to indirect access memory. The results show that the optimizations implemented in the numerical code remarkably reduce the execution time of the simulations.
The scientific gateway BioinfoPortal for bioinformatics applications is hosted in the National Laboratory for Scientific Computing (LNCC) and is coupled to the Santos Dumont (SDumont) supercomputer environment. BioinfoPortal offers a catalog of bioinformatics software that benefits from the parallel and distributed architecture offered by LNCC. Task submissions consume SDumont nodes shared by other users of the supercomputer; thus, it is important they use the best configuration, which is defined as the best choice of the number of threads/nodes to be allocated for every task submission. This article presents an analysis using neural networks to estimate the computational time required to execute bioinformatics software in several scenarios using a pre-configured number of nodes and threads. Our goal is to demonstrate the performance behavior of software such as RAxML in Bioinfoportal, and which computational scenario can be chosen to efficiently execute software in SDumont. Results support that the neural networks are adequate to predict the variable elapsed time, Elapsed, to evaluate the relationships between input parameters, number of bootstraps (RAxML), number of threads, and number of nodes, and to identify the fastest configuration. The goal is to make BioinfoPortal a smart, efficient, and green gateway. In future studies, we propose to study more variables and predictors as well as other bioinformatics software in BioinfoPortal.
High-Performance Computing (HPC) platforms are required to solve the most diverse large-scale scientific problems in various research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific softwares, which have different requirements. These include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle mixed workload when storing data from the applications. Understanding the set of applications and their performance running in a supercomputer is paramount to understanding the storage system's usage, pinpointing possible bottlenecks, and guiding optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, identifying inefficient usage, problematic performance factors, and providing guidelines on how to tackle those issues.
O artigo traz discussões sobre a eleição de modificações no formato de execução do workflow ParslRNA-Seq, que levam a melhora do desempenho e escalabilidade computacional, baseado em redução de gastos com operações de E/S com o uso de SSD em relação ao sistema de arquivos paralelos Lustre no supercomputador Santos Dumont.