Understanding the capabilities and limitations of modern time series models is essential for their effective use in domain-specific contexts. This study investigates the zero-shot forecasting performance of state-of-the-art Time Series Foundation Models (TSFMs) using real-world epidemiological data. Guided by a central hypothesis and three research questions, we benchmark several TSFMs, including TimesFM, ChronosT5Base, and Moirai variants, on monthly time series of notifiable diseases in Brazil, considering different forecast horizons and context lengths. Through robust evaluation using MASE and CRPS metrics, along with statistical significance testing, we find that TimesFM and ChronosT5Base consistently outperform classical statistical baselines, while others demonstrate architectural limitations. These findings contribute to a deeper methodological understanding of TSFMs in health-related forecasting scenarios and suggest that performance differences may arise from model design and training data characteristics. The results provide valuable insights to guide future model development and support more informed use of forecasting tools in public health decision-making.
QSAR models capable of predicting biological, toxicity, and pharmacokinetic properties were widely used to search lead bioactive molecules in chemical databases. The dataset’s preparation to build these models has a strong influence on the quality of the generated models, and sampling requires that the original dataset be divided into training (for model training) and test (for statistical evaluation) sets. This sampling can be done randomly or rationally, but the rational division is superior. In this paper, we present MASSA, a Python tool that can be used to automatically sample datasets by exploring the biological, physicochemical, and structural spaces of molecules using PCA, HCA, and K-modes. The proposed algorithm is very useful when the variables used for QSAR are not available or to construct multiple QSAR models with the same training and test sets, producing models with lower variability and better values for validation metrics. These results were obtained even when the descriptors used in the QSAR/QSPR were different from those used in the separation of training and test sets, indicating that this tool can be used to build models for more than one QSAR/QSPR technique. Finally, this tool also generates useful graphical representations that can provide insights into the data.
A comunicação animal é essencial para a sobrevivência, e os pesquisadores têm se dedicado ao estudo dos padrões de sinais emitidos por lagartos, com destaque para a comunicação visual por meio dos movimentos de cabeceios (headbobs). Nesse contexto, este trabalho propõe o uso de técnicas de aprendizado de máquina para identificar e analisar padrões de comunicação em lagartos do gênero Tropidurus. A metodologia consiste em extrair os sinais dos vídeos por meio de algoritmos de aprendizado profundo e, em seguida, aplicar algoritmos não supervisionados para identificar padrões nos sinais extraídos. Os resultados obtidos demonstram a presença de padrões claros nos sinais analisados.
The drug design process has been evolving since its existence, and several knowledge fields have contributed to the discovery and development of novel drugs. The emergence of artificial intelligence and related algorithms was requested by medicinal chemists to improve the success rates and to speed up drug discovery and optimization processes. In this sense, methods such as principal component analysis, hierarchical clustering analysis, K-means, k-nearest neighbors, decision trees and random forests, support vector machines, and artificial neural networks have been widely employed to predict physicochemical and biological properties of chemicals as part of rational drug design campaigns. This chapter covers the theoretical background and applications of methods as mentioned earlier and other relevant ones in the medicinal chemistry area.
As the amount of financial data generated grows yearly, there is a growing need to leverage this data to develop customized financial products to meet individual users' unique needs and preferences. This study proposes a method for identifying potential spending patterns based on categorized financial transactions. Different clustering and outlier detection algorithms are compared using various internal validation metrics and empirical analysis of cluster balancing. A visualization of the spending patterns is created from the proposed method and validated by an expert in the domain in order to extract more insights based on user behavior. The visualization was found to be helpful when analyzing for insights into spending pattern.
Evolutionary and Swarm algorithms show great effectiveness when performing feature selection, classification problems, and other optimization tasks. These scenarios highlight several algorithms such as Genetic Algorithm, Particle Swarm Optimization, Artificial Bee Colony, and Genetic Bee Colony (GBC). The last one is a combination of Genetic Algorithm and Artificial Bee Colony. Our study proposes an improvement over the GBC algorithm, including the Chaos Theory behavior in its foundation. In addition, we modified the Onlooker Bee and Scout Bee phase to improve the proposed model exploration and exploitation capabilities. Our novel proposal, Chaos-GBC, showed a very competitive performance compared to GBC and other metaheuristics. CGBC achieved the highest classification accuracy and the lowest average number of selected genes, showing the importance of our new proposal.
Evolutionary and Swarm algorithms show great effectiveness when performing feature selection, classification problems, and other optimization tasks. These scenarios highlight several algorithms such as Genetic Algorithm, Particle Swarm Optimization, Artificial Bee Colony, and Genetic Bee Colony (GBC). The last one is a combination of Genetic Algorithm and Artificial Bee Colony. Our study proposes an improvement over the GBC algorithm, including the Chaos Theory behavior in its foundation. In addition, we modified the Onlooker Bee and Scout Bee phase to improve the proposed model exploration and exploitation capabilities. Our novel proposal, Chaos-GBC, showed a very competitive performance compared to GBC and other metaheuristics. CGBC achieved the highest classification accuracy and the lowest average number of selected genes, showing the importance of our new proposal.
The enoyl-[acyl-carrier-protein] reductase (FabI) is an important enzyme in the fatty acid metabolism of Gram-positive bacteria, such as Staphylococcus aureus. FabI is also a potential target for the development of novel antibacterials. Several machine learning-driven studies were reported to develop FabI inhibitors, describing robust and predictive models. Herein, the authors applied the kGCN, a graph convolutional network framework, to generate classification models to select potential S. aureus FabI inhibitors. The most predictive model showed robustness for both active and inactive class prediction, according to statistical validation. Finally, the chemical interpretation of the model was consistent with prior experimental and theoretical works. The SAR analysis highlighted the importance of the occupation of hydrophobic pockets and polar interactions with Tyr-156 and NADPH cofactor present in the FabI catalytic site by potential inhibitors. A density functional theory study endorsed the SAR, where the electrostatic surfaces were consistent with the expected interactions with the pocket.
Since the emergence of the new severe acute respiratory syndrome-related coronavirus 2 (SARS-CoV-2) at the end of December 2019 in China, and with the urge of the coronavirus disease 2019 (COVID-19) pandemic, there have been huge efforts of many research teams and governmental institutions worldwide to mitigate the current scenario. Reaching more than 1,377,000 deaths in the world and still with a growing number of infections, SARS-CoV-2 remains a critical issue for global health and economic systems, with an urgency for available therapeutic options. In this scenario, as drug repurposing and discovery remains a challenge, computer-aided drug design (CADD) approaches, including machine learning (ML) techniques, can be useful tools to the design and discovery of novel potential antiviral inhibitors against SARS-CoV-2. In this work, we describe and review the current knowledge on this virus and the pandemic, the latest strategies and computational approaches applied to search for treatment options, as well as the challenges to overcome COVID-19. ABSTRACT Since the emergence of the new severe acute respiratory syndrome-related coronavirus 2 (SARS-CoV-2) at the end of December 2019 in China, and with the urge of the coronavirus disease 2019 (COVID-19) pandemic, there have been huge efforts of many research teams and governmental institutions worldwide to mitigate the current scenario. Reaching more than 1,377,000 deaths in the world and still with a growing number of infections, SARS-CoV-2 remains a critical issue for global health and economic systems, with an urgency for available therapeutic options. In this scenario, as drug repurposing and discovery remains a challenge, computer-aided drug design (CADD) approaches, including machine learning (ML) techniques, can be useful tools to the design and discovery of novel potential antiviral inhibitors against SARS-CoV-2. In this work, we describe and review the current knowledge on this virus and the pandemic, the latest strategies and computational approaches applied to search for treatment options, as well as the challenges to overcome
Introduction: Drug design and discovery of new antivirals will always be extremely important in medicinal chemistry, taking into account known and new viral diseases that are yet to come. Although machine learning (ML) have shown to improve predictions on the biological potential of chemicals and accelerate the discovery of drugs over the past decade, new methods and their combinations have improved their performance and established promising perspectives regarding ML in the search for new antivirals.Areas covered: The authors consider some interesting areas that deal with different ML techniques applied to antivirals. Recent innovative studies on ML and antivirals were selected and analyzed in detail. Also, the authors provide a brief look at the past to the present to detect advances and bottlenecks in the area.Expert opinion: From classical ML techniques, it was possible to boost the searches for antivirals. However, from the emergence of new algorithms and the improvement in old approaches, promising results will be achieved every day, as we have observed in the case of SARS-CoV-2. Recent experience has shown that it is possible to use ML to discover new antiviral candidates from virtual screening and drug repurposing.
Since the emergence of the new severe acute respiratory syndrome-related coronavirus 2 (SARS-CoV-2) at the end of December 2019 in China, and with the urge of the coronavirus disease 2019 (COVID-19) pandemic, there have been huge efforts of many research teams and governmental institutions worldwide to mitigate the current scenario. Reaching more than 1,377,000 deaths in the world and still with a growing number of infections, SARS-CoV-2 remains a critical issue for global health and economic systems, with an urgency for available therapeutic options. In this scenario, as drug repurposing and discovery remains a challenge, computer-aided drug design (CADD) approaches, including machine learning (ML) techniques, can be useful tools to the design and discovery of novel potential antiviral inhibitors against SARS-CoV-2. In this work, we describe and review the current knowledge on this virus and the pandemic, the latest strategies and computational approaches applied to search for treatment options, as well as the challenges to overcome COVID-19.
Machine Learning (ML) algorithms have been increasingly applied to problems from several different areas. Despite their growing popularity, their predictive performance is usually affected by the values assigned to their hyperparameters (HPs). As consequence, researchers and practitioners face the challenge of how to set these values. Many users have limited knowledge about ML algorithms and the effect of their HP values and, therefore, do not take advantage of suitable settings. They usually define the HP values by trial and error, which is very subjective, not guaranteed to find good values and dependent on the user experience. Tuning techniques search for HP values able to maximize the predictive performance of induced models for a given dataset, but have the drawback of a high computational cost. Thus, practitioners use default values suggested by the algorithm developer or by tools implementing the algorithm. Although default values usually result in models with acceptable predictive performance, different implementations of the same algorithm can suggest distinct default values. To maintain a balance between tuning and using default values, we propose a strategy to generate new optimized default values. Our approach is grounded on a small set of optimized values able to obtain predictive performance values better than default settings provided by popular tools. After performing a large experiment and a careful analysis of the results, we concluded that our approach delivers better default values. Besides, it leads to competitive solutions when compared to tuned values, making it easier to use and having a lower cost. We also extracted simple rules to guide practitioners in deciding whether to use our new methodology or a HP tuning approach.
Semi-supervised learning is drawing increasing attention in the era of big data, as the gap between the abundance of cheap, automatically collected unlabeled data and the scarcity of labeled data that are laborious and expensive to obtain is dramatically increasing. In this paper, we first introduce a unified view of density-based clustering algorithms. We then build upon this view and bridge the areas of semi-supervised clustering and classification under a common umbrella of density-based techniques. We show that there are close relations between density-based clustering algorithms and the graph-based approach for transductive classification. These relations are then used as a basis for a new framework for semi-supervised classification based on building-blocks from density-based clustering. This framework is not only efficient and effective, but it is also statistically sound. In addition, we generalize the core algorithm in our framework, HDBSCAN*, so that it can also perform semi-supervised clustering by directly taking advantage of any fraction of labeled data that may be available. Experimental results on a large collection of datasets show the advantages of the proposed approach both for semi-supervised classification as well as for semi-supervised clustering.
Extracting a flat solution from a clustering hierarchy, as opposed to deriving it directly from data using a partitional clustering algorithm, is advantageous as it allows the hierarchical relationships between clusters and sub-clusters as well their stability across different hierarchical levels to be revealed before any decision on what clusters are more relevant is made. Traditionally, flat solutions are obtained by performing a global, horizontal cut through a clustering hierarchy (e.g. a dendrogram). This problem has gained special importance in the context of density-based hierarchical algorithms, because only sophisticated cutting strategies, in particular non-horizontal local cuts, are able to select clusters at different density levels. In this paper, we propose an adaptation of a variant of the Modularity Q measure, widely used in the realm of community detection in complex networks, so that it can be applied as an optimization criterion to the problem of optimal local cuts through clustering hierarchies. Our results suggest that the proposed measure is a competitive alternative, especially for high-dimensional data.
Semi-supervised classification is drawing increasing attention in the era of big data, as the gap between the abundance of cheap, automatically collected unlabeled data and the scarcity of labeled data that are laborious and expensive to obtain is dramatically increasing. In this paper, we introduce a unified framework for semi-supervised classification based on building-blocks from density-based clustering. This framework is not only efficient and effective, but it is also statistically sound. Experimental results on a large collection of datasets show the advantages of the proposed framework.
Dipeptidyl peptidase-4 (DPP-4) is an important biological target related to the treatment of diabetes as DPP-4 inhibitors can lead to an increase in the insulin levels and a prolonged activity of glucagon-like peptide-1 (GLP-1) and gastric inhibitory polypeptide (GIP), being effective in glycemic control. Thus, this study analyses the main molecular interactions between DPP-4 and a series of bioactive ligands. The methodology used here employed molecular modeling methods, such as HQSAR (Hologram Quantitative Structure-Activity) analyses and molecular docking, with the aim of understanding the main structural features of the compound series that are essential for the biological activity. Analyses of the main interactions in the active site of DPP-4, in particular, the contribution of the hydroxyl coordination between Tyr547 and Ser630 by the water molecule, which is described in the literature as important for the coordinated interactions in the active site, were performed. Significant correlation coefficients of the best 2D model (r(2) = 0.942 and q(2) = 0.836) were obtained, indicating the predictive power of this model for untested compounds. Therefore, the final model constructed in this study, along with the information from the contribution maps, could be useful in the design of novel DPP-4 ligands with improved activity.
Introduction: Pharmacokinetics involves the study of absorption, distribution, metabolism, excretion and toxicity of xenobiotics (ADME-Tox). In this sense, the ADME-Tox profile of a bioactive compound can impact its efficacy and safety. Moreover, efficacy and safety were considered some of the major causes of clinical failures in the development of new chemical entities. In this context, machine learning (ML) techniques have been often used in ADME-Tox studies due to the existence of compounds with known pharmacokinetic properties available for generating predictive models.Areas covered: This review examines the growth in the use of some ML techniques in ADME-Tox studies, in particular supervised and unsupervised techniques. Also, some critical points (e. g., size of the data set and type of output variable) must be considered during the generation of models that relate ADME-Tox properties and biological activity.Expert opinion: ML techniques have been successfully employed in pharmacokinetic studies, helping the complex process of designing new drug candidates from the use of reliable ML models. An application of this procedure would be the prediction of ADME-Tox properties from studies of quantitative structure-activity relationships or the discovery of new compounds from a virtual screening using filters based on results obtained from ML techniques.
Activin-like kinase 5 (ALK-5) receptor represents an attractive object to treat cancer. Analyses on the quantitative structure-activity relationship were performed to explore the relationship between the molecular structure of 1,5-naphthyridine, pyrazole and quinazoline derivatives and the inhibition of the activin-like kinase 5. From a data set containing 59 compounds, various electronic descriptors were calculated using density functional theory (DFT) method; stereochemical descriptors (as molecular volume and area), polar surface area (PSA), log P and dragon descriptors were also calculated. The ordered predictor selection (OPS) algorithm, weighted principal component analysis (PCA) and Fisher's weights (FW), combined with sequential forward selection, were employed to select the most relevant descriptors to be employed in all partial least square regressions. Using this procedure, we selected the most informative descriptors and significant correlation coefficients were achieved (r(2) = 0.74, q(2) = 0.83). Additional validation tests were carried out, indicating that the obtained model is robust and reliable and, consequently, it can be used to predict the biological activity of new compounds.
Researches in Medicinal Chemistry's area have focused on the search of methods that accelerate the process of drug discovery.Among several steps related to the process of discovery of bioactive substances there is the analysis of the relationships between chemical structure and biological activity of compounds.In this process, researchers of medicinal chemistry analyze data sets that are characterized by high dimensionality and small number of observations.Within this context, this work presents a computational approach that aims to contribute to the analysis of chemical data and, consequently, the discovery of new drugs for the treatment of chronic diseases.Approaches used in exploratory data analysis, employed in this work, combine techniques of dimensionality reduction and clustering for detecting natural structures that reflect the biological activity of the analyzed compounds.Among several existing techniques for dimensionality reduction, we have focused the Fisher's score, principal component analysis and sparse principal component analysis.For the clustering procedure, this study evaluated k-means, fuzzy c-means and enhanced ICA mixture model.In order to perform experiments, we used four data sets, containing information of bioactive substances.Two sets are related to the treatment of diabetes mellitus and metabolic syndrome, the third set is related to cardiovascular disease and the latter set has substances that can be used in cancer treatment.In the experiments, the obtained results suggest the use of dimensionality reduction techniques along with clustering algorithms for the task of clustering chemical data, since from these experiments, it was possible to describe different levels of biological activity of the studied compounds.Therefore, we conclude that the techniques of dimensionality reduction and clustering can be used as guides in the process of discovery and development of new compounds in the field of Medicinal Chemistry.
Diabetes affects approximately 4% of world's population and metabolic syndrome has been directly related to obesity. There is a class of nuclear receptors, peroxisome proliferator-activated receptors (PPARs), which controls the metabolism of carbohydrates and lipids. It has been considered an attractive target to treat diabetes and metabolic syndrome. Accordingly, the primary objective of this study was to employ molecular modelling techniques to understand the factors involved in PPARδ activation. The QSAR models obtained showed good internal and external consistency and presented good validation coefficients (QSAR: q(2) = 0.83, r(2) = 0.87; HQSAR: q(2) = 0.73, r(2) = 0.90; CoMFA: q(2) = 0.88, r(2) = 0.94). The selected properties and the contour maps described the possible interactions between the PPARδ receptor and its agonists. From these findings, it is possible to propose molecular modifications to design new compounds with improved biological properties.