
Simulation studies were performed to assess the performances in realistic estimation problems of B-splines (BS) and restricted cubic splines (RCS) within flexible models for survival data. Several theoretical distributions, designed to match the estimates from clinical cancer studies, were adopted for data generation, including a proportional hazards (PH) model and a non-PH one. The simulation plan included sample sizes equal to 500 and 2000, with 10 ^2 ), and the ability to identify the correct number of peaks of the target hazard, along with the ability to detect peak times and peak heights within clinically relevant ranges. We found reduced performances in the pattern identification task, even for the simplest distributional models, for which estimation error was optimal. Overall, BS slightly outperformed RCS in pattern identification, while estimation error was equivalent. These results will be compared to further simulations concerning a more extensive set of smoothing methods.
Furthermore, accelerating the process of interpretation of forensics DNA analysis can be crucial for faster and more precise identification, especially in cases where time and accuracy are critical, and improving forensic analysis and decision-making.
This study aims to evaluate haematological parameters of patients who have contracted the COVID-19 infection. In particular, we considered patients who were already hospitalised at the time of nasopharyngeal swab (NF) testing, and patients at the Emergency Department and/or who required hospitalisation following a positive result from the NF swab. The collected data are defined as longitudinal data (i.e. constituted by measurements accumulated sequentially over time), mainly characterised by observation times that are irregular, different for each patient, and distributed non-uniformly across the observation interval. In light of these considerations, we exploit CONNECTOR, a data-driven framework designed for longitudinal data, which returns a grouping of the curves of the haemochromocytometric parameters, based on a functional clustering algorithm. These clusters are analysed in terms of disease outcome and survival, finding a good correlation to mortality rates. Finally, the CONNECTOR clusters are exploited to stratify the patients based on profiles of the haemochromocytometric parameters evolution over time, showing that comorbidities appear to have an impact on mortality independently of the outcome of the monitoring carried out through the laboratory tests considered.
In medical imaging, several deep learning models, such as vision transformers (ViT), have shown improved capability in recognizing image patterns efficiently by enhancing model efficiency in parameter optimization and sample effectiveness. Our research introduces a fusion approach that leverages multi-backbone pre-trained models (ResNet, EfficientNet, VGG) as feature extractors and the Caenorhabditis elegans' pyramid connectome ViT in the tail. In most cases, the proposed model demonstrates superior performance on MEDMNIST2D V2.0 classification challenges. Moreover, our approach maintains comparable performance while utilizing fewer training parameters than conventional state-of-the-art models. This indicates a significant step toward more efficient deep-learning architectures in medical diagnostics.
Feature (variable) selection methods are used to detect the most important features (variables) within high-dimensional data. Tree-based models, such as Random Forests, are often exploited for this purpose, as they provide a built-in mechanism to quantify feature importance. However, the stochastic sampling strategies used in these models can lead to unstable feature importance rankings, particularly when the number of trees is low. In our study, we investigate the extent to which these unstable feature rankings can be consolidated through rank aggregation and consensus signal techniques. We propose to compute consensus values from multiple feature selection runs, where each run generates a ranked list of features. We have evaluated our approach while varying a spectrum of hyperparameters such as the number of trees and the number of features available for splitting a node. Our results suggest that consensus ranks provide a more accurate and robust selection of features compared to single-run feature selection procedures. The proposed approach is especially relevant for biomarker discovery, as it can improve the accuracy and reliability of feature selection, leading to the identification of the most informative and relevant biomarkers associated with a particular disease or condition. By consolidating the results from multiple feature selection runs, our approach may help to overcome problems associated with noisy or complex data because it can provide more robust and accurate estimates of the feature importance rankings.
In the medical field, image classification is crucial for identifying respiratory diseases. Researchers propose combining the NIH Chest X-rays dataset with the Chest X-Ray Images dataset, creating a 14-class dataset including conditions like pneumonia, viral infections, and COVID-19. The goal is image classification. Given that an image can have multiple labels, the problem is treated as multilabel classification, transformed into a binary classification where each sample can have one or more labels. In AutoML, this task is framed as multiclass classification. The primary techniques employed are transfer learning and AutoML. Transfer learning fine-tunes pre-trained CNN models on the target dataset, while AutoML optimizes the entire process using automated pipelines. Existing studies differ in their focus, some employing transfer learning on all 14 classes, while others using AutoML for only three. Evaluation metrics like AUROC, precision, F1-score, and recall are used to assess performance, providing insights into accuracy and prediction quality. By comparing results from transfer learning and AutoML, researchers aim to determine the most effective approach for accurately classifying respiratory diseases in medical images.
We have ideated a novel mechanism of how odorants with differing and electronic-structural features can excite the same olfactory receptor. We explored the conformational space of an odorant and identified its likely "activating" or "inhibiting" conformers. The eventual challenge is to provide a molecular glimpse into how a few hundred olfactory receptors can discern several thousand odorants and odors-of various structures and functional groups. For a functionally characterized mouse olfactory receptor, mOR-EG, we tested odorant molecules with experimentally determined activating strengths: strong and weak. Vanillin-an odorant that strongly activates mOR-EG was used as a reference. The conformational space of each test odorant was explored and superimposed over the structure of the reference odorant. By narrowing our studies to mOR-EG's strong activator, eugenol and weak activator, allyl benzene, we identified specific molecular features among conformations that were reproduced even for disparate-looking odorants. It is likely then that a specific conformation catalyzes OR activation leading to the perception of an odor. We explore activating conformers as opposed to an odorant's gross chemical features such as aromatic rings, aliphatic chains, degree of bond-saturation, etc. Our studies can be extended to drug-design, primarily because olfactory receptors have a putative structure that is typical of type-A G- Protein Coupled Receptor. To accomplish the tasks represented in this work the BioSolveIT suite of software was used.
Single variant Genome-Wide Association Studies (GWAS) are the most common data-driven approach to discover genetic variants associated to phenotypes. However, in complex diseases, single variants often have no effect unless they coexist with other variants. Conversely, machine learning (ML) can model potential interactions among variants, potentially addressing the missing heritability problem in these diseases. Nevertheless, the curse of dimensionality must be considered, given the tremendous number of variants in genomic datasets, requiring feature selection techniques to reduce the number of features. This study aims at a preliminary benchmark of the Relevance-Redundancy assessment (ReRa) feature selection method using a public genetic dataset of a Parkinson's cohort. Obtained results demonstrated that ReRa can achieve performances comparable to common filter-based feature selection techniques, with the benefit of building simpler models with fewer features.
Fuzzy granulation is a technique that allows us to represent complex data in a more interpretable and simple way. Granules provide a higher level of abstraction by capturing the essence of the data within specific intervals or ranges. In this work we propose Electrodermal activity (EDA) feature extraction through fuzzy granulation to classify four different scenarios of academic stress. EDA is a physiological signal controlled by the sympathetic nervous system, for this reason it is considered an important component of animical states as stress. According to the research carried out, the detection of stress using EDA is based on the use of morphological attributes or statistical characteristics obtained both in the frequency domain and in time domain. As a time series EDA data often exhibit complex patterns, traditional feature extraction methods may not fully capture these variations, and fuzzy granulation offers an alternative. We have experimented with four different techniques of fuzzy granulation, our results allow us to verify that fuzzy granulation is a useful approach to characterize EDA signals to adequately distinguish the effects of listening to music while performing a stressful task reaching classification accuracy up to 98.9%.
RNAs are single-stranded molecules that fold into themselves, determining a complex shape to perform their biological functions. Considering the chemical bonds established, such shapes can be abstracted into secondary structures, which are tractable from a computational point of view and encode valuable biological information. The analysis of such structures, including comparison and classification, plays a fundamental role in different biological studies. Unfortunately, the available tools take secondary structures as input using different formats, making the translation among different them a necessary step in every analysis. In this work, we propose TARNAS, a software that permits the translation of secondary structure formats, including BPSEQ, CT, Dot-Bracket, RNAML, FASTA (only primary structure) and Arc-annotated Sequence. TARNAS also allows the abstraction of RNA secondary structures into three views, namely Core, Core Plus and Shape. Finally, TARNAS permits to delete or retain comments, blank lines and headers of the files. TARNAS is developed as a standalone desktop application and as a web app. The tool, developed in Java, is available as a standalone application at https://github.com/bdslab/TARNAS or as a web application at https://bdslab.unicam.it/tarnas/. The standalone version allows the processing of large sets of RNA secondary structures in a batch fashion, whereas the web version translates one molecule at a time.
A gene fusion is a chromosomal aberration from juxtaposing separate genes. Since some gene fusions are involved in tumorigenesis, proper gene fusion investigation and analysis are crucial in the literature. After DNA/RNA sample extraction, detecting gene fusions requires first gene fusion detection tools, which usually provide many false positives. Given the high experimental costs in wet lab validation of a single fusion, gene fusion prioritization tools were made available over the years to significantly narrow down candidate gene fusions for validation (e.g., Oncofuse, Pegasus, DEEPrior, ChimerDriver). Although a few reviews about gene fusion detection tools are available, a benchmark on prioritization tools is not available yet in the literature. The aim of this paper is twofold: 1. to provide a curated dataset for a fair gene fusion prioritization tool evaluation. 2. to develop a proper comparison based on time, resources, and tool confidence on selected gene fusions. Based on this benchmark, it can be stated that ChimerDriver is the most reliable tool for prioritizing oncogenic fusions.
SARS-CoV-2 has become an endemic disease, and we will have to face the continuous rise of new variants. Designing and evaluating the effects of new containment policies is of primary importance to keep social activities going as safely as possible according to the different stages of the pandemic. Therefore, we propose an Agent-Based Model to study the evolution of SARS-CoV-2 spread in a well-defined environment (of small/medium size, like a shop, a restaurant, an office, a school, with fewer than a hundred or a few hundred people) to assess the efficacy of different non-pharmaceutical interventions and vaccination strategies. Specifically, we focused on schools, given that the COVID-19 quarantine has resulted in substantial disruptions to education, leading to a transition to remote learning and worsening educational inequalities. We consider using face masks and several real-world testing protocols combined with quarantine policies. All protocols/policies have been evaluated at various stages of the pandemic evolution. Results show that testing campaigns are effective as far as the testing process is faster than the virus diffusion. Also, vaccination campaigns covering less than 40
We propose a rigorous and efficient method for evaluating homophily and heterophily in edge-weighted networks. In a network with nodes partitioned into classes, homophily (resp., heterophily) is defined as the tendency to have edges between nodes in the same class (resp., in different classes). Assuming a suitable null model, we provide a closed formula for the z-score of the total weight of homophilic/heterophilic edges for each class/pair of classes. The z-score directly measures how much this weight deviates from its expected value under the null model. In addition, we also propose a global homophily measure, that gives a significant score of how the set of all classes at a glance tend to be homophilic. The proposed statistics can be computed for very large networks since, as we show, they can be efficiently computed in a data streaming setting. For a network with n nodes and m edges, our algorithm only needs O(n) internal memory space, optimal O(m) worst case time, and a single scan of the m input edges, in any order, is required. Experimental results are shown on ten Protein-Protein Interaction networks, reporting homophily w.r.t. protein functional classes.
Reducing the high dimensionality of the original feature space through the use of feature selection algorithms is crucial in gene-expression-based predictive tasks to potentially improve performance and provide a better understanding of each feature's power and biological meaning. Feature selection approaches like LASSO and other embedded techniques select small subsets of relevant features based solely on their quantitative contribution and predictive power, often leading to the selection of features with limited biological relevance. This work aims to provide a wide exploratory analysis of LASSO feature selection to assess the effects of different hyper-parameters on the selection of the most relevant features and their corresponding biological significance. Then, it introduces a new approach that can guide LASSO in the selection of the features by considering their predictive power as well as their biological relevance. With this intention, this work proposes a novel Gene Information Score to estimate each gene's biological relevance and shows its use in enhancing the feature selection.
Ordinary Differential Equations (ODEs) and Agent-Based Models (ABMs) represent nowadays the two main approaches for Immune System (IS) modeling. While the former approach does not allow for representing aleatory variations, the latter lacks a clear well-defined semantics, entailing possible biases on simulation results. We present here the application of our modeling pipeline, that has been designed to cope with these shortcomings, to a case-study about the competition between cancer and IS under the administration of a pre-clinical vaccine in transgenic mice. The pipeline involves the use of Extended Stochastic Symmetric Nets (ESSN) for a formal definition of the conceptual model, and allows to study the domain problem from a macro-perspective by means of the Stochastic Simulation Algorithm (SSA) or from a micro-perspective through an Agent Based Model with a clear defined semantics. The numerical results obtained in this study using SSA are presented and global sensitivity analysis is performed using Latin Hypercube Sampling - Partial Rank Correlation Coefficients (LHS-PRCC) to analyze and improve vaccine dosages and timings.
Nested named entity recognition (NER) is crucial in processing Chinese electronic medical records (EMRs). Recently, the BERT-based model using CNN and a multi-head Biaffine decoder has shown promising results in nested NER on news datasets. However, this model faces difficulties in dealing with the complex and unevenly distributed entities in Chinese EMRs, resulting in prediction errors. This paper proposes an MC-BERT-CGC model based on MC-BERT semantic features comprising Context-Gated Convolution and multi-head Biaffine decoder. Our model initially incorporates Chinese medical language knowledge by leveraging MC-BERT to represent medical descriptions as sentence vectors. We then use Context-Gated Convolution to accurately define the boundaries of nested entities by learning overlapping relationships between different entities. Finally, we use Focal Loss to classify difficult-to-distinguish entities. Experimental results tested on our Chinese EMRs and the CMeEE-V2 dataset show that our model performs better than existing baseline models in Chinese medical NER tasks. The impacts of this study on the life of patients are significant, as more accurate and detailed medical information can be extracted from EMRs, potentially leading to improved diagnoses, personalized treatment recommendations, and proactive identification of health risks. Our code is available at https://github.com/ymlmorning/MC-BERT-CGC.
Quality control (QC) is fundamental in single-cell RNA sequencing (scRNA-seq) data analysis pipelines to ensure data reliability. A critical QC step involves identifying damaged cells using quality metrics like the percentage of mitochondrial genes or the total number of reads. However, automatically determining the threshold of these metrics for filtering damaged cells can be challenging. Moreover, using this metric alone may result in the removal of biologically meaningful cells. This study aims to find alternative biomarkers to improve the identification of damaged cells, focusing on gene lists other than mitochondrial genes. We hypothesized that genes localized within other organelles, particularly the nucleus, would exhibit similar enrichment patterns as mitochondrial genes. To test this hypothesis, we used a public scRNA-seq dataset where damaged cells were labelled via optical inspection. We considered as potential descriptors the percentage of genes from various lists, in particular lists of transcripts detected within the nucleus. We built a binary logistic regression model to differentiate damaged cells from good cells and evaluated its performance. Our results showed that the traditional criteria, such as mitochondrial genes, number of genes, and total counts, successfully identified damaged cells but tended to overestimate damage. Our findings suggest that although standard features are effective, their poor precision can be problematic. Incorporating other gene lists, particularly those related to nuclear transcripts, into classification models can improve the prediction of damaged cells. Further investigation is needed to understand the underlying mechanisms driving these relationships.
Given the significant improvement in flow cytometry technologies, a massive dataset has been acquired, measuring a larger number of cellular markers. Managing and analysing these data by hand, based on the sole expertise of human operators, is no longer feasible. Recent literature suggests exploiting machine learning algorithms to automate data analysis for extracting and counting sub-populations of cells and supporting diagnosis. In this paper, we applied Support Vector Machine, XGBoost, Decision Tree, Logistic Regression, and Multi-Layer Perceptron to identify three specific cellular types: Lymphocyte T, B, and T cytotoxic. Performances are promising across all models and experiments, with a balanced accuracy above 0.85. Moreover, when looking at the Recall and the F1 score, the Decision Tree is the unique model with values below 0.8 in the classification B Lymphocyte. Moreover, to improve the interpretability of the trained models, we computed the SHAP-based explanations for the XGBoost, Decision Tree and Multi-Layer Perceptron, obtaining a set of extracted features that domain experts recognised as significant for the three classification tasks, thus emphasising the viability of this approach in automating the gating process in flow citometry.
Stem cells play a central role in the development of organisms; hence, studying their gene regulatory networks (GRNs) is of great importance. The Reasoning Engine for Interaction Networks (RE:IN) is a toolset that supports modeling of GRNs to investigate their dynamics systematically and efficiently and make new predictions. Here we constructed a RE:IN model of the GRN which describes the regulation of gene expression in purple sea urchin stem cells. It consists of a constrained abstract Boolean network - a collection of Boolean networks, each corresponding to a possible structure and logic of the GRN consistent with experiments. We examined the model's compatibility with observed behavior and explored its robustness. To this end, we developed several new methods for modeling GRNs in RE:IN. These include methods for handling cases where models don't behave in accordance with observed behavior, tools for fast and efficient RE:IN modeling, and tools for synthesizing complex conditions in which models can be tested. Our results show that the current model cannot behave according to the entirety of the expected behavior. Moreover, we show that the model is robust to perturbations in a subset of key genes in the network. These results suggest that there is still work to be done to better capture the intricacies of this GRN in RE:IN. Furthermore, the tools we developed proved to be useful and may serve future research on GRNs within the RE:IN framework.
The integration of bioinformatics and molecular biology has advanced the search for biomarkers in several diseases. In this study, we searched for possible microRNA biomarkers related to Ulcerative Colitis, a chronic and progressive immune-mediated inflammatory condition characterized by inflammation of the gastrointestinal tract. Starting from a set of public datasets, we collect a set of miRNAs and analyze the possible interactions between miRNAs, their target genes, and the associated enriched pathways, related in particular to inflammation and oxidative stress.