
This article observes fish acoustic telemetry RMS acceleration data as stochastic process system. The system is tested for stationarity and differentiability. In order to access continuation of the stationary stochastic system, it is considered Lagrange mean theorem and Wiener-Levy process. The continuation is tested on transition matrix of probability distributions between cause and product state levels of acceleration. The Wiener-Levy process estimates the diffusion system as the transition distributions with finer resolution, while preserving the mean value and standard deviation of the original distribution. The approach may help to simulate and understate the behavior of fish in captive.
This study evaluates the capabilities of artificial intelligence (AI) in scientific research, focusing on the prediction of breast cancer recurrence using MRI data. Here, ChatGPT 4o was employed to propose complementary analytical approaches, demonstrating potential in generating valid statistical methods such as PCA, K-means, Random Forest, and Cox models. However, significant limitations were observed in its ability to perform numerical calculations accurately. For instance, discrepancies in silhouette coefficients and variable importance rankings highlighted the AI’s unreliability in executing complex analyses without supervision. While ChatGPT 4o proved useful for hypothesis generation and experimental design, its lack of transparency and consistency underscores the need for human oversight. These findings suggest that AI can improve scientific research but is not yet ready to replace human expertise in rigorous, unsupervised analysis. The study calls for a balanced integration of AI tools, emphasizing ethical and methodological considerations in the evolving landscape of automated research.
Scientific research and practice of sports selection confirm the importance of peak height velocity (PHV) and maturity offset (MO) for the development of physical fitness and training programs for young athletes. This requires a method for assessing these indicators that combines maximum validity, ease of use, and confirmed results of testing in wide groups of athletes. In this paper, we aimed to test the interchangeability of two empirical equations for assessing maturity: Mirwald’s equation, created in 2002 and requiring measurements of standing height, body weight, sitting height, and leg length, and Moore’s equation, introduced in 2015 and requiring only standing height to be measured. The study involved 314 healthy non-athletic boys aged 13 to 17 years and 403 football players aged 8 to 18 years. Athletes were divided into 3 age groups: school childhood, puberty, and adolescence. All volunteers had their standing height, sitting height, leg length, and body weight measured. Based on the obtained data, peak height velocity (PHV) and maturity offset were calculated using Mirwald’s equation and Moore’s equation. Bland-Altman analysis was used to assess the consistency of these maturation parameters. The use of the simpler Moore’s equation leads to underestimation of MO and overestimation of age of PHV by an average of 2 years in subjects from all groups. The maximum error (3 years) of this method was found in adolescents aged 15 to 17 years. Replacing Mirwald’s equation with Moore’s equation for the purpose of simplification may lead to false positive predictions regarding age of PHV and overestimation of physical capabilities in a given athlete.
Brain gliomas, including astrocytomas, oligodendrogliomas, and mixed gliomas, present diagnostic and prognosis challenges. This study leverages RNA-Seq data obtained from TCGA database and machine learning approaches to distinguish between these glioma subtypes and tumor grades. By combining differential expression analysis with feature selection methods such as minimum Redundancy Maximum Relevance (mRMR) and Random Forest, implemented via the KnowSeq package, we identified compact and informative gene expression signatures. These signatures demonstrated strong potential for improving the distinction between glioma types and grades, yielding promising results. The best binary classification model for glioma subtype achieved an average accuracy of 84.15
Metagenomics has revolutionized the study of microbial communities by facilitating direct genetic analysis of environmental samples. This article proposes a novel metagenomic pipeline that integrates alignment-free k-mer screening using Mash Screen with precise approximate mapping via MashMap. The pipeline dynamically constructs targeted reference databases, reducing computational overhead while enhancing representativeness. Taxonomic assignments are strengthened by a weighted strategy based on alignment coverage and confidence scoring. Results demonstrate excellent F1 scores, particularly for prokaryotes and viruses, with scores exceeding 0.95 across all taxonomic levels. The pipeline processes most datasets in under 60 min with a memory footprint of approximately 5.82 GB, making it suitable for high-throughput analyses. By addressing the limitations of existing tools, this pipeline offers a scalable, accurate and user-friendly solution for the analysis of microbial communities, advancing our understanding of complex ecosystems and their functional potential.
Tyrosinase plays a crucial role in the biosynthetic pathway of melanogenesis, the process through which melanin, the pigment responsible for the color of skin, hair, and eyes, is produced in humans. A series of chemical compounds were assessed through computational methods to predict the best chemical hits that might potentially inhibit the tyrosinase activity in the enzyme screening assay. Multiple computational tools and servers were used to check the chemoinformatic and drug-like behavior. Furthermore, molecular docking studies were utilized to check the interaction profile of chemical compounds against tyrosinase through docking energy values and bonding patterns. Finally, the MD simulations were performed to check the stability behavior of docked complexes by computing root mean square deviation/fluctuation (RMSD/RMSF), solvent accessible surface area (SASA), and radius of gyration (Rg). Overall, our results indicate that the compound #8 N-(4-fluorophenyl)-2-(5-(2-fluorophenyl)-4-(4-fluorophenyl)-4H-1,2,4-triazol-3-ylthio) acetamide showed promising results in all evaluations and could be used as a novel therapeutic molecule for developing drugs against melanogenesis.
Accurate classification of skin lesions into benign or malignant categories is crucial for early diagnosis and treatment of dermatological conditions. In this study, we present a comprehensive evaluation of Vision Transformer (ViT) models for binary classification tasks using a curated subset of the DermNet dataset. By leveraging a pre-trained ViT model fine-tuned on domain-specific data, we achieve a test accuracy of 94
Personalized dietary planning became the state of art method for weight loss and management. Recent efforts to compute personalized meal-plans focused on the use of machine learning and large language models. These tools achieved a good level of accuracy and could serve as a starting point if it is complemented by physician review. However, these tools suffer from different drawbacks that limit their use for personalized planning. These briefly include accurate calory computation, compliance with USDA guidelines, and do not take the culture specific components in terms of food choices, eating habits, or budget into consideration. In this paper, we introduce PMP-LLM (Personalized Meal Planner – Large Language Models) as a culturally specific AI-based system to generate safe, individualized meal plans. It is based on leveraging large language models (LLMs) combined with structured nutritional intelligence. Our system has been trained and evaluated over datasets composed of real and virtual cases. Its performance is comparable to dieticians and could cover the cultural aspects more than the general purpose LLMs tools.
Accurate prediction of Drug–Target Interactions (DTIs) plays a central role in accelerating drug discovery and improving therapeutic interventions. In this paper, we introduce a novel deep learning framework that integrates specialized Convolutional Neural Networks (CNNs) and Fully Connected Neural Networks (FCNNs) to extract and learn complex molecular features from both SMILES representations and protein amino acid sequences. Unlike conventional methods that rely primarily on one-hot encoding, our approach leverages dense, learned embeddings to capture nuanced structural and physicochemical relationships. We benchmark our model on widely used datasets, including BindingDB and DAVIS, both of which focus on drug–target binding affinity. The results demonstrate high specificity and accuracy across diverse data splits, albeit with some variation in sensitivity and precision depending on the dataset and encoding strategy. These findings suggest that while embedding-based representations offer significant potential for capturing subtle interaction patterns, carefully tuning hyperparameters and considering dataset characteristics remain pivotal for optimal performance. Altogether, our study highlights the promise of advanced deep learning techniques in streamlining virtual screening, guiding in vitro validation, and ultimately expediting the drug design pipeline.
High-throughput sequencing technologies continue to generate massive amounts of genomic variant data, often stored in the widely adopted Variant Call Format (VCF). Despite the availability of powerful comprehensive frameworks such as BCFtools or GATK, researchers and developers may require smaller, more flexible tools to tackle specialized tasks or quickly prototype novel pipelines. In this paper, we present VCFX, a modular suite of C++ utilities designed to support the entire lifecycle of variant data analysis, from filtering and annotation to merging, phasing, and structural variant manipulation. Adhering to a minimalist philosophy reminiscent of Unix pipelines, VCFX components stream VCF records via standard input/output, minimizing intermediate disk usage and simplifying integration with existing genomic workflows. By combining minimal software dependencies, concise code, and an approach tailored to HPC-friendly pipelines, our approach seeks to lower barriers to large-scale variant processing while granting researchers granular control over custom analyses. The code of the toolkit is available at https://github.com/ieeta-pt/VCFX .
Infertility treatment begins with an interview, during which all medical information related to the patient’s history and diagnostic test results are collected. Next, a physician may determine the most appropriate direction for further treatment. We hypothesize that a task-specific deep learning (DL) model could support clinical decisions. However, there is a lack of Polish medical language corpus and language-specific medical DL models. Thus, we tested three general-use Natural Language Processing (NLP) models trained on Polish word collections: (i) bert-base-polish-cased-v1; (ii) herbert-base-cased; (iii) polish-roberta-base-v2. We evaluated the efficacy of the models using 80 Electronic Health Records (EHRs) with information obtained during the medical interview to predict the future stages of infertility treatment. The best Polish medical text classification model was the bert-base-polish-cased-v1 trained on pre-processed data. The lack of a Polish medical corpus decreased the results’ interpretability. Nevertheless, we believe that thanks to the developed model, it will be possible to estimate the pregnancy and decisions on the choice of treatment method more precisely.
The accumulation of amyloid-beta (A β ) plays a critical role in Alzheimer’s Disease progression. We present a mathematical framework using stochastic Smoluchowski-type models to investigate the aggregation and degradation dynamics of A β in the brain. The framework allows us to capture the evolution of monomers, dimers, and higher-order polymers, emphasizing their interplay under physiological and pathological conditions. To mitigate the pathological effects of A β aggregation, we formulate an optimal control problem incorporating control inputs designed to reduce dimer and polymer densities. The objective function minimizes polymer density over the treatment period and dimer density at the final time, reflecting immediate and long-term therapeutic goals. Numerical simulations explore the effects of key parameters on A β dynamics, including clearance rates, weight factors, and control input limits. Our results demonstrate the effectiveness of optimal control strategies in reducing pathological aggregates while preserving system stability.
The escalating cost and protracted timelines of traditional drug discovery have spurred the adoption of computational strategies, especially those leveraging artificial intelligence (AI). In particular, deep learning (DL)-based de novo design has shown promise in generating novel drug-like molecules. However, rigorously validating these AI-generated compounds remains a significant hurdle. To address this challenge, we present an integrative workflow that combines molecular docking and molecular dynamics (MD) simulations for the systematic evaluation of deep learning-generated ligands targeting two pharmacologically critical proteins: the Adenosine A _2 A Receptor (A2aR) and Ubiquitin-Specific Protease 7 (USP7). We first benchmarked six docking tools (AutoDock 4, AutoDock FR, AutoDock Vina, LeDock, PLANTS, and rDock) against reference ligands to identify those with the highest predictive accuracy. After selecting top-performing tools, we screened AI-generated compounds and applied exponential consensus scoring to refine hit prioritization. Finally, the most promising candidate complexes were subjected to all-atom MD simulations to assess binding stability and interaction fidelity. Our results indicate that AutoDock FR and AutoDock Vina consistently outperform other software in pose prediction, and that consensus scoring substantially improves hit enrichment. By providing a robust, multi-step validation pipeline, this study offers an efficient means to benchmark AI-generated molecules, ultimately facilitating more reliable drug candidate selection in both early-stage discovery and repurposing efforts.
Mandibular reconstruction, particularly with temporomandibular joint (TMJ) involvement, is a challenging procedure where precise evaluation of condylar displacement is critical for assessing functional and anatomical outcomes. The complexity of this evaluation stems from anatomical variability, surgical techniques, and post-surgical changes, making consistent and reliable assessments difficult. To address this issue, we developed an innovative, standardized procedure for the evaluation of condylar displacement following mandibular reconstruction. The proposed procedure starts from biomedical image analysis and integrates precise segmentation with an innovative standardized alignment technique, ensuring reproducibility across different reconstructive approaches, regardless of the surgical method used. This procedure was validated through a retrospective study of 18 patients, employing a consistent 3D model reconstruction and analysis workflow. The results demonstrate that the developed procedure provides a reliable framework for comparing different reconstructive methods, regardless of whether they employ CAD/AM techniques or conventional approaches, offering a repeatable and adaptable approach for assessing mandibular reconstruction outcomes. This methodology lays the groundwork for future improvements in surgical planning and post-operative evaluation, ultimately contributing to the optimization of both functional and aesthetic results in mandibular reconstruction.
Sickle cell anemia (SCA), a genetic disorder caused by the deformation of red blood cells due to hemoglobin S, constitutes a critical global public health challenge, particularly in areas with a high incidence of malaria and Afro-descendant populations such as the Colombian Caribbean. Traditional diagnostic methods, such as hemoglobin electrophoresis, face economic and logistical obstacles in rural and low-income areas. This work proposes an innovative system that combines medical ontologies with knowledge-based models, such as artificial intelligence, to improve the diagnosis and management of the disease. The developed system uses an ontological model structured into domains such as clinical phenotypes, treatments, genetics (HBB gene mutations), and quality of life. The architecture consists of five layers: ontology processing (OWLAPI), triple database storage (RDF), inference engines (SWRL rules), query interfaces, and data visualization. The results demonstrate the usefulness of semantic rules for alerting patients about disease-modifying factors, monitoring patients’ quality of life, and adjusting treatments. The proposal prioritizes accessibility in underserved regions, minimizing dependence on expensive equipment and specialized personnel, which could transform care in resource-limited settings.
Understanding disease relationships is vital for uncovering causal mechanisms and improving prevention and treatment strategies. Traditional disease networks, often based on Pearson correlation, are limited by their inability to account for confounding effects, making them unsuitable for causal inference. In this study, we present a novel approach using partial correlation to construct a putative causal disease network from comorbidity data in the Human Disease Network (HuDiNe), which includes over 291,000 associations among 995 diseases derived from Medicare claims. Using a James–Stein shrinkage estimator, we computed partial correlations while controlling for the influence of other diseases. This method identified 7,894 statistically significant associations among 697 diseases, capturing both positive and negative relationships. Compared to Pearson-based networks, the partial correlation network was sparser and less modular, highlighting its specificity for direct associations. We validated key findings, particularly the high connectivity of hypertension, through Mendelian randomization studies. In addition to recovering established links (e.g., hypertension with obesity and cardiovascular diseases), the method uncovered novel associations with conditions such as breast cancer, endometriosis, and glaucoma. Negative correlations—such as between diabetes and aortic aneurysm—further demonstrated the method’s ability to detect inverse relationships. Our results show that partial correlation analysis applied to summary-level data offers a promising tool for causal disease network construction. This approach enables hypothesis generation without the need for individual-level data, laying the groundwork for integrative analyses that incorporate genomic and clinical information to refine disease prediction and intervention strategies.
The performance of Single Cell RNA-Sequencing (scRNA-seq) protocols has yet to be evaluated. In this paper, we analyze the performance of scRNA-seq protocols in detecting gene abundance by investigating their dropout, that is the event of inability to detect gene abundance. In our research, we use 33 datasets that contain 689 million gene abundance reads obtained using the most commonly used protocols for scRNA-seq: Tang, SMARTer, and Smart-Seq. Our findings show that there is an increase in dropout rate detected in all datasets when going further in scanning the genes of the samples. Some datasets showed sharp increase in dropout while some showed a slight increase. Notably, the rise in dropout rate starts after passing the first third of the gene groups in the sample.
Objective. The aim of this study was to explore the relationship between the performance in the six-minute walk test and the trabecular bone score (TBS) in a group of older sarcopenic women. Summary of facts and results. – A total of 36 Lebanese older women whose ages range from 60 to 82 years participated in this study. The participants underwent multiple clinical and physical tests including Dual Energy X-ray absorptiometry (DXA) scan, six-minute walk test (6 MWT), handgrip, vertical jump test and physical activity questionnaire. The participants were diagnosed as sarcopenic based on the skeletal muscle mass index results (<5.5 kg/m2) and handgrip test (<20 kg) as determined by the requirements of European Working Group on Sarcopenia in Older People (EWGSOP) to diagnose sarcopenia. L1-L4 TBS was positively correlated with 6MWT (r = 0.42; p = 0.01), vertical jump height (r = 0.35; p = 0.03) and physical activity level (r = 0.38; p = 0.02). In addition, age presented a significant positive correlation with 6MWT performance (r = 0.33; p = 0.04), physical activity level (r = 0.32; p = 0.05) and L1-L4 TBS (r = 0.35; p = 0.033). However, age did not present a significant correlation with vertical jump height (r = -0.11; p = 0.51). Moreover, the positive association between L1-L4 TBS and 6MWT/vertical jump test remained significant after controlling for age. However, the positive association between L1-L4 TBS and physical activity level disappeared after controlling for age. Conclusion. – The present study suggests that the performance in the 6MWT and vertical jump level are predictors of L1-L4 TBS in older sarcopenic women.
This article explores the combined use of immunoglobulin heavy chain variable (IGHV) gene mutation status and DNA sequence entropy analysis to assess survival outcomes in leukemia patients. Leveraging a chronic lymphocytic leukemia (CLL) patient database, we calculated various entropy metrics for individual DNA sequences. A randomized entropy divergence (EnRD) after as a consequence of sequence randomization, focusing on the disruption of neighborhood interrelation information, is employed to estimate patient-specific DNA statistical properties, enabling the categorization of patients into two distinct groups. Survival analyses integrate IGHV mutation subtypes with entropy measures, revealing statistically significant differences between groups with high and low randomized entropy divergence levels, as well as among IGHV subtypes. The findings suggest that patients with mutated IGHV genes, high DNA sequence entropy and high DNA sequence randomized entropy divergence exhibit notably improved survival prospects (Hazard Ratio HR = 5.23, p = 0.001). These results are validated through robust methodologies, including Kaplan-Meier survival curves, Cox regression analyses, and Linear Mixed Model (LMM). The study underscores the utility of combining IGHV mutational status with DNA entropy analysis as a prognostic tool, highlighting the potential to refine patient stratification and improve predictions of disease outcomes.
This paper presents a distributed web architecture for real-time patient monitoring that integrates customer relationship management (CRM) tools to improve clinical management. The proposed framework uses scalable, cloud-based technologies to ensure high availability and seamless data exchange between healthcare professionals, patients, and administrative staff. By incorporating CRM functionalities, the system provides a unified platform for tracking patient interactions, personalizing care, and managing clinical workflows. The approach emphasizes data interoperability and security through standardized communication protocols, facilitating compatibility with existing electronic health records. The results suggest that real-time monitoring, combined with CRM analytics, enables clinical teams to make more informed decisions, streamline resource allocation, and improve patient satisfaction. The flexibility of the architecture makes it adaptable to different healthcare settings, providing a scalable and secure solution for modern telemedicine and future-proof healthcare services.