
Unlabelled:JMIR Bioinformatics and Biotechnology announced a strategic partnership with the MidSouth Computational Biology and Bioinformatics Society (MCBIOS) in late 2025; this partnership establishes JMIR Bioinformatics and Biotechnology as the official journal of MCBIOS. This collaboration reflects a shared commitment to advancing computational biology, bioinformatics, and biotechnology through open science, interdisciplinary collaboration, and real-world data. By connecting MCBIOS' community-building and professional development initiatives with JMIR Publications' open-access publishing platform, this partnership aims to support emerging researchers, accelerate the dissemination of innovative computational methods and artificial intelligence applications, and strengthen data-driven research across medicine and biology.
BackgroundCattle are among the most important livestock resources in Ethiopia, contributing significantly to the agricultural economy and rural livelihoods. They provide meat, milk, hides, draft power for crop production, and serve as a major source of income for farmers. Despite their vital role, cattle productivity is often constrained by various diseases, particularly parasitic diseases. One of the most significant of these is bovine fasciolosis, a condition caused by ingestion of metacercariae of liver flukes belonging to the genus Fasciola. ObjectiveThis study aimed to assess the prevalence and associated risk factors of bovine fasciolosis in Bahir Dar, Ethiopia. MethodsA cross-sectional study was conducted from November 2021 to April 2022. A total of 384 cattle were randomly selected from different locations within the study area. Animals of all age groups and both sexes were included. Fecal samples were collected directly from the rectum of each animal using clean, labeled containers. The samples were examined using standard coprological techniques, specifically the sedimentation method, to detect liver fluke eggs. All findings were recorded, and the data were analyzed using descriptive statistical methods. ResultsThe overall prevalence of fasciolosis was 49.21% (n=189). Based on origin, Sebatamit had the most incidence at 61.84% (n=47), followed by Kebele 11 at 59.37% (n=57), Tikurit at 50% (n=59), and Latammba at 27.65% (n=26). Statistical analysis revealed significant disparities in occurrence among areas. Cattle in poor condition had the largest prevalence (n=80, 64%), followed by medium condition (n=85, 50%) and fat cattle (n=24, 26.96%). This variation was statistically significant. Age-group analysis revealed comparable prevalence rates, with young cattle at 50.38% (n=65), adults at 47.33% (n=71), and elderly cattle at 50.47% (n=53), with no significant differences found. There were no significant sex-related variations in prevalence, with males exhibiting a prevalence of 49.73% (n=93) and females 48.73% (n=96). Local cattle had a slightly higher prevalence (n=111, 51.62%) than crossbreeds (n=78, 46.15%), although the difference was not statistically significant (P=.29). ConclusionsThese findings underscore the need for targeted, location-specific control strategies and highlight the importance of improved nutritional and health management practices to reduce the burden of fasciolosis in cattle populations.
Background:Non-small cell lung cancer (NSCLC) is one of the leading causes of cancer-related mortality. Programmed cell death receptor-1 (PD-1) immunotherapy has shown results in the treatment of NSCLC; however, not all patients respond effectively to it. Identifying predictive biomarkers for PD-1 therapy response is critical to improving patient outcomes and treatment strategies. Traditional methods of biomarker discovery often fall short in terms of accuracy and comprehensiveness. Recent advancements in deep learning provide a powerful approach to analyze complex genomic data to resolve this issue. Objective:This study aims to leverage deep neural networks (DNNs) to identify genomic biomarkers predictive of patient responses to PD-1 immunotherapy in NSCLC. DeepImmunoGene is a model designed using a reduced feature set to identify the most critical biomarkers. We use feature selection to reduce the space and apply deep learning to identify the highly predictive gene subset. Methods:Differentially expressed genes were identified in RNA-seq data from 355 patients with NSCLC using the LIMMA package in R, followed by preprocessing with log2 transformation, removing outliers, and detecting easily identified genes. Machine learning models, including support vector machines, extreme gradient boosting (XGBoost), and DNNs, were applied to gene expression data to predict patient responses to immunotherapy. Key predictive genes were identified through model interpretation techniques, and differences in model performance were assessed for statistical significance. Primarily, the metric used identifies which genes serve as key biomarkers in regard to immunotherapy detection. Results:Initially, we identified 1093 differentially expressed genes from RNA-seq data of 355 patients. We then trained models using SVM, XGBoost, and DNN to predict immunotherapy response. The DNN model outperformed both SVM and XGBoost with an accuracy of 82%, an area under the curve of 90%, and recall of 85%. To identify key biomarkers, we performed a permutation importance analysis, narrowing down the gene set to 98 genes. DeepImmunoGene, trained on these 98 genes, showed superior results, with an accuracy of 87% and an area under the curve of 95%. The top 36 upregulated genes in responders and 62 upregulated genes in nonresponders were identified, which could serve as potential biomarkers for predicting response to PD-1 inhibitors. These findings suggest that DeepImmunoGene can reliably forecast immunotherapy outcomes and aid in biomarker discovery, supporting the development of more personalized treatment strategies in NSCLC. Conclusions:The DeepImmunoGene predictive model identified 36 upregulated genes that may represent candidate genomic biomarkers associated with response to PD-1 immunotherapy in patients with NSCLC. Notably, the 10 most significant genes offer valuable insights into the underlying mechanisms of treatment responses. These biomarkers may not only aid in predicting which patients are more likely to respond to PD-1 immunotherapy but also offer insights into the molecular differences associated with nonresponse.
Background:Plant-derived exosome-like nanovesicles (P-ELNs) effectively deliver bioactive compounds due to their high biocompatibility and low immunogenicity. While liquid chromatography-mass spectrometry (LC-MS) profiles compounds in complex samples, its analysis of large datasets remains limited by traditional methods. Recent advances in large language models (LLMs) and domain-specific systems have enhanced Chinese biomedical data processing and cross-modal pharmaceutical research. Objective:This study aimed to create a multimodal framework of LC-MS combined with DeepSeek models for data mining of compounds with wound-healing properties from exosome-like nanovesicles derived from Cayratia japonica (CJ-ELNs). Methods:LC-MS identified compounds enriched in CJ (n=3) and CJ-ELNs (n=3), and then compounds specifically enriched in CJ-ELNs were filtered via a four-step filtering workflow. The CJ-ELNs-specific compounds were processed by DeepSeek models for screening naturally active compounds with targeted functions of antioxidation, anti-inflammation, anticellular damage, antiapoptosis, wound healing and tissue regeneration, and cell proliferation. Results:A multimodal framework of LC-MS combined with the DeepSeek-DF model was created. With the assistance of artificial intelligence (AI), a total of 46 naturally active compounds derived from CJ-ELNs with targeted functions were identified. Conclusions:A self-designed multimodal framework of LC-MS, combined with DeepSeek models, rapidly and accurately identifies naturally active compounds from CJ-ELNs. This AI-powered system innovatively integrates the traditional analytical technique with modern LLMs, thus greatly favoring data mining of active ingredients in traditional Chinese medicine herbs.
Background:The manual abstraction of unstructured clinical data is often necessary for granular clinical outcomes research but is time consuming and can be of variable quality. Large language models (LLMs) show promise in medical data extraction yet integrating them into research workflows remains challenging and poorly described. Objective:This study aimed to develop and integrate an LLM-based system for automated data extraction from unstructured electronic health record (EHR) text reports within an established clinical outcomes database. Methods:We implemented a generative artificial intelligence pipeline (UODBLLM) utilizing a flexible language model interface that supports various LLM implementations, including Health Insurance Portability and Accountability Act-compliant cloud services and local open-source models. We used extensible markup language (XML)-structured prompts and integrated using an open database connectivity interface to generate structured data from clinical documentation in the EHR. We evaluated the UODBLLM's performance on the completion rate, processing time, and extraction capabilities across multiple clinical data elements, including quantitative measurements, categorical assessments, and anatomical descriptions, using sample magnetic resonance imaging (MRI) reports as test cases. System reliability was tested across multiple batches to assess scalability and consistency. Results:Piloted against MRI reports, UODBLLM processed 1800 clinical documents with a 100% completion rate and an average processing time of 8.90 seconds per report. The token utilization averaged 2692 tokens per report, with an input-to-output ratio of approximately 13:2, resulting in a processing cost of US $0.009 per report. UODBLLM had consistent performance across 18 batches of 100 reports each and completed all processing in 4.45 hours. From each report, UODBLLM extracted 16 structured clinical elements, including prostate volume, prostate-specific antigen values, Prostate Imaging Reporting and Data System scores, clinical staging, and anatomical assessments. All extracted data were automatically validated against predefined schemas and stored in standardized JSON format. Conclusions:We demonstrated the successful integration of an LLM-based extraction system within an existing clinical outcomes database, achieving rapid, comprehensive data extraction at minimal cost. UODBLLM provides a scalable, efficient solution for automating clinical data extraction while maintaining protected health information security. This approach could significantly accelerate research timelines and expand feasible clinical studies, particularly for large-scale database projects.
Background:Autosomal dominant nonsyndromic hearing loss (ADNSHL) is highly heterogeneous, with more than 64 genes implicated in its etiology. This complexity limits the diagnostic power of clinical examinations and audiometry alone, while existing computational approaches have achieved only moderate accuracy and often lack interpretability. As precision medicine increasingly emphasizes genotype-phenotype correlations, there is a recognized need for diagnostic tools that provide clinicians with transparent, interpretable outputs. Objective:This study aimed to develop and evaluate the AudioGene Translational Dashboard, an interpretable clinical informatics tool that integrates machine learning models and interactive visualizations to enhance genotype-phenotype correlations and support diagnostic decision-making in ADNSHL. Methods:We developed the AudioGene Translational Dashboard, integrating 2 machine learning models (AudioGene version 4 and AudioGene version 9.1) with 6 interactive visualization tools. AudioGene version 4 uses a multi-instance support vector machine classifier for patients with multiple audiograms, while AudioGene version 9.1 combines adaptive boosting, k-nearest neighbors, random forest models, and logistic regression for patients with a single audiogram. Visualizations include audiometric profile plots, audioprofile surfaces, clustering analyses, and data distribution charts designed to facilitate clinical interpretation. Results:The AudioGene Translational Dashboard was developed to address the "70/30" phenomenon, indicating a 74% likelihood that the causative gene is among the top 3 predicted genes, thereby providing clinicians with a clear confidence indicator ("green flag") or a caution alert ("red flag") during diagnosis. While this level of performance is well suited for hypothesis generation, the remaining uncertainty underscores the need for interpretive context in clinical decision-making. Visualization tools enhanced clinicians' ability to interpret and correlate phenotypic data with predicted genetic outcomes, improving diagnostic confidence and interpretability. Conclusions:The AudioGene Translational Dashboard advances clinical informatics in genetic diagnosis of ADNSHL by integrating explainable artificial intelligence with interactive visualizations, enhancing clinical interpretability and diagnostic accuracy. This approach facilitates informed clinical decision-making, highlights the translational potential of genotype-phenotype computational models, and supports precision medicine in hearing loss diagnostics. Future enhancements will target improving class balance and incorporating additional user-customizable features to further optimize clinical applicability.
The integration of artificial intelligence (AI) into personalized medicine is revolutionizing drug delivery by transitioning from the traditional “one-size-fits-all” approach to patient-specific therapeutic strategies. This review aims to explore the transformative role of artificial intelligence (AI) in personalizing drug delivery systems by leveraging genomic, proteomic, and metabolic data. A comprehensive literature review was conducted using electronic databases such as PubMed, Scopus, and Web of Science to examine studies related to AI in pharmacogenomics, smart drug delivery, biosensing technologies, and drug repurposing. Inclusion and exclusion criteria were applied, and findings were thematically synthesized. AI facilitates real-time data interpretation and personalized therapy through smart nanoformulations, biosensor-enabled monitoring, and deep learning-based pharmacogenomic modeling. Applications in oncology, diabetes, and neurodegenerative disorders show improved treatment outcomes. However, challenges include data security, regulatory constraints, and interpretability of AI models. AI bridges genomics and pharmaceutics, driving precision medicine. Innovations like AI-driven 3D printing and federated learning promise a new era in personalized healthcare.
Background Adalimumab, a monoclonal antibody targeting tumor necrosis factor α, treats autoimmune diseases but induces antidrug antibodies in 30% to 60% of patients, reducing its efficacy. Objective This study aims to investigate molecular mimicry as a mechanism behind this immunogenicity, where bacterial immunoglobulin domains structurally resemble adalimumab’s light chain, triggering immune responses. Methods Using PSI-BLASTp (National Center for Biotechnology Information) and PRALINE (Center for Integrative Bioinformatics), there are 40 bacterial antigens homologous to adalimumab, with 8 clinically relevant strains. Results Structural analysis revealed 94% amino acid identity between the immunoglobulin domain of Escherichia coli strain B1 and adalimumab’s light chain, and 89.67% similarity with Corynebacterium pyruviciproducens. Root mean square deviation values confirmed strong structural homology. Additionally, 5 cross-reactive B-cell epitopes were predicted, suggesting overlapping surfaces that may promote immune cross-reactivity and antidrug antibody development. Conclusions This study represents a first step toward identifying a potential microbial factor driving antiadalimumab antibody formation. The predicted cross-reactive regions provide specific candidates for further in vitro validation to confirm molecular mimicry and refine epitope mapping. Understanding these mechanisms may ultimately inform the design of less immunogenic biologics and guide clinical strategies to predict and prevent antidrug antibody formation.
Background:Bladder cancer is a disease characterized by complex perturbations in gene networks and is heterogeneous in terms of histology, mutations, and prognosis. Advances in high-throughput sequencing technologies, genome-wide association studies, and bioinformatics methods have revealed greater insights into the pathogenesis of complex diseases. Network biology-based approaches have been used to identify complex protein-protein interactions (PPIs) that can lead to potential drug targets. There is a need to better understand PPIs specific to urothelial carcinoma. Objective:This study aimed to elucidate PPIs specific to papillary and nonpapillary urothelial carcinoma and identify the most connected or "hub" proteins, as these are potential drug targets. Methods:A novel PPI analysis tool, Proteinarium, was used to analyze RNA sequencing data from 132 patients with papillary and 270 patients with nonpapillary urothelial carcinoma from the TCGA Cell 2017 dataset and 39 patients with papillary and 88 patients with nonpapillary urothelial carcinoma from the TCGA Nature 2014 dataset. Hub proteins were identified in distinct PPI networks specific to papillary and nonpapillary urothelial carcinoma. Statistical significance of clusters was assessed using the Fisher exact test (P<.001), and network separation was quantified using the interactome-based separation score. Results:RPS27A, UBA52, and VAMP8 were the most connected or "hub" proteins identified in the network specific to the papillary urothelial carcinoma. In the network specific to the nonpapillary carcinoma, GNB1, RHOA, UBC, and FPR2 were found to be the hub proteins. Notably, GNB1 and FPR2 were among the proteins that have existing drugs targeting them. Conclusions:We identified distinct PPI networks and the hub proteins specific to papillary and nonpapillary urothelial carcinomas. However, these findings are limited by the use of transcriptomic data and require experimental validation to confirm the functional relevance of the identified targets.
Background:Sensitivity-expressed as percent positive agreement (PPA) with a reference assay-is a primary metric for evaluating lateral-flow antigen tests (ATs), typically benchmarked against a quantitative reverse transcription polymerase chain reaction (qRT-PCR). In SARS-CoV-2 diagnostics, ATs detect nucleocapsid protein, whereas qRT-PCR detects viral RNA copy numbers. Since observed PPA depends on the underlying viral load distribution (proxied by the number of cycle thresholds [Cts], which is inversely related to load), study-specific sampling can bias sensitivity estimates. Cohort differences-such as enrichment for high- or low-Ct specimens-therefore complicate cross-test comparisons, and real-world datasets often deviate from regulatory guidance to sample across the full concentration range. Although logistic models relating test positivity to Ct are well described, they are seldom used to reweight results to a standardized reference viral load distribution. As a result, reported sensitivities remain difficult to compare across studies, limiting both accuracy and generalizability. Objective:The aim of this study was to develop and validate a statistical methodology that estimates the sensitivity of ATs by recalibrating clinical performance data-originally obtained from uncontrolled viral load distributions-against a standardized reference distribution of target concentrations, thereby enabling more accurate and comparable assessments of diagnostic test performance. Methods:AT sensitivity is estimated by modeling the PPA as a function of qRT-PCR Ct values (PPA function) using logistic regression on paired test results. Raw sensitivity is the proportion of AT positives among PCR-positive samples. Adjusted sensitivity is calculated by applying the PPA function to a reference Ct distribution, correcting for viral load variability. This enables standardized comparisons across tests. The method was validated using clinical data from a community study in Chelsea, Massachusetts, demonstrating its effectiveness in reducing sampling bias. Results:Over a 2-year period, paired ATs and qRT-PCR-positive samples were collected from 4 suppliers: A (n=211), B (n=156), C (n=85), and D (n=43). Ct value distributions varied substantially, with suppliers A and D showing lower Ct (high viral load) values in the samples, and supplier C skewed toward higher Ct values (low viral load). These differences led to inconsistent raw sensitivity estimates. To correct for this, we used logistic regression to model the PPA as a function of Cts and applied these models to a standardized reference Ct distribution. This adjustment reduced bias and enabled more accurate comparisons of test performance across suppliers. Conclusions:We present a distribution-aware framework that models PPA as a logistic function of Ct and reweights results to a standardized reference Ct distribution to produce bias-corrected sensitivity estimates. This yields fairer, more consistent comparisons across AT suppliers and studies, strengthens quality control, and supports regulatory review. Collectively, our results provide a robust basis for recalibrating reported sensitivities and underscore the importance of distribution-aware evaluation in diagnostic test assessment.
Background:Integrating clinical, genomic, and social determinants of health (SDOH) data is essential for advancing precision medicine and addressing cancer health disparities. However, existing bioinformatics tools often lack the flexibility to perform equity-driven analyses or require significant programming expertise. Objective:We developed AI-HOPE-PM (Artificial Intelligence Agent for High-Optimization and Precision Medicine in Population Metrics), a conversational artificial intelligence system designed to enable natural language-driven, multidimensional cancer analysis. This study describes the development, implementation, and application of AI-HOPE-PM to support hypothesis testing that integrates genomic, clinical, and SDOH data. Methods:AI-HOPE-PM leverages large language models and Python-based statistical scripts to convert user-defined natural language queries into executable workflows. It was evaluated using curated colorectal cancer datasets from The Cancer Genome Atlas and cBioPortal, enriched with harmonized SDOH variables. Accuracy of natural language interpretation, run time efficiency, and usability were benchmarked against cBioPortal and UCSC Xena. Results:AI-HOPE-PM successfully supported case-control stratification, survival modeling, and odds ratio analysis using natural language prompts. In colorectal cancer case studies, the system revealed significant disparities in progression-free survival and treatment access based on financial strain, health care access, food insecurity, and social support, demonstrating the importance of integrating SDOH in cancer research. Benchmark testing showed faster task execution compared to existing platforms, and the system achieved 92.5% accuracy in parsing biomedical queries. Conclusions:AI-HOPE-PM lowers technical barriers to integrative cancer research by enabling real-time, user-friendly exploration of clinical, genomic, and SDOH data. It expands on prior work by incorporating equity metrics into precision oncology workflows and offers a scalable tool for supporting disparities-focused translational research. Five videos are included as multimedia appendices to demonstrate platform functionality in real-world scenarios.
Background The systemic treatment of cancer typically requires the use of multiple anticancer agents in combination or sequentially. Clinical narrative texts often contain extensive descriptions of the temporal sequencing of systemic anticancer therapy (SACT), setting up an important task that may be amenable to automated extraction of SACT timelines. Objective We aimed to explore automatic methods for extracting patient-level SACT timelines from clinical narratives in the electronic medical records (EMRs). Methods We used two datasets from two institutions: (1) a colorectal cancer (CRC) dataset including the entire EMR of the 199 patients in the THYME (Temporal Histories of Your Medical Event) dataset and (2) the 2024 ChemoTimelines shared task dataset including 149 patients with ovarian cancer, breast cancer, and melanoma. We explored finetuning smaller language models trained to attend to events and time expressions, and few-shot prompting of large language models (LLMs). Evaluation used the 2024 ChemoTimelines shared task configuration—Subtask1 involving the construction of SACT timelines from manually annotated SACT event and time expression mentions provided as input in addition to the patient’s notes and Subtask2 requiring extraction of SACT timelines directly from the patient’s notes. Results Our task-specific finetuned EntityBERT model achieved 93% F1-score, outperforming the best results in Subtask1 of the 2024 ChemoTimelines shared task (90%). It ranked second in Subtask2. LLM (LLaMA2, LLaMA3.1, and Mixtral) performance lagged the task-specific finetuned model performance for both the THYME and shared task datasets. On the shared task datasets, the best LLM performance was 77% macro F1-score, 16% points lower than the task-specific finetuned system (Subtask1). Conclusions In this paper, we explored approaches for patient-level timeline extraction through the SACT timeline extraction task. Our results and analysis add to the knowledge of extracting treatment timelines from EMR clinical narratives using language modeling methods.
Background Deep learning (DL) shows promise for automated lung cancer diagnosis, but limited clinical data can restrict performance. While data augmentation (DA) helps, existing methods struggle with chest computed tomography (CT) scans across diverse DL architectures. Objective This study proposes Random Pixel Swap (RPS), a novel DA technique, to enhance diagnostic performance in both convolutional neural networks and transformers for lung cancer diagnosis from CT scan images. Methods RPS generates augmented data by randomly swapping pixels within patient CT scan images. We evaluated it on ResNet, MobileNet, Vision Transformer, and Swin Transformer models, using 2 public CT datasets (Iraq-Oncology Teaching Hospital/National Center for Cancer Diseases [IQ-OTH/NCCD] dataset and chest CT scan images dataset), and measured accuracy and area under the receiver operating characteristic curve (AUROC). Statistical significance was assessed via paired t tests. Results The RPS outperformed state-of-the-art DA methods (Cutout, Random Erasing, MixUp, and CutMix), achieving 97.56% accuracy and 98.61% AUROC on the IQ-OTH/NCCD dataset and 97.78% accuracy and 99.46% AUROC on the chest CT scan images dataset. While traditional augmentation approaches (flipping and rotation) remained effective, RPS complemented them, surpassing the performance findings in prior studies and demonstrating the potential of artificial intelligence for early lung cancer detection. Conclusions The RPS technique enhances convolutional neural network and transformer models, enabling more accurate automated lung cancer detection from CT scan images.
Background:The protein A disintegrin and metalloprotease (ADAM) domain containing 17, also called tumor necrosis factor alpha-converting enzyme, is mainly responsible for cleaving a specific sequence Pro-Leu-Ala-Gln-Ala-/-Val-Arg-Ser-Ser-Ser in the membrane-bound precursor of tumor necrosis factor alpha. This cleavage process has significant implications for inflammatory and immune responses, and recent research indicates that genetic variants of ADAM17 may influence susceptibility to and severity of SARS-CoV-2 infection. Objective:The aim of the study is to identify the most deleterious missense variants of ADAM17 that impact protein stability, structure, and function and to assess specific variants potentially involved in SARS-CoV-2 infection. Methods:A bioinformatics approach was used on 12,042 single-nucleotide polymorphisms using tools including SIFT (Sorting Intolerant From Tolerant), PolyPhen2.0, PROVEAN (Protein Variation Effect Analyzer), PANTHER (Protein Analysis Through Evolutionary Relationships), SNP&GO (Single Nucleotide Polymorphisms and Gene Ontology), PhD-SNP (Predictor of Human Deleterious Single Nucleotide Polymorphisms), Mutation Assessor, SNAP2 (Screening for Non-Acceptable Polymorphisms 2), MUpro, I-Mutant, iStable, InterPro, Sparks-x, PROCHECK (Programs to Check the Stereochemical Quality of Protein Structures), PyMol, Project HOPE (Have (y)Our Protein Explained), ConSurf, and SWISS-MODEL. Missense variants of ADAM17 were collected from the Ensembl database for analysis. Results:In total, 7 nonsynonymous single-nucleotide polymorphisms (P556L, G550D, V483A, G479E, G349E, T339P, and D232E) were identified as high-risk pathogenic by all prediction tools, and these variants were found to potentially have deleterious effects on the stability, structure, and function of the ADAM17 protein, potentially destroying the entire cleavage process. Additionally, 4 missense variants (Q658H, D657G, D654N, and F652L) in positions related to SARS-CoV-2 infection exhibited high conservation scores and were predicted to be deleterious, suggesting that they play an important role in SARS-CoV-2 infection. Conclusions:Specific missense variants of ADAM17 are predicted to be highly pathogenic, potentially affecting protein stability and function and contributing to SARS-CoV-2 pathogenesis. These findings provide a basis for understanding their clinical relevance, aiding in early diagnosis, risk assessment, and therapeutic development.
Background:The COVID-19 pandemic requires a deep understanding of SARS-CoV-2, particularly how mutations in the spike receptor-binding domain (RBD) chain E affect its structure and function. Current methods lack comprehensive analysis of these mutations at different structural levels. Objective:This study aims to analyze the impact of specific COVID-19-associated point mutations (N501Y, L452R, N440K, K417N, and E484A) on the SARS-CoV-2 spike RBD structure and function using predictive modeling, including a graph-theoretic model, protein modeling techniques, and molecular dynamics simulations. Methods:The study used a multitiered graph-theoretic framework to represent protein structure across 3 interconnected levels. This model incorporated 19 top-level vertices, connected to intermediate graphs based on 6-angstrom proximity within the protein's 3D structure. Graph-theoretic molecular descriptors or invariants were applied to weigh vertices and edges at all levels. The study also used Iterative Threading Assembly Refinement (I-TASSER) to model mutated sequences and molecular dynamics simulation tools to evaluate changes in protein folding and stability compared to the wildtype. Results:A total of 3 distinct predictive modeling and analytical approaches successfully identified structural and functional changes in the SARS-CoV-2 spike RBD (chain E) resulting from point mutations. The novel graph-theoretic model detected notable structural changes, with N501Y and L452R showing the most pronounced effects on conformation and stability compared to the wildtype. K147N and E484A mutations demonstrated less significant impacts compared to the severe mutations, N501Y and L452R. Ab initio modeling and molecular simulation dynamics findings corroborated the results from graph-theoretic analysis. The multilevel analytical approach provided a comprehensive visualization of mutation effects, deepening our understanding of their functional consequences. Conclusions:This study advanced our understanding of SARS-CoV-2 spike RBD mutations and their implications. The multifaceted approach characterized the effects of various mutations, identifying N501Y and L452R as having the most substantial impact on RBD conformation and stability. The findings have important implications for vaccine development, therapeutic design, and variant monitoring. Our research underscores the power of combining multiple predictive analytical approaches in virology, contributing valuable knowledge to ongoing efforts against the COVID-19 pandemic and providing a framework for future studies on viral mutations and their impacts on protein structure and function.
Background:Cancer is one of the leading causes of disease burden globally, and early and accurate diagnosis is crucial for effective treatment. This study presents a deep learning-based model designed to classify 5 common types of cancer in Saudi Arabia: breast, colorectal, thyroid, non-Hodgkin lymphoma, and corpus uteri. Objective:This study aimed to evaluate whether integrating RNA sequencing, somatic mutation, and DNA methylation profiles within a stacking deep learning ensemble improves cancer type classification accuracy relative to the current state-of-the-art multiomics models. Methods:Using a stacking ensemble learning approach, our model integrates 5 well-established methods: support vector machine, k-nearest neighbors, artificial neural network, convolutional neural network, and random forest. The methodology involves 2 main stages: data preprocessing (including normalization and feature extraction) and ensemble stacking classification. We prepared the data before applying the stacking model. Results:The stacking ensemble model achieved 98% accuracy with multiomics versus 96% using RNA sequencing and methylation individually, 81% using somatic mutation data, suggesting that multiomics data can be used for diagnosis in primary care settings. The models used in ensemble learning are among the most widely used in cancer classification research. Their prevalent use in previous studies underscores their effectiveness and flexibility, enhancing the performance of multiomics data integration. Conclusions:This study highlights the importance of advanced machine learning techniques in improving cancer detection and prognosis, contributing valuable insights by applying ensemble learning to integrate multiomics data for more effective cancer classification.
Background:National and ethnic mutation frequency databases (NEMDBs) play a crucial role in documenting gene variations across populations, offering invaluable insights for gene mutation research and the advancement of precision medicine. These databases provide an essential resource for understanding genetic diversity and its implications for health and disease across different ethnic groups. Objective:The aim of this study is to systematically evaluate 42 NEMDBs to (1) quantify gaps in standardization (70% nonstandard formats, 50% outdated data), (2) propose artificial intelligence/linked open data solutions for interoperability, and (3) highlight clinical implications for precision medicine across NEMDBs. Methods:A systematic approach was used to assess the databases based on several criteria, including data collection methods, system design, and querying mechanisms. We analyzed the accessibility and user-centric features of each database, noting their ability to integrate with other systems and their role in advancing genetic disorder research. The review also addressed standardization and data quality challenges prevalent in current NEMDBs. Results:The analysis of 42 NEMDBs revealed significant issues, with 70% (29/42) lacking standardized data formats and 60% (25/42) having notable gaps in the cross-comparison of genetic variations, and 50% (21/42) of the databases contained incomplete or outdated data, limiting their clinical utility. However, databases developed on open-source platforms, such as LOVD, showed a 40% increase in usability for researchers, highlighting the benefits of using flexible, open-access systems. Conclusions:We propose cloud-based platforms and linked open data frameworks to address critical gaps in standardization (70% of databases) and outdated data (50%) alongside artificial intelligence-driven models for improved interoperability. These solutions prioritize user-centric design to effectively serve clinicians, researchers, and public stakeholders.
Artificial intelligence (AI) is poised to become an integral component in health care research and delivery, promising to address complex challenges with unprecedented efficiency and precision. However, many clinicians lack training and experience with AI, and for those who wish to incorporate AI into research and practice, the path forward remains unclear. Technical barriers, institutional constraints, and lack of familiarity with computer and data science frequently stall progress. In this tutorial, we present a transparent account of our experiences as a newly established interdisciplinary team of clinical oncology researchers and data scientists working to develop a natural language processing model to identify symptomatic adverse events during pediatric cancer therapy. We outline the key steps for clinicians to consider as they explore the utility of AI in their inquiry and practice, including building a digital laboratory, curating a large clinical dataset, and developing early-stage AI models. We emphasize the invaluable role of institutional support, including financial and logistical resources, and dedicated and innovative computer and data scientists as equal partners in the research team. Our account highlights both facilitators and barriers encountered spanning financial support, learning curves inherent with interdisciplinary collaboration, and constraints of time and personnel. Through this narrative tutorial, we intend to demystify the process of AI research and equip clinicians with actionable steps to initiate new ventures in oncology research. As AI continues to reshape the research and practice landscapes, sharing insights from past successes and challenges will be essential to informing a clear path forward.