Autism spectrum disorder (ASD) encompasses a range of neurodevelopmental conditions characterized by impairments in social interaction, communication, and behavior. Early detection is critical for timely intervention and improved outcomes; however, conventional approaches remain subjective, labor-intensive, and dependent on highly trained professionals. Facial image-based ASD research also faces challenges such as inter-class resemblance, intra-class variation, and pose differences. To address these issues, we propose the Multi-Aspect Cross-Fusion Transformer (MACFUT), a network that integrates facial landmark and image features through a multi-level cross-fusion mechanism for early ASD screening. The architecture incorporates spatial attention to highlight salient facial regions and a global-local cross-fusion transformer encoder to capture detailed feature relationships. We further analyze facial landmark distances, vectors, and angles to reveal developmental differences and enhance clinical interpretability. The proposed model is trained on the ASD Facial Image Dataset (AFID) using 10-fold cross-validation, achieving 89.56% accuracy and an area under the receiver operating characteristic curve (AUC) of 95.36%. To evaluate cross-dataset generalization, the trained model is tested on an independent dataset from Bangladesh (BACFED), yielding 86.51% accuracy and an AUC of 94.78%. The model also highlights diagnostic-group and sex-based differences in facial features among children. MACFUT provides an interpretable and accessible early ASD screening tool that demonstrates robustness across heterogeneous datasets. This method outperforms existing approaches and may help identify children who could benefit from further clinical evaluation in resource-limited and telehealth settings.
Accurate cell segmentation remains challenging due to morphological variations, diverse imaging modalities, and unclear cellular boundaries. Existing deep learning (DL) methods struggle to extract features adaptively across heterogeneous cellular environments, thereby limiting generalization capacity. To address these challenges, we propose FlexiCell, a novel adaptive segmentation framework that integrates a learnable adaptive filter with dual attention mechanisms. The core innovation lies in the FlexiFilter approach, which combines standard convolution with adaptive residual learning through learnable mixing parameters. These parameters dynamically balance input preservation and feature enhancement. FlexiCell employs multi-scale FlexiFilter blocks with varying kernel sizes, channel and spatial attention networks, and a dedicated boundary extractor for precise edge detection. Extensive experiments demonstrate superior performance compared to benchmark models, achieving 3.8% improvement in detection accuracy and 5.5 % in segmentation quality on our newly developed induced pluripotent stem (iPS) cell datasets. Further evaluation on standardized Cell Tracking Challenge (CTC) benchmarks confirms state-of-the-art performance on mesenchymal stem cells and glioblastoma datasets, outperforming established CTC methods. The framework demonstrates robust generalization across fluorescence, phase contrast, and differential interference contrast microscopy, without requiring dataset-specific optimization. Codes are available at https://github.com/jovialniyo93/FlexiCell.
Neuroblastoma is a prevalent pediatric tumor with a low 5-year survival rate among high-risk patients, and the prognosis remains poor despite available therapeutic interventions. Therefore, identifying novel and effective therapeutic targets is critical for improving outcomes in these patients. In this study, we performed an integrative analysis of two neuroblastoma circRNA sequencing datasets to identify potential drug targets. By comparing circRNA expression levels between neuroblastoma tissues and adjacent normal tissues, we identified differentially expressed circRNAs and subsequently predicted 30 hub circRNAs through Weighted Gene Co-expression Network Analysis. To elucidate the functional roles of these circRNAs, we investigated their interactions with RNA-binding proteins. The results suggest that hsa_circ_0051680 and hsa_circ_0006107 may influence neuroblastoma progression through interactions with FUS and IGF2BP1, respectively. Furthermore, we analyzed the translational potential of the hub circRNAs, revealing that six circRNAs encode proteins with complex secondary structures. Molecular docking analysis identified five high-affinity complexes between circRNA-encoded proteins (hsa_circ_0000786, hsa_circ_0005087, hsa_circ_0006867) and their corresponding ligands. These circRNA-derived proteins present promising novel drug targets for both the diagnosis and treatment of neuroblastoma.
OBJECTIVES:Beta-lactamase is a bacterial enzyme that deactivates beta-lactam antibiotics, and it is one of the leading causes of antibiotic resistance problems globally. In current drug discovery research, molecular simulation, like molecular docking, has been routinely integrated to virtually screen an enzyme inhibitory effect. However, a commonly known limitation of molecular docking is a low percent success rate. Previously, we reported a proof-of-concept of combining machine learning with a quantitative structure-activity relationship (QSAR) model that overcame this limitation ( https://doi.org/10.1186/s13065-024-01324-x ). Here, we presented and navigated the dataset used in our previous report, including sixty trained models (thirty for random forest and another thirty for logistic regression). DATA DESCRIPTION:This data note has three essential parts. The first part is an in vitro beta-lactamase inhibitory screening of eighty-nine bioactive molecules. The second part consisted of three molecular docking approaches (AutoDock Vina, DOCK6, and consensus docking). The last part is machine learning integrated with QSAR models. Therefore, this data note is vital for further model development to increase performance.
Sequence clustering software is essential in bioinformatics. However, selecting the appropriate one can be challenging due to its diverse algorithms and targeted applications. This paper analyzes and evaluates eight representative softwares (algorithms) in terms of precision, sensitivity, speed, scale of running time, and memory consumption. Furthermore, this paper examines the effects of sequence count, sequence length, identity, thread count, and GPU on the above aspects. Sequence length and identity significantly impact clustering efficiency (speed and memory consumption), with fluctuation amplitudes exceeding an order of magnitude and non-monotonic effects observed. The evaluation results are analyzed and summarized in tables for users' reference.
In recent years, numerous studies have demonstrated that circRNAs play crucial biological roles through their capacity to encode functional proteins. Computational methods have become essential for investigating circRNA translation. In this review, we first outline circRNA biogenesis and translation mechanisms to establish the rationale for developing specialized computational strategies. We then summarize experimental techniques and existing databases that support computational method development. Subsequently, we provide a systematic introduction to existing circRNA translation analysis tools and their underlying algorithms, with emphasis on benchmarking the performance of sequence-based methods using a unified dataset. Our benchmarking revealed that: (1) cirCodAn achieved superior predictive accuracy while maintaining user accessibility; (2) the training data selection during method development critically impacts model performance. This review serves as a comprehensive reference for the selection and application of circRNA translation analysis methods and provides foundational guidance for the development and refinement of future computational tools.
Molecular dynamics simulation is a crucial research domain within the life sciences, focusing on comprehending the mechanisms of biomolecular interactions at atomic scales. Protein simulation, as a critical subfield, often utilizes MD for implementation, with trajectory data play a pivotal role in drug discovery. The advancement of high-performance computing and deep learning technology becomes popular and critical to predict protein properties from vast trajectory data, posing challenges regarding data features extraction from the complicated simulation data and dimensionality reduction. Simultaneously, it is essential to provide a meaningful explanation of the biological mechanism behind dimensionality. To tackle this challenge, we propose a new unsupervised model named RevGraphVAMP to intelligently analyze the simulation trajectory. This model is based on the variational approach for Markov processes (VAMP) and integrates graph convolutional neural networks and physical constraint optimization to enhance the learning performance. Additionally, we introduce attention mechanism to assess the importance of key interaction region, facilitating the interpretation of molecular mechanism. In comparison to other VAMPNets models, our model showcases competitive performance, improved accuracy in state transition prediction, as demonstrated through its application to two public datasets and the Shank3-Rap1 complex, which is associated with autism spectrum disorder. Moreover, it enhanced dimensionality reduction discrimination across different substates and provides interpretable results for protein structural characterization.
Autism spectrum disorder (ASD) encompasses a range of neurodevelopmental conditions marked by challenges in social interaction, communication, and behavior. Currently, there are no specific medications for ASD; instead, early identification and intervention are crucial for enhancing brain function and reducing the disorder's negative impacts through specialized education and rehabilitation. Identification of ASD can be achieved through two primary methods. The first is a manual approach, which involves identifying the disorder by observing the individual or interviewing a parent or caregiver. This approach is laborintensive and subjective. The second method employs automated techniques using both traditional machine learning (ML) and advanced deep learning (DL) strategies, which analyze facial images. The human face, reflecting brain conditions, serves as a practical biomarker for early identification. However, identification of ASD from facial images presents significant challenges due to inter-class similarities, intra-class differences, and variations in facial poses. No existing research comprehensively addresses all these issues within a unified model. In this paper, we introduce a novel multi-feature, multi-level cross-fusion transformer network (MFMLCFT), designed to effectively address these challenges. The proposed method enables effective fusion between facial landmark features and multiple image features. Additionally, it focuses on global and local attention features, learning the multi-level relationships among local regions, and combines global and local features to efficiently discern ASD-related facial features. Specifically, our model first applies spatial attention to multiple image features to accentuate salient facial regions, followed by a multi-level cross-fusion transformer encoder that fuses landmark and image features, capturing the internal dynamics of local facial regions and the interactions between global and local features. The outputs from multiple transformer encoders yield a detailed profile of ASD-related facial features. Furthermore, a vision transformer network is integrated to further enhance feature assimilation. Comprehensive experimental results on an ASD facial image dataset demonstrate that our model significantly outperforms state-of-the-art methods, achieving high metrics of accuracy (96.79%), precision (97.12%), recall (96.43%), F1 score (96.77%), and AUC (99.49). This model offers a valuable tool for healthcare professionals to confirm initial screenings and identify children with ASD more effectively.
Autism Spectrum Disorder (ASD) is a highly disabling mental disease that brings significant impairments of social interaction ability to the patients, making early screening and intervention of ASD critical. With the development of the machine learning and neuroimaging technology, extensive research has been conducted on machine classification of ASD based on structural Magnetic Resonance Imaging (s-MRI). However, most studies involve with datasets where participants' age are above 5 and lack interpretability. In this paper, we propose a machine learning method for ASD classification in children with age range from 0.92 to 4.83 years, based on s-MRI features extracted using Contrastive Variational AutoEncoder (CVAE). 78 s-MRIs, collected from Shenzhen Children's Hospital, are used for training CVAE, which consists of both ASD-specific feature channel and common-shared feature channel. The ASD participants represented by ASD-specific features can be easily discriminated from Typical Control (TC) participants represented by the common-shared features. In case of degraded predictive accuracy when data size is extremely small, a transfer learning strategy is proposed here as a potential solution. Finally, we conduct neuroanatomical interpretation based on the correlation between s-MRI features extracted from CVAE and surface area of different cortical regions, which discloses potential biomarkers that could help target treatments of ASD in the future.
Recent studies have shed light on the potential of circular RNA (circRNA) as a biomarker for disease diagnosis and as a nucleic acid vaccine. The exploration of these functionalities requires correct circRNA full-length sequences; however, existing assembly tools can only correctly assemble some circRNAs, and their performance can be further improved. Here, we introduce a novel feature known as the junction contig (JC), which is an extension of the back-splice junction (BSJ). Leveraging the strengths of both BSJ and JC, we present a novel method called JCcirc (https://github.com/cbbzhang/JCcirc). It enables efficient reconstruction of all types of circRNA full-length sequences and their alternative isoforms using splice graphs and fragment coverage. Our findings demonstrate the superiority of JCcirc over existing methods on human simulation datasets, and its average F1 score surpasses CircAST by 0.40 and outperforms both CIRI-full and circRNAfull by 0.13. For circRNAs below 400 bp, 400-800 bp, 800 bp-1200 bp and above 1200 bp, the correct assembly rates are 0.13, 0.09, 0.04 and 0.03 higher, respectively, than those achieved by existing methods. Moreover, JCcirc also outperforms existing assembly tools on other five model species datasets and real sequencing datasets. These results show that JCcirc is a robust tool for accurately assembling circRNA full-length sequences, laying the foundation for the functional analysis of circRNAs.
In virtual drug screening, consensus docking is a standard in-silico approach consisting of a combined result from optimized docking experiments, a minimum of two results combination. Therefore, consensus docking is subjected to a lower success rate than the best docking method due to its mathematical nature, an unavoidable limitation. This study aims to overcome this drawback via random forest, an ensemble machine learning model. First, in vitro beta-lactamase inhibitory screening was performed using an in-house chemical library. The in vitro results were later used as a validation. Consequently, we optimized docking protocols for AutoDock Vina and DOCK6 programs. With an appropriate scoring function, we found that DOCK6 could identify up to 70% of all active molecules, double the inappropriate. Further consensus analysis reduced the success rate to 50%. Simultaneously, a false positive rate was down to 16%, which was experimentally favorable for a drug search. Finally, we trained two quantitative structure-activity relationship (QSAR) models using logistic regression as a reference model and a random forest as a test model. After combining consensus docking results, random forest-based QSAR outperformed a logistic regression by restoring the success rate to 70% and maintaining a low false positive rate of around 21%. In conclusion, this study demonstrated the benefit of using a random forest (machine learning)-based QSAR model to overcome a standard consensus docking limitation in beta-lactamase inhibitor search as a proof-of-concept.
Sequence clustering software is essential in bioinformatics, yet selecting the most suitable one poses a challenge due to its diverse algorithm design and targeted bioinformatics applications. This paper comprehensively reviewed the developments of most representative sequence clustering software and evaluated 8 representative software based on criteria such as precision, speed, scalability, and memory consumption. This paper divides the clustering software into four aspects: NMI scores greater than 0.95, running time less than 1min/h, 64 core acceleration exceeding 30 times, and memory consumption less than 3 times the dataset, and summarizes them into a table for user querying. Finally, taking OTU, tree of life building, and metagenomic analysis, as examples, this paper demonstrates how to analyze the requirements of scenarios for clustering software and provides recommendations for selecting the most suitable one based on evaluation results.
The automatic detection of cells in microscopy image sequences is a significant task in biomedical research. However, routine microscopy images with cells, which are taken during the process whereby constant division and differentiation occur, are notoriously difficult to detect due to changes in their appearance and number. Recently, convolutional neural network (CNN)-based methods have made significant progress in cell detection and tracking. However, these approaches require many manually annotated data for fully supervised training, which is time-consuming and often requires professional researchers. To alleviate such tiresome and labor-intensive costs, we propose a novel weakly supervised learning cell detection and tracking framework that trains the deep neural network using incomplete initial labels. Our approach uses incomplete cell markers obtained from fluorescent images for initial training on the Induced Pluripotent Stem (iPS) cell dataset, which is rarely studied for cell detection and tracking. During training, the incomplete initial labels were updated iteratively by combining detection and tracking results to obtain a model with better robustness. Our method was evaluated using two fields of the iPS cell dataset, along with the cell detection accuracy (DET) evaluation metric from the Cell Tracking Challenge (CTC) initiative, and it achieved 0.862 and 0.924 DET, respectively. The transferability of the developed model was tested using the public dataset FluoN2DH-GOWT1, which was taken from CTC; this contains two datasets with reference annotations. We randomly removed parts of the annotations in each labeled data to simulate the initial annotations on the public dataset. After training the model on the two datasets, with labels that comprise 10% cell markers, the DET improved from 0.130 to 0.903 and 0.116 to 0.877. When trained with labels that comprise 60% cell markers, the performance was better than the model trained using the supervised learning method. This outcome indicates that the model’s performance improved as the quality of the labels used for training increased.
Neuroblastoma is a prevalent solid tumor affecting children, with a low 5-year survival rate in high-risk patients. Previous studies have shed light on the involvement of specific circRNAs in neuroblastoma development. However, there is still a pressing need to identify novel therapeutic targets associated with circRNAs. In this study, we performed an integrated analysis of two circRNA sequencing datasets, the results revealed dysregulation of 36 circRNAs in neuroblastoma tissues, with their parental genes likely implicated in tumor development. In addition, we identified three specific circRNAs, namely hsa_circ_0001079, hsa_circ_0099504, and hsa_circ_0003171, that exhibit interaction with miRNAs, modulating the expression of genes associated with neuroblastoma. Additionally, by analyzing the translational potential of differentially expressed circRNAs, we uncovered seven circRNAs with the potential capacity for polypeptide translation. Notably, structural predictions suggest that the protein product derived from hsa_circ_0001073 belongs to the TGF-beta receptor protein family, indicating its potential involvement in promoting neuroblastoma occurrence.
基于结构磁共振影像的孤独症分类对孤独症疾病的早期筛查和精准诊断具有重要意义,但受数据噪声及样本不足的影响,基于结构磁共振影像的孤独症分类模型的准确率并不理想。该文提出一种新的数据增强模型,并采用孤独症脑成像交换数据库I中的密西根大学样本库1数据集进行模型测试,随机选取密西根大学样本库1数据集中的78个样本进行实验,以对孤独症分类准确率进行评估。实验结果显示:该方法在500次实验的494次(98%以上的实验)中能够将分类准确率提升10%~20%,在不增加数据量的情况下显著提升了孤独症的分类准确率。通过分析准确率提升和标注变化比例之间的关系,该文进一步对数据标注噪声的问题进行了探讨。
Bioinformatics analysis has been playing a vital role in identifying potential genomic biomarkers more accurately from an enormous number of candidates by reducing time and cost compared to the wet-lab-based experimental procedures for disease diagnosis, prognosis, and therapies. Cervical cancer (CC) is one of the most malignant diseases seen in women worldwide. This study aimed at identifying potential key genes (KGs), highlighting their functions, signaling pathways, and candidate drugs for CC diagnosis and targeting therapies. Four publicly available microarray datasets of CC were analyzed for identifying differentially expressed genes (DEGs) by the LIMMA approach through GEO2R online tool. We identified 116 common DEGs (cDEGs) that were utilized to identify seven KGs (AURKA, BRCA1, CCNB1, CDK1, MCM2, NCAPG2, and TOP2A) by the protein–protein interaction (PPI) network analysis. The GO functional and KEGG pathway enrichment analyses of KGs revealed some important functions and signaling pathways that were significantly associated with CC infections. The interaction network analysis identified four TFs proteins and two miRNAs as the key transcriptional and post-transcriptional regulators of KGs. Considering seven KGs-based proteins, four key TFs proteins, and already published top-ranked seven KGs-based proteins (where five KGs were common with our proposed seven KGs) as drug target receptors, we performed their docking analysis with the 80 meta-drug agents that were already published by different reputed journals as CC drugs. We found Paclitaxel, Vinorelbine, Vincristine, Docetaxel, Everolimus, Temsirolimus, and Cabazitaxel as the top-ranked seven candidate drugs. Finally, we investigated the binding stability of the top-ranked three drugs (Paclitaxel, Vincristine, Vinorelbine) by using 100 ns MD-based MM-PBSA simulations with the three top-ranked proposed receptors (AURKA, CDK1, TOP2A) and observed their stable performance. Therefore, the proposed drugs might play a vital role in the treatment against CC.
Residue distance prediction from the sequence is critical for many biological applications such as protein structure reconstruction, protein-protein interaction prediction, and protein design. However, prediction of fine-grained distances between residues with long sequence separations still remains challenging. In this study, we propose DuetDis, a method based on duet feature sets and deep residual network with squeeze-and-excitation (SE), for protein inter-residue distance prediction. DuetDis embraces the ability to learn and fuse features directly or indirectly extracted from the whole-genome/metagenomic databases and, therefore, minimize the information loss through ensembling models trained on different feature sets. We evaluate DuetDis and 11 widely used peer methods on a large-scale test set (610 proteins chains). The experimental results suggest that 1) prediction results from different feature sets show obvious differences; 2) ensembling different feature sets can improve the prediction performance; 3) high-quality multiple sequence alignment (MSA) used for both training and testing can greatly improve the prediction performance; and 4) DuetDis is more accurate than peer methods for the overall prediction, more reliable in terms of model prediction score, and more robust against shallow multiple sequence alignment (MSA).
The development of noninvasive brain imaging such as resting-state functional magnetic resonance imaging (rs-fMRI) and its combination with AI algorithm provides a promising solution for the early diagnosis of Autism spectrum disorder (ASD). However, the performance of the current ASD classification based on rs-fMRI still needs to be improved. This paper introduces a classification framework to aid ASD diagnosis based on rs-fMRI. In the framework, we proposed a novel filter feature selection method based on the difference between step distribution curves (DSDC) to select remarkable functional connectivities (FCs) and utilized a multilayer perceptron (MLP) which was pretrained by a simplified Variational Autoencoder (VAE) for classification. We also designed a pipeline consisting of a normalization procedure and a modified hyperbolic tangent (tanh) activation function to replace the classical tanh function, further improving the model accuracy. Our model was evaluated by 10 times 10-fold cross-validation and achieved an average accuracy of 78.12%, outperforming the state-of-the-art methods reported on the same dataset. Given the importance of sensitivity and specificity in disease diagnosis, two constraints were designed in our model which can improve the model's sensitivity and specificity by up to 9.32% and 10.21%, respectively. The added constraints allow our model to handle different application scenarios and can be used broadly.