
Actin filaments are fundamental components of the cytoskeleton, essential for maintaining cell shape, enabling motility, and facilitating intracellular transport. Cryo-electron tomography (cryo-ET) enables nanometer-resolution visualization of filament networks in situ; however, low signal-to-noise ratios, missing wedge artifacts, and complex 3D architectures present significant challenges for accurate analysis. Manual filament annotation is highly resource-intensive and prone to variability, underscoring the need for automated approaches. In this study, we develop and evaluate deep learning-based semantic segmentation architectures for accurate segmentation of actin filament networks in cryo-ET data. We systematically assess several architectures-3D U-Net, Attention 3D U-Net, TransUNet 3D, and UNETR-on simulated tomograms with known ground truth. Model performance is quantified using Dice scores and Intersection over Union (IoU) to evaluate segmentation performance under challenging imaging conditions. Our results show that no single deep learning architecture consistently outperforms others, highlighting the importance of accounting for filament arrangement and noise characteristics. By providing a comparative evaluation of these architectures, we demonstrate their effectiveness in detecting filamentous structures and offer guidance for future efforts to improve segmentation performance.
Gut microbiome-based disease diagnosis holds significant promise but remains challenging due to the high data dimensionality, typically small sample sizes, and the necessity of incorporating biological knowledge. Due to these challenges, traditional machine learning approaches often tend to overfit the data and fail to capture true biological relationships, resulting in inaccurate diagnoses. To fill in the gap, we propose a two-phase, knowledge-guided large language model (LLM) framework for disease diagnosis that integrates biomedical expertise with in-context learning. In Phase 1, an LLM is employed to identify disease-associated taxa from hundreds of microbial families and to infer their biological relationships with the disease outcome. This process reduces the feature space dimensionality through biologically-informed feature selection and acquires essential domain knowledge. In Phase 2, we employ few-shot prompting to guide the LLM in disease outcome classification based on the domain knowledge acquired in Phase 1. Thanks to the universal applicability of LLM and our two-phase approach, this is a generic framework that can be applied to a wide range of microbiome-based disease diagnostic tasks. We demonstrate the superiority of our framework using inflammatory bowel disease (IBD) as a representative case study, where our approach achieves an accuracy of 73.91%, significantly outperforming an optimized XGBoost classifier. Overall, our knowledge-guided framework provides a powerful and generalizable strategy for leveraging LLMs in microbiome-based disease diagnosis, and opens a new avenue for disease diagnosis in the era of LLM.
With the growing prevalence of cognitive impairment, early detection has become increasingly critical. Prior studies have examined the association between neuropsychiatric symptoms (NPS) and cognitive impairment, identifying potential predictive relationships. However, they hardly evaluated the heterogeneous relationships between serial patterns of NPS and evolving cognition status of the patients. To address this limitation, we investigate the statistical causal relationship between NPS and cognitive impairment, as well as the dynamic changes in their predictive effects over time, with a specific focus on sex differences. Our approach accounts for the fluctuating nature of NPS and varying follow-up durations across participants by implementing a bootstrap strategy that repeatedly samples a fixed number of visits per participant in a temporal order. Then, we apply causal discovery techniques and counterfactual framework-based causal inference methods to estimate the independent effects of NPS over time. Our findings highlight apathy as a key predictive symptom of cognitive impairment. Moreover, its predictive effect peaks earlier in females than in males, indicating that early-stage tracking is particularly informative in female participants. This suggests sex-specific monitoring strategies may improve early detection and intervention of cognitive impairment.
Integrating novel medical concepts and relationships into existing ontologies can significantly enhance their coverage and utility for both biomedical research and clinical applications. Clinical notes, as unstructured documents rich with detailed patient observations, offer valuable context-specific insights and represent a promising yet underutilized source for ontology extension. Despite this potential, directly leveraging clinical notes for ontology extension remains largely unexplored. To address this gap, we propose CLOZE, a novel framework that uses large language models (LLMs) to automatically extract medical entities from clinical notes and integrate them into hierarchical medical ontologies. By capitalizing on the strong language understanding and extensive biomedical knowledge of pre-trained LLMs, CLOZE effectively identifies disease-related concepts and captures complex hierarchical relationships. The zero-shot framework requires no additional training or labeled data, making it a cost-efficient solution. Furthermore, CLOZE ensures patient privacy through automated removal of protected health information (PHI). Experimental results demonstrate that CLOZE provides an accurate, scalable, and privacy-preserving ontology extension framework, with strong potential to support a wide range of downstream applications in biomedical research and clinical informatics.
Biomedical image segmentation plays a critical role in clinical applications such as disease diagnosis and surgical planning. While deep learning, especially U-Net and its variants, have achieved impressive performance in this domain, their success heavily relies on manual architectural design, which limits scalability and generalizability. Neural architecture search (NAS) offers a promising solution by automating model design, but existing NAS-based methods still suffer from high computational costs and limited architectural diversity. In this work, we propose a novel NAS framework, termed ranking-aware predictorassisted architecture design (RPA2D), for efficient and accurate biomedical image segmentation. RPA2D introduces a lightweight and expressive search space comprising diverse convolution and pooling operations tailored for medical images. Furthermore, we design a contrastive learning-based ranking-aware performance predictor that enables accurate architecture ranking under limited supervision. Extensive experiments on public biomedical datasets demonstrate that RPA2D consistently outperforms both handcrafted and NAS-based baselines in segmentation accuracy while significantly reducing search overhead.
With the increasingly widespread application of omics data and the rapid emergence of artificial intelligence techniques, it has become possible to study biomedical problems at a more system levels. Studying these issues not only requires larger datasets and more advanced technologies, but more importantly, a change in the way we think. Taking medical research as an example, traditional reductionist thinking-driven omic data analysis has yet yielded true breakthroughs in disease research, including cancer, Alzheimer's disease, and diabetes. This is because the roots of these problems may lie in persistent imbalances at the chemical homeostasis level, and metabolic reprogramming that cells deploy to restore the balances for survival. The discovery and study of these issues require a more holistic way of thinking. In this regard, bioinformaticians can play a more important role than traditional technology developers. Through system-level analyses guided by holistic way of thinking, they could identify imbalanced chemical homeostasis and the efforts to regain balances made by the cells as stress responses, potentially leading to the establishment of entirely new paradigms in medical research, thereby advancing medical science.
Brain atlas is an indispensable tool for studying the relationship between brain structure and cognitive function. We proposed to create a new brain atlas – the Brainnetome atlas, using brain connectivity profiles. The Brainnetome atlas lays the foundation for research in brain science and brain-inspired intelligence, and opens a new avenue not only for the study of brain science and brain diseases, but also for brain-inspired intelligence. In this lecture, we first introduce the research background and content of the Brainnetome, including the definition and the main research directions of the Brainnetome, the idea of creating the Brainnetome atlas, and the essential differences from existing brain atlas. Then, we will introduce the application of the Brainnetome atlas in elucidating brain cognitive mechanisms and precise diagnosis of brain diseases. We will also introduce the challenges and solutions of neuromodulation robots guided by the Brainnetome atlas for precise treatment of brain diseases. Finally, a summary and perspective on future research directions are provided.
Genome-wide association studies (GWAS) have uncovered tens of thousands of genetic variants associated with complex traits and diseases. However, translating these associations into mechanistic insights and therapeutic opportunities remains a fundamental challenge. In this talk, I will present several recent studies that bridge this gap by integrating GWAS discoveries with large-scale functional and computational genomics. We combine experimental approaches-including Massively Parallel Reporter Assays (MPRA), single-cell and bulk RNA sequencing, long read sequencing and CRISPR-based perturbations-with advanced computational frameworks that leverage statistical modeling, causal inference, and machine learning models. By jointly analyzing multi-omics data at scale, we are able to identify regulatory variants, resolve cell-type-specific effects, and prioritize causal genes underlying complex phenotypes. I will also discuss emerging strategies for building predictive models of disease risk and cellular function, illustrating how close integration of computation and experimentation is reshaping our ability to connect genetic variation with molecular mechanisms and clinical outcomes.
The artificial intelligence (AI)-assisted design of synthetic, optimized RNA elements provides a new innovative strategy for the development of cancer/infectious disease vaccines and new RNA drugs. Recently, many related landmark studies and patents on the artificial features have been filed and occupied by a few companies and academia. In this talk, I will discuss our recent progress in AI-assisted design of synthetic RNA vaccines with some ongoing studies and the future prospects for intelligent design of new RNA drugs for disease prevention. Our new deep neural network model for designing next-generation RNAi for RNA drugs as well as advanced CRISPR sgRNAs for genome editing will be presented. In addition, recent research trends in the field of advanced biotechnology and the Advanced Biotechnology Initiative announced by the government will be briefly introduced.
Medical contrastive Vision-Language Pre-training (VLP) has emerged as a promising approach, enabling models to learn joint representations from paired medical images and radiology reports. Despite existing methods exploring local visual representation learning techniques, they often fall short in local alignment and knowledge infusion, e.g., uniform token treatment and isolated knowledge assignment. To address these issues, we propose a novel Fine-grained Knowledge-Guided Alignment (FKGA) framework for medical VLP. Specifically, we propose a Fine-grained Disease Knowledge Integration (FDKI) module to inject detailed disease descriptions into corresponding disease tokens in reports. Based on these semantic-enriched tokens, we introduce global instance-wise and local token-wise contrastive learning to further align the semantically related visual and textual modalities. In contrast to previous local visual representation learning methods, our design of semantic-enriched token alignment and context-preserved knowledge infusion enhances the semantic understanding of diseases. Extensive experimental results on five downstream tasks demonstrate that our proposed method outperforms other state-of-the-art methods across seven datasets.
Motor Imagery Brain-Computer Interface (MIBCI) is one of the most widely used BCI paradigms. However, due to large inter-individual variability in EEG signals, EEGbased pattern recognition models face significant challenges in cross-subject generalization. In this study, we posit that although cross-subject MI-BCI is fundamentally a cross-domain task, the population is not homogeneous; instead, latent subpopulations exist in which subjects share more consistent classification boundaries. To exploit this structure, we propose a Mixture-of-Experts (MoE) framework that automatically partitions training trials into latent groups and trains expert classifiers specialized for each group, while simultaneously learning a gating network that assigns test trials to the most suitable expert(s). This design enables the system to adapt to subpopulation structure, mitigate negative transfer from dissimilar subjects, and better model inter-subject heterogeneity. Evaluations on two public MI-EEG datasets (EEGMMIDB and OpenBMI) using k-fold cross-subject protocols demonstrate that our MoE approach significantly improves classification accuracy compared to most baseline models.
Domain shifts in computational pathology, caused by variations in staining, imaging devices, and tissue morphology, challenge model performance in segmentation tasks. Existing domain generalization methods, such as style transfer and feature alignment, often fail to account for organ-level morphological differences. In this paper, we propose PathVLG, a visionlanguage model designed to improve domain generalization for adenocarcinoma segmentation. PathVLG leverages a CONCHbased encoder with three key innovations: the Text-informed Content Query Reformer (TCQR), Text-driven Style Augmentor (TSA), and Style Regeneration Decoder (SRD). These components help the model adapt across domains by incorporating text embeddings, generating diverse styles, and combining source and target domain features. Experimental results show that PathVLG outperforms existing methods in cross-domain generalization.
Accurate instance segmentation of nuclei is a foundational task in computational pathology, critical for quantitative analysis in disease diagnosis and research. While Vision Transformers (ViTs) have shown promise in capturing global context, their direct application to high-resolution histopathology images is hindered by quadratic computational complexity and representational redundancy in their attention mechanisms, which limit the computational power required to capture nuances in nuclei segmentation. To address these limitations, we propose the Regularized-Efficient Transformer U-Net (RET-UNet), a novel hybrid architecture that combines a convolutional encoder-decoder with a redesigned transformer bottleneck. The core of our contribution is the RET-UNet, which introduces three key innovations: 1) an Efficient Pooling Attention (EPA) mechanism that drastically reduces computational cost while maintaining a global receptive field; 2) a self-tuning, projection-based orthogonality regularizer that forces attention heads to learn diverse and complementary feature subspaces; and 3) a Convolutional Feed-Forward Network (CFFN) that reintroduces spatial inductive biases to the transformer. We evaluated our model on the 2018 MoNuSeg dataset, where it achieved a Dice score of 0.6881, outperforming recent transformer-based methods such as LVIT-T. Critically, this superior accuracy was achieved with a computational cost of only 16.32 GFLOPs, making our model over 3.3 times more efficient. These results demonstrate that by strategically addressing the core limitations of ViTs, it is possible to create a model that is not only more accurate but also significantly more practical for real-world applications in computational pathology.
Accurate segmentation and classification of colorectal polyps are critical for the early diagnosis and prevention of colorectal cancer. However, high variability in polyp appearance, along with object-background and inter-class visual similarities, presents significant challenges for automated analysis. In this paper, we propose SC-UMamba, a unified deep learning architecture that integrates a CNN backbone with a Mamba-based UNet-like architecture to jointly address segmentation and classification tasks. Specifically, SC-UMamba leverages shared feature representations between the SSM encoder and CNN backbone, and further incorporates semantic guidance from the segmentation mask output. Additionally, pretrained weights and a step-based curriculum learning strategy are adopted to enable robust learning across both tasks. Furthermore, the architecture offers modular flexibility to adapt to different clinical requirements by supporting backbone substitution. Extensive experiments on the SUN-SEG and our proprietary clinical dataset demonstrate that SC-UMamba achieves state-of-the-art performance, outperforming existing methods in both segmentation accuracy and classification robustness. These results highlight the potential of SC-UMamba to enhance real-time, intelligent assistance during colonoscopy procedures and support reliable clinical decision-making.
Optical coherence tomography (OCT) has emerged as a clinically viable tool for detecting early cervical diseases due to its non-invasive, high-resolution imaging capabilities. While computer-aided OCT diagnosis systems have shown promising potential, they confront three primary challenges: limited annotated data, class imbalance, and cross-center bias. These factors collectively compromise model generalizability across medical centers. Therefore, we propose a joint structure-texture representation learning framework to enhance the generalization of cervical OCT diagnosis in multi-center, small-sample scenarios by leveraging histomorphological invariance. The framework synergizes hierarchical structure features from tissue segmentation with histomorphology-enhanced texture representations learned through lesion-focused contrastive reconstruction. We design a semi-supervised segmentation strategy that achieves promising segmentation performance for layered tissue architecture with minimal annotation burden. An attention-based feature fusion module dynamically combines complementary structural and textural features, generating enriched representations optimized for classification robustness. External cross-center validations demonstrate the superior generalization performance and interpretability of our approach over competitive baseline methods. Moreover, low model complexity and high inference efficiency make our approach well-suited for adoption in low-resource clinical settings.
Sleep staging is essential for assessing sleep quality and diagnosing disorders. Current deep learning models exhibit limitations in spatial feature modeling, cross-frequency coupling awareness, contextual modeling, and stage transition dynamics. We propose a novel deep neural architecture integrating multi-band and multimodal modeling. It includes Frequency-Time-Spatial Convolution (FTSC) with wavelet decompositions for multi-scale frequency features, $1 \times 1$ convolutions for crossfrequency coupling, and spatiotemporal convolutions for local patterns. Enhanced by Multimodal Bottleneck Transformer (MBT) and Epoch Transformer (ET) for cross-modal semantics and a Conditional Random Field (CRF) for physiologically plausible transitions, our approach demonstrates state-of-the-art performance and robustness in experiments on multiple public datasets.
Automatic coronary artery labeling is essential for accurate vascular identification and the diagnosis of coronary disease. The task requires delineating the full vasculature and classifying each segment; however, preserving global topology and local demarcation line precision is difficult due to complex anatomy and blurry contours. We propose a coarse-to-fine ensemble framework with two modules: a Coarse-to-fine Topology Extraction (CTE) network using topology priors for global continuity, and a Progressive Vessel Labeling (PVL) module with multibranch fusion for segmentation and classification. Experiments on the ARCADE dataset achieve a mean F1-score of 0.6028, outperforming state-of-the-art methods and enhancing topological integrity and labeling accuracy. Code: https://github.com/IPMINWU/PGSMODEL.