The automated discovery of structural patterns in macromolecular complexes remains a central challenge in cryo-electron tomography, particularly in highly heterogeneous datasets. Although fully unsupervised clustering methods have shown promise in grouping subtomograms by structural similarity, they often ignore a crucial source of information: the partial ground truth routinely available to structural biologists from prior studies or manual annotations. In this work, we propose a semi-supervised structural discovery framework that utilizes partial supervision to guide clustering without compromising the ability to uncover previously unknown structures. At the core of our method is a label-anchored probabilistic clustering mechanism that seeds the latent space using a small subset of labeled examples and refines it through a multi-resolution consensus strategy based on PCA-space voting. This is complemented by an entropy-based confidence scoring scheme that attenuates the influence of ambiguous samples, as well as a feature propagation procedure that extends structural labels to low-confidence regions using local similarity in feature space. Together, these components create a stable and adaptive pipeline capable of discovering both known and novel structures. Our approach is efficient, requires as little as 1% of labeled data per class, and consistently produces clearer, more interpretable feature embeddings compared to fully unsupervised methods, with well-separated clusters from the very first iterations. Extensive experiments on simulated and realistic tomographic datasets demonstrate that this semi-supervised strategy significantly improves clustering performance, robustness, and biological relevance in cryo-electron tomography analysis. These methods are integrated as extensions to the existing Deep Iterative Subtomogram Clustering Approach pipeline, enhancing its capability for guided structural discovery.
Cryo-electron tomography (cryo-ET) is an essential tool in structural biology, uniquely capable of visualizing three-dimensional macromolecular complexes within their native cellular environments, thereby providing profound molecular-level insights. Despite its significant promise, cryo-ET faces persistent challenges in the systematic localization, identification, segmentation, and structural recovery of three-dimensional subcellular components, necessitating the development of efficient and accurate large-scale image analysis methods. In response to these complexities, this paper introduces AITom, an open-source artificial intelligence platform tailored for cryo-ET researchers. AITom integrates a comprehensive suite of public and proprietary algorithms, supporting both traditional template-based and template-free approaches, alongside state-of-the-art deep learning methodologies for cryo-ET data analysis. By incorporating diverse computational strategies, AITom enables researchers to more effectively tackle the complexities inherent in cryo-ET, facilitating precise analysis and interpretation of complex biological structures. Furthermore, AITom provides extensive tutorials for each analysis module, offering valuable guidance to users in utilizing its comprehensive functionalities.
Subtomogram alignment is a critical task in cryo-electron tomography (cryo-ET) analysis, essential for achieving high-resolution reconstructions of macromolecular complexes. However, learning effective positional representations remains challenging due to limited labels and high noise levels inherent in cryo-ET data. In this work, we address this challenge by proposing a self-supervised learning approach that leverages intrinsic geometric transformations as implicit supervisory signals, enabling robust representation learning despite data scarcity. We introduce BOE-ViT, the first Vision Transformer (ViT) framework for 3D subtomogram alignment. Recognizing that traditional ViTs lack equivariance and are therefore suboptimal for orientation estimation, we enhance the model with two innovative modules that introduce equivariance include 1) the Polyshift module for improved shift estimation and 2) Multi-Axis Rotation Encoding (MARE) for enhanced rotation estimation. Experimental results demonstrate that BOE-ViT significantly outperforms state-of-the-art methods. Notably, at SNR 0.01 dataset, our approach achieves a 77.3% reduction in rotation estimation error and a 62.5% reduction in translation estimation error, effectively overcoming the challenges in cryo-ET subtomogram alignment.
Cryo-electron tomography (cryo-ET) is an important technique used to explore the structure and position of macromolecular complexes in a cellular environment. However, it is extremely difficult to extract sufficient information due to missing-wedge effects, low signal-to-noise ratio (SNR), and the compactness of the particles. Currently, unsupervised cryoET analysis methods struggle to achieve the desired accuracy, while supervised methods require a large amount of labeled data, which costs a lot of manpower, material resources, and financial resources. Even so, Cryo-ET simulation is beneficial as it provides an effective solution to assist supervised analysis methods by providing a large amount of labeled data to train, test, and optimize the models. However, most simulation methods only focus on a single structure, which can only provide limited help. Few methods simulate macromolecular crowding, but they either difficult to achieve a sufficient crowding level or cannot avoid overlap. Therefore, we proposed a new cryo-ET simulation method that packs macromolecules based on molecular dynamics to generate realistic and compact macromolecular crowding without overlap. We also simplified each macromolecule into a small ball to accelerate the simulation process. Our mechanics-based model computes the non-specific interaction, electrostatic interaction, and an external force to obtain the packed macromolecule crowding. From there, the obtained simulated cryo-ET is more realistic according to the imaging principle based on the coordinates of macromolecules under an equilibrium state. In conclusion, our experiments show that our method is able to pack the structures more tightly without overlap, and the simulated cryo-ET is more biologically realistic. Our simulation data can be used to expand the real data set and test methods such as particle picking, protein classification, and protein segmentation. The experimental results show that using the simulation method in this paper, the accuracy of training with all real data can be achieved when only 30% of the real data is used. This simulation method solves the scarcity, expense, and difficult-to-label problems of cryo-ET data, and provides a large amount of simulation data for researchers in the field of computational biology.
Model organisms are instrumental substitutes for human studies to expedite basic, translational, and clinical research. Despite their indispensable role in mechanistic investigation and drug development, molecular congruence of animal models to humans has long been questioned and debated. Little effort has been made for an objective quantification and mechanistic exploration of a model organism’s resemblance to humans in terms of molecular response under disease or drug treatment. We hereby propose a framework, namely Congruence Analysis for Model Organisms (CAMO), for transcriptomic response analysis by developing threshold-free differential expression analysis, quantitative concordance/discordance scores incorporating data variabilities, pathway-centric downstream investigation, knowledge retrieval by text mining, and topological gene module detection for hypothesis generation. Instead of a genome-wide vague and dichotomous answer of “poorly” or “greatly” mimicking humans, CAMO assists researchers to numerically quantify congruence, to dissect true cross-species differences from unwanted biological or cohort variabilities, and to visually identify molecular mechanisms and pathway subnetworks that are best or least mimicked by model organisms, which altogether provides foundations for hypothesis generation and subsequent translational decisions.
Mitochondria adapt to changing cellular environments, stress stimuli, and metabolic demands through dramatic morphological remodeling of their shape, and thus function. Such mitochondrial dynamics is often dependent on cytoskeletal filament interactions. However, the precise organization of these filamentous assemblies remains speculative. Here, we apply cryogenic electron tomography to directly image the nanoscale architecture of the cytoskeletal-membrane interactions involved in mitochondrial dynamics in response to damage. We induced mitochondrial damage via membrane depolarization, a cellular stress associated with mitochondrial fragmentation and mitophagy. We find that, in response to acute membrane depolarization, mammalian mitochondria predominantly organize into tubular morphology that abundantly displays constrictions. We observe long bundles of both unbranched actin and septin filaments enriched at these constrictions. We also observed septin-microtubule interactions at these sites and elsewhere, suggesting that these two filaments guide each other in the cytosolic space. Together, our results provide empirical parameters for the architecture of mitochondrial constriction factors to validate/refine existing models and inform the development of new ones.
Central to active learning (AL) is what data should be selected for annotation. Existing works attempt to select highly uncertain or informative data for annotation. Nevertheless, it remains unclear how selected data impacts the test performance of the task model used in AL. In this work, we explore such an impact by theoretically proving that selecting unlabeled data of higher gradient norm leads to a lower upper-bound of test loss, resulting in a better test performance. However, due to the lack of label information, directly computing gradient norm for unlabeled data is infeasible. To address this challenge, we propose two schemes, namely expected-gradnorm and entropy-gradnorm. The former computes the gradient norm by constructing an expected empirical loss while the latter constructs an unsupervised loss with entropy. Furthermore, we integrate the two schemes in a universal AL framework. We evaluate our method on classical image classification and semantic segmentation tasks. To demonstrate its competency in domain applications and its robustness to noise, we also validate our method on a cellular imaging analysis task, namely cryo-Electron Tomography subtomogram classification. Results demonstrate that our method achieves superior performance against the state of the art. We refer readers to https://arxiv.org/pdf/2112.05683.pdf for the full version of this paper which includes the appendix and source code link.
Cryo-electron tomography, combined with subtomogram averaging (STA), can reveal three-dimensional (3D) macromolecule structures in the near-native state from cells and other biological samples. In STA, to get a high-resolution 3D view of macromolecule structures, diverse macromolecules captured by the cellular tomograms need to be accurately classified. However, due to the poor signal-to-noise-ratio (SNR) and severe ray artifacts in the tomogram, it remains a major challenge to classify macromolecules with high accuracy. In this paper, we propose a new convolutional neural network, named 3D-Dilated-DenseNet, to improve the performance of macromolecule classification. In 3D-Dilated-DenseNet, there are two key strategies to guarantee macromolecule classification accuracy: 1) Using dense connections to enhance feature map utilization (corresponding to the baseline 3D-C-DenseNet); 2) Adopting dilated convolution to enrich multi-level information in feature maps. We tested 3D-Dilated-DenseNet and 3D-C-DenseNet both on synthetic data and experimental data. The results show that, on synthetic data, compared with the state-of-the-art method in the SHREC contest (SHREC-CNN), both 3D-C-DenseNet and 3D-Dilated-DenseNet outperform SHREC-CNN. In particular, 3D-Dilated-DenseNet improves 0.393 of F1 metric on tiny-size macromolecules and 0.213 on small-size macromolecules. On experimental data, compared with 3D-C-DenseNet, 3D-Dilated-DenseNet can increase classification performance by 2.1 percent.
Macromolecular structure classification from cryo-electron tomography (cryo-ET) data is important for understanding macro-molecular dynamics. It has a wide range of applications and is essential in enhancing our knowledge of the sub-cellular environment. However, a major limitation has been insufficient labelled cryo-ET data. In this work, we use Contrastive Self-supervised Learning (CSSL) to improve the previous approaches for macromolecular structure classification from cryo-ET data with limited labels. We first pretrain an encoder with unlabelled data using CSSL and then fine-tune the pretrained weights on the downstream classification task. To this end, we design a cryo-ET domain-specific data-augmentation pipeline. The benefit of augmenting cryo-ET datasets is most prominent when the original dataset is limited in size. Overall, extensive experiments performed on real and simulated cryo-ET data in the semi-supervised learning setting demonstrate the effectiveness of our approach in macromolecular labeling and classification.
MOTIVATION Cryo-Electron Tomography (cryo-ET) is a 3D imaging technology that enables the visualization of subcellular structures in situ at near-atomic resolution. Cellular cryo-ET images help in resolving the structures of macromolecules and determining their spatial relationship in a single cell, which has broad significance in cell and structural biology. Subtomogram classification and recognition constitute a primary step in the systematic recovery of these macromolecular structures. Supervised deep learning methods have been proven to be highly accurate and efficient for subtomogram classification, but suffer from limited applicability due to scarcity of annotated data. While generating simulated data for training supervised models is a potential solution, a sizeable difference in the image intensity distribution in generated data as compared with real experimental data will cause the trained models to perform poorly in predicting classes on real subtomograms. RESULTS In this work, we present Cryo-Shift, a fully unsupervised domain adaptation and randomization framework for deep learning-based cross-domain subtomogram classification. We use unsupervised multi-adversarial domain adaption to reduce the domain shift between features of simulated and experimental data. We develop a network-driven domain randomization procedure with 'warp' modules to alter the simulated data and help the classifier generalize better on experimental data. We do not use any labeled experimental data to train our model, whereas some of the existing alternative approaches require labeled experimental samples for cross-domain classification. Nevertheless, Cryo-Shift outperforms the existing alternative approaches in cross-domain subtomogram classification in extensive evaluation studies demonstrated herein using both simulated and experimental data. AVAILABILITYAND IMPLEMENTATION https://github.com/xulabs/aitom. SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.
Cryo-electron tomography (cryo-ET) is an imaging technique that allows three-dimensional visualization of macro-molecular assemblies under near-native conditions. Cryo-ET comes with a number of challenges, mainly low signal-to-noise and inability to obtain images from all angles. Computational methods are key to analyze cryo-electron tomograms. To promote innovation in computational methods, we generate a novel simulated dataset to benchmark different methods of localization and classification of biological macromolecules in tomograms. Our publicly available dataset contains ten tomographic reconstructions of simulated cell-like volumes. Each volume contains twelve different types of complexes, varying in size, function and structure. In this paper, we have evaluated seven different methods of finding and classifying proteins. Seven research groups present results obtained with learning-based methods and trained on the simulated dataset, as well as a baseline template matching (TM), a traditional method widely used in cryoET research. We show that learning-based approaches can achieve notably better localization and classification performance than TM. We also experimentally confirm that there is a negative relationship between particle size and performance for all methods.
In many real-life image analysis applications, particularly in biomedical research domains, the objects of interest undergo multiple transformations that alters their visual properties while keeping the semantic content unchanged. Disentangling images into semantic content factors and transformations can provide significant benefits into many domain-specific image analysis tasks. To this end, we propose a generic unsupervised framework, Harmony, that simultaneously and explicitly disentangles semantic content from multiple parameterized transformations. Harmony leverages a simple cross-contrastive learning framework with multiple explicitly parameterized latent representations to disentangle content from transformations. To demonstrate the efficacy of Harmony, we apply it to disentangle image semantic content from several parameterized transformations (rotation, translation, scaling, and contrast). Harmony achieves significantly improved disentanglement over the baseline models on several image datasets of diverse domains. With such disentanglement, Harmony is demonstrated to incentivize bioimage analysis research by modeling structural heterogeneity of macromolecules from cryo-ET images and learning transformation-invariant representations of protein particles from single-particle cryo-EM images. Harmony also performs very well in disentangling content from 3D transformations and can perform coarse and fast alignment of 3D cryo-ET subtomograms. Therefore, Harmony is generalizable to many other imaging domains and can potentially be extended to domains beyond imaging as well.
Cryo-electron tomography (Cryo-ET) has been regarded as a revolution in structural biology and can reveal molecular sociology. Its unprecedented quality enables it to visualize cellular organelles and macromolecular complexes at nanometer resolution with native conformations. Motivated by developments in nanotechnology and machine learning, establishing machine learning approaches such as classification, detection and averaging for Cryo-ET image analysis has inspired broad interest. Yet, deep learning-based methods for biomedical imaging typically require large labeled datasets for good results, which can be a great challenge due to the expense of obtaining and labeling training data. To deal with this problem, we propose a generative model to simulate Cryo-ET images efficiently and reliably: CryoETGAN. This cycle-consistent and Wasserstein generative adversarial network (GAN) is able to generate images with an appearance similar to the original experimental data. Quantitative and visual grading results on generated images are provided to show that the results of our proposed method achieve better performance compared to the previous state-of-the-art simulation methods. Moreover, CryoETGAN is stable to train and capable of generating plausibly diverse image samples.
Cryo-electron tomography (cryo-ET) is an imaging technique that allows three-dimensional visualization of macro-molecular assemblies under near-native conditions. Cryo-ET comes with a number of challenges, mainly low signal-to-noise and inability to obtain images from all angles. Computational methods are key to analyze cryo-electron tomograms. To promote innovation in computational methods, we generate a novel simulated dataset to benchmark different methods of localization and classification of biological macromolecules in tomograms. Our publicly available dataset contains ten tomographic reconstructions of simulated cell-like volumes. Each volume contains twelve different types of complexes, varying in size, function and structure. In this paper, we have evaluated seven different methods of finding and classifying proteins. Seven research groups present results obtained with learning-based methods and trained on the simulated dataset, as well as a baseline template matching (TM), a traditional method widely used in cryo-ET research. We show that learning-based approaches can achieve notably better localization and classification performance than TM. We also experimentally confirm that there is a negative relationship between particle size and performance for all methods.
Estimating the number of clusters ( K ) is a critical and often difficult task in cluster analysis. Many methods have been proposed to estimate K , including some top performers using resampling approach. When performing cluster analysis in high-dimensional data, simultaneous clustering and feature selection is needed for improved interpretation and performance. To our knowledge, little has been studied for simultaneous estimation of K and feature sparsity parameter in a high-dimensional exploratory cluster analysis. In this paper, we propose a resampling method to bridge this gap and evaluate its performance under the sparse K -means clustering framework. The proposed target function balances between sensitivity and specificity of clustering evaluation of pairwise subjects from clustering of full and subsampled data. Through extensive simulations, the method performs among the best over classical methods in estimating K in low-dimensional data. For high-dimensional simulation data, it also shows superior performance to simultaneously estimate K and feature sparsity parameter. Finally, we evaluated the methods in four microarray, two RNA-seq, one SNP, and two nonomics datasets. The proposed method achieves better clustering accuracy with fewer selected predictive genes in almost all real applications.
Cryo-electron tomography (Cryo-ET) is an emerging technology for three-dimensional (3D) visualization of macromolecular structures in the near-native state. To recover structures of macromolecules, millions of diverse macromolecules captured in tomograms should be accurately classified into structurally homogeneous subsets. Although existing supervised deep learning–based methods have improved classification accuracy, such trained models have limited ability to classify novel macromolecules that are unseen in the training stage. To adapt the trained model to the macromolecule classification of a novel class, massive labeled macromolecules of the novel class are needed. However, data labeling is very time-consuming and labor-intensive. In this work, we propose a novel few-shot learning method for the classification of novel macromolecules (named FSCC). A two-stage training strategy is designed in FSCC to enhance the generalization ability of the model to novel macromolecules. First, FSCC uses contrastive learning to pre-train the model on a sufficient number of labeled macromolecules. Second, FSCC uses distribution calibration to re-train the classifier, enabling the model to classify macromolecules of novel classes (unseen class in the pre-training). Distribution calibration transfers learned knowledge in the pre-training stage to novel macromolecules with limited labeled macromolecules of novel class. Experiments were performed on both synthetic and real datasets. On the synthetic datasets, compared with the state-of-the-art (SOTA) method based on supervised deep learning, FSCC achieves competitive performance. To achieve such performance, FSCC only needs five labeled macromolecules per novel class. However, the SOTA method needs 1100 ∼ 1500 labeled macromolecules per novel class. On the real datasets, FSCC improves the accuracy by 5% ∼ 16% when compared to the baseline model. These demonstrate good generalization ability of contrastive learning and calibration distribution to classify novel macromolecules with very few labeled macromolecules.
Cryo-Electron Tomography (cryo-ET) is an emerging 3D imaging technique which shows great potentials in structural biology research. One of the main challenges is to perform classification of macromolecules captured by cryo-ET. Recent efforts exploit deep learning to address this challenge. However, training reliable deep models usually requires a huge amount of labeled data in supervised fashion. Annotating cryo-ET data is arguably very expensive. Deep Active Learning (DAL) can be used to reduce labeling cost while not sacrificing the task performance too much. Nevertheless, most existing methods resort to auxiliary models or complex fashions (e.g. adversarial learning) for uncertainty estimation, the core of DAL. These models need to be highly customized for cryo-ET tasks which require 3D networks, and extra efforts are also indispensable for tuning these models, rendering a difficulty of deployment on cryo-ET tasks. To address these challenges, we propose a novel metric for data selection in DAL, which can also be leveraged as a regularizer of the empirical loss, further boosting the task model. We demonstrate the superiority of our method via extensive experiments on both simulated and real cryo-ET datasets. Our source Code and Appendix can be found at this URL.
3D subtomogram image alignment, clustering, and segmentation are vital to macromolecular structure recognition in cryo-electron tomography (cryo-ET). However, acquiring ground-truth labels to train a unified deep learning model that can simultaneously deal with these tasks is unaffordable. To this end, we propose an end-to-end unified multi-task learning framework to simultaneously complete the three tasks, where models are trained in an unsupervised manner without using any labels. In particular, we have three parallel branches. In the alignment branch, we adopt a two-stage training scheme, i.e., self-supervised pretraining and constrained unsupervised training using our proposed skip correlation attention layer and constrained loss. Synchronously, in the clustering branch, the learned deep cluster features are utilized to iteratively cluster subtomograms into groups using pseudo-labels from an image-wise Gaussian Mixture Model (GMM). Meanwhile, in the segmentation branch, we use rough pseudo-labels generated from a voxel-wise GMM as supervision signals, and prior knowledge from humans is utilized to jointly learn how to correct these labels as well as predict reliable segmentation results. Benefiting from the end-to-end unified network architecture, our method achieves overall state-of-the-art performance on both simulated and real subtomogram processing benchmarks.
MOTIVATION Cryo-Electron Tomography (cryo-ET) is a 3D bioimaging tool that visualizes the structural and spatial organization of macromolecules at a near-native state in single cells, which has broad applications in life science. However, the systematic structural recognition and recovery of macromolecules captured by cryo-ET are difficult due to high structural complexity and imaging limits. Deep learning based subtomogram classification have played critical roles for such tasks. As supervised approaches, however, their performance relies on sufficient and laborious annotation on a large training dataset. RESULTS To alleviate this major labeling burden, we proposed a Hybrid Active Learning (HAL) framework for querying subtomograms for labelling from a large unlabeled subtomogram pool. Firstly, HAL adopts uncertainty sampling to select the subtomograms that have the most uncertain predictions. This strategy enforces the model to be aware of the inductive bias during classification and subtomogram selection, which satisfies the discriminativeness principle in AL literature. Moreover, to mitigate the sampling bias caused by such strategy, a discriminator is introduced to judge if a certain subtomogram is labeled or unlabeled and subsequently the model queries the subtomogram that have higher probabilities to be unlabeled. Such query strategy encourages to match the data distribution between the labeled and unlabeled subtomogram samples, which essentially encodes the representativeness criterion into the subtomogram selection process. Additionally, HAL introduces a subset sampling strategy to improve the diversity of the query set, so that the information overlap is decreased between the queried batches and the algorithmic efficiency is improved. Our experiments on subtomogram classification tasks using both simulated and real data demonstrate that we can achieve comparable testing performance (on average only 3% accuracy drop) by using less than 30% of the labeled subtomograms, which shows a very promising result for subtomogram classification task with limited labeling resources. AVAILABILITY https://github.com/xulabs/aitom.
Computing dense pixel-to-pixel image correspondences is a fundamental task of computer vision. Often, the objective is to align image pairs from the same semantic category for manipulation or segmentation purposes. Despite achieving superior performance, existing deep learning alignment methods cannot cluster images; consequently, clustering and pairing images needed to be a separate laborious and expensive step.Given a dataset with diverse semantic categories, we propose a multi-task model, Jim-Net, that can directly learn to cluster and align images without any pixel-level or image-level annotations. We design a pair-matching alignment unsupervised training algorithm that selectively matches and aligns image pairs from the clustering branch. Our unsupervised Jim-Net achieves comparable accuracy with state-of-the-art supervised methods on benchmark 2D image alignment dataset PF-PASCAL. Specifically, we apply Jim-Net to cryo-electron tomography, a revolutionary 3D microscopy imaging technique of native subcellular structures. After extensive evaluation on seven datasets, we demonstrate that Jim-Net enables systematic discovery and recovery of representative macromolecular structures in situ, which is essential for revealing molecular mechanisms underlying cellular functions. To our knowledge, Jim-Net is the first end-to-end model that can simultaneously align and cluster images, which significantly improves the performance as compared to performing each task alone.