The recent large vision-language models (LVLMs), such as CLIP and mPlug-Owl2 have shown great success in various vision processing tasks, demonstrating strong capabilities in understanding high-level information from visual inputs. A worth-researching question is whether these pre-trained LVLMs can be effectively transferred to the domain of electron microscopy images, providing new insights for volume alignment quality assessment (VAQA) methods. To fill this gap, in this paper, we propose SAVIOR, a novel VAQA model based on LVLM and rich quality-aware features. Specifically, SAVIOR leverages three parallel branches: (1) a large vision-language model-based branch to extract high-level quality-aware features from cross-view samples; (2) a Swin-Transformer branch employed on section image sequences to derive spatiotemporal local features; (3) a membrane affinity map-guided motion analysis branch implemented on section-wise optical flow fields to integrate biological priors into motion feature extraction. The complementary features extracted by the branch models are regressed into sub-scores using multi-layer perceptrons, and the final quality score is obtained through a feature-based weight adaptive learner that dynamically adjust contribution of each branch model. Experimental results demonstrate that the proposed method can be effectively fine-tuned on small-scale datasets, achieving state-of-the-art performance on synthetic datasets, and can accurately assess the alignment quality of real-world data. All data and code will be publicly available at https://github.com/HaoranChen-CASIA/SAVIOR.
The registration of serial section electron microscopy (ssEM) images allows to restore the three-dimensional volumes of biological tissues. Automated registration quality assessment is crucial because it can detect potential registration errors, which can impact the accuracy of neural-circuit reconstruction in large-scale micro-connectomics studies. However, due to substantial natural variations in morphology of neural structures across adjacent sections, correctly aligned image contents may still appear dissimilar, causing traditional image similarity-based metrics to yield misleading assessments. In this paper, we first analyze the impact of neural structure size on registration quality using a spherical deformation model. Theoretical analysis and experimental validation indicate that it is more challenging to accurately assess registration quality in image regions with large neural structure size. To address the above challenge, an automated method is proposed for the local registration quality assessment of ssEM images using a feature extraction network that is insensitive to variation of neural structure across adjacent sections. Results show that the proposed method achieves higher stability on the in-situ focused ion beam scanning electron microscopy (FIB-SEM) dataset compared with image similarity-based metrics, and can accurately distinguish local registration quality of images with various neural structure size.
Diabetic foot ulcers represent a severe complication of diabetes, associated with high rates of amputation and mortality risk. Consequently, early identification of high-risk foot lesions is crucial for prevention. Since abnormal plantar temperature changes often precede ulcer formation, infrared thermography has emerged as a promising screening tool. However, its clinical application is limited by uneven temperature distribution and a primary focus on established ulcers. This study integrates the deep features from infrared images, and plantar temperature characteristics to construct a machine learning-based classification model for high-risk diabetic foot lesions. After evaluating multiple classifiers and feature fusion strategies, the XGBoost model trained using fused ResNet-101 Global Average Pooling features and temperature features demonstrated optimal performance. This confirms the potential of multimodal infrared thermography models for clinical early screening and preventive management.
Large-scale connectomics requires the nanometer-scale resolution of electron microscopy to resolve ultrastructural details, but its acquisition is time- and labor-intensive. In contrast, high-throughput light microscopy offers high throughput but lacks the spatial resolution required for fine-grained analysis. To bridge this gap, we propose a computational framework for cross-modal ultrastructural inference, which maps high-throughput light microscopy data to an electron microscopy-like structural representation to accelerate downstream workflows. Rather than performing de novo synthesis, our framework recovers ultrastructural representation from diffraction-limited optical signals by enhancing latent morphological cues within the light microscopy data. This workflow is powered by a novel deep learning framework, the Content-Decoupled Schrödinger Bridge, which disentangles modality-invariant physical content from imaging-specific attributes and incorporates a physics-informed perceptual loss to ensure structural plausibility. Our approach accelerates the connectomics pipeline, demonstrated in three key applications. First, the enhanced clarity of the generated images reduced expert miss rates for region-of-interest selection. Second, the generated images are inherently aligned with the source light microscopy data while matching the appearance of target electron microscopy data, streamlining multi-modal registration. Third, they improve segmentation accuracy by allowing pre-trained electron microscopy models to be applied directly to light microscopy data. In summary, this work provides a computational solution for bridging the gap between imaging speed and resolution. By enhancing the analytical value of light microscopy data, our workflow accelerates key stages of large-scale connectomics mapping, from targeted acquisition to quantitative analysis.
Volume microscopy, including electron and light microscopy, suffers from severe anisotropic resolution due to physical axial sectioning. Existing self-supervised axial super-resolution (ASR) methods face a trilemma bounded by overly smoothed regression textures, structural hallucinations of pure diffusion models, and prohibitive inference latency. In this paper, we propose Skeleton-refinE Microscopy (SkelEM), a self-supervised framework that decouples ASR at the training-signal level: a frozen topological network and a diffusion refiner are optimized by disjoint objectives, separating low-frequency topology formulation from high-frequency detail enhancement. Building on this deterministic skeleton, we exploit a unified cycle-consistent mechanism on input sparse slices to simultaneously extract a real-domain residual prior and bidirectionally align the diffusion refiner, washing away cross-plane artifacts without synthetic bias. By truncating the reverse diffusion process with this physical prior, SkelEM achieves high-fidelity detail restoration in merely ≤ 5 steps. To rigorously assess cross-instrument generalization, we further introduce BRAVE-ASR, a new benchmark of co-aligned anisotropic and isotropic volumes acquired on a Plasma-FIB instrument. Across public benchmarks, SkelEM achieves the most favorable balance across the fidelity-perception trade-off among self-supervised methods, with state-of-the-art downstream membrane segmentation performance and robust zero-shot generalization across distinct modalities.
IntroductionMitochondrial networks exhibit striking heterogeneity in their morphology and distribution across different neuronal compartments, reflecting the diverse metabolic demands of these structures.MethodsIn this study, we used automated tape-collecting ultramicrotome scanning electron microscopy (ATUM-SEM) to reconstruct and quantify mitochondrial networks in the somata and neurites of neurons in the rat prefrontal cortex (PFC) and hippocampus (HPC; CA1 stratum radiatum). We developed an automated segmentation pipeline based on an attention-enhanced 3D U-Net to extract all mitochondria from volumetric EM data.ResultsOur quantitative analyses revealed pronounced regional and subcellular heterogeneity. In the PFC, the mitochondrial volume fraction was higher in neurites (7.2%) than in somata (2.9%; 7.1% when nucleus was excluded). Mean individual mitochondrial volume was 0.11 μm³ for neuritic and 0.33 μm³ for somatic mitochondria in the PFC, with similar results observed in the HPC (0.13 μm³ in neurites, 0.31 μm³ in somata). In both regions, the vast majority of mitochondria (~91%) assumed an oval or rod shape, with few displaying branched or donut-shaped structures (~1%). Notably, elongated linear mitochondria (~8%) were mostly confined to neurites, and approximately 90% of these comprised up to 120 nanotunnels—thin segments (<220 nm) connecting enlarged, oval-shaped structures (>350 nm) in tandem.ConclusionThese data provide a detailed quantitative characterization of mitochondrial network architecture in the adult rat cortex and hippocampus, revealing significant regional and subcellular differences in mitochondrial morphology and distribution.
Accurate measurement of biological slice thickness is essential for the precise reconstruction of neuronal connections in microscale connectomics. Conventional metrology often faces performance bottlenecks, struggling to balance measurement accuracy with system complexity and operational efficiency. This study proposes a novel, low-cost, and high-precision method for biological slice thickness measurement. The proposed data-driven approach eliminates the need for high-precision physical modeling and complex instrumentation. By utilizing a simple optical setup based on a standard RGB camera, the proposed method achieves a mean absolute error (MAE) of 0.77 nm over a wide measurement range of 30-1000 nm. To address the limited generalization capability of conventional deep learning (Conv. DL) methods under sparse training data conditions, a physics-inspired three-wavelength interferometric neural network (TWI-NN) framework is proposed. The architecture incorporates physics-based structural constraints, including task decoupling and input-output inversion, enabling the network to learn accurate mappings from limited data. Experimental results indicate that, under sparse training data conditions, the prediction accuracy of the TWI-NN is significantly better than that of Conv. DL models. Uncertainty analysis confirms that the total expanded uncertainty (k = 2) is less than 0.947 nm. Furthermore, compared with existing interferometric color measurement methods, the proposed method eliminates the need for repeated calibration and features a simpler implementation. Thickness distribution measurement at an effective resolution of 200 & times; 200 pixels can be completed in 0.73 s. Overall, the proposed method provides a scalable, cost-effective, and high-precision solution, demonstrating strong potential for high-throughput thickness measurement.
Multi-animal tracking (MAT) is critical for wildlife monitoring and behavioral analysis, yet remains challenging due to uniform appearance, high density, and irregular motion. Existing methods typically follow heuristic- or query-based paradigms: the former relies on handcrafted geometric associations without end-to-end optimization, whereas the latter enables joint optimization but relies heavily on appearance embeddings. In such conditions, continuous geometric embeddings can be unstable, as small coordinate perturbations may disproportionately alter cross-frame attention weights, degrading identity association performance. To address this limitation, we propose HieDG, a Hierarchical Discrete Geometry-guided tracking framework that reformulates geometric dynamics as structured discrete representations within a query-based tracker. Instead of directly using raw geometric signals, HieDG employs a two-stage residual codebook to discretize position, scale, and velocity cues, transforming unstable continuous geometry into structured, stable discrete tokens. These tokens are aligned with visual embeddings and integrated into the tracking queries to enhance identity consistency. Extensive experiments on animal-specific benchmarks (AnimalTrack, BFT, and BuckTales) demonstrate state-of-the-art association performance with significant improvements in HOTA, AssA, and IDF1. Additional evaluations on generic multi-object tracking benchmarks, including DanceTrack and SportsMOT, show competitive performance, indicating the broader applicability of discretized geometric modeling beyond animal-specific scenarios.
Volume electron microscopy (VEM) enables nanometer-resolution three-dimensional (3D) visualization of biological specimens via serial sectioning and imaging. Owing to limitations of downstream analysis, VEM datasets are often acquired at slow speeds and high resolutions, thereby limiting achievable imaging throughput. By systematically searching for optimal VEM acquisition conditions, we find that sufficient spatial resolution effectively counteracts high image noise in preserving 3D structural information. To further verify that denoising is more effective in restoring volumetric datasets than axial interpolation, we compared machine learning-based methods, including a newly developed 3D context-based denoising model, through various tasks on VEM datasets acquired simultaneously. Our volumetric approach not only outperforms other baseline methods in faithful feature recovery but also facilitates robust serial block-face cutting down to 20 nm by allowing fast imaging. This work provides both an optimized acquisition strategy and volumetric denoising methods as actionable guidelines for maximizing VEM throughput.
Multiscale information integration is essential for a comprehensive understanding of brain structure and function [1]. Optical microscopy provides mesoscopic information on brain region distribution, neuronal projections, and functional activity patterns, while electron microscopy offers microscopic details, such as cell morphology and synaptic connectivity. Aligning cell nucleus information from these two modalities is therefore critical for linking global tissue organization with cellular-level mechanisms. However, fundamental differences in imaging principles and data characteristics make cross-modal cell nucleus point clouds registration highly challenging [2]. Variations in point density, inconsistent noise levels, and complex nonlinear deformations often prevent traditional methods from achieving the accuracy and biological plausibility required in neuroscience research. We propose a non-rigid registration strategy for cross-modal cell nucleus point clouds. The method adopts a multi-level block correspondence scheme to achieve consistent alignment across local and global scales, while integrating neighborhood constraints and a bidirectional weighting mechanism to mitigate the influence of outliers and data sparsity. Experimental results demonstrate that this strategy significantly improves registration accuracy and robustness while preserving structural continuity. This work provides technical support for cross-modal neural tissue mapping and demonstrates the potential of multiscale information integration to advance brain connectomics and neurological disease research. [1] Shapson-Coe A, Januszewski M, Berger D R, et al. ‘A petavoxel fragment of human cerebral cortex reconstructed at nanoscale resolution’ [J]. Science, 2024, 384(6696). [2] Huang X, Mei G, Zhang J. ‘Cross-source point cloud registration: Challenges, progress and prospects’ [J]. Neurocomputing, 2023, 548: 126383.
Advancements in animal behavior quantification methods have driven the development of computational ethology, enabling fully automated behavior analysis. Existing multi-animal pose estimation workflows rely on tracking-by-detection frameworks for either bottom-up or top-down approaches, requiring retraining to accommodate diverse animal appearances. This study introduces InteBOMB, an integrated workflow that enhances top-down approaches by incorporating generic object tracking, eliminating the need for prior knowledge of target animals while maintaining broad generalizability. InteBOMB includes two key strategies for tracking and segmentation in laboratory environments and two techniques for pose estimation in natural settings. The "background enhancement" strategy optimizes foreground-background contrastive loss, generating more discriminative correlation maps. The "online proofreading" strategy stores human-in-the-loop long-term memory and dynamic short-term memory, enabling adaptive updates to object visual features. The "automated labeling suggestion" technique reuses the visual features saved during tracking to identify representative frames for training set labeling. Additionally, the "joint behavior analysis" technique integrates these features with multimodal data, expanding the latent space for behavior classification and clustering. To evaluate the framework, six datasets of mice and six datasets of non-human primates were compiled, covering laboratory and natural scenes. Benchmarking results demonstrated a 24% improvement in zero-shot generic tracking and a 21% enhancement in joint latent space performance across datasets, highlighting the effectiveness of this approach in robust, generalizable behavior analysis.
Volume electron microscopy (VEM) enables three-dimensional (3D) visualization of thick specimens through serial sectioning and nanometer-resolution imaging. In an effort to overcome daunting challenges in subsequent data analysis, people tend to employ excessively slow imaging and high resolution, both of which significantly prolong the acquisition. Here, using authentic and synthetic VEM datasets through image stack re-registration, we demonstrate the implications of 3D oversampling in maintaining structural information fidelity against high background noise. This provides a fact-based argument for a higher priority of improving the spatial sampling rate, particularly in the most difficult axial direction, than suppressing the image noise during a VEM acquisition. To accelerate VEM acquisition by leveraging the oversampled 3D context, we explore distinct machine learning-based methods for restoring serial images that contain either low-contrast snapshots or skipped sections. On the datasets of similar acquisition consumptions, 3D context-based (volumetric) denoising outperforms state-of-the-art interpolation and 2D denoising approaches. Furthermore, the volumetric denoising models can proceed in a self-supervised manner, thereby no longer relying on specific training datasets. This work is not only instructive for planning efficient large-scale acquisition on commercial setups but also benchmarks the methodology of optimizing VEM acquisition to facilitate automated image processing. ### Competing Interest Statement The authors have declared no competing interest. Scientific Research Instrument and Equipment Development Project of Chinese Academy of Sciences, PTYQ2025TD0002 National Natural Science Foundation of China, 32171461, 82171133 Innovative Research Team of High-level Local Universities in Shanghai, SHSMU-ZLCX20211700
We have developed a parallel ion beam thinning device with low incident ion beam energy, enabling simultaneous 20nm thickness reduction for biological sections which are collected on 4-inch wafer. Together with SEM imaging and volume stitching, it is trustworthy and efficient to achieve three-dimensional electron microscopy imaging of millimeter-scale samples for ultra large scale connectomics. ### Competing Interest Statement The authors have declared no competing interest.
Anisotropic resolution remains a fundamental challenge in 3D microscopy, where axial resolution is significantly lower than lateral resolution due to physical limitations. To address this, we propose a self-supervised volume super-resolution (VSR) framework named Diffusion to Resolution (D2R), which leverages 2D diffusion priors to enhance axial resolution without requiring high-resolution (HR) volume as supervision. D2R consists of three stages: (1) learning biological priors via a 2D diffusion model trained on high-resolution XY slices, (2) generating pseudo-HR lateral (XZ/YZ) volumes through cross-plane fusion, and (3) performing stable structure distillation to train a 3D VSR network. To further improve VSR quality, we introduce Axial Enhancement Network (AENet), a 3D VSR model incorporating lightweight channel attention to enhance fine details while maintaining inter-slice continuity. Extensive experiments on FIB-SEM datasets demonstrate that D2R-AENet outperforms state-of-the-art self-supervised methods in both image similarity and membrane segmentation accuracy, achieving performance close to supervised approaches. These results validate the effectiveness of our framework in high-fidelity volumetric reconstruction under practical conditions where HR references are unavailable. Codes are available at https://github.com/hmzawz2/D2R-models.
Volume electron microscopy (VEM) enables three-dimensional (3D) visualization of specimens through serial sectioning and nanometer-resolution imaging. To alleviate daunting difficulties in subsequent data analysis, people tend to employ excessively slow imaging and high resolution, limiting the acquisition throughput. Here, we titrated the combinations of acquisition parameters for preserving key structural information. This revealed a prior role of sufficient spatial sampling over high image contrast. To save acquisition time, we then attempted to restore serial images that contained either low-contrast snapshots or skipped sections using machine learning. Owing to constraints of cytoarchitecture rationality, 3D context-based (volumetric) denoising not only preserved more structural features but also generated fewer artifacts than other methods. Moreover, volumetric denoising lowered the requirement of imaging dwell time and thereby allowed serial block-face removal down to 20 nm because of reduced radiation damage. This work demonstrated how machine learning-based image processing enabled the optimization of VEM acquisition parameters.
The brain is complementarily assembled by the sensorimotor and neuromodulatory (NM) pathways[1][1]. While connectome mapping is crucial for elucidating the synaptic organization principles underlying this bi-pathway architecture, most electron microscopy (EM) reconstructions provide limited information about cell types, particularly NM neurons[2][2]–[7][3]. Here we present Fish-X, a synapse-level, multiplexed NM-type-annotated reconstruction of an intact larval zebrafish brain, identifying noradrenergic, dopaminergic, serotonergic, hypocretinergic, and glycinergic neurons via subcellular localization of peroxidase APEX2, with glutamatergic/GABAergic neurons inferred through morphology comparison with the Zebrafish Mesoscopic Atlas. NM neurons with varying indegree scale-distinctly innervate various sensory– motor brain regions that exhibit heterogeneity in synapse number and strength. Whereas some NM neurons receive localized unitary inputs, individual locus coeruleus noradrenergic (LC-NE) neurons integrate brain-wide inputs with modality-specific spatial organization: Motor inputs converge proximally, while sensory inputs favor distal dendrites of these neurons. LC-NE neurons receive distinct yet overlapping inputs — while most presynaptic neurons establish exclusive one-to-one connections, a specialized subset broadcasts divergent projections to multiple LC-NE targets, creating a hierarchical control architecture. Notably, these shared input patterns extend across monoaminergic systems, suggesting a structural basis for coordinated neuromodulation. Based on the first vertebrate brain-wide EM reconstruction with multiplexed NM-type annotation, our study demonstrate the synaptic organization principles of the LC-NE system’s inputome. Integrated with multi-modal mesoscopic data[8][4], Fish-X offers a valuable resource with various precisely identified NM types, and a critical reference for elucidating synaptic organization principles of bi-pathway architecture in vertebrate brains. ### Competing Interest Statement The authors have declared no competing interest. [1]: #ref-1 [2]: #ref-2 [3]: #ref-7 [4]: #ref-8
Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames. Existing adaptations, such as RoPE-3D, attempt to encode spatial and temporal dimensions separately but suffer from two major limitations: positional bias in attention distribution and disruptions in video-text transitions. To overcome these issues, we propose Video Rotary Position Embedding (VRoPE), a novel positional encoding method tailored for Video-LLMs. Specifically, we introduce a more balanced encoding strategy that mitigates attention biases, ensuring a more uniform distribution of spatial focus. Additionally, our approach restructures positional indices to ensure a smooth transition between video and text tokens. Extensive experiments on different models demonstrate that VRoPE consistently outperforms previous RoPE variants, achieving significant improvements in video understanding, temporal reasoning, and retrieval tasks. Code is available at https://github.com/johncaged/VRoPE.
Volume electron microscopy (vEM) imaging technology was rapidly developed in recent years. It has been the advanced technology to solve high-resolution three-dimensional structures of biological samples. Much wonderful work has revealed the fine structure and interactions of intracellular organelles, the ultrastructure of tissues, and even the three-dimensional structure of entire small biological organisms. With the continuous improvement of resolution, scale and throughput, vEM is becoming more and more widely used in medicine, life sciences, clinical diagnostics and other fields. As a result, this technology has been rated by Nature as one of the seven most noteworthy frontier technologies to watch in 2023. However, the development and application of vEM-related technologies started late in China and need to be further promoted. We write this review to introduce all related vEM technologies, covering the development history of vEM, technology classification, sample preparation, data collection, image processing, etc., which is convenient for people in various fields to understand, learn, apply and further develop this technology.