Photoacoustic computed tomography (PACT) combines optical functional contrast with ultrasonic spatial resolution, enabling noninvasive structural and functional imaging of deep tissues. However, PACT has inherent limitations, including low soft tissue anatomical contrast, depth-dependent quantitative accuracy, and limited penetration depth. By integrating PACT with ultrasound imaging, magnetic resonance imaging (MRI), optical imaging (fluorescence imaging, optical coherence tomography, Raman imaging), and radiological imaging (CT, SPECT/PET), complementary and enhanced information across modalities can be achieved. This review systematically outlines the principles and mechanisms of multimodal PACT techniques, including hardware integration, image registration, probe development, and dual-modality system design. It also introduces practical applications across fields, including tumor diagnosis, neuroscience, vascular disease assessment, and skin lesion detection. Research indicates that multimodal PACT can integrate structural anatomy, functional metabolism, and molecular specificity information, leading to more comprehensive imaging. This provides a powerful tool for preclinical research and clinical translation, offering broad application prospects.
Photoacoustic tomography (PAT) is a powerful imaging modality for visualizing tissue physiology and exogenous contrast agents. However, PAT faces challenges in visualizing deep-seated vascular structures due to light scattering, absorption, and reduced signal intensity with depth. Optical coherence tomography angiography (OCTA) offers high-contrast visualization of vasculature networks, yet its imaging depth is limited to a millimeter scale. Herein, we propose OCPA-Net, a novel unsupervised deep learning method that utilizes the rich vascular feature of OCTA to enhance PAT images. Trained on unpaired OCTA and PAT images, OCPA-Net incorporates a vessel-aware attention module to enhance deep-seated vessel details captured from OCTA. It leverages a domain-adversarial loss function to enforce structural consistency and a novel identity invariant loss to mitigate excessive image content generation. We validate the structural fidelity of OCPA-Net on simulation experiments, and then demonstrate its vascular enhancement performance on in vivo imaging experiments of tumor-bearing mice and contrast-enhanced pregnant mice. The results show the promise of our method for comprehensive vessel-related image analysis in preclinical research applications.
This paper presents a novel method based on leveraging physics-informed neural networks for magnetic resonance electrical property tomography (MREPT). MREPT is a noninvasive technique that can retrieve the spatial distribution of electrical properties (EPs) of scanned tissues from measured transmit radiofrequency (RF) in magnetic resonance imaging (MRI) systems. The reconstruction of EP values in MREPT is achieved by solving a partial differential equation derived from Maxwell’s equations that lacks a direct solution. Most conventional MREPT methods suffer from artifacts caused by the invalidation of the assumption applied for simplification of the problem and numerical errors caused by numerical differentiation. Existing deep learning-based (DL-based) MREPT methods comprise data-driven methods that need to collect massive datasets for training or model-driven methods that are only validated in trivial cases. Hence we proposed a model-driven method that learns mapping from a measured RF, its spatial gradient and Laplacian to EPs using fully connected networks (FCNNs). The spatial gradient of EP can be computed through the automatic differentiation of FCNNs and the chain rule. FCNNs are optimized using the residual of the central physical equation of convection-reaction MREPT as the loss function ( L ). To alleviate the ill condition of the problem, we added multiconstraints, including the similarity constraint between permittivity and conductivity and the ℓ 1 norm of spatial gradients of permittivity and conductivity, to the L . We demonstrate the proposed method with a three-dimensional realistic head model, a digital phantom simulation, and a practical phantom experiment at a 9.4T animal MRI system.
Photoacoustic tomography (PAT) is a promising imaging technique that can visualize the distribution of chromophores within biological tissue. However, the accuracy of PAT imaging is compromised by light fluence (LF), which hinders the quantification of light absorbers. Currently, model-based iterative methods are used for LF correction, but they require extensive computational resources due to repeated LF estimation based on differential light transport models. To improve LF correction efficiency, we propose to use Fourier neural operator (FNO), a neural network specially designed for estimating partial differential equations, to learn the forward projection of light transport in PAT. Trained using paired finite-element-based LF simulation data, our FNO model replaces the traditional computational heavy LF estimator during iterative correction, such that the correction procedure is considerably accelerated. Simulation and experimental results demonstrate that our method achieves comparable LF correction quality to traditional iterative methods while reducing the correction time by over 30 times.
Photoacoustic tomography (PAT) enables non-invasive cross-sectional imaging of biological tissues, but it fails to map the spatial variation of speed-of-sound (SOS) within tissues. While SOS is intimately linked to density and elastic modulus of tissues, the imaging of SOS distribution serves as a complementary imaging modality to PAT. Moreover, an accurate SOS map can be leveraged to correct for PAT image degradation arising from acoustic heterogeneities. Herein, we propose a method for SOS imaging using scanned photoacoustic beacons excited by short laser pulse with inversion reconstruction. Our method is based on photoacoustic reversal beacons (PRBs), which are small light-absorbing targets with strong photoacoustic contrast. We excite and scan a number of PRBs positioned at the periphery of the target, and the generated photoacoustic waves propagate through the target from various directions, thereby achieve spatial sampling of the internal SOS. By picking up the PRB signal using a graph-based dynamic programing algorithm, we formulate a linear inverse model for pixel-wise SOS reconstruction and solve it with iterative optimization technique. We validate the feasibility of the proposed method through simulations, phantoms, and ex vivo biological tissue tests. Experimental results demonstrate that our approach can achieve accurate reconstruction of SOS distribution. Leveraging the obtained SOS map, we further demonstrate significantly enhanced PAT image reconstruction with acoustic correction.
Cylindrical organs, e.g., blood vessels, airways, and intestines, are ubiquitous structures in biomedical optical imaging analysis. Image segmentation of these structures serves as a vital step in tissue physiology analysis. Traditional model-driven segmentation methods seek to fit the structure by constructing a corresponding topological geometry based on domain knowledge. Classification-based deep learning methods neglect the geometric features of the cylindrical structure and therefore cannot ensure the continuity of the segmentation surface. In this paper, by treating the cylindrical structures as a 3D graph, we introduce a novel contour-based graph neural network for 3D cylindrical structure segmentation in biomedical optical imaging. Our proposed method, which we named CylinGCN, adopts a novel learnable framework that extracts semantic features and complex topological relationships in the 3D volumetric data to achieve continuous and effective 3D segmentation. Our CylinGCN consists of a multiscale 3D semantic feature extractor for extracting inter-frame multiscale semantic features, and a residual graph convolutional network (GCN) contour generator that combines the semantic features and cylindrical topological priors to generate segmentation contours. We tested the CylinGCN framework on two types of optical tomographic imaging data, small animal whole body photoacoustic tomography (PAT) and endoscopic airway optical coherence tomography (OCT), and the results show that CylinGCN achieves state-of-the-art performance. Code will be released at https://github.com/lzc-smu/CylinGCN.git.
Objective: To comprehensively capture intra-tumor heterogeneity in head and neck cancer (HNC) and maximize the use of valid information collected in the clinical field, we propose a novel multi-modal image–text fusion strategy aimed at improving prognosis. Method: We have developed a tailored diagnostic algorithm for HNC, leveraging a deep learning-based model that integrates both image and clinical text information. For the image fusion part, we used the cross-attention mechanism to fuse the image information between PET and CT, and for the fusion of text and image, we used the Q-former architecture to fuse the text and image information. We also improved the traditional prognostic model by introducing time as a variable in the construction of the model, and finally obtained the corresponding prognostic results. Result: We assessed the efficacy of our methodology through the compilation of a multicenter dataset, achieving commendable outcomes in multicenter validations. Notably, our results for metastasis-free survival (MFS), recurrence-free survival (RFS), overall survival (OS), and progression-free survival (PFS) were as follows: 0.796, 0.626, 0.641, and 0.691. Our results demonstrate a notable superiority over the utilization of CT and PET independently, and exceed the result derived without the clinical textual information. Conclusions: Our model not only validates the effectiveness of multi-modal fusion in aiding diagnosis, but also provides insights for optimizing survival analysis. The study underscores the potential of our approach in enhancing prognosis and contributing to the advancement of personalized medicine in HNC.
In gradient echo MRI, quantitative susceptibility mapping (QSM) quantifies the magnetic susceptibility distributions of tissues, which has great potential in detecting brain diseases. However, QSM reconstruction is an ill-conditional inversion problem because of the zeros in the frequency domain of the dipole kernel. The intrinsic nature of the ill-posedness would affect the accuracy of quantifying tissue susceptibility. Recently, deep learning-based methods have been proposed to improve accuracy by suppressing the streaking artifacts. In this work, we proposed a hybrid architecture to enforce data consistency by involving numerical optimization blocks within convolutional neural networks (CNN), which aimed to reconstruct high-quality QSM images, referred to as NoQSM-net. The Calculation of Susceptibility through Multiple Orientation Sampling (COSMOS) QSM maps were used as labels for training. The performance of the proposed method was evaluated on two healthy volunteers and brain images of patients with diseases. Our experiments showed that the proposed method achieved good performance in terms of quantitative metrics and could effectively suppress artifacts in reconstructed QSM images, demonstrating its potential for future applications. For experiments on patients with multiple sclerosis (MS), the proposed method could better detect lesion regions in the results of NoQSM-net.
Deep learning methods show great potential for the efficient and precise estimation of quantitative parameter maps from multiple magnetic resonance (MR) images. Current deep learning-based MR parameter mapping (MPM) methods are mostly trained and tested using data with specific acquisition settings. However, scan protocols usually vary with centers, scanners, and studies in practice. Thus, deep learning methods applicable to MPM with varying acquisition settings are highly required but still rarely investigated. In this work, we develop a model-based deep network termed MMPM-Net for robust MPM with varying acquisition settings. A deep learning-based denoiser is introduced to construct the regularization term in the nonlinear inversion problem of MPM. The alternating direction method of multipliers is used to solve the optimization problem and then unrolled to construct MMPM-Net. The variation in acquisition parameters can be addressed by the data fidelity component in MMPM-Net. Extensive experiments are performed on R2 mapping and R1 mapping datasets with substantial variations in acquisition settings, and the results demonstrate that the proposed MMPM-Net method outperforms other state-of-the-art MR parameter mapping methods both qualitatively and quantitatively.
Photoacoustic tomography (PAT), as a novel medical imaging technology, provides structural, functional, and metabolism information of biological tissue in vivo. Sparse Sampling PAT, or SS-PAT, generates images with a smaller number of detectors, yet its image reconstruction is inherently ill-posed. Model-based methods are the state-of-the-art method for SS-PAT image reconstruction, but they require design of complex handcrafted prior. Owing to their ability to derive robust prior from labeled datasets, deep-learning-based methods have achieved great success in solving inverse problems, yet their interpretability is poor. Herein, we propose a novel SS-PAT image reconstruction method based on deep algorithm unrolling (DAU), which integrates the advantages of model-based and deep-learning-based methods. We firstly provide a thorough analysis of DAU for PAT reconstruction. Then, in order to incorporate the structural prior constraint, we propose a nested DAU framework based on plug-and-play Alternating Direction Method of Multipliers (PnP-ADMM) to deal with the sparse sampling problem. Experimental results on numerical simulation, in vivo animal imaging, and multispectral un-mixing demonstrate that the proposed DAU image reconstruction framework outperforms state-of-the-art model-based and deep-learning-based methods.
Photoacoustic tomography (PAT), as a novel biomedical imaging technique, is able to capture temporal, spatial and spectral tomographic information from organisms. Organ-level multi-parametric analysis of continuous PAT images are of interest since it enables the quantification of organ specific morphological and functional parameters in small animals. Accurate organ delineation is imperative for organ-level image analysis, yet the low contrast and blurred organ boundaries in PAT images pose challenge for their precise segmentation. Fortunately, shared structural information among continuous images in the time-space-spectrum domain may be used to enhance segmentation. In this paper, we introduce a structure fusion enhanced graph convolutional network (SFE-GCN), which aims at automatically segmenting major organs including the body, liver, kidneys, spleen, vessel and spine of abdominal PAT image of mice. SFE-GCN enhances the structural feature of organs by fusing information in continuous image sequence captured at time, space and spectrum domains. As validated on large-scale datasets across different imaging scenarios, our method not only preserves fine structural details but also ensures anatomically aligned organ contours. Most importantly, this study explores the application of SFE-GCN in multi-dimensional organ image analysis, including organ-based dynamic morphological analysis, organ-wise light fluence correction and segmentation-enhanced spectral un-mixing. Code will be released at https://github.com/lzc-smu/SFEGCN.git.
High censoring phenomenon usually occurs in cancer prognosis analysis, which, however, would introduce bias for model construction and limit generalization performance. In this paper, we first explore and identify an appropriate censoring range for cancer prognosis evaluation, upon which we present a mix-supervised multiset learning framework to cope with high-censoring data. Specifically, we construct multiple subsets with the specified censoring proportion, followed by a multiset representation learning method to learn subset-specific representations, which equips with adversary integrality preservation and dependency limitation constraints to ensure the unbiasedness of subsets and eliminate the redundancy among subsets, respectively. Furthermore, a mix-supervised multiset fusion model is proposed to estimate the relative survival risk, in which teacher model can make full use of the survival time of uncensored samples and the prognosis-related attributes of censored ones to generate reliable pseudo-labels and latent-space for student model. We evaluate the proposed method on three public datasets, and extensive experimental results demonstrate its superiority.
Photoacoustic tomography (PAT) and magnetic resonance imaging (MRI) are two advanced imaging techniques widely used in pre-clinical research. PAT has high optical contrast and deep imaging range but poor soft tissue contrast, whereas MRI provides excellent soft tissue information but poor temporal resolution. Despite recent advances in medical image fusion with pre-aligned multimodal data, PAT-MRI image fusion remains challenging due to misaligned images and spatial distortion. To address these issues, we propose an unsupervised multi-stage deep learning framework called PAMRFuse for misaligned PAT and MRI image fusion. PAMRFuse comprises a multimodal to unimodal registration network to accurately align the input PAT-MRI image pairs and a self-attentive fusion network that selects information-rich features for fusion. We employ an end-to-end mutually reinforcing mode in our registration network, which enables joint optimization of cross-modality image generation and registration. To the best of our knowledge, this is the first attempt at information fusion for misaligned PAT and MRI. Qualitative and quantitative experimental results show the excellent performance of our method in fusing PAT-MRI images of small animals captured from commercial imaging systems.
Multispectral photoacoustic tomography (PAT) is an imaging modality that utilizes the photoacoustic effect to achieve non-invasive and high-contrast imaging of internal tissues but also molecular functional information derived from multi-spectral measurements. However, the hardware cost and computational demand of a multispectral PAT system consisting of up to thousands of detectors are huge. To address this challenge, we propose an ultra-sparse spiral sampling strategy for multispectral PAT, which we named U3S-PAT. Our strategy employs a sparse ring-shaped transducer that, when switching excitation wavelengths, simultaneously rotates and translates. This creates a spiral scanning pattern with multispectral angle-interlaced sampling. To solve the highly ill-conditioned image reconstruction problem, we propose a self-supervised learning method that is able to introduce structural information shared during spiral scanning. We simulate the proposed U3S-PAT method on a commercial PAT system and conduct in vivo animal experiments to verify its performance. The results show that even with a sparse sampling rate as low as 1/30, our U3S-PAT strategy achieves similar reconstruction and spectral unmixing accuracy as non-spiral dense sampling. Given its ability to dramatically reduce the time required for three-dimensional multispectral scanning, our U3S-PAT strategy has the potential to perform volumetric molecular imaging of dynamic biological activities.
Quantitative imaging has been very useful in neuroscientific and clinical applications, including glioma, tumor diagnosis and prognosis, brain maturation, and Alzheimer's disease. EPI is a powerful tool for quantitative imaging owing to its extremely fast acquisition. This work aims to develop a distortion-free, blip-up/down acquisition (BUDA) 3D-EPI with controlled aliasing in parallel imaging (CAIPI) sampling and joint Hankel low-rank image reconstruction for fast and robust multi-contrast high-resolution whole-brain imaging. The developed technique could generate distortion-free high-resolution whole-brain T2* mapping and quantitative susceptibility mapping in 47s at 1.1×1.1×1 mm3 resolution.
Magnetic resonance imaging (MRI) and photoacoustic tomography (PAT) offer two distinct image contrasts. To integrate these two modalities, we present a comprehensive hardware-software solution for the successive acquisition and co-registration of PAT and MRI images in in vivo animal studies. Based on commercial PAT and MRI scanners, our solution includes a 3D-printed dual-modality imaging bed, a 3-D spatial image co-registration algorithm with dual-modality markers, and a robust modality switching protocol for in vivo imaging studies. Using the proposed solution, we successfully demonstrated co-registered hybrid-contrast PAT-MRI imaging that simultaneously displays multi-scale anatomical, functional and molecular characteristics on healthy and cancerous living mice. Week-long longitudinal dual-modality imaging of tumor development reveals information on size, border, vascular pattern, blood oxygenation, and molecular probe metabolism of the tumor micro-environment at the same time. The proposed methodology holds promise for a wide range of pre-clinical research applications that benefit from the PAT-MRI dual-modality image contrast.
The Cox proportional hazard model has been widely applied to cancer prognosis prediction. Nowadays, multi-modal data, such as histopathological images and gene data, have advanced this field by providing histologic phenotype and genotype information. However, how to efficiently fuse and select the complementary information of high-dimensional multi-modal data remains challenging for Cox model, as it generally does not equip with feature fusion/selection mechanism. Many previous studies typically perform feature fusion/selection in the original feature space before Cox modeling. Alternatively, learning a latent shared feature space that is tailored for Cox model and simultaneously keeps sparsity is desirable. In addition, existing Cox-based models commonly pay little attention to the actual length of the observed time that may help to boost the model's performance. In this article, we propose a novel Cox-driven multi-constraint latent representation learning framework for prognosis analysis with multi-modal data. Specifically, for efficient feature fusion, a multi-modal latent space is learned via a bi-mapping approach under ranking and regression constraints. The ranking constraint utilizes the log-partial likelihood of Cox model to induce learning discriminative representations in a task-oriented manner. Meanwhile, the representations also benefit from regression constraint, which imposes the supervision of specific survival time on representation learning. To improve generalization and alleviate overfitting, we further introduce similarity and sparsity constraints to encourage extra consistency and sparseness. Extensive experiments on three datasets acquired from The Cancer Genome Atlas (TCGA) demonstrate that the proposed method is superior to state-of-the-art Cox-based models.