Optical-resolution photoacoustic microscopy (OR-PAM) provides the superior penetration depth and enables the visualization of optical absorption distributions within diverse biological tissues. However, achieving high spatial resolution across large tissue areas within a short timeframe remains a significant challenge, primarily due to the limitations of traditional scanning methods and laser repetition rate. While various optimization strategies have been proposed for OR-PAM, the integration of the scanning techniques with advanced deep learning-based models to address the spatiotemporal resolution trade-off remains underexplored. In this study, we present a rapid photoacoustic microscopy framework that combines rotary-scanning imaging with an enhanced denoising diffusion probabilistic model, termed as RotDiffTransPAM. The raw radiofrequency (RF) signals acquired through rotary-scanning are directly undersampled and used as inputs to the model. RotDiffTransPAM introduces a hybrid attention mechanism comprising anchored stripe attention, channel attention, and shifted windowbased self-attention within the 16 x 16 resolution layers, replacing the conventional self-attention mechanism in DDPM. Experimental validation on both ex vivo bovine cancellous bone and in vivo mouse datasets demonstrates the superior performance of RotDiffTransPAM over several classic image denoising models. The experimental results show that with a downsampling ratio of [2,2], RotDiffTransPAM achieved an 11.41% improvement in peak signal-to-noise ratio (PSNR) (2.7705 dB) and a 9.48% increase in structural similarity index (SSIM) (0.0773) compared to bilinear interpolation on the cancellous bone test set. Moreover, in vivo experiments further confirmed that RotDiffTransPAM enable to effectively reconstruct high-resolution OR-PAM images of vascular structures in critical regions such as the mouse brain and kidney.
Pneumonia is an acute respiratory infection, posing a serious threat to health and lives. Lung ultrasound (LUS), as a non-invasive and rapid imaging technique, can monitor real-time changes in lung, providing valuable assistance in clinical diagnosis. However, most LUS studies are limited to frame-level analysis and ignore respiratory cycle changes, leading to diagnostic errors. To address these problems, we propose a cross-dimensional spatial-temporal feature integration model for LUS video analysis. Specifically, the sliding window and feature difference analysis are first utilized to preprocess the original LUS videos for eliminating invalid and highly similar frames and implementing abstract video. Subsequently, a cross-dimensional feature fusion backbone integrates an improved temporal-C3D network and a self-designed recursive inception-meet-transformer (IMT) network to extract features from different dimensions for fusion. Thereby, comprehensive features can be obtained for characterizing LUS videos. Finally, the Longformer is employed to analyze the temporal dependencies of cross-dimensional features, supplemented by a classification head for evaluating LUS videos. 3018 LUS video clips were collected from 119 patients in three hospitals for the evaluation of the proposed LUS video scoring model. By dividing at the patient level, the training and testing set consist of 2652 clips from 104 patients and 366 clips from 15 patients, respectively. Experimental results of 5-fold cross validation demonstrate that the proposed model achieves outstanding scoring performance, with an accuracy, precision, recall, specificity, F1-score, and AUC of $91.78~\pm ~0.52$ %, $92.19~\pm ~0.76$ %, $91.81~\pm ~0.62$ %, $97.17~\pm ~0.20$ %, $91.94~\pm ~0.44$ %, and $97.83~\pm ~0.19$ %, respectively. The independent testing set also shows the superior generalization capability with a scoring accuracy of $87.65~\pm ~1.12$ %. Moreover, ablation studies confirm that each designed module contributes significantly to the model's performance, and comparative experiments further confirm the superiority of the proposed model compared to previous models. These robust findings highlight the proposed LUS video scoring model's strong potential for clinical deployment.
The 2-D ultrasonography has shown great potential for evaluating periodontal structures due to its portability, flexibility, and nonionizing nature. However, for clinical applications, 3-D ultrasound imaging of oral anatomy is essential. This study presents the development of a freehand 3-D intraoral imaging system using the high-frequency ultrasound for reconstructing the 3-D anatomical structures of teeth. The proposed hand-held system integrates an in-house designed ultrasound transducer with positioning markers and four optical cameras to acquire the 2-D ultrasound images along with corresponding 3-D tracking poses. A forward-mapping algorithm was employed for 3-D data fusion to enable the real-time imaging. Both dimensional measurements and 3-D visualized difference colormaps were generated for a total of 24 teeth from maxillary and mandibular phantoms to compare the proposed system with an optical scanner and cone-beam computed tomography (CBCT). The mean distance differences ranged from -0.13 to 0.18 mm versus the optical scanner and from -0.18 to 0.15 mm versus CBCT. The overall mean absolute difference (MAD) across all dimensional measurements of two raters was 0.17 mm for the optical scanner and 0.16 mm for CBCT, respectively. The correlation coefficient with both reference modalities exceeded 0.98. The average model-to-model distance between the 3-D volumes acquired from the proposed system and the corresponding reference volumes was approximately 0.20 mm for the optical scanner and 0.28 mm for CBCT. The proposed 3-D intraoral imaging system demonstrated the feasibility of achieving high-quality 3-D intraoral imaging in real time, showing the potential to significantly accelerate the clinical adoption of the ultrasound technology in dentistry.
Effective clutter filtering is essential for ultrafast ultrasound small blood flow imaging, where weak blood flow signals are often obscured by tissue clutter and noise. While singular value decomposition (SVD)-based methods are widely used, their performance is limited by overlapping singular value distributions between blood flow, clutter, and noise. This study proposes a Cauchy-norm-based sparse SVD clutter filtering method that integrates a sparsity penalty into the classical SVD framework. By leveraging the structural sparsity of blood flow, the method employs a Cauchy norm regularization to suppress noise while preserving small vessel details. A randomized spatial downsampling strategy is further used to increase computational efficiency. The proposed method is validated through simulated and in vivo rat brain experiments, demonstrating superior performance in power Doppler imaging and functional ultrasound imaging compared to conventional SVD and robust principal component analysis. Results show significant improvements in signal-to-noise ratio by about 10 dB and contrast-to-noise ratio by about 9 dB, with enhanced visualization of small vascular networks and dynamic cerebral blood volume responses to whisker stimulation. Future work will focus on accelerating optimization and integrating deconvolution for further resolution enhancement.
As a noninvasive medical imaging modality, ultrasound offers the advantages of safety, convenience, and affordability. However, bone imaging has long been a challenge due to the high acoustic impedance. Inversion algorithms have shown promise in advancing ultrasound computed tomography (UCT) as a viable technique for bone imaging. Traditional physics-driven inversion algorithms, such as full-waveform inversion (FWI), are computationally intensive and tend to be trapped in local minima. In this study, we propose a deep-learning (DL) inversion method to achieve fast bone imaging. This data-driven approach establishes a mapping from ultrasound data to bone sound speed (SoS) maps. Moreover, a physics-informed prior knowledge enhancement module is designed to extract and fuse travel-time and low-frequency information from the data. This enables the network to focus on key features based on physical principles and improves its generalization capability under waveform-domain variations. In simulation tests, the DL inversion can reconstruct a cortical bone image within two seconds and achieves a mean structural similarity index measure (SSIM) of 0.9823. Although trained on data from a specific source, the proposed network remains robust to waveform -domain mismatches, such as changes in pulselength, center frequency, measurement noise, and propagation effects introduced by soft tissues. In laboratory validation, despite being trained on simulated datasets, the network is able to reconstruct bone phantom images using experimental data, with an SSIM of 0.9528. This demonstrates the robustness of the prior-informed DL inversion method for bone imaging, which offers a rapid, accurate, and adaptable solution that overcomes the limitations of traditional inversion methods.
ABSTRACT Closed‐loop bioelectronic systems that adapt stimulation to real‐time physiological feedback hold transformative potential for treating neurological and cardiac disorders and are emerging as key components of future ultrasonic brain–machine interfaces (uBMIs). Realizing this requires the simultaneous achievement of millimeter‑scale deep‐tissue targeting, artifact‐free physiological feedback, and robust wireless power and data transfer, which remain elusive with current methods. Here, we present an integrated ultrasonic platform engineered to overcome these fundamental limitations. We propose a physics‐constrained metasurface design framework to enable high‐resolution multifocal ultrasound energy delivery through highly aberrating biological barriers such as the skull and ribs, achieving improved experimental targeting accuracy (e.g., ±6.5% intensity uniformity across multiple foci). We demonstrate the platform's adaptive stimulation capabilities through two distinct paradigms: attention‐based ultrasound stimulation and cardiac‐synchronized ultrasound stimulation. Furthermore, we introduce a novel dual‐channel acoustic link that enables continuous wireless power and wireless data streaming through the skull with a single acoustic metasurface, demonstrating robustness even with a 400‐fold power differential. This integrated ultrasonic framework, providing seamless integration of precise spatial targeting through biological barriers, adaptive physiological feedback, and untethered operation, contributes to the development of next‐generation uBMIs and closed‐loop bioelectronic therapies.
Objective: To improve image quality in multispectral optoacoustic tomography (MSOT) under conditions of strong noise and extreme sparse sampling. Impact Statement: This work provides a practical solution to enhance MSOT image quality while reducing hardware requirements, which may help expand its use in preclinical and clinical imaging. Introduction: MSOT provides useful biochemical and molecular contrast in tissues. However, image quality is often limited by system noise and sparse sampling. Methods: We propose the optoacoustic universal denoising network (OA-UDNet), a hybrid diffusion-based framework for sparse MSOT data. The model is trained on more than 250,000 in vivo images. It performs joint denoising and high-fidelity image restoration by combining an edge-aware module with diffusion-based generation to preserve structural boundaries. Results: With data from only 32 detectors, the method improves image quality and increases peak signal-to-noise ratio (PSNR) by nearly 14 dB. Evaluated against the standard 256-detector reference, the framework consistently preserves structural fidelity under extreme undersampling. Validation on whole-body mouse imaging, tumor models, and human samples shows reduced artifacts and improved recovery of anatomical and functional features. Conclusion: OA-UDNet improves MSOT image quality under sparse conditions. It offers a simple and effective way to accelerate imaging while reducing hardware complexity.
Ultrasound (US) imaging is emerging as a promising radiation-free modality for three-dimensional (3D) visualization of the midpalatal suture (MPS), a key anatomical structure in rapid maxillary expansion treatment. However, current approaches are largely confined to laboratory settings, rely on mechanical scanning systems, and require extensive manual image processing. Additionally, interpreting and processing of intraoral US images remain challenging and timeconsuming for clinicians. To overcome these challenges and facilitate clinical translation, this study proposes a fully automated, end-to-end framework for freehand 3D US visualization of the MPS. The framework integrates optical tracking-based freehand US acquisition, automated deep-learning-based anatomical segmentation, and 3D image reconstruction within a single workflow. Feasibility was evaluated using ex-vivo pig jaw samples with simulated MPS opening. Quantitative validation against gold-standard microcomputed tomography demonstrated strong agreement for MPS separation measurements $\left(\boldsymbol{R}^{\mathbf{2}}=\mathbf{0 . 9 2}\right)$, with a mean absolute difference of $\mathbf{0 . 2 2} \pm \mathbf{0 . 1 6 ~ m m}$ and 97% of measurements within the clinically acceptable tolerance of 0.5 mm. These findings demonstrated the feasibility of automated freehand 3D US for objective, repeatable, and radiation-free assessment of MPS opening.
Cerebrospinal fluid circulation through the glymphatic system plays a crucial role in removing metabolic waste from the central nervous system. However, the mechanism underlying the brain-wide glymphatic dynamics is not yet fully understood, in part due to the lack of glymphatic imaging technologies on deep brains. Here, we report a hybrid imaging technology that integrates three-dimensional photoacoustic tomography and ultrasound localization microscopy (3D-PAULM), enhanced by a photoacoustic dye with strong optical absorption in the second near-infrared window (NIR-II). 3D-PAULM allows for continuous, noninvasive, whole-brain imaging in mice through intact skull, providing superresolution mapping of the brain vasculature and highly sensitive tracing of the NIR-II dye in the glymphatic system. Using 3D-PAULM, we investigated the glymphatic function impaired by ischemic stroke, aging, and anesthesia. Our results provide insights into glymphatic transport under various physiological as well as pathological conditions and establish 3D-PAULM as a valuable tool for preclinical glymphatic research.
Ischemic stroke represents a leading cause of global mortality and long-term disability. Consequently, rapid quantification of cerebral damage severity, coupled with timely neuromodulatory intervention, is imperative for improving clinical prognosis. In this study, we present a label-free quantitative monitoring framework utilizing optical-resolution photoacoustic microscopy (OR-PAM) to assess microvascular alterations and evaluate the therapeutic efficacy of low-intensity transcranial ultrasound stimulation (LITUS). A graded ischemic stroke model was established by modulating photothrombotic irradiation duration (3, 5, and 10 min) to validate system sensitivity. Subsequently, a Hessian filter-based segmentation pipeline was employed to extract quantitative vascular metrics. Five key morphological and functional parameters, including the photoacoustic (PA) signal intensity, vessel area fraction (VAF), average vessel diameter (AVD), branch-point number (BN), and perfused vessel density (PVD), were extracted to characterize the severity of ischemic injury. Our analysis revealed a strong negative correlation between irradiation duration and vascular perfusion-related metrics, which was further confirmed by ex vivo 2,3,5-triphenyltetrazolium chloride (TTC) staining showing duration-dependent infarct expansion. Applying this quantitative framework to the established model, we further demonstrated that acute LITUS treatment significantly alleviated microvascular hypoperfusion and promoted rapid hemodynamic recovery by restoring both functional perfusion and vascular morphology via vasodilation. These findings highlight the potential of a PAM-based quantitative framework for accurate grading of ischemic severity and evaluation of ultrasound-based neuromodulation.
This study developed and validated a deep learning model for diagnosing lymphadenopathy (LA) using B-mode ultrasound (BUS) and color Doppler flow imaging (CDFI) videos. A retrospective and prospective study was conducted from January 2016 to August 2025, including 7371 patients (3824 male [51.9
Acoustic imaging is widely used in biomedicine and industrial nondestructive testing, yet its spatial resolution remains fundamentally limited by acoustic diffraction. Despite significant progress in nonlinear harmonic imaging, localization microscopy, and acoustic metamaterials, these approaches face intrinsic trade-offs between efficiency, invasiveness, structural complexity, and scalability. Here, we introduce a far-field super-resolution ultrasound imaging strategy based on composite orbital angular momentum (OAM). By synthesizing vortex beams spanning multiple topological charges, the collective angular spectrum converges toward a Dirac delta function, enabling the recovery of high spatial frequencies from far-field scattered fields. Numerical simulations and underwater experiments demonstrate robust subdiffraction imaging of metallic wires, complex steel structures, and blood vessels, achieving resolutions down to 0.2λ in simulation and 0.24λ in experiment at 1 MHz, well beyond the Rayleigh limit. Moreover, tailored OAM spectra enable functional imaging capabilities such as chirality discrimination, offering sensitivity to structural handedness. Our results establish composite OAM as a powerful new degree of freedom for ultrasound imaging.
Obtaining information on bone metabolism through intraoperative or non-invasive examination remains a challenge in medical practice. Photoacoustic (PA) spectroscopy offers a method for identifying molecules within biological tissues by exploiting the contrast in their optical absorption properties. However, difficulties arise when analyzing bone tissue, which comprises a complex mixture of organic and inorganic components. The overlapping optical absorption peaks of various chemical constituents in bone tissue can significantly hinder the accuracy of PA absorption spectra decoupling inversion. In this study, we developed a decoupling technique that integrates PA correlation spectra (PCS) with physics-guided machine learning to analyze the chemical components of bone tissue quantitatively. The feasibility of using PCS and machine learning for bone metabolism was evaluated through numerical simulations and experimental studies on bone models with varying chemical compositions. The calculated quantitative parameters for the chemical components closely matched the ground truth values, thus allowing for the characterization of changes in bone composition. This non-invasive, radiation-free PA technology holds promise for advancing the diagnosis and monitoring of bone diseases.
Abdominal trauma with bleeding is a leading cause of post-traumatic death, and detecting free fluid in the abdomen or hemoperitoneum can provide critical guidance for clinical management. Rapid and accurate diagnosis of abdominal bleeding using ultrasound is significant for making decisions regarding the need for surgical intervention. This study introduces a multi-task network for the segmentation and classification of ascites in ultrasound images. The network utilizes a U-Net backbone with a ResNext encoder as the basic architecture for the segmentation and classification models. The segmentation network includes a Frequency Channel Attention (FCA) attention module, which effectively broadens the range of captured information and enhances the robustness of channel representation. Furthermore, an Enhanced Channel Attention Multi Feature Fusion (EMFF) was used to extract the interdependencies between feature channels by combining high-order and low-order feature mappings, thereby improving segmentation accuracy. Lastly, a classification branch was created to classify ascites by sharing encoder features. Experiments on the collected ascites ultrasound dataset demonstrated that the proposed method achieved a segmentation Dice of 85.28% and a classification accuracy of 86.18%. It outperformed the leading multi-task SOTA method by 0.7% in Dice and 2.03% in accuracy, establishing a new benchmark for simultaneous ascites assessment. This study showed that the proposed network is valuable for the preliminary diagnosis of ascites in ultrasound and can serve as a potential auxiliary tool for clinical ascites examination in emergency situations.
Accurate orientation estimation using miniature inertial measurement units (IMUs) is critical for portable biomedical tracking applications, particularly for guiding freehand ultrasound probes. However, the combined effect of sampling rate, motion speed, and filtering algorithm on estimation performance remains insufficiently characterized. This study presents a systematic evaluation of orientation accuracy using a custom-built six-axis IMU module ( 2 & times;2.5cm) across varying sampling rates and controlled rotational speeds. Static experiments assessed accelerometer-based tilt estimation from 0 degrees to +/- 90 degrees in 10 degrees increments, using a 0.1 degrees resolution digital protractor as ground truth, yielding mean absolute errors (MAEs) below 1.5 degrees. Dynamic experiments evaluated six commonly used filters: complementary filter, Madgwick filter, Mahony filter, Kalman filter (KF), extended Kalman filter (EKF), and unscented Kalman filter (UKF), under four sampling rates (12.5, 26, 52, and 104 Hz) and eight angular velocities (15 degrees-120 degrees/s at 15 degrees/s intervals). A robotic arm executed controlled roll and pitch motions within +/- 160 degrees, while an optical tracking system provided reference orientations. Results show that orientation accuracy strongly depends on both sampling frequency and rotational speed. At low speeds (<= 30 degrees/s), all filters achieved MAE < 2 degrees even at low-sampling rates (12.5 and 26 Hz). At medium speeds (45 degrees-75 degrees/s), Kalman-based filters achieved MAE < 2 degrees at 52 and 104 Hz. At high speeds ( >= 90 degrees), the EKF achieved the best performance, with MAE < 1 degrees at 26 Hz. Overall, the EKF operating at 26 Hz provided the best balance between accuracy, robustness, and computational efficiency, supporting its suitability for low-power biomedical tracking applications. This work provides the first systematic characterization of how sampling frequency and motion speed jointly influence IMU orientation accuracy in portable biomedical tracking systems.
Computer-aided diagnosis (CAD) technology has become an integral part of early breast cancer diagnosis in medical ultrasound imaging. Nonetheless, the accurate segmentation and classification of breast tumor images continue to pose significant challenges due to the uneven intensity distribution, indistinct boundaries, and the irregular shapes of tumors. A primary reason is that many CAD methods fail to capitalize on the interrelation between tumor segmentation and classification. Furthermore, most existing methods focus on classifying global images or tumor regions of interest (ROIs), often disregarding the interplay between global and local features. This study proposes a segmentation knowledge based global-local attention classification network (SGLA-Net). First, the segment anything model (SAM) is utilized to obtain high-quality segmentation masks from a limited number of annotated samples. Then, global and local feature representations are derived from enhanced images obtained through the segmentation mask. Moreover, a global-local feature interaction (GLFI) block is designed to adaptively integrate global and local information for classification. The proposed method achieved segmentation Dice coefficients of 81.239 % and 80.516 % on the internal and external datasets, respectively. In terms of classification, the method obtained Area Under the Curve (AUC) values of 0.9532 and 0.8521 on the internal and external datasets, surpassing five state-of-the-art breast ultrasound classification methods.
Plane wave (PW) transmission enables ultrahigh-frame-rate imaging by insonifying the entire field of view with a full-array excitation. However, because PW transmission lacks transmit focusing, the echo signals suffer from low signal-to-noise ratio, which in turn degrades the overall image quality. In contrast, synthetic transmit aperture (STA) imaging uses single-element transmission, which reduces interference and noise in the measured echoes. Combined with full-aperture reception and dynamic focusing during image reconstruction, STA enables high-quality imaging. However, STA requires as many sequential transmissions as the number of probe elements, resulting in a significantly reduced frame rate due to the prolonged acquisition time. In this context, we propose designing a convolutional neural network (CNN) to map the apodized PW measurement signals to their STA-equivalent counterparts, aiming to preserve the high frame rate of PW imaging while achieving image quality comparable to that of STA. Specifically, based on the superposition principle of linear acoustics, we first establish a mathematical measurement relationship between PW and STA signals. This theoretical model then guides the design of a CNN for measurement signal transformation. Finally, high-resolution B-mode ultrasound images are reconstructed using the delay-and-sum (DAS) beamforming algorithm. We trained the network via self-supervised learning on a simulated dataset generated by Field II and assessed its performance through both in vitro and in vivo experiments. The results demonstrate the feasibility and effectiveness of using a CNN to generate STA measurement signals from apodized PW measurement signals, thereby advancing ultrasound imaging by combining high frame rates with improved image quality. Notably, the proposed method exhibits excellent performance in both lateral and axial (longitudinal) resolution. Overall, this study presents a promising approach for achieving high spatiotemporal resolution in ultrasound imaging based on apodized PW acquisition.
Abstract Full waveform inversion (FWI) enables high-resolution reconstruction of the speed of sound (SOS) in medical ultrasound imaging, but conventional FWI does not quantify reconstruction uncertainty. We propose a Stein variational gradient descent (SVGD)-based FWI framework for posterior approximation and uncertainty quantification, in which the posterior distribution of SOS is represented by interacting particles. The particle updates are driven by the log-posterior gradient, combining an adjoint-state likelihood term with a Gaussian prior. The converged particle ensemble provides both the posterior mean SOS and the posterior standard deviation. The method is validated on synthetic models and a realistic breast-tissue model. Numerical results show that SVGD yields accurate SOS reconstructions, with mean estimates comparable to or slightly better than those of conventional FWI, while also providing informative uncertainty maps. Reliable mean reconstructions can be obtained with relatively few particles, whereas uncertainty estimates become more stable as the particle number increases. In non-ideal cases with source-signature mismatch and source–receiver positioning error, the posterior standard deviation remains useful for identifying regions associated with inversion failure or acquisition error. The uncertainty estimates produced by SVGD are consistent with the Metropolis–Hasting–Markov Chain Monte Carlo benchmark along selected profiles. Overall, SVGD provides a practical compromise between reconstruction accuracy, computational feasibility, and uncertainty quantification for medical ultrasound FWI.
Ultrasound computed tomography (UCT) via full waveform inversion (FWI) enables high-resolution quantitative imaging for tissue characterization and disease diagnosis. However, UCT suffers from large computational burden and severe convergence issues due to highly nonlinear optimization. Deep learning can accelerate UCT reconstruction, but supervised training requires large-scale labeled datasets difficult to obtain in vivo. To address these limitations, we propose SDA-UCT, a two-stage self-supervised domain-adaptive framework for rapid and accurate UCT imaging of musculoskeletal tissues. SDA-UCT employs an attention-enhanced network (AttUCT) pre-trained on simulation datasets and transfers to in-vivo data via physics-informed self-supervised learning, effectively bridging the simulation-to-real domain gap. A Low-Rank Adaptation (LoRA) mechanism is integrated to enable efficient adaptation across diverse clinical scenarios. Results showed that AttUCT achieved high-quality SOS reconstruction for simulated human forearm with a PSNR of 29.23 dB and SSIM of 0.928, outperforming conventional FWI and existing deep learning methods. Validated on in-vivo data, SDA-UCT successfully reconstructed SOS images revealing complex anatomical structures (skin, fat, muscle, tendon, bone and bone marrow) for human forearm, in high concordance with MRI references. The LoRA mechanism adjusting only 3
Weiqi Wang (王威琪)合作论文数Institute of Biomedical Engineering and Technology, Fudan University95