Ultrasound computed tomography (UCT) via full waveform inversion (FWI) enables high-resolution quantitative imaging for tissue characterization and disease diagnosis. However, UCT suffers from large computational burden and severe convergence issues due to highly nonlinear optimization. Deep learning can accelerate UCT reconstruction, but supervised training requires large-scale labeled datasets difficult to obtain in vivo. To address these limitations, we propose SDA-UCT, a two-stage self-supervised domain-adaptive framework for rapid and accurate UCT imaging of musculoskeletal tissues. SDA-UCT employs an attention-enhanced network (AttUCT) pre-trained on simulation datasets and transfers to in-vivo data via physics-informed self-supervised learning, effectively bridging the simulation-to-real domain gap. A Low-Rank Adaptation (LoRA) mechanism is integrated to enable efficient adaptation across diverse clinical scenarios. Results showed that AttUCT achieved high-quality SOS reconstruction for simulated human forearm with a PSNR of 29.23 dB and SSIM of 0.928, outperforming conventional FWI and existing deep learning methods. Validated on in-vivo data, SDA-UCT successfully reconstructed SOS images revealing complex anatomical structures (skin, fat, muscle, tendon, bone and bone marrow) for human forearm, in high concordance with MRI references. The LoRA mechanism adjusting only 3
Maximum-likelihood (ML) direction-of-arrival (DOA) estimation for vector hydrophone arrays provides strong statistical performance. However, multidimensional nonlinear search still causes a heavy computational burden. To address this problem, this paper proposes an ML-DOA estimation method based on an improved velvet worm optimization (IVWO) algorithm. The original velvet worm optimization (VWO) algorithm balances global exploration and local exploitation, but its direct application to complex ML-DOA objective functions exhibits insufficient initial search coverage, an abrupt exploration–exploitation transition, and limited late-stage local search. Accordingly, IVWO introduces randomized Hammersley initialization, proposes an iteration-dependent three-stage smooth quota allocation strategy, and incorporates logarithmic-spiral local search to improve initial coverage, transition smoothness, and local search, respectively. Simulations under different SNRs, numbers of snapshots, population sizes, and numbers of sources show that IVWO-ML achieves faster convergence than the Improved Secretary Bird Optimization Algorithm (IMSBOA), Improved Invasive Weed Optimization (IIWO), Genetic Algorithm (GA), Particle Swarm Optimization (PSO), the original VWO, and Grey Wolf Optimizer (GWO), while also attaining lower RMSE and better run-to-run stability, especially in closely spaced multi-source scenarios. Compared with the traditional Multiple Signal Classification (MUSIC) and Alternating Projection Maximum Likelihood (AP-ML) baselines, IVWO-ML also provides higher estimation accuracy and stronger robustness under low-SNR and closely spaced multi-source scenarios.
Medical ultrasound full-waveform inversion (FWI) is typically formulated as a deterministic optimization problem and provides limited uncertainty information. This study presents a Bayesian FWI framework for quantitative speed-of-sound (SOS) imaging. The medium is parameterized by squared slowness, and a Gaussian additive-noise model yields a posterior distribution proportional to a data-misfit term and a prior term. Posterior statistics are approximated via Metropolis-Hastings Markov chain Monte Carlo (MH-MCMC) with Metropolis-adjusted Langevin algorithm (MALA) proposals, enabling efficient gradient-informed sampling. Numerical experiments on a two-dimensional synthetic musculoskeletal phantom demonstrate that the posterior mean SOS reconstruction closely matches the reference anatomy, while additionally producing spatially resolved uncertainty maps. The posterior standard deviation is elevated primarily within bone and in localized regions near the transducer ring, coinciding with larger reconstruction errors and indicating reduced confidence in these areas. These results highlight the potential of MALA-based Bayesian FWI to provide uncertainty-aware SOS imaging for medical ultrasound.
Full waveform inversion (FWI) based on ultrasonic guided waves has been widely employed as a tomographic approach for nondestructive evaluation (NDE) of structural defects. In practical monitoring scenarios, however, noise and weak prior information can degrade reconstruction fidelity and drive the inversion toward local minima. To address these challenges, we propose an FWI framework that incorporates a dual-decoder network (DD-Net) as a deep image prior (DIP). Through its encoder-dual-decoder architecture, DD-Net provides implicit regularization via network structure and feature fusion, which helps alleviate cycle-skipping effects in time-domain inversion. By feeding the DD-Net output back as the initial model for subsequent FWI iterations, the proposed method yields more stable model updates and more reliable defect imaging. Numerical simulations and experiments show that, under weak priors, the proposed method improves the peak signal-to-noise ratio (PSNR) from 7.16 dB to 28.91 dB and the structural similarity index measure (SSIM) from 0.62 to 0.85. In addition, it tolerates initial-velocity deviations of up to +/- 16%, compared with approximately +/- 6% for conventional approaches. These results indicate that the proposed framework improves the robustness and practical applicability of ultrasonic guided-wave FWI for structural health monitoring (SHM) of safety-critical structures.
Accurate tissue classification of chronic prostatitis is critical for pathological research and clinical evaluation. Traditional diagnostic methods are either highly invasive or lack sufficient specificity, making non-invasive and precise quantitative characterization challenging. Photoacoustic imaging (PAI), which enables label-free and high-contrast detection of tissue biochemical properties, provides a novel approach for the non-invasive quantification of prostate inflammation. In this study, 11 inflamed and 9 normal rat prostate tissue sections were used to extract the power spectrum slope of photoacoustic signals at 36 wavelengths, and its efficacy for inflammation classification was evaluated. Results showed the most significant intergroup differences at 720 nm and $770 \mathrm{~nm}(P=0.005)$. The range from 650 to 890 nm contained numerous characteristic wavelengths that effectively distinguished the two tissue types, with negative slopes in the normal group and positive slopes in the inflamed group. This study clarifies the discriminative value of the power spectrum slope parameter, providing reliable quantitative evidence and technical support for the non-invasive detection of prostatitis, with important theoretical and practical significance.
Full waveform inversion(FWI)has showed great potential in the detection of musculoskeletal disease.However,FWI is an ill-posed inverse problem and has a high requirement on the initial model during the imaging process.An inaccurate initial model may lead to local minima in the inversion and unexpected imaging results caused by cycle-skipping phenomenon.Deep learning methods have been applied in musculoskeletal imaging,but need a large amount of data for training.Inspired by work related to generative adversarial networks with physical informed constrain,we proposed a method named as bone ultrasound imaging with physics informed generative adversarial network(BUIPIGAN)to achieve unsupervised multi-parameter imaging for musculoskeletal tissues,focusing on speed of sound(SOS)and density.In the in-silico experiments using a ring array transducer,conventional FWI methods and BUIPIGAN were employed for multi-parameter imaging of two musculoskeletal tissue models.The results were evaluated based on visual appearance,structural similarity index measure(SSIM),signal-to-noise ratio(SNR),and relative error(RE).For SOS imaging of the tibia-fibula model,the proposed BUIPIGAN achieved accurate SOS imaging with best performance.The specific quantitative metrics for SOS imaging were SSIM 0.9573,SNR 28.70 dB,and RE 5.78%.For the multi-parameter imaging of the tibia-fibula and human forearm,the BUIPIGAN successfully reconstructed SOS and density distributions with SSIM above 94%,SNR above 21 dB,and RE below 10%.The BUIPIGAN also showed robustness across various noise levels(i.e.,30 dB,10 dB).The results demonstrated that the proposed BUIPIGAN can achieve high-accuracy SOS and density imaging,proving its potential for applications in musculoskeletal ultrasound imaging.
To address the significant degradation of Space-Time Adaptive Processing (STAP) performance when the array elements have mutual coupling and gain/phase errors, a STAP algorithm with adaptive calibration for the above two array errors is proposed in this article. First, based on a defined error matrix that simultaneously considers both array mutual coupling and gain/phase errors, a STAP signal model including these errors is given. Then, utilizing the defined signal model, it is demonstrated that the estimation of the defined error matrix can be formulized as a standard convex optimization problem with the low-rank structure of the clutter covariance matrix and the subspace projection theory. Once the defined error matrix is estimated by solving the convex optimization problem, it is illustrated that a STAP method with adaptive calibration of the mutual coupling and gain/phase errors is coined. Analyses also show that the proposed adaptive calibration algorithm only needs one training sample to construct the adaptive weight vector. Therefore, it can achieve a good detection performance even with severe non-homogeneous clutter environments. Finally, the simulation experiments verify the effectiveness of the proposed algorithm and the correctness of the analytical results.
Owing to changes in the spatial position of Autonomous Aerial Vehicle (AAV) aerial images and limited platform resources, most existing AAV aerial image detection models have low accuracy, and it is difficult to achieve a good balance between detection performance and lightweight. To solve the above problems, an object detection model of AAV aerial images based on YOLOv8s, called BSDS-YOLOv8s, is proposed. In the proposed model the Dynamic Head (DyHead) is used to replace the detection head of YOLOv8s firstly, which can improve the spatial perception ability of the detection head of the model. Second, to improve the detection performance of DyHead and make the model lightweight, a new feature pyramid network (SDI-MBiFPN) in the Neck of YOLOv8s is proposed, which contains a Multiscale Bidirectional Feature Pyramid Network (BiFPN), called MBiFPN, and a feature fusion method based on the redesigned Semantic and Detail Fusion (SDI). Finally, a method that integrates Soft Non-Maximum Suppression (Soft-NMS) with Generalized Intersection over Union (GIoU), called GIoU-Soft-NMS, is proposed to enhance the model’s post-processing capability and reduce the missed and false detection rates, thereby further improving the detection accuracy of the model. The experimental results showed that the mAP0.5 and mAP0.5:0.95 of BSDS-YOLOv8s reached 49.6% and 34.1%, respectively, on the VisDrone2019 dataset. This is an improvement of 9.2 and 9.8 percentage points over YOLOv8s, respectively. The number of parameters and the floating point operations (FLOPs) were 8.03M and 26.6G, 27.8% and 7.6% less than those of YOLOv8s, respectively.
In order to address the challenging issue where model parameters of Thevenin equivalent circuit models of Li(thium)-ion batteries cannot be identified online with constant current, voltage and/or power charge/discharge, a new equivalent time-varying model is reported in this paper. Under various operating conditions, it is demonstrated that the open-circuit voltage (OCV) and internal resistance of Li-ion batteries can be respectively viewed as the time-varying ones with evolutions described by coulomb counting and the random walk model. Although the time-varying self-discharge transient state evolution of batteries cannot be represented by coulomb counting and the random walk model, it can be illustrated by observation noise with a long tail probability density function (pdf) of unknown parameters. In this way, the novel equivalent time-varying model described in terms of a state space equation is given. By a presented adaptive Bayesian learning method, both the parameters of the state space equation and model can be inferred online in sense of maximum a posterior probability. It also implies that Li-ion batteries can be robustly characterized/estimated by our model online. Both the datasets available in internet and our experiments with the Li-ion ternary and Li(thium)-FePO4 batteries verify the effectiveness of our model, analyses, and algorithm.
In order to enhance the array aperture and improve the quality of phased array Lamb wave imaging without adding physical array elements, an iterative adaptive Lamb wave imaging method using a virtual aperture extension array is proposed in this article. Analyses show that the virtual aperture extension of an array with a linear prediction method will make the rank of its sample covariance matrix (SCM) deficient and the noise between the physical and virtual elements correlated with one another. It is illustrated that the reconstruction of an interference plus noise covariance matrix (INCM) via the iterative adaptive approach (IAA) can be used to defeat such two problems without any array aperture loss. By implementing adaptive beamforming with weights derived from the reconstructed INCM, the proposed Lamb wave imaging method achieves enhanced damage localization accuracy and image resolution without requiring additional physical transducers. Numerical simulation and imaging experiments on aluminum plate damage validate the feasibility of the proposed method.
Multimodal ultrasonic guided wave (UGW) signal reconstruction technology can accurately separate individual modes, providing more comprehensive and precise information for material nondestructive testing. However, the accuracy of existing reconstruction techniques heavily depends on the precision and completeness of time-frequency (TF) ridge extraction. To address this challenge, this paper proposes a TF energy segmentation reconstruction method without relying on complete TF ridge extraction, as traditionally required. This approach introduces an adaptive noise variance estimation Bayesian filter to extract the TF ridges under unknown noise distribution, particularly in regions where TF ridges intersect or overlap. By using the extracted TF ridges as references, the energy segmentation method directly separates and reconstructs UGW modes from the TF representation even when the extracted TF ridges are incomplete. This is because the proposed method can automatically retrieve the energy of each mode with a region growing algorithm from the time domain and frequency domain so that both modes with rapidly changing instantaneous frequency or group delay can be recovered, while the traditional method can only separate modes from a single domain. Numerical simulations and photoacoustic-guided wave experiments validate the effectiveness of the proposed method, achieving reconstruction accuracies of 96.9% and 92.5% for the simulated and experimental signals, respectively.
In Lamb wave imaging based on a phased array, higher frequencies narrowband excitation pulses enable more precise damage detection and localization. However, due to the size constraints of individual transducer elements, the spacing between array elements may exceed half the wavelength of the excitation signal. This can lead to a grating lobe effect. To overcome this limitation, a Lamb wave imaging method via dual-frequency fusion for grating lobe effect compensation is proposed in this study. Analyses indicate that the grating lobe effect may introduce artifacts or distortions in the imaging results. This method utilizes two frequencies of narrowband excitation pulses for imaging and subsequently fuses the results. By doing so, the imaging artifacts caused by the grating lobes produced by high-frequency narrowband excitation pulses are effectively compensated. The proposed method is validated through simulations and experiments on an aluminum plate, showing superior accuracy, contrast, and imaging quality.
Recent advances in machine vision have played an important role in addressing the challenging problem of motion blur. However, most deep learning–based deblurring methods operate in the RGB domain, rely on recursive strategies, and are often trained on unrealistic synthetic data. In this paper, we introduce a preventive solution from a new perspective, leveraging the opportunity to operate directly in the RAW domain on high-bit sensor data. Since no publicly available high–frame rate RAW-based blur prevention dataset exists, we construct Blurry-RAW, a novel dataset containing paired blurry and sharp frames in both RAW and RGB formats. We further propose 3D-ISPNet, a CNN–Transformer hybrid architecture, trained exclusively on RAW sensor data. This model achieves superior quantitative and qualitative performance compared to RGB-based counterparts. Moreover, by fine-tuning on data from different camera sensors, 3D-ISPNet demonstrates strong generalization across diverse hardware. Ultimately, the introduction of RAW-driven blur prevention and the new dataset paves the way for further research in this emerging direction.
The focus on integrated sensing and communication (ISAC) is growing with explosive growth in the low-altitude economy. However, research on multi-site BS collaboration ISAC, which is vital for future UAV positioning and surveillance, is scarce. Therefore, this paper proposes a dynamic sub-graph (DSG) to analyze the performance and optimize resource allocation of the collaborative ISAC system. It constructs a DSG model that adjusts BS relationships based on resource schemes. The goal of resource allocation is to minimize system costs while meeting key space (KS) coverage and interference constraints. Simulations validate the approach's effectiveness in optimizing resource allocation for such systems.
BACKGROUND AND OBJECTIVE:The full waveform inversion (FWI) method plays a significant role in bone quantitative imaging. It is shown that even a small deviation in transducer positions can lead to a considerable variation in frequency-domain signals, and result in a marked decline in the performance of frequency-domain full waveform inversion (FDFWI). To address this limitation, a multi-parameter time-domain full waveform inversion algorithm based on a recurrent neural network (RNN-MPTDFWI) is proposed for bone quantitative imaging. METHODS:In the proposed method, a variable-density acoustic wave equation which takes sound velocity and bone density as parameters is solved in the time domain within RNN cells as the forward model. Multiscale inversion is conducted with filtered signals iteratively from low to high frequency bands via minimizing a misfit function between the simulated and observed data. Automatic differentiation for multi-parameter gradient calculation and adaptive momentum estimation (Adam) algorithm are utilized in the optimization process. After the estimated velocity and density are obtained, bone images are generated based on these parameters. RESULTS:Numerical simulation results demonstrate that RNN-MPTDFWI reduces the mean relative errors (MREs) of the reconstructed velocity and density by at least 54.56% and 71.64%, respectively, compared to FDFWI. CONCLUSIONS:These improvements highlight that RNN-MPTDFWI offers more accurate representations of bone geometry and microarchitecture, along with enhanced robustness against transducer position errors.
Background Voice disorders are common otolaryngological conditions that significantly impair patients’ communication ability and quality of life. Most existing studies rely on two-dimensional Mel-spectrograms, which, while practical, have inherent limitations in feature extraction. This study demonstrates that three-dimensional Mel-spectrograms can capture more comprehensive pathological voice features, thereby enabling more accurate diagnosis. Methods This study was conducted using the voice database from Fudan University Eye, Ear, Nose, and Throat Hospital (1,839 cases covering six categories). Focusing on six voice conditions—healthy voice (NV), spasmodic dysphonia (SD), uncompensated unilateral vocal fold paralysis (UVFP), vocal cord sulcus (VFS), benign proliferative lesions (BPLVF), and malignant vocal fold tumors (MVFT)—we propose a classification method termed Voice-3D, which integrates three-dimensional Mel-spectrograms (3D-Mel) with a deep learning framework. By mapping voice signals into a three-dimensional “time-frequency-energy” space and incorporating a multi-view convolutional feature fusion structure, the model comprehensively captures pathological voice characteristics. Results Compared with 2D-Mel, 3D-Mel demonstrated superior overall classification performance; under identical data partitions and training settings, the overall accuracy increased from 68.0% to 81.1%. The most notable improvements were observed in malignant vocal fold tumors (from 57.2% to 84.3%) and vocal cord sulcus (from 86.7% to 93.5%), while normal voices also showed a modest but statistically significant gain (from 89.8% to 93.6%). Confusion matrix analysis further revealed that 3D-Mel substantially reduced cross-class misclassifications, particularly in clinically challenging categories. It is worth noting that the overall accuracy for spasmodic dysphonia slightly decreased, although its F1-score improved, indicating a better balance between precision and recall. Conclusions The deep learning framework based on 3D Mel-spectrograms substantially outperforms traditional 2D methods in multiclass classification of voice disorders, enabling noninvasive, objective, and automated auxiliary diagnosis. This method holds promise for clinical decision support and remote voice health screening. Future work will focus on large-scale multicenter validation and the exploration of multimodal fusion with laryngoscopic images and clinical records to enhance generalizability and clinical applicability.
Addressing the shortcomings of the Sparrow Search Algorithm (SSA), such as low accuracy of convergence and tendency of falling into local optimum, a Multi-strategy Integrated Sparrow Search Algorithm (MISSA) is proposed. In this method, by improving the black-winged kite algorithm and applying it to the producer’s position update formula, an improved search strategy (ISS) is firstly proposed to enhance search ability. Secondly, a new strategy inspired by the Coot algorithm, called the group follow strategy (GFS), is proposed to improve the ability to jump out of the local optimum. Finally, a proposed random opposition-based learning strategy (ROBLS) is applied to the population after each iteration to enhance its diversity. To verify MISSA’s effectiveness, extensive testing is conducted on 24 benchmark functions as well as CEC 2017 functions. The experimental results, complemented by Wilcoxon rank-sum tests, conclusively demonstrate that MISSA outperforms SSA and other advanced optimization algorithms, exhibiting superior overall performance.
The Lamb wave imaging approach based on adaptive beamforming method, notably the minimum variance distortionless response beamforming algorithm, is analyzed to explain its serious performance degradation or even complete invalidation, when the scattering signals generated by damages are coherent. It is demonstrated that in such case, the covariance matrix of the received scattering signals becomes rank-deficient, and when decomposed into the description form of the array manifold matrix left and conjugate right multiplied by a matrix P, P is a non-diagonal matrix with off-diagonal elements representing the coherences among scattering signals. These factors result in the failure of the conventional adaptive beamforming algorithm. To break the above limitations, a covariance matrix fitting method using a weighted least squares criterion is proposed by constructing P as a diagonal matrix. Based on the fitted covariance matrix, a robust adaptive beamforming Lamb wave imaging approach for coherent scattering signals is presented. Extensive numerical simulations and imaging experiments conducted on an aluminum plate corroborate the effectiveness of the proposed algorithm. The results affirm that the proposed method preserves the performance levels achieved by non-coherent scattering signals Lamb wave imaging approaches.
BACKGROUND AND OBJECTIVE:It is a challenging task to use ultrasound for bone imaging, as the bone tissue has a complex structure with high acoustic impedance and speed-of-sound (SOS). Recently, full waveform inversion (FWI) has shown promising imaging for musculoskeletal tissues. However, the FWI showed a limited ability and tended to produce artifacts in bone imaging because the inversion process would be more easily trapped in local minimum for bone tissue with a large discrepancy in SOS distribution between bony and soft tissues. In addition, the application of FWI required a high computational burden and relatively long iterations. The objective of this study was to achieve high-resolution ultrasonic imaging of bone using a deep learning-based FWI approach. METHOD:In this paper, we proposed a novel network named CEDD-Unet. The CEDD-Unet adopts a Dual-Decoder architecture, with the first decoder tasked with reconstructing the SOS model, and the second decoder tasked with finding the main boundaries between bony and soft tissues. To effectively capture multi-scale spatial-temporal features from ultrasound radio frequency (RF) signals, we integrated a Convolutional LSTM (ConvLSTM) module. Additionally, an Efficient Multi-scale Attention (EMA) module was incorporated into the encoder to enhance feature representation and improve reconstruction accuracy. RESULTS:Using the ultrasonic imaging modality with a ring array transducer, the performance of CEDD-Unet was tested on the SOS model datasets from human bones (noted as Dataset1) and mouse bones (noted as Dataset2), and compared with three classic reconstruction architectures (Unet, Unet++, and Att-Unet), four state-of-the-art architecture (InversionNet, DD-Net, UPFWI, and DEFE-Unet). Experiments showed that CEDD-Unet outperforms all competing methods, achieving the lowest MAE of 23.30 on Dataset1 and 25.29 on Dataset2, the highest SSIM of 0.9702 on Dataset1 and 0.9550 on Dataset2, and the highest PSNR of 30.60 dB on Dataset1 and 32.87 dB on Dataset2. Our method demonstrated superior reconstruction quality, with clearer bone boundaries, reduced artifacts, and improved consistency with ground truth. Moreover, CEDD-Unet surpasses traditional FWI by producing sharper skeletal SOS reconstructions, reducing computational cost, and eliminating the reliance for an initial model. Ablation studies further confirm the effectiveness of each network component. CONCLUSION:The results suggest that CEDD-Unet is a promising deep learning-based FWI method for high-resolution bone imaging, with the potential to reconstruct accurate and sharp-edged skeletal SOS models.
Weiqi Wang (王威琪)合作论文数Institute of Biomedical Engineering and Technology, Fudan University4