Objective emotion recognition is significantly important for some fields such as physiological health, healthcare and education. From the perspectives of signal acquisition difficulty, cost and user acceptance, the electrocardiogram (ECG) signals are appropriate biomarkers for achieving objective emotion recognition. As it is difficult for deep-learning based methods to extract and fuse the spatio-temporal features of one-dimensional ECG signals, a 1D-2D signal transformation method based on wavelet packet decomposition was proposed firstly, which converted one-dimensional ECG signals to "two-dimensional images". Subsequently, the ResNet18 was used as the backbone network and the "two-dimensional images" were used as its input, where a Fusion Block module was designed to improve the network's spatio-temporal feature extraction and fusion capabilities. Finally, extensive experiments were implemented for the emotion recognition task on the WESAD and SWELL-KW datasets. The experimental results demonstrated that in comparison with the suboptimal method, the emotion recognition method proposed in this paper improved the performance in both metrics of average accuracy and F1 scores by 2.19 and 4.48 percentage points respectively, which may provide the technical support for objective emotion recognition.
Accurate detection of surgical instruments is critical for both routine surgical procedures and surgical robotics research. To the best of our knowledge, there is a notable lack of datasets and dedicated detection studies specifically addressing orthopedic surgical instruments. Detecting orthopedic surgical instruments presents particular challenges including significant size variations, highly similar shapes, and frequent, severe occlusions due to instrument intersections. To address these issues, we propose an orthopedic surgical instrument detection method (OrthoDetNet) incorporating three specialized modules. The FilterUnit mitigates occlusion effects via an adaptive feature filtering mechanism, that dynamically adjusts its filtering strategy based on context, prioritizing features from key regions while suppressing distracting interference features. The DEUnit enhances fine-grained feature discrimination in local regions to distinguish instruments with high shape similarity, and the BDFusion module improves multi-scale detection performance through bi-directional feature fusion between deep and shallow-level feature maps. A dataset for orthopedic surgical instrument detection is created, which is based on the proximal femoral nail antirotation (PFNA) instrument package manufactured by Shenzhen Mindray Bio-Medical Electronics Co., Ltd. Images were captured in a controlled, simulated experimental environment, ensuring no patient privacy or ethical concerns. We obtained explicit authorization from the manufacturer for instrument use. Experimental results on this dataset demonstrate the effectiveness of the OrthoDetNet and its constituent modules.
Calibration is an initial but crucial step in reliable light-field 3D reconstruction researches. Traditional calibration methods typically require capturing multiple images with the calibration board at different positions and poses, which inevitably makes the calibration time-consuming. To address this issue, this paper proposes a one-shot calibration method that only requires capturing a single image of a well-designed 3D calibration target composed of three mutually orthogonal planes with checkerboard patterns. In this method, based on the vanishing point principle, we first deduce a positive definite linear system to solve the intrinsic parameters, then adopt the method previously proposed by us to solve the extrinsic parameters. Furthermore, since the re-projection error of traditional calibration methods can only ensure the position consistency between feature points, to improve calibration performance, we propose an optimization objective function that can ensure both position consistency and angular consistency based on the inherent relationships between the three planes of the 3D calibration target, then we propose an ADMM-based joint optimization algorithm to address the non-smooth optimization problem introduced by the new objective function. Extensive experiments show that the typical RMS of re-projection error and angular error are about 0.004 mm and 2 degrees, respectively, which demonstrate the one-shot light field calibration method is effective and reliable.
Within the brachial plexus, the C6&C7 nerves play crucial roles in peripheral nerve block. The precise identification of nerve locations ultrasound images is critical for drug injections in anesthesia. However, their segmentation remains challenging due to limited annotated datasets, small target size in noisy images, and underutilized anatomical context. In this paper, we proposed LA-BPNet for the segmentation of the brachial plexus C6&C7 nerves in ultrasound images, which is able to use the anatomy information as the indicator of their segmentation task. We first introduce the Hybrid Spatial Perception (HSP) Module to solve the small target (nerves) segmentation problem in the low signal-to-noise ultrasound images, which consists of an Axial Spatial Attention (ASA) module and a Multi-scale Asymmetric Local Spatial Attention (MA-LSA) Module. To use the anatomical priors of nerves in the segmentation task adequately, we present the Anatomical Information Guidance (AIG) Module consisting of two parallel task branches to supervise the nerve centers and the coarse segmentation, additionally, the Interactive Guidance Mechanism (IGM) is integrated the information interaction in these two branches. In this paper, a private dataset is constructed, which contains 863 ultrasound images of the brachial plexus C6&C7 nerves. Extensive experiments with state-of-the-art (SOTA) models on our private dataset demonstrates that the LA-BPNet achieves better performance. The code will be available at https://github.com/zhankn/LA-BPNet.
Hand-held light field (LF) cameras often exhibit low spatial resolution due to the inherent trade-off between spatial and angular dimensions. Existing supervised learning-based LF spatial super-resolution (SR) methods, which rely on pre-defined image degradation models, struggle to overcome the domain gap between the training phase-where LFs with natural resolution are used as ground truth-and the inference phase, which aims to reconstruct higher-resolution LFs, especially when applied to real-world data. To address this challenge, this paper introduces a novel self-supervised learning-based method for LF spatial SR, which can produce higher spatial resolution LF images than originally captured ones without pre-defined image degradation models. The self-supervised method incorporates a hybrid LF imaging prototype, a real-world hybrid LF dataset, and a self-supervised LF spatial SR framework. The prototype makes reference image pairs between low-resolution central-view sub-aperture images and high-resolution (HR) images. The self-supervised framework consists of a well-designed LF spatial SR network with hybrid input, a central-view synthesis network with an HR-aware loss that enables side-view sub-aperture images to learn high-frequency information from the only HR central view reference image, and a backward degradation network with an epipolar-plane image gradient loss to preserve LF parallax structures. Extensive experiments on both simulated and real-world datasets demonstrate the significant superiority of our approach over state-of-the-art ones in reconstructing higher spatial resolution LF images without pre-defined degradation.
Surgical workflow recognition (SWR) stands as a pivotal component in computer-assisted surgery and is dedicated to identifying phases from surgical videos. Many deep learning-based methods have been proposed for this task and achieved acceptable SWR results. However, these methods usually implicitly extract and aggregate spatio-temporal features, so that it is challenging for these methods to adequately use some spatial information that is strongly relevant to surgical phase in SWR task, such as the information from the surgical instruments. To address this issue, an Explicit Surgical Instrument Prompting (ESIP) approach is proposed for SWR task. ESIP leverages surgical instrument segmentation to generate instrument-specific visual prompts, which explicitly guide the extraction of crucial intra-frame spatial features through a frozen pre-trained backbone, then enable effective inter-frame spatio-temporal feature extraction and aggregation. Unlike multi-task approaches that jointly perform SWR with auxiliary tasks within a shared network framework, ESIP is a single-task SWR approach dedicated to optimize framework itself for more adequate feature extraction. Furthermore, to accomplish the segmentation prompting efficiently, this paper presents SAM-based segmentation with prompt tuning strategy to explicitly integrate segmentation features into spatial features. Experimental results on Cholec80, M2CAI and AutoLaparo datasets demonstrate that our ESIP method achieves the best performance in comparison with 16 SOTA methods, with a Precision of 91.8%, 89.5% and 89.6%, Recall of 92.2%, 89.5% and 76.9%, Jaccard of 83.3%, 77.0% and 67.3%, respectively.
Lung ultrasound (LUS), valued for its portability and AI-driven analysis, has become an essential di agnostic tool in emergency medicine, critical care, and the screening of infectious diseases. However, the scarcity of annotated data and limited expert availability remain major barriers to large-scale deployment. In this study, we pro pose a Similarity-Guided Multi-Source Few-Shot Learning (SM-FSL) paradigm that shifts the focus from domain alignment to domain selection, enabling robust feature learning from multiple semantically relevant domains under limited data conditions. Based on this paradigm, we develop the Global-Local Multi-Source Domain Network (GLMD-Net) for LUS analysis in data-scarce scenarios, which first employs the proposed Average Nearest Neighbor Set Distance (ANNSD) to prioritize relevant source domains for pre-training. A similarity-guided optimization mechanism then harmonizes multi-source gradient updates, while a local feature enhancement module and a self-supervised auxiliary task improve robustness against noise and artifacts. Finally, a few-shot adaptation strategy fine-tunes only the classification head for efficient and stable knowledge transfer. Extensive experiments on two public COVID-19 ultrasound datasets demonstrate that our method surpasses state-of-the-art approaches, validating its effectiveness in enhancing generalization and robustness under few-shot settings.
Semi-supervised medical image segmentation (SSMIS) methods predominantly rely on consistency regularization to reinforce invariant feature learning under perturbations. However, the reliance on uniform perturbation strategies makes SSMIS models susceptible to overconfident pseudo-labeling, resulting in confirmation bias issues. Moreover, the inefficient knowledge propagation from labeled to unlabeled data exacerbates this bias on pseudo-labeling, restricting the model’s performance and generalization capability. In this study, we propose a novel channel-spatial hierarchical adversarial perturbation scheme to employ a diverse-driven perturbation learning and coupling learning of labeled and unlabeled data for SSMIS. Specifically, we present a discrepancy-aware spatial adversarial perturbation strategy, injecting adversarial noise into uncertain regions identified by high divergence between two decoders and enforcing a cross-consistency constraint on their outputs under perturbations. This mechanism encourages the model to confront latent ambiguities and refine its decision boundary through adversarial learning. Moreover, an innovative gradient-guided channel perturbation mechanism is designed to selectively inhibit channels exhibiting low deviation by quantifying the similarity between supervised and unsupervised loss gradients, facilitating bridging the learning from labeled to unlabeled data to enhance the learning efficiency of hard samples. This channel-spatial dual-targeted perturbation collaboratively reinforces the perturbation diversity for consistency regularization, enforcing the discriminative features learning to achieve unbiased prediction on unannotated data. We have extensively evaluated our method with state-of-the-art semi-supervised approaches on three widely recognized SSMIS benchmarks. The experimental results, obtained across various labeled data ratios, demonstrate the superiority of our proposed method over existing techniques, suggesting its effectiveness in SSMIS. Our code is available at https://github.com/gardnerzhou/CHAP.
With rapid developments in light field imaging, a great deal of attention has been given to its applications in industrial, medical and other fields due to its ability to perform three-dimensional reconstruction in single-shot. In these applications, Light field Laparoscope (LFL) is an important one, but it often suffers severe micro-lens image deformations that lead to incorrect LFL decoding, calibration and three-dimensional reconstruction. Based on the micro-lens image non-deformation constraint presented by us before, we propose the flexible aperture- angular plane to analyze the LFL imaging model, the modified microlens image non-deformation constraint for 3D LFL system and an advanced two-step calibration method to compute 3D LFL imaging parameters. Moreover, a 3D LFL imaging prototype is designed and calibrated. Experimental results show that microlens image deformations are avoided in this 3D LFL prototype, and the typical RMS re-projection error is about 0.06 pixels.
As a fundamental aspect of bone fracture treatment, fracture reduction plays a decisive role in restoring the structural integrity and function of bones. At present, fracture reduction techniques mostly rely on semi-automatic interaction methods or healthy-side bone templates for registration, which have many limitations in clinical practice. In order to enhance treatment efficiency and accuracy, an automatic fracture reduction algorithm is proposed. This algorithm utilizes the similarity of fracture cross-sections for registration, thereby reducing the workload of physicians and eliminating the need for a healthy-side bone template. Initially, the closed edge is identified and extracted by analyzing the differences in the fracture surface and the calorific value diagram of the roughness distribution. Next, the fracture section is determined by using the identified closed edge as a guideline for regional expansion and similarity matching. During the registration phase, the iterative closest point (ICP) algorithm is highly sensitive to distance. Therefore, the geometric features of point clouds are incorporated into the objective function of the registration algorithm to mitigate the influence of noise, and fracture section registration is implemented one by one. Finally, the algorithm is tested and compared on 180 simulated datasets and 16 publicly available datasets. The results show that the proposed algorithm significantly improves the registration accuracy, and the registration error of clinical bone fracture cases is controlled within 1.7 mm.
One of the core challenges in ultrasound image quality assessment (IQA) is the entanglement of semantic content and quality-related information, such as blurring and shadows. Insufficient attention to the latter can easily lead to biased IQA results. Furthermore, fine-grained quality inconsistencies, i.e., subtle variations in ultrasound images that can impact quality interpretations, may further complicate the IQA tasks. To address these challenges, we propose a novel degradation-aware model (DAM) for the ultrasound IQA, which effectively perceives various and subtle variations of quality patterns, accurately assessing the quality of ultrasound images. The advanced degradation-derived augmentation (DDA) in DAM incorporates degradations that clinicians may focus on during IQA into the synthesis of appearance changes, promoting the disentanglement of quality-related representations from semantic contents. Subsequently, we present fine-grained degradation learning (FGDL), which encourages distinctions between image versions with diminishing quality inconsistencies, boosting the awareness of quality nuances from easy to hard for better ultrasound IQA performance. A universal boundary acquisition operator (UBAO) is also developed to suppress interferences from redundant information, achieving the standardization of ultrasound images from various devices. Extensive experimental results on an in-house ultrasound dataset demonstrate that DAM outperforms 14 baseline methods, achieving a PLCC of 0.760 and an SROCC of 0.766. The code can be available at this URL.
Ultrasound imaging has emerged as an effective tool for aiding diagnosis. The automatic segmentation of ultrasound images is crucial in identifying the lesion target and evaluating clinical indicators for accurate diagnosis and prognosis. However, the segmentation problems are challenging due to the inherent speckle noise interference and low contrast of ultrasound images. The complex-value-based neural network can directly deal with the phase components, offering a potential solution in a better-perceiving structure for ultrasound image segmentation. In this study, we develop a Frequency Phase-Guided Attention Network (FPGANet) for ultrasound image segmentation by exploring the properties of the complex-valued model under the guide of phase and frequency perspectives. First, our proposed method transforms images into a complex domain as the input to an advanced complex-value model consisting of pure complex-value convolutions and operations. Especially this model can then effectively scrutinize phase information to distinguish target areas from similar backgrounds better. Moreover, we introduce a complex hybrid attention module following complex convolution to selectively adjust the perception of phase components and the model's bias. Also, we designed a frequency-adaptive separation module to emphasize frequency features prioritized by the encoder and decoder using a combination of wavelet decomposition and frequency channel attention. We evaluate the proposed FPGANet on three publicly available ultrasound datasets of breast, cardiac and thyroid nodules and a private abdominal effusion ultrasound dataset. Comparative experiments were also conducted with state-of-the-art methods. The results demonstrate the superior performance of FPGANet, implying its potential for advancing ultrasound image segmentation.
Colonoscopy is an important technical means for screening early colorectal cancer lesions. Accurate segmentation of intestinal polyps helps improve the accuracy of screening. Early screening for lesions is of great significance for the prevention of colorectal cancer, and the segmentation of intestinal polyps is an important research direction. Although intestinal polyp segmentation based on deep learning has achieved acceptable performance, the color variation among intestinal endoscopic images significantly affects it. Based on the ResNet architecture, this study proposes an advanced PE-ResNet in which histogram equalization is used to reduce color influence. Experimental results on five datasets, including ClinicDB, demonstrate that the PE-ResNet model achieves improved performance in intestinal polyp segmentation.
Constructing the health indicator (HI) and predicting the remaining useful life (RUL) are essential steps in bearing health management. Some prediction methods depend on prior information about HIs, especially when these indicators are generated by deep learning models. However, acquiring such prior information can be challenging in practical applications. This paper introduces a novel unsupervised adaptive density-based clustering filter (UADCF) for RUL prediction of bearings, which operates without the need for prior knowledge. Firstly, a post-hoc interpretation HI model (PIHIM) is proposed to characterize the deep learning constructed HIs from the perspective of what the deep learning has done. Then, leveraging the classical density-based clustering algorithm, we introduce the UADCF for unsupervised estimation of model parameters, which can dynamically adjust density parameters based on the current conditions. Finally, we develop a prediction framework combining PIHIM and UADCF, enabling unsupervised RUL prediction of bearings. The experimental studies validate the effectiveness of the proposed method.
Depth estimation is a fundamental problem in light field processing. Epipolar-plane image (EPI)-based methods often encounter challenges such as low accuracy in slope computation due to discretization errors and limited angular resolution. Besides, existing methods perform well in most regions but struggle to produce sharp edges in occluded regions and resolve ambiguities in texture-less regions. To address these issues, we propose the concept of stitched-EPI (SEPI) to enhance slope computation. SEPI achieves this by shifting and concatenating lines from different EPIs that correspond to the same 3D point. Moreover, we introduce the half-SEPI algorithm, which focuses exclusively on the non-occluded portion of lines to handle occlusion. Additionally, we present a depth propagation strategy aimed at improving depth estimation in texture-less regions. This strategy involves determining the depth of such regions by progressing from the edges towards the interior, prioritizing accurate regions over coarse regions. Through extensive experimental evaluations and ablation studies, we validate the effectiveness of our proposed method. The results demonstrate its superior ability to generate more accurate and robust depth maps across all regions compared to state-of-the-art methods. The source code will be publicly available at https://github.com/PingZhou-LF/Light-Field-Depth-Estimation-Based-on-Stitched-EPIs.
Coherent plane-wave compounding (CPWC) is a commonly applied beamforming technique for realizing ultrafast ultrasound imaging, though it often involves a trade-off in image quality. To address this limitation, the minimum variance distortionless response (MVDR) is advanced for the adaptive combination of ultrasound signals obtained from different directions. This process effectively suppresses the artifacts and improves the imaging quality. While previous studies have proved that MVDR can be approximated to a minimum mean square error (MMSE) algorithm after combining it with a Wiener post-filter (WPF), a systematic survey of optimized MVDR beamformers utilizing different estimation of covariance matrices and corresponding post-filter has not been undertaken. In this study, we scrutinize the impact on the image resolution and contrast ratios (CRs) by varying beamforming combination strategies related to MVDR beamformers for CPWC imaging. Moreover, we propose a novel approach to effectively integrate the multidimensional signals from both the transmission and reception stages, improving overall image quality for MVDR beamformers. Our comprehensive evaluation, including the simulation, phantom, and in vivo experiments suggests that the spatial-smoothing MVDR in the reception and transmission stages have inherent strengths in lateral and axial resolution, respectively. Also, WPF further benefits the improvement of the image contrast and lateral resolution. By comparison, the results of our proposed approach offer valuable insights into the future of research in optimized ultrafast ultrasound imaging related to MVDR and coherence-based factors, paving a better way to combine the correlated signals in the transmission and reception stages, applicable not only to CPWC but also other two-way focusing systems.
Stiffer cages provide sufficient mechanical support but fail to promote bone ingrowth due to stress shielding. It remains challenging for fusion cage to satisfy both bone bridging and mechanical stability. Here we designed a fusion cage based on twist metamaterial for improved bone ingrowth, and proved its superiority to the conventional diagonal-based cage in silico. The fusion process was numerically reproduced via an injury-induced osteogenesis model and the mechano-driven bone remodeling algorithm, and the outcomes fusion effects were evaluated by the morphological features of the newly-formed bone and the biomechanical behaviors of the bone-cage composite. The twist-based cages exhibited oriented bone formation in the depth direction, in comparison to the diagonal-based cages. The axial stiffness of the bone-cage composites with twist-based cages was notably higher than that with diagonal-based cages; meanwhile, the ranges of motion of the twist-based fusion segment were lower. It was concluded that the twist metamaterial cages led to oriented bone ingrowth, superior mechanical stability of the bone-cage composite, and less detrimental impacts on the adjacent bones. More generally, metamaterials with a tunable displacement mode of struts might provide more design freedom in implant designs to offer customized mechanical stimulus for osseointegration.
Light field cameras have a wide range of uses due to their ability to simultaneously record light intensity and direction. The angular resolution of light fields is important for downstream tasks such as depth estimation, yet is often difficult to improve due to hardware limitations. Conventional methods tend to perform poorly against the challenge of large disparity in sparse light fields, while general CNNs have difficulty extracting spatial and angular features coupled together in 4D light fields. The light field disentangling mechanism transforms the 4D light field into 2D image format, which is more favorable for CNN for feature extraction. In this paper, we propose a Deep Disentangling Mechanism, which inherits the principle of the light field disentangling mechanism and further develops the design of the feature extractor and adds advanced network structure. We design a light-field reconstruction network (i.e., DDASR) on the basis of the Deep Disentangling Mechanism, and achieve SOTA performance in the experiments. In addition, we design a Block Traversal Angular Super-Resolution Strategy for the practical application of depth estimation enhancement where the input views is often higher than 2x2 in the experiments resulting in a high memory usage, which can reduce the memory usage while having a better reconstruction performance.
Porous cages with lower global stiffness induce more bone ingrowth and enhance bone-implant anchorage. However, it's dangerous for spinal fusion cages, which usually act as stabilizers, to sacrifice global stiffness for bone ingrowth. Intentional design on internal mechanical environment might be a promising approach to promote osseointegration without undermining global stiffness excessively. In this study, three porous cages with different architectures were designed to provide distinct internal mechanical environments for bone remodeling during spinal fusion process. A design space optimization-topology optimization based algorithm was utilized to numerically reproduce the mechano-driven bone ingrowth process under three daily load cases, and the fusion outcomes were analyzed in terms of bone morphological parameters and bone-cage stability. Simulation results show that the uniform cage with higher compliance induces deeper bone ingrowth than the optimized graded cage. Whereas, the optimized graded cage with the lowest compliance exhibits the lowest stress at the bone-cage interface and better mechanical stability. Combining the advantages of both, the strain-enhanced cage with locally weakened struts offers extra mechanical stimulus while keeping relatively low compliance, leading to more bone formation and the best mechanical stability. Thus, the internal mechanical environment can be well-designed via tailoring architectures to promote bone ingrowth and achieve a long-term bone-scaffold stability.