Ex situ conservation in breading centers is one of the key strategies for saving giant pandas (Ailuropoda melanoleuca). Abnormal behaviors (e.g., inappetence) are key symptoms of potential health issues (e.g., Klebsiella pneumoniae) for the captives. Therefore, monitoring their normal activity patterns could set a baseline to detect these abnormalities for implementing timely interventions. However, traditional monitoring methods are labor-intensive, which often rely on manual observations. Here, we proposed a deep learning framework, termed as DeepPanda, for automatically recognizing four essential behaviors (i.e., eating, walking, resting and drinking) of giant pandas based on videos from common surveillance cameras. Experimental results demonstrated that the DeepPanda model achieved high performance on the self-established APanda dataset, with the testing mean average precision at an IoU threshold of 0.5 (mAP@0.5) of 98.8%. This methodology provides a powerful tool for monitoring the captive giant panda’s behaviors.
Glass-insulated terminals (GITs) are widely used in high-reliability microelectronic systems, where glass fall-offs in the sealing region may seriously degrade the reliability of the microelectronic component and further degrade the device reliability. Automatic inspection of such defects is challenging due to strong light reflection, irregular defect appearances, and limited defective samples. To address these issues, a coarse-to-fine machine-learning framework is proposed for glass fall-off detection in GIT images. By exploiting the circular-ring geometric prior of GITs, an adaptive sector partition scheme is introduced to divide the region of interest into sectors. Four categories of sector features, including color statistics, gray-level variations, reflective properties, and gradient distributions, are designed for coarse classification using a gradient boosting decision tree (GBDT). Furthermore, a sector neighbor (SN) feature vector is constructed from adjacent sectors to enhance fine classification. Experiments on real industrial GIT images show that the proposed method outperforms several representative inspection approaches, achieving an average IoU of 96.85%, an F1-score of 0.984, a pixel-level false alarm rate of 0.55%, and a pixel-level missed alarm rate of 35.62% at a practical inspection speed of 32.18 s per image.
Accurate assessment of axillary lymph node metastatic burden (ALNMB) is crucial for breast cancer staging and treatment planning. Current deep learning approaches primarily focus on single lymph nodes for ALNMB assessment, failing to capture the inter-node correlations and holistic medical hints required to evaluate the patient-level metastatic burden. To address this limitation, we propose a Multi-branch Dual-modal Fusion Transformer (MDF-former), which is an advanced multi-node assessment framework by simultaneously integrating B-mode ultrasound (BUS) and shear wave elastography (SWE) images of the three most suspicious ALNs. This multi-branch architecture mimics the comprehensive clinical diagnostic workflow to effectively capture spatial heterogeneity and inter-node correlations, thereby delivering a truly holistic and patient-level assessment (non-metastatic, low-burden, or high-burden). Specifically, the MDF-former employs a Conditional Key Dual-Modal Attention Fusion Module (CK-DMAFM) to align dual-modal features based on clinical logics. Meanwhile, a Geometric Consistency-based Balancing Loss (GCBLoss) is introduced to maintain branch equilibrium. Subsequently, the Multi-branch Soft Voting Aggregation (MSVA) strategy aggregates individual nodal probabilities to produce a patient-level diagnostic consensus. Through comprehensive ablation studies and comparative evaluations against state-of-the-art methods, the MDF-former demonstrated superior diagnostic performance on the independent test set. Specifically, the model achieved a higher AUC of 0.9629, an accuracy of 0.9268, a sensitivity of 0.8712, and a specificity of 0.9485, significantly outperforming all baseline models. We further demonstrated the clinical interpretability of MDF-former through heatmap visualization.
The inspection of voids in Metal-Oxide-Semiconductor Field-Effect Transistors (MOSFETs) is crucial for enhancing product quality and optimizing manufacturing processes. In this paper, an integrated adaptive framework is proposed to inspect internal voids in MOSFETs. First, a position-aware vertex location scheme is proposed to acquire the transistor region, which can filter out four MOSFET’s vertices based on the grayscale characteristics of their neighborhoods. To address the grayscale unevenness in the acquired transistor region, the region is partitioned into different patches. Then, a coarse-to-fine void identification scheme is proposed to identify whether the patches involve voids with different sizes, which is based on KL divergence and a high-order gray-level difference. Next, an adaptive thresholding scheme is proposed to localize the voids at the pixel level, which is based on the interval cumulative frequency strategy. Finally, the void ratio of a MOSFET is calculated by the combination of the areas of voids and the transistor. Experimental results indicate that the proposed integrated adaptive framework outperforms some existing inspection methods at a reasonable inspection speed of 1.152s/image, with the inspection performance of 0.9213 Dice coefficient, 0.8597 mIoU, 98.62% Accuracy, 98.97% Precision, 99.31% Recall, and 0.9914 F1-score.
Small space object detection (SSOD) plays a crucial role in orbital debris monitoring and spacecraft defense, providing critical support for space situational awareness. To adapt complex background interference and lightweight deployment, a Mamba-driven multilevel feature fusion network is designed based on you only look once (YOLO) architecture. Specifically, an adaptive feature enhancement module is introduced to effectively incorporate multilevel information. To consolidate valid information while suppressing redundant noise, a Mamba-driven dual-branch feature extraction module is developed for spatial-aware and channel-aware selective scanning and merging. Comparative experiments demonstrate that the proposed SSOD-YOLO achieves superior performance on a public small space object dataset, with 95.21% recall, 97.94% mAP50, and 96.12% F1-score, outperforming both the YOLO series models and the space object detection methods.
Automatic detection of safety helmets and harnesses is crucial for industrial safety supervision. Existing methods face challenges in detecting small objects, handling complex environments, and capturing fine-grained features. This paper presents a deep learning-based approach that incorporates a fine-grained feature extraction (FFE) branch for fusing low-level details with high-level semantics, and a dynamic label assignment (DLA) strategy to optimize positive sample selection during training. A comprehensive real-world dataset, GDUT-HHD, is created for model development and evaluation. Experiments demonstrate robust detection performance, achieving an mAP@50 of 92.2 320× 320 , confirming its effectiveness and practical applicability for on-site safety monitoring.
Low-slow-small (LSS) target monitoring is critical for airport safety management, particularly when LSS objects such as unmanned aerial vehicles (UAVs) and birds unexpectedly enter airport airspace, posing significant risks to regular flights and airport operations. Recent advances have taken advantage of deep learning for the recognition of LSS radar targets, achieving promising classification accuracy. However, existing LSS radar target recognition approaches often rely on statistical correlations, including unstable spurious correlations, which can undermine the generalization performance of classification networks, limiting their effectiveness in all-time radar recognition. To address this, we propose a causality-inspired single-source domain generalization method for radar LSS target recognition. Our method introduces a Causal-Symmetric Transformation (CST) module for data augmentation, combining Non-Causal Augmentation for global perturbations and Symmetric Transformation for local motion reversal, enhancing data diversity and reducing bias. Additionally, we propose a Causal Mining (CM) module with a Causal Consistency loss to extract causal features that boost generalization. A Fourier-Aware Attention (FAA) module leverages frequency-domain information to strengthen feature representation and preserve causal information. Extensive experiments on four real-world datasets validate the effectiveness of our approach.
Retinopathy of Prematurity (ROP) recurrence is significant for the prognosis of ROP treatment. In this paper, corrected gestational age at treatment is involved as an important risk factor for the assessment of ROP recurrence. To reveal the complementary information from fundus images and risk factors, a dual-modal deep learning framework with two feature extraction streams, termed as ROPRNet, is designed to assist recurrence prediction of ROP after anti-vascular endothelial growth factor (Anti-VEGF) treatment, involving a stacked autoencoder (SAE) stream for risk factors and a cascaded deep network (CDN) stream for fundus images. Here, the specifically-designed CDN stream involves several novel modules to effectively capture subtle structural changes of retina in the fundus images, involving enhancement head (EH), enhanced ConvNeXt (EnConvNeXt) and multi-dimensional multi-scale feature fusion (MMFF). Specifically, EH is designed to suppress the variations of color and contrast in fundus images, which can highlight the informative features in the images. To comprehensively reveal the inherent medical hints submerged in the fundus images, an adaptive triple-branch attention (ATBA) and a special ConvNeXt with a rare-class sample generator (RSG) were designed to compose the EnConvNeXt for effectively extracting features from fundus images. The MMFF is designed for feature aggregation to mitigate redundant features from several fundus images from different shooting angles, involving a designed multi-dimensional and multi-sale attention (MD-MSA). The designed ROPRNet is validated on a real clinical dataset, which indicate that it is superior to several existing ROP diagnostic models, in terms of 0.894 AUC, 0.818 accuracy, 0.828 sensitivity and 0.800 specificity.
The three-dimensional (3D) reconstruction of the hepatic duct tree is significant for the minimally invasive surgery of hepatobiliary stone disease, which can be influenced by the quantity of the CT scans of hepatic ducts. If insufficient CT scans with low inter-slice resolution are directly utilized for 3D reconstruction, discontinuities and gaps will emerge in the reconstructed hepatic duct tree. In this paper, a novel end-to-end deep learning framework is designed for the inter-slice super-resolution segmentation of the CT slices of hepatic ducts, which can improve the 3D reconstruction performance in the inter-slice dimension. Specifically, the framework cascades into an inter-slice super-resolution subnetwork and a segmentation subnetwork. A deep learning network is introduced as the inter-slice super-resolution subnetwork to generate an intermediate slice between two adjacent CT slices in the simulated CT scans with low inter-slice resolution. To capture the spatiotemporal correlation existing in the CT scans of hepatic ducts, the ConvLSTM is introduced into the U-Net-like segmentation subnetwork in the high-dimensional feature space. To further suppress the problems of discontinuities and gaps, a structure-aware loss function is proposed by incorporating the structural similarity index measure (SSIM) as a regulator to dynamically assign the contribution of the generated CT slice to the total loss of the designed framework. Experimental results demonstrate that our proposed framework performs better segmentations for hepatic ducts than several existing deep learning networks with the performance of 0.7690 DICE and a 0.7712 F1-score, which is beneficial for the 3D reconstruction of the hepatic duct tree.
Correct cable connection is critical for safe and reliable operation of the low-voltage switchgear but currently relies on time-consuming and labor-intensive manual inspection. To improve inspection accuracy and efficiency, a novel multiscale edge-enhanced deep learning (MEDL) framework is designed to visually inspect cable connections in a dense cable scenario. Specifically, the MEDL detects the keypoints at the cable-terminal junctions through an encoder-decoder architecture with an edge enhancement (EE) module and a multiscale feature extraction (MSFE) module, followed by a matching stage. The EE module is designed to highlight the edges of the cables, which can, to some extent, suppress environmental interferences. The MSFE module is designed to extract multiscale features at the cable-terminal junctions while guiding the MEDL model to focus on the target regions. In the matching stage, the HDBSCAN is combined with a shared nearest neighbor (SNN) distance metric to cluster candidate keypoints for keypoint matching. The experimental results on cable connection images acquired in real-world scenarios demonstrate the superiority of the MEDL to some existing deep learning methods, achieving a matching accuracy (MA) of 0.9463 at an acceptable inspection speed.
In the domain of micro-nano electronic printing technology, the integration of the operational characteristics with piezoelectric inkjet printing (PIJP) and electrohydrodynamic printing overcomes the limitations of single-mode printing. This integration facilitates the advancement and enhancement of hybrid printing technology, encompassing piezoelectric and electric forces. A study is conducted on jet printing utilizing a hybrid force as the driving power. The hybrid printing system is initially developed and assembled. The study focuses on the printing mechanism of piezoelectric deforming units and electrohydrodynamic generating units, leading to the development and integration of hybrid printing equipment. Secondly, three types of printing models are established using simulation software, including piezoelectric, electrohydrodynamic, and hybrid models. The jetting mechanism of the single force and hybrid force is described, and the feasibility of the process is assessed. The experiment is designed to assess the performance of hybrid printing in comparison to other single-force printing methods. The experimental results indicate that hybrid printing demonstrates superior precision and consistency compared to single-force printing. The average diameters of the printed droplet dots for piezoelectric-driven and composite-driven methods are 225.9 mu m and 181.4 mu m, respectively. Compared to PIJP, hybrid printing ensures uniform droplet formation and reduces satellite formation through electrohydrodynamic stabilization. Compared to electrohydrodynamic methods, hybrid printing improves deposition accuracy, refines jetting control, minimizes unintended spreading, and achieves higher resolution.
Accurate predictions of the state of health (SOH) and remaining useful life (RUL) of lithium-ion batteries are crucial to ensure their efficient and safe operations. Existing deep learning methods are hindered by the limited quantity and diversity of datasets, leading to model overfitting and poor generalization. To this end, a bimodal large-small model collaborative network (BLSCN) is designed in this article, which seamlessly combines a pretrained large vision model (LVM) with a designed small model (SM) for joint prediction of SOH and RUL of lithium-ion batteries. Specifically, the BLSCN utilizes the LVM to extract features from the bimodal images generated by the battery data, followed by an SM for feature fusion and prediction. To promote the generalization ability of LVM on the battery data, an attention mask strategy is proposed to guide the LVM to focus on more features of interest in the images. In the SM, an elaborately designed cross-fusion module (CFM) is employed to interactively fuse bimodal image features for subsequent joint prediction of SOH and RUL. Experimental results on the public dataset demonstrate the superiority of the proposed BLSCN on joint prediction of SOH and RUL, with coefficients of determination (R-2) of 0.983 and 0.942 for SOH and RUL prediction tasks, respectively.
Accurately predicting the state of health (SOH) and remaining useful life (RUL) of lithium-ion batteries is crucial for battery management systems. Most of existing methods rely on complete charging-discharging data, which are possibly limited in the real scenarios since discharging data are easily affected by different operating conditions. To this end, a deep learning framework is designed for joint prediction of SOH and RUL of lithium-ion batteries using only charging data, involving a novel differential scheme for charging data structuring and an elaborately-designed multi-strategy attention regression network (MARN). The MARN is a hybrid dual-branch deep learning network with several self-attentions, consisting of the stages of shared feature extraction, task-specific feature extraction, and prediction. To comprehensively extract temporal features inherent existing in charging data, a multi-strategy attention mechanism is designed to construct three residual modules with different attentions to extract temporal features in different domains. Experiments on the public dataset demonstrate that our proposed method can perform well in joint prediction of SOH and RUL with the performance of coefficients of determination (R2) of 0.99 and 0.94 for SOH and RUL tasks, respectively, which is superior to some existing single-state models and a joint prediction model.
Melt electrospinning writing (MEW) technology has demonstrated its utility in additive manufacturing 3D micrometer structures, such as biodevices and flexible electronic devices. It needs to achieve high-precision and micrometer-scale deposition through the 3D jet formed by the Taylor cone. However, jet lag effects during high-speed direct printing present a significant challenge in realizing high-precision deposition, especially at 90 degrees angles. This study first analyzes the generation of jet lag effects. Subsequently, a detection model for the jet deflection angle theta is constructed using a real-time charge coupled device camera. The exploration of collector acceleration and the jet-deposited trajectory is also described. Finally, the influence and relationship of various MEW parameters-such as nozzle height, heating temperature, and acceleration-on achieving high-precision jet deposition are examined. This study provides a new perspective and ideas for adjusting collector acceleration to enhance MEW high-precision jet deposition.
During reflow soldering, Metal-Oxide- Semiconductor Field-Effect Transistors (MOSFETs) are prone to internal voids. A high void ratio for a MOSFET significantly affects its electrical performance to further decrease the reliability and lifespan of the electronic device. In this article, a multi-scale adaptive inspection framework is proposed to inspect internal voids in MOSFETs through X-ray imaging. First, a location-aware vertex detection scheme is utilized to detect the transistor region possibly involving the voids, which makes full advantages of the neighborhood grayscale distributions of the transistor region vertices. Then, after the detected transistor region is preprocessed and divided into several patches, coarse-to-fine void inspection is performed on the patches for void inspection. Specifically, a multi-scale scheme is proposed for coarse void inspection to identify whether the patches involve the potential voids, in which an identification function is defined by integrating the constructed multi-scale radial decaying weight matrix and a fitting metric to characterize the patches. An adaptive thresholding scheme is proposed to finely inspect the voids in the MOSFETs, in which the segmentation threshold model is constructed based on grayscale distributions of the MOSFET X-ray images. Finally, the void ratio is calculated according to the detected transistor and the voids. Experimental results indicate that the proposed framework achieves better inspection performance for voids in MOSFETs than some existing methods at the speed of 0.812s per MOSFET, with an average Dice of 0.8852, an Accuracy of 97.11%, and an F1-score of 0.9819.
Automatic intention recognition in financial service scenarios faces challenges such as limited corpus size, high colloquialism, and ambiguous intentions. This paper proposes a hybrid intention recognition framework for financial customer service, which involves semi-supervised learning data augmentation, label semantic inference, and text classification. A semi-supervised learning method is designed to augment the limited corpus data obtained from the Chinese financial service scenario, which combines back-translation with BERT models. Then, a K-means-based semantic inference method is introduced to extract label semantic information from categorized corpus data, serving as constraints for subsequent text classification. Finally, a BERT-based text classification network is designed to recognize the intentions in financial customer service, involving a multi-level feature fusion for corpus information and label semantic information. During the multi-level feature fusion, a shallow-to-deep (StD) mechanism is designed to alleviate feature collapse. To validate our hybrid framework, 2977 corpus texts about loan service are provided by a financial company in China. Experimental results demonstrate that our hybrid framework outperforms existing deep learning methods in financial customer service intention recognition, achieving an accuracy of 89.06%, precision of 90.27%, recall of 90.40%, and an F1 score of 90.07%. This study demonstrates the potential of the hybrid framework to automatic intention recognition in financial customer service, which is beneficial for the improvement of the financial service quality.
Electrohydrodynamic (EHD) technology is renowned for its significant advantages in high resolution and micro-nanoscale printing, demonstrating an immense potential in the development of micro-nano devices. During the printing process, it is inevitably influenced by different interferences, which result in printing errors that influence its printing precision. This article summarizes several research topics on printing errors of EHD printing technology, involving the sources, and correction of different types of printing errors. First, the induced factors of printing errors are summarized in details, which are used to categorize the error correction methods. Then, the existing correction methods are comprehensively summarized and analyzed according to the types of printing errors. Finally, the conclusions are provided, involving some potential research topics.
To address the issue of error compensation for optical encoders in various measurement environments, a compensation model based on an improved deep reinforcement learning (DRL) framework has been designed. This marks the first attempt to introduce DRL into the compensation problem of optical encoders. To achieve this, an improved deep Q-learning network (DQN) has been developed to construct the compensation model. Specifically, the network enhances measurement accuracy by first compensating the input grating images and then cascading with a positioning and decoding network. Experimental results demonstrate that the designed compensation network is suitable for various speed measurement environments, and the measurement accuracy achieved using this compensation network surpasses that of existing compensation methods.
Outdoor images often suffer from severe degradation due to rain, haze, and noise, impairing image quality and challenging high-level tasks. Current image restoration methods struggle to handle complex degradation while maintaining efficiency. This paper introduces a novel image restoration architecture that combines multi-dimensional dynamic attention and self-attention within a U-Net framework. To leverage the global modeling capabilities of transformers and the local modeling capabilities of convolutions, we integrate sole CNNs in the encoder-decoder and sole transformers in the latent layer. Additionally, we design convolutional kernels with selected multi-dimensional dynamic attention to capture diverse degraded inputs efficiently. A transformer block with transposed self-attention further enhances global feature extraction while maintaining efficiency. Extensive experiments demonstrate that our method achieves a better balance between performance and computational complexity across five image restoration tasks: deraining, deblurring, denoising, dehazing, and enhancement, as well as superior performance for high-level vision tasks. The source code will be available at https://github.com/House-yuyu/MDDA-former.
Online inspection of flexographic printing labels (FPLs) is significant for preventing batch misprints and ensuring mass production quality. In this article, an online inspection system is designed for fabric FPLs, in which a multiscale generative adversarial network (GAN) framework with several region adaptive schemes is proposed to inspect the fabric FPL images acquired by the designed hardware system. In order to address the scarcity of defective samples in real production, the framework only employs qualified FPL samples to train the GAN template generators for generating the templates. Different from previous template generators, the framework incorporates skip connections and a multiscale architecture to improve the template representation ability. In order to deal with tensions in the FPL substrates and to capture subtle defects, several region adaptive schemes are designed for the inspection stage, involving statistical region subtracter (SRS), region adaptive thresholding (RAT), and region defect assessment (RDA). Specifically, during inspection, the inspected FPL is adaptively divided into the character and background regions, which are dealt with by these region adaptive schemes. Experimental results demonstrate that the proposed framework achieves the fabric FPL inspection performance of 0.1% omission rate and 1.8% error rate at a reasonable speed of 45 ms per sample, which can meet the online inspection requirements for real production.