Compound Facial Expressions (CFEs) are composed of basic expressions with different contributions. However, existing approaches to compound facial expression recognition (CFER) often overlook the distinction between primary and secondary contributions, leading to ambiguous margins and inter-class overlap among similar CFEs. To address this challenge, we propose a novel framework, Differential-Guided Tri-Cyclic Suppression (DGCS3), which explicitly disentangles the contribution differences within CFEs to enhance classification separability. DGCS3 comprises two key modules: the Tri-Cyclic Expression Suppression (CES3) module and the Differential Contribution Detector (DCD) module. The CES3 module is designed to disentangle different expressions by employing a tri-cyclic suppression strategy to decouple the contributions of CFEs. Additionally, we introduce Restoration Consistency Loss and Cycle Consistency Loss to enforce cyclicity and restorability, ensuring more precise expression disentanglement. To further mitigate inter-class overlap, the DCD module leverages the differences in expression contributions to refine classification centers and compress classification margins. Specifically, based on the prior knowledge that the primary expression is closer to the classification center than the secondary, the DCD module optimizes the placement of classification centers accordingly. Meanwhile, the classification margin is constrained via a regularization mechanism that accounts for contribution differences. Extensive experiments on both in-the-lab and in-the-wild datasets demonstrate that DGCS3 surpasses state-of-the-art methods in terms of effectiveness and robustness, offering a more structured and discriminative approach to CFER.
Facial expression is a key signal in social human robot interaction, but natural facial motion is still difficult to achieve on physical robots when the system must rely on camera sensing and low-cost embedded control. The main difficulty is not only to reproduce a correct target face. The robot must also move with a reasonable temporal process, react to new visual changes in time, and remain stable on limited hardware. To address this problem, this work presents a camera-based sensing to actuation framework for humanoid facial expression control. A camera first captures facial image sequences, and frame-wise Action Unit intensities are then estimated from the observed face. Based on offline AU sequences, an AU-specific canonical velocity prior is built to describe the relation between AU intensity and motion speed. During online control, the latest three frames are used to sense the local motion trend, and the sensed velocity is used to anchor the canonical prior for short-term trajectory generation. A dead zone is further introduced to suppress false motion near the static state. Finally, a quantized velocity matching strategy is designed for microcontroller-based execution. Experiments on local velocity sensing, online trajectory generation, abrupt motion change, and embedded runtime show that the proposed method improves temporal stability, reduces overshoot and static chatter, and remains suitable for real-time physical deployment. The results demonstrate the feasibility of the proposed sensor-driven control framework on the physical YU-F02 facial robot.
Defect detection in photovoltaic (PV) cells is essential for maintaining the stability and efficiency of PV power generation systems. However, existing detection models often struggle to achieve high detection accuracy while maintaining a lightweight design for deployment on Unmanned Aerial Vehicle (UAV) or other edge devices. To address this challenge, we propose a new efficient PV defect detection model termed YOLO-PDD (YOLO for PV Defect Detection). This model achieves accurate and fast detection with a lightweight architecture. Specifically, we introduce a novel Curve Deformable Convolution (CDC) that incorporates a learnable curvature-guided offset mechanism into standard convolution, enabling the sampling grid to adaptively deform along spatial directions. This design allows the model to more effectively capture geometric variations in feature maps, thereby improving its capacity to represent irregular and small-scale defect regions. Moreover, an innovative Multi-Kernel Depthwise Convolution (MKDWC) is proposed. It uses channel-splitting and heterogeneous kernel designs to simultaneously reduce model parameters and computational cost while preserving high detection accuracy. The proposed model is evaluated on the public PVEL-AD dataset and ELDDS1400C5 dataset, achieving an mAP50 of 0.89, mAP50:95 of 0.609, with only 2.56 million parameters, 6.8G FLOPs, and a real-time inference speed of 435 FPS on the PVEL-AD dataset. It further achieves a mAP50 of 0.75, mAP50:95 of 0.539 on ELDDS1400C5 dataset. These results demonstrate that the proposed model simultaneously achieves high accuracy and lightweight design, making it well-suited for deployment on UAVs or other edge devices in PV inspection tasks, and showing strong potential for real-world photovoltaic defect detection applications.
In immersive panoramic video (PV) encoding scenarios, PV exhibits larger flat regions compared to traditional videos. The existing intra-rate control methods in Versatile Video Coding (VVC) allocate the bitrate based on the construction of weights according to the energy distribution characteristics of different encoding blocks. However, in reality, the human visual system (HVS) is not sensitive to blocks with many flat regions in perception, which leads to excessive allocation of bitrate in insensitive regions. On the contrary, insufficient bitrate allocation in sensitive regions leads to the inability to achieve better reconstruction quality. To address this challenge, we propose a Perception-Driven Rate Control (PD-RC) strategy for panoramic video encoding based on energy distribution optimization, which makes the intra-rate control closer to the perception habits of HVS. Firstly, we propose a low-complexity filtering method guided by rate-distortion performance to optimize the energy distribution of I-frame features. Subsequently, leveraging the optimized perception features of the energy distribution, a perception-driven intra-mode coding-tree-unit-level rate control strategy is proposed to improve the coding performance for PV. Extensive evaluations show the performance of PD-RC over the stateof-the-art rate control methods of VVC. Specifically, in all-intra encoding mode, the average bitrate savings of PD-RC is-5.002%, while the average gain in weighted-spherically quality is 0.239 dB, with a rate exceeding the upper limit as low as 2%, and a reduction in encoding complexity gain of-0.163%. PD-RC effectively improves the rate control performance of PV intra-frame coding while saving computational overhead. It is significant for optimizing data transmission efficiency, enhancing video quality, and reducing storage costs. The source code will be available at https://github.com/liulinyun324/PD-RC.
Compound Facial Expressions consist of multiple basic expressions and present increased complexity in recognition tasks. Enhancing the representation of primary and secondary expressions is critical for improving the classification performance. However, when features are enhanced indiscriminately, irrelevant information may also be amplified, leading to the ‘Indiscriminate Enhancement Trap' and robustness degradation. To address these issues, we propose a novel Bi-Enhancement based Disentangled Dual-Attention (BED2A) Framework. The bi-enhancement strategy is devised to achieve signal complementarity between channel and spatial features. By transferring semantic information from channel attention to spatial attention in the latent space, the proposed method enhances the ability to localize salient regions. Subsequently, primary and secondary expression features are disentangled across tri-branch feature enhancement architecture, thereby enabling more effective dual-attention based classification. Specifically, the Differential Cross-Consistency mechanism is introduced to disentangle inter-expression features and enhance the precision of expression representations. Dual-Attention mechanism leverages attention to primary and secondary expressions to optimize the classification center and improve inter class separability. Experimental results demonstrate that BED2A effectively optimizes feature enhancement to improve robustness, achieving state-of-the-art performance.
Existing facial robotic systems often suffer from high costs, limited flexibility in expression imitation, and a lack of adaptability to diverse facial identities, resulting in suboptimal replication of human-like expressions. This work designs YU-F01, a low-cost soft-skin facial robot capable of real-time emotion-driven actuation through visual perception. The system integrates computer vision and a multi-actuator facial mechanism to enable dynamic expression imitation. The main innovations include: (1) a landmark-based hierarchical mapping framework that converts facial keypoints into coordinated servo movements, allowing for adaptive expression replication across different facial structures; (2) a Polling-based Multi-actuator Synchronous Control Framework (PMSCF) that enables pseudo-parallel control of 16 servos on a microcontroller, reducing flash and memory usage by 21
Underwater object detection plays a significant role in marine ecosystem research and marine species conservation. The improvement of related technologies holds practical significance. Although existing object-detection algorithms have achieved an excellent performance on land, they are not satisfactory in underwater scenarios due to two limitations: the underwater objects are often small, densely distributed, and prone to occlusion characteristics, and underwater embedded devices have limited storage and computational capabilities. In this paper, we propose a high-precision, lightweight underwater detector specifically optimizing for underwater scenarios based on the You Only Look Once Version 8 (YOLOv8) model. Firstly, we replace the Darknet-53 backbone of YOLOv8s with FasterNet-T0, reducing model parameters by 22.52%, FLOPS by 23.59%, and model size by 22.73%, achieving model lightweighting. Secondly, we add a Prediction Head for Small Objects, increase the number of channels for high-resolution feature map detection heads, and decrease the number of channels for low-resolution feature map detection heads. This results in a 1.2% improvement in small-object detection accuracy, while the remaining model parameters and memory consumption are nearly unchanged. Thirdly, we use Deformable ConvNets and Coordinate Attention in the neck part to enhance the accuracy in the detection of irregularly shaped and densely occluded small targets. This is achieved by learning convolution offsets from feature maps and emphasizing the regions of interest (RoIs). Our method achieves 52.12% AP on the underwater dataset UTDAC2020, with only 8.5 M parameters, 25.5 B FLOPS, and 17 MB model size. It surpasses the performance of large model YOLOv8l, at 51.69% AP, with 43.6 M parameters, 164.8 B FLOPS, and 84 MB model size. Furthermore, by increasing the input image resolution to 1280 × 1280 pixels, our model achieves 53.18% AP, making it the state-of-the-art (SOTA) model for the UTDAC2020 underwater dataset. Additionally, we achieve 84.4% mAP on the Pascal VOC dataset, with a substantial reduction in model parameters compared to previous, well-established detectors. The experimental results demonstrate that our proposed lightweight method retains effectiveness on underwater datasets and can be generalized to common datasets.
The growing demand for mobile devices has generated interest in lightweight human pose estimation. Currently, lightweight estimation generally uses heatmap-based methods, which has demonstrated exceptional performance. However, their use of non-differentiable post-processing imposes considerable inference latencies. Conversely, integral-based approaches expedite the inference process by employing a soft-argmax operation but compromise in accuracy. Integrating explicit heatmap knowledge learned using the heatmap-based method into the implicit heatmap generated by the integral-based method, thereby combining the best of both worlds, offers a promising avenue. However, owing to the disparities in supervision and inference processes, the explicit and implicit heatmaps are heterogeneous. Consequently, direct transfer of knowledge presents difficulties in ensuring consistencies in heat value and location. In this paper, we propose a novel Heterogeneous Heatmap Distillation (HHD) framework that effectively tackles these challenges. The framework seamlessly integrates explicit heatmap knowledge that contains high-precision localization information into implicit heatmaps. The framework revolves around an unbiased heatmap alignment scheme encompassing two steps: heterogeneous heatmap normalization and unbiased cropping. Heterogeneous heatmap normalization separately normalizes the output feature maps of both the teacher and student models, alleviating potential heat value bias during the knowledge transfer. Unbiased cropping applies closed-form computation on the normalized teacher and student heatmap to eliminate location bias. Additionally, mirror expansion is implemented to handle potential cases wherein the cropped region extends beyond the image boundary. Extensive experiments demonstrate the efficiency and effectiveness of our methods on the MSCOCO and MPII datasets compared to other integral-based lightweight networks. Our source codes and pre-trained models are available at https://github.com/ducongju/HHD.
2D heatmap-based human pose estimation has exhibited remarkable performance. However, constrained by quantization error, heatmap-based methods heavily rely on high-resolution heatmaps and intricate post-processing to enhance detection accuracy, thereby incurring substantial computational costs. To pursue more effective keypoint representation, we propose a novel scheme named Offset-based Disentangled Representation (ODR). ODR conducts coordinate classification and offset prediction simultaneously using 1D vectors and aggregates their outputs to precisely pinpoint keypoint positions, thus freeing from dependence on high-resolution heatmaps. To eliminate the impact of long-range offsets, we propose a scale-aware eraser that generates noticeable intervals based on the relative scale of different keypoints, directing the regression task to focus on short-range offsets. By doing so, upsampling layers are no longer necessary, enabling a more concise and effective architecture for human pose estimation. Extensive experiments conducted over COCO and MPII datasets validate the superiority of ODR over counterparts based on 2D or 1D heatmaps. Our source codes are available at the link.
Facial action unit (AU) detection constitutes precise measurements for facial appearance variances, holding great significance within the realms of affective computing, human-computer interaction, and negotiation. Subject-invariant AU detection remains a challenge primarily due to the distribution variations among individuals. More importantly, the inherent subtlety and localized nature of facial action frequently give rise to the dominance of interference factors, particularly those related to individual identity. To tackle these issues, we propose a novel knowledge-driven hierarchical feature alignment (KHFA) framework, which aims to investigate the multifaceted consistency within representations of facial actions. AUs unambiguously define facial appearance variations induced by specific groups of facial muscle movements. At the same time, the intrinsic physiological interconnections between these facial muscles impose substantial constraints on the correlations between different AUs. Therefore, KHFA presents a dual classwise alignment scheme to ensure a harmonious balance between consistency within the same class and coherence across different categories. Furthermore, the similarity in the sample-level AU combinations reflects the semantic proximity of global features within the feature space. KHFA integrates an intersample relationship to enhance the coherence of semantic information across samples via a multilabel alignment scheme. Finally, a hybrid attention mechanism equipped with an importance-aware feature fusion layer is proposed to capture nuanced spatial features that are specific to individual AUs and to adeptly embed AU correlations. Extended experiments conducted on two benchmark datasets, BP4D and DISFA, reveal that KHFA outperforms state-of-the-art methods, underscoring the effectiveness and superiority of our approach.
To implement a metaverse exhibition interaction system, the instability problem of high-quality avatar facial reenactment must be considered. How to void identity limitations and eliminate artifacts are key challenges for avatar reenactment. It also lacks the support of the application system it is implemented in. We propose a metaverse system architecture oriented to emotional interaction. And we propose a novel method for avatar expression reenactment named Overall-Local Feature Warping Fusion Model based on Optical-Flow field prediction. We solve the identity limitation by overall optical-flow estimation and local optical-flow estimation and eliminate artifacts by illumination consistency. We compare with the mainstream optical flow face reenactment methods and outperform them in identity similarity, structural similarity, and facial action unit recognition ratio. We experimentally compared our method improves by an average improvement of 3.79%. And we also implement our method in the metaverse exhibition system. Although we satisfy most of the interaction scenarios, our method is still insufficient in some side-face cases.
Micro-expressions are spontaneous, rapid and subtle facial movements that can neither be forged nor suppressed. They are very important nonverbal communication clues, but are transient and of low intensity thus difficult to recognize. Recently deep learning based methods have been developed for micro-expression (ME) recognition using feature extraction and fusion techniques, however, targeted feature learning and efficient feature fusion still lack further study according to the ME characteristics. To address these issues, we propose a novel framework Feature Representation Learning with adaptive Displacement Generation and Transformer fusion (FRL-DGT), in which a convolutional Displacement Generation Module (DGM) with self-supervised learning is used to extract dynamic features from onset/apex frames targeted to the subsequent ME recognition task, and a well-designed Transformer Fusion mechanism composed of three Transformer-based fusion modules (local, global fusions based on AU regions and full-face fusion) is applied to extract the multi-level informative features after DGM for the final ME prediction. The extensive experiments with solid leave-one-subject-out (LOSO) evaluation results have demonstrated the superiority of our proposed FRL-DGT to state-of-the-art methods.
Background: Cashew (Anacardium occidentale L.) is a commercially important plant. Cashew nuts are a popular food source that belong to the tree nut family. Tree nuts are one of the eight major food allergens identified by the Food and Drug Administration in the USA. Allergies to cashew nuts cause severe and systemic immune reactions. Tree nut allergies are frequently fatal and are becoming more common. Aim: We aimed to identify the key allergenic epitopes of cashew nut proteins by correlating the phage display epitope prediction results with bioinformatics analysis. Design: We predicted and experimentally confirmed cashew nut allergen antigenic peptides, which we named Ana o 2 (cupin superfamily) and Ana o 3 (prolamin superfamily). The Ana o 2 and Ana o 3 epitopes were predicted using DNAstar and PyMoL (incorporated in the Swiss-model package). The predicted weak and strong epitopes were synthesized as peptides. The related phage library was built. The peptides were also tested using phage display technology. The expressed antigens were tested and confirmed using microtiter plates coated with pooled human sera from patients with cashew nut allergies or healthy controls. Results: The Ana o 2 epitopes were represented by four linear peptides, with the epitopes corresponding to amino acids 108–111, 113–119, 181–186, and 218–224. Furthermore, the identified Ana o 3 epitopes corresponding to amino acids 10–24, 13–27, 39–49, 66–70, 101–106, 107–114, and 115–122 were also screened out and chosen as the key allergenic epitopes. Discussion: The Ana o 3 epitopes accounted for more than 40% of the total amino acid sequence of the protein; thus, Ana o 3 is potentially more allergenic than Ana o 2. Conclusions: The bioinformatic epitope prediction produced subpar results in this study. Furthermore, the phage display method was extremely effective in identifying the allergenic epitopes of cashew nut proteins. The key allergenic epitopes were chosen, providing important information for the study of cashew nut allergens.
Facial action unit (AU) detection is a hot topic in computer vision, but it remains challenging due to individual characteristics. Facial action features are a vital informative factor to explain facial anatomical variations but are often entangled with other facial attribution information leading to representation inconsistency within one category. We propose a novel Contrastive Disentangled Representation Autoencoder (CDAE) to learn discriminative identity-invariant representation for AU detection by factorizing face images into temporally varying action parts and stationary facial attributions components. Facial image space is mapped onto the facial action subspace and action-independent identity subspace to disentangle facial action information from identity information. In addition, we design a contrastive learning scheme to obtain a semantic-aware AU manifold by mapping the facial action features onto the continuous space of the latent variables, thus minimizing the misalignment between subjects and reducing the dimension of the facial action features. Experiment results show that CDAE outperforms or is comparable to previous AU detection methods on the challenging BP4D and DISFA benchmarks, demonstrating that the learned facial action representation is discriminative for AU detection.
Human expression often happens simultaneously with head posture in real-time video, facial Action Units (AUs) is the key factor in facial expression detection. Optical flow can effectively capture weak motion displacements, that is, it can capture facial AUs caused by facial expressions. But the optical flow would produce a lot of noise, which would adverse the detection performance. To achieve a better facial AU detection performance, we propose a novel Optical Flow Synthesis Generative Adversarial Network (OFS-GAN). Firstly, we calculate the optical flow vector of the source frame and the target frame pair that randomly selected from video clips of facial expressions to promote the robustness of OFS-GAN. Secondly, in the generator of OFS-GAN, representation input into the encoder is yield by the optical flow vector concatenated with the source frame. The feature map output from the encoder is fed into the decoder to synthesize a target frame named generated target frame. In the end, through the comparison of the generated target frame and the target frame randomly selected from the target stage, OFS-GAN learns discriminative and robust facial action feature for facial AU detection following the principle of adversarial learning. Our novel OFS-GAN has been tested on DISFA+ and CK+ dataset with LOSO evaluation method. Qualitative results of our experiment demonstrate that OFS-GAN approaches or exceeds existing optical flow or deep learning algorithms.
The emotion of human beings tends to be complex in real conditions, generating compound expressions in human faces. Compound expression recognition is an important challenge for the assessment of human complex emotion. The recognition system based on six basic expressions cannot meet the demand of compound expressions recognition. The recognition performance of models learned from the basic expressions is poor due to the small number of compound expression datasets with highly accurate labels and insufficient sample diversity. Making full use of domains outside of compound expressions in small sample datasets will help promote diversity. We propose the Multi-Domain Fusion Generative Adversarial Network (MDFGAN), which innovatively fuses the face domain, compound expression domain and basic expression domain to obtain rich expression generation capability and high accuracy recognition. Pairing the face domain and the contour-unrelated compound expression domain in the generator will expand the sample diversity. The contour-related compound expression domain and the basic expression domain will jointly improve the expression recognition accuracy of the discriminator. Finally, we conducted comprehensive experiments on CFEE-26, CFEE-7 and CK+. In the experiments, the results of MDFGAN improved 6.79% on UF1 and 8.5% on UAR.
In the past decade, with the deepening of the aging of the population and the strengthening of the health consciousness of the whole society, the Internet healthcare service has grown to be an inevitable trend of current society. We propose a collaborative adaptive architecture named Trusted Healthcare Smart Brain (THSB) for cross-blockchain intelligent collaboration of multi-institution healthcare services. THSB is an interdisciplinary system with Healthcare Internet of Things (H-IoT), blockchain, Artificial Intelligence, Cloud Computing, Big Data, and Internet. The participants of THSB include patients, rehabilitation institutions, medical service institutions, healthcare content service institutions, medical regulatory institutions, scientific research institutions, and government institutions. In addition, we propose a medical resources service balance method to maximize the utilization of medical service resources to solve the contradiction between random medical events and the normal distribution of medical resources.
Facial Action Unit (AU) detection is a challenging task for the reason that AU features extracted from videos always entangle other inevitable variations, including head posture motion characteristics, which is even much more intense than facial actions, and individual facial features because of race, age, gender or face shape. These AU-unrelated features would adverse the AU detection performance. To achieve better performance of AU detection, we proposed a novel Feature Disentangled Autoencoder (FDAE) to learn more discriminative facial action representation from large amounts of videos. Different from previous approaches, FDAE disentangled AU-related features, head posture motion characteristics and identity code characteristics to eliminate the impact of irrelevant factors. At the same time, we added two classifiers to discriminate the AU embedding and identity code embedding respectively to accelerate training process and make the model more stable and robust. Experiments on BP4D and DISFA demonstrated that the learned representation is discriminative, where FDAE outperformed or was comparable with existing representation learning method for AU detection.
Chronic diseases are a growing concern worldwide, with nearly 25% of adults suffering from one or more chronic health conditions, thus placing a heavy burden on individuals, families, and healthcare systems. With the advent of the "Smart Healthcare" era, a series of cutting-edge technologies has brought new experiences to the management of chronic diseases. Among them, smart wearable technology not only helps people pursue a healthier lifestyle but also provides a continuous flow of healthcare data for disease diagnosis and treatment by actively recording physiological parameters and tracking the metabolic state. However, how to organize and analyze the data to achieve the ultimate goal of improving chronic disease management, in terms of quality of life, patient outcomes, and privacy protection, is an urgent issue that needs to be addressed. Artificial intelligence (AI) can provide intelligent suggestions by analyzing a patient's physiological data from wearable devices for the diagnosis and treatment of diseases. In addition, blockchain can improve healthcare services by authorizing decentralized data sharing, protecting the privacy of users, providing data empowerment, and ensuring the reliability of data management. Integrating AI, blockchain, and wearable technology could optimize the existing chronic disease management models, with a shift from a hospital-centered model to a patient-centered one. In this paper, we conceptually demonstrate a patient-centric technical framework based on AI, blockchain, and wearable technology and further explore the application of these integrated technologies in chronic disease management. Finally, the shortcomings of this new paradigm and future research directions are also discussed.