Research has shown that facial expressions can effectively infer the severity of depression. Recently, depression assessment methods based on images or facial landmarks have gained widespread attention. However, existing methods have not fully considered the complementary and synergistic effects between facial landmark features and images in extracting depression cues. Therefore, this article proposes a bidirectional depression detection network, aimed at comprehensively capturing facial depression cues. We employ a differential feature enhancement module to extract the differences between features and inject them into the corresponding features. Simultaneously, the common feature enhancement module utilizes facial landmark-based structured positional information as a guiding signal, directing the model to focus on depression-related information within the facial image features and fuse shared features. Furthermore, through adaptive feature fusion, we enhance the expressive power of depression-related features to improve the accuracy of assessment. We conducted extensive experiments on the AVEC2014 depression dataset and the RAF-DB facial expression dataset. The experimental results demonstrate that our method outperforms the current state-of-the-art methods, providing strong evidence of its effectiveness.
PurposeThis study aims to tackle the primary challenges in human-robot-environment interaction (HREI) within unknown environments. The key issues include recognizing human motion intention and managing force impacts during transitions from free space to constrained space. Addressing these challenges is critical for improving compliance, enhancing force control accuracy, and ensuring the safety and performance of HREI systems.Design/methodology/approachFirst, the energy equation of the second-order system is presented, and variable admittance control laws are designed for both free space and constraint space based on the energy equation. Then, a smooth switching method based on selection matrix is developed. Subsequently, the admittance-based overall control system is discussed. Finally, comparative simulations and experiments are conducted to verify the efficacy of the variable admittance control method using the 7-degree-of-freedom (7-DOF) manipulator Panda arm.FindingsThe simulation and experiment results demonstrate that the proposed variable admittance control method outperforms the traditional method in terms of force overshoot and accuracy.Research limitations/implicationsThis study does not account for the shape of unknown surfaces in the formulation of the variable admittance control law.Originality/valueThis paper proposes an energy-based variable admittance control method that uses energy considerations and uses a smooth switching technique to deduce human intentions and mitigate the effects of impact force.
Diffusion models have achieved remarkable success in image generation and, more recently, have shown strong capabilities in restoring clean images from various degradations. However, existing approaches typically adapt diffusion models for image restoration through supervised learning, which requires large amounts of paired training data for a predefined set of degradations. This not only incurs high data-labeling costs but also limits the trained models to closed-set scenarios. To overcome these limitations, we propose a zero-shot image restoration framework based on a pretrained latent diffusion model. Since such models are originally trained for image generation, they inherently capture rich low-level image statistics. We observe that by keeping the pretrained model frozen and optimizing only with respect to the input noise, we can effectively reconstruct any degraded images without using additional training data. Furthermore, when conditioning the model on both the degraded image and a suitable text prompt, we can significantly accelerate the reconstruction process and reduce the tendency to overfit to artifacts, resulting in cleaner restorations. Extensive experiments demonstrate the effectiveness of our method on image denoising and dehazing tasks.
The advancement of cable-driven robot technology broadens their application scope, showcasing significant developmental promise. This study addresses the kinematic control precision deficit in quaternion joint cable-driven continuum robots by reconceptualizing the quaternion joint as a dual universal joint. Through systematic error analysis and modeling, it introduces a kinematic calibration method leveraging optimized trajectories to refine algorithmic precision. This approach streamlines calibration data acquisition, ensuring data quality and enhancing computational accuracy. Both simulations and experiments verify the proposed calibration algorithm’s efficacy in augmenting cable-driven robots’ kinematic accuracy.
Remote photoplethysmography (rPPG) is a promising non-contact method for measuring heart rate (HR) and physiological signals. However, current deep learning approaches in this field primarily focus on extracting subtle rPPG cues using convolutional neural networks with limited spatio-temporal receptive fields. This approach overlooks the importance of long-distance temporal perception and interaction, which are crucial for accurate rPPG modeling. In this paper, we present an end-to-end temporal hierarchical quick spatial attention (THQSA) network to adaptive aggregate the local and global temporal features through multi-scale temporal slice perception and interaction. Specifically, the Temporal Hierarchical Block (THB) is firstly constructed to incorporate the global periodic rPPG attention features across multi-scale temporal slices attention mechanism. And then THB refines local temporal slice representations to account for disturbances. Quick spatial attention is further developed to ensure the model obtains better global and local feature representation by avoiding irrelevant features. Additionally, we apply the rPPG signal to a 1-D convolutional network for emotion recognition. Promising results are achieved in the emotion test. To demonstrate the effectiveness of our approach, extensive experiments conducted on two benchmark datasets showcase an outperformance over existing state-of-the-art methods.
It is very challenging for robots to perform grinding and polishing tasks on surfaces with unknown geometry. Most existing methods solve this problem by modeling the relationship between the force sensing information and surface normal vectors by analyzing the forces on special end tools such as spherical tools and cylindrical tools and simplified friction model. In this paper, we propose a normal vectors learning method to simultaneously control end-effector force and direction on unknown surfaces. First, the relation that mapping the force sensing information to the surface normal vectors is learned from the demonstrated data on the known plane using locally weighted regression. Next, the learned relation is used to estimate surface normal vectors on the unknown surface. To improve the force control precision on the unknown geometry surface, the adaptive force control is developed. To improve the direction control precision due to friction, the iterative learning control is developed. The proposed method is verified by comparative simulations and experiments using the Franka robot. Results show that the end-effector can be controlled perpendicular to the surface with a certain force.
In order to perform accurate physical analysis of digital core, the reconstruction of high-quality digital core image has become a problem to be resolved at present. In this paper, a digital core image reconstruction method based on the residual self-attention generative adversarial networks is proposed. In the process of digital core image reconstruction, the traditional generative adversarial networks (GANs) can obtain high resolution detail features only by the spatial local point generation in low resolution details, and the far away dependency can only be processed by multiple convolution operations. In view of this, in this paper the residual self-attention block is introduced in the traditional GANs, which can strengthen the correlation learning between features and extract more features. In order to analyze the quality of generated shale images, in this paper the Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) are used to evaluate the consistency of Gaussian distribution between reconstructed shale images and original ones, and the two-point covariance function is used to evaluate the structural similarity between reconstructed shale images and original ones. Plenty experiments show that the reconstructed shale images by the proposed method in the paper are closer to the original images and have better effect, compared to those of the state-of-art methods.
The issue of prescribed tracking error fixed-time control of stochastic nonlinear systems is investigated in this article. Different from the conventional quartic Lyapunov function (LF) on tracking error, a novel LF based on two important tuning functions is constructed. By means of the fixed-time command filtered dynamic surface control (DSC) technique with the newly error compensating signals (ECSs), the designed controller can ensure that the tracking error and state tracking errors can be predicted in advance without any state transformation. Meanwhile, the problem of “curse of dimensionality” is avoided, and the filtering errors are effectively compensated for. Furthermore, an improved event-triggering mechanism (ETM) is designed to save network resources. A simulation result verifies the scheme developed.
Fractional-order calculus is an extension of integer order calculus. In signal processing, fractional-order calculus can non-linearly enhance the low-frequency signal and suppress the high-frequency signal. In this paper, a new fractional-order local minimum pixel prior (FOLMP) is proposed by combining fractional-order calculus with the local minimum pixel prior. The FOLMP of the sharp images includes fewer non-zero pixels than the blur images. A new blur kernel estimation algorithm is proposed by combining L0 regularized FOLMP with the maximum posterior probability. Furthermore, the kernel similarity is employed to adjust the iteration times to accelerate the computational efficiency. Comparative experiments show that the proposed algorithm can perform better on different types of datasets than the most advanced algorithms. In addition, non-overlapping image patches are adopted to compute the FOLMP, and the kernel similarity is used to suppress excessive iterations. Therefore, the proposed algorithm is several times or even tens of times more efficient than the classical prior-based methods.
Neural networks are increasingly used widely in the solution of partial differential equations (PDEs). This letter proposes 3D-PDE-Net to solve the three-dimensional PDE. We give a mathematical derivation of a three-dimensional convolution kernel that can approximate any order differential operator within the range of expressing ability and then conduct 3D-PDE-Net based on this theory. An optimum network is obtained by minimizing the normalized mean square error (NMSE) of training data, and L-BFGS is the optimized algorithm of second-order precision. Numerical experimental results show that 3D-PDE-Net can achieve the solution with good accuracy using few training samples, and it is of highly significant in solving linear and nonlinear unsteady PDEs.
In this paper, the event-triggered fixed-time control scheme is developed for a class of stochastic nonlinear systems with prescribed boundary constraints and actuator faults. For the controlled systems with unknown nonlinear functions, the neural networks are employed to reestablish the system model. Based on the event-triggered control technology, an original adaptive fixed-time control strategy with prescribed performance and the fault state is proposed by using the backstepping technique. Under the developed adaptive controller, the tracking error satisfies the predefined boundary functions and all the closed-loop system signals are bounded in probability in a fixed time, and the convergence time is irrelevant to the initial states of the system. The practicability of the designed controller is illustrated by a simulation example.
The prior-based blind image deblurring methods have recently achieved good performance. However, many state-of-art algorithms are time-consuming since some nonlinear operators are involved. Presented in this paper is a fast blind image deblurring algorithm which uses the salience map and gradient cepstrum. The inspiration for this work comes from the fact that the extreme values of the salience map of the clear image are more sparse than those of the blurred one. By enforcing the L-0 norm constraint to the terms involving salience map and incorporating them into the traditional deblurring framework, an effective optimization scheme is explored. Furthermore, gradient cepstrum is used to adjust the number of iterations in each scale and determine the size of the initial kernel. Experimental results illustrate that our algorithm outperforms the state-of-art deblurring algorithms in both benchmark datasets and real blur scenes. Besides, this algorithm greatly shortens the running time since it restrains excessive iterations and does not involve any nonlinear operators.
In order to solve the problems of highly redundant spatial information and motion noise in the heart rate (HR) estimation from facial videos based on remote photoplethysmography (rPPG), this article proposes a novel HR estimation method based on spatial–temporal attention model. First, to reduce the redundant information and strengthen the association relationships of long-range videos, the spatial–temporal facial features are extracted by the 2-D convolutional neural network (2DCNN) and 3-D convolutional neural network (3DCNN), respectively. The aggregation function is adopted to incorporate feature maps into short segment spatial–temporal feature maps. Second, the spatial–temporal strip pooling is designed in the spatial–temporal attention module to reduce head movement noises. Then, via the two-part loss function, the model can focus more on the rPPG signal rather than the interference. We conduct extensive experiments on two public data sets to verify the effectiveness of our model. The experimental results show that the proposed method achieves significantly better performances than the state-of-the-art baselines: The mean absolute error could be reduced by 11% on the PURE data set, and by 25% on the COHFACE data set.
Well test analysis is a crucial technique to monitor reservoir performance, which is based on the theory of seepage mechanics, through the study of well test data, to identify reservoir models and estimate reservoir parameters. Reservoir model recognition is the first and essential step of well test analysis. It is usually judged by professionals’ experience, which results in low efficiency and accuracy. This paper is devoted to applying convolutional neural network (CNN) to well test analysis and proposes a new intelligent reservoir model identification method. Eight reservoir models studied in this paper include homogenous reservoirs with different outer boundaries such as infinite acting boundary, circular, single, angular, channel, U-shaped and rectangular sealing fault boundaries, and a radial composite reservoir with infinite acting boundary. Well testing data used in this paper, including actual field data and theoretical data, are generated by analytical solutions. To improve the classification accuracy of actual field data, noise processing was carried out on the data before training. The CNN that is most suitable for model recognition has been obtained through trial-and-error procedures. The availability of proposed CNN is proved with actual field cases of Daqing oil field, China. The method realizes the automatic identification of reservoir model with the total classification accuracy (TCA) of test data set of 98.68% and 95.18% for original data and noisy data, respectively.
Robot dynamic model has been applied in many fields. The traditional dynamic parameter identification is often based on the linearized model of inverse dynamics. To acquire a more accurate model, it is necessary to reduce acceleration noise and optimize excitation trajectory. In this paper, we propose a differential momentum parameter identification model which avoid the acceleration in inverse dynamics based method and the accumulation errors in momentum based method. The model of differential momentum parameter identification is elaborated in detail. The excitation trajectory is produced by optimizing Hadamard's inequality using genetic algorithm. In order to verify the proposed method, we carry out simulations and experiments of dynamic parameter identification using Franka robot. Results show that the differential momentum parameter identification model can achieve more accurate dynamics compared with traditional inverse dynamics based parameter identification model.
Remote photoplethysmography (rPPG) technology is widely used to measure heart rate (HR) from facial video. However, the accuracy of rPPG signal extraction is affected by the slow content changes in long-range facial videos, and the extraction process is also susceptible to interference from illumination variation and head movement noise. To address the above problems, in this article, we propose an end-to-end effective time-domain attention network (ETA-rPPGNet). First, in order to overcome video redundancy information, we construct the time-domain segment subnet. The video is divided into several segments, which are fed into the subspace networks to extract important spatial facial features and aggregate temporal information, respectively. Then, the time-domain attention mechanism is designed in the backbone net. In this mechanism, the 1-D convolution is used to effectively model the information association in the local time domain, so as to improve the antinoise ability of the model. Finally, via the two-part loss function, the model can reduce the interference of other physiological signals. We conduct extensive experiments to verify the effectiveness of our model on the public PURE, COHFACE, UBFC-rPPG, and MMSE-HR data sets. Compared with other models, our model shows better performance in the measurement accuracy.
目的 单幅图像超分辨率重建的深度学习算法中,大多数网络都采用了单一尺度的卷积核来提取特征(如3×3的卷积核),往往忽略了不同卷积核尺寸带来的不同大小感受域的问题,而不同大小的感受域会使网络注意到不同程度的特征,因此只采用单一尺度的卷积核会使网络忽略了不同特征图之间的宏观联系.针对上述问题,本文提出了多层次感知残差卷积网络(multi-level perception residual convolutional network,MLP-Net,用于单幅图像超分辨率重建).方法 通过特征提取模块提取图像低频特征作为输入.输入部分由密集连接的多个多层次感知模块组成,其中多层次感知模块分为浅层多层次特征提取和深层多层次特征提取,以确保网络既能注意到图像的低级特征,又能注意到高级特征,同时也能保证特征之间的宏观联系.结果 实验结果采用客观评价的峰值信噪比(peak signal to noise ratio,PSNR)和结构相似性(structural similarity,SSIM)两个指标,将本文算法其他超分辨率算法进行了对比.最终结果表明本文算法在4个基准测试集上(Set5、Set14、Urban100和BSD100(Berkeley Segmentation Dataset))放大2倍的平均峰值信噪比分别为37.851 1 dB,33.933 8 dB,32.219 1 dB,32.148 9 dB,均高于其他几种算法的结果.结论 本文提出的卷积网络采用多尺度卷积充分提取分层特征中的不同层次特征,同时利用低分辨率图像本身的结构信息完成重建,并取得不错的重建效果.
As a well-known ill-conditional problem in the image processing field, image deblurring has become a hot topic recently. The prior-based blind image deblurring methods have recently shown promising effectiveness. A lot of advanced algorithms such as dark channel prior, bright channel prior, and local maximum gradient prior are time-consuming since nonlinear operators are involved. Presented in this paper is a fast blind image deblurring algorithm which uses the simplified extreme channel prior (SECP) and gradient cepstrum. The inspiration for this work comes from the fact that the simplified bright channel prior (SBCP) of the clear image has fewer non-one elements than the blurred one. We propose a novel SECP based on the proposed SBCP and the simplified dark channel prior (SDCP). By enforcing the $$L_{0}$$ norm constraint to the terms involving SECP and incorporating them into the traditional deblurring framework, an effective optimization scheme is explored. Furthermore, gradient cepstrum is used to determine the size of the initial kernel and restrain excessive iterations in each scale. Experimental results illustrate that our algorithm outperforms the state-of-the-art deblurring algorithms in terms of computational efficiency and deblurring effect on both benchmark datasets and real-world blur scenes.