In recent years, large visual language models (LVLMs) have shown impressive performance and promising generalization capability in multi-modal tasks, thus replacing humans as receivers of visual information in various application scenarios. In this paper, we pioneer to propose a variable bitrate image compression scheme consisting of a pre-editing module and an end-to-end codec to achieve promising rate-accuracy performance for different LVLMs. In particular, instead of optimizing an adaptive pre-editing network towards a particular task or several representative tasks, we propose a new optimization strategy tailored for LVLMs, which is designed based on the representation and discrimination capability with token-level distortion and rank. The pre-editing module and the variable bitrate end-to-end image codec are jointly trained by the losses based on semantic tokens of the large model, which introduce enhanced generalization capability for various data and tasks. Experimental results demonstrate that the proposed framework could efficiently achieve much better rate-accuracy performance compared to the state-of-the-art coding standard, Versatile Video Coding. Meanwhile, experiments with multi-modal tasks have revealed the robustness and generalization capability of the proposed framework.
The widespread application of power-assist exoskeletons in physical labor and daily activities has increased the demand for robust control strategies to address challenges in human-exoskeleton interaction. Factors such as collisions and friction introduce uncertain disturbances, making it difficult to establish an accurate human-exoskeleton interaction model, thereby limiting the applicability of current model-based control methods. To overcome these problems, this study proposes an improved data-driven model-free adaptive control method (IMFAC) for the upper extremity power-assist exoskeleton. The stability and convergence of the closed-loop system are rigorously proven. To optimize the initial conditions of IMFAC, we propose an improved snake optimizer (ISO) algorithm incorporating opposition-based learning. The proposed ISO-IMFAC method is evaluated in two scenarios: a nonlinear Hammerstein model benchmark and a physical exoskeleton platform. Experimental results demonstrate that ISO-IMFAC outperforms other popular data-driven control methods across six metrics: integrated absolute error (4.756), mean integral of time-weighted absolute error (0.457), maximum error (1.167), minimum error (0), mean error (0.032), and error standard deviation (0.169). Additionally, the ISO-IMFAC method effectively drives the exoskeleton without relying on its dynamic model. In two load-bearing experiments conducted with five subjects wearing the exoskeleton, the proposed method reduces average muscle exertion per unit time by over 50 https://github.com/Shurun-Wang/ISO-IMFAC .
In this paper, we propose a novel framework for Interactive Face Video Coding (IFVC), which allows humans to interact with the intrinsic visual representations instead of the signals. The proposed solution enjoys several distinct advantages, including ultra-compact representation, low delay interaction, and vivid expression/headpose animation. In particular, we propose the Internal Dimension Increase (IDI) based representation, greatly enhancing the fidelity and flexibility in rendering the appearance while maintaining reasonable representation cost. By leveraging strong statistical regularities, the visual signals can be effectively projected into controllable semantics in the three dimensional space (e.g., mouth motion, eye blinking, head rotation, head translation and head location), which are compressed and transmitted. The editable bitstream, which naturally supports the interactivity at the semantic level, can synthesize the face frames via the strong inference ability of the deep generative model. Experimental results have demonstrated the performance superiority and application prospects of our proposed IFVC scheme. In particular, the proposed scheme not only outperforms the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes in terms of rate-distortion performance for face videos, but also enables the interactive coding without introducing additional manipulation processes. Furthermore, the proposed framework is expected to shed lights on the future design of the digital human communication in the metaverse. The project page can be found at https://github.com/Berlin0610/Interactive Face Video Coding.
Accurate finger gesture recognition with surface electromyography (sEMG) is essential and long-challenge in the muscle-computer interface, and many high-performance deep learning models have been developed to predict gestures. For these models, problem-specific tuning of network architecture is essential for improving the performance, yet it requires substantial knowledge of network architecture design and commitment of time and effort. This process thus imposes a major obstacle to the widespread and flexible application of modern deep learning. To address this issue, we present an auto-learning search framework (ALSF) to generate the integrated block-wised neural network (IBWNN) for sEMG-based gesture recognition. IBWNN contains several feature extraction blocks and dimensional reduction layers, and each feature extraction block integrates two sub-blocks (i.e., multi-branch convolutional block and triplet attention block). Meanwhile, ALSF generates optimal models for gesture recognition through the reinforcement learning method. The results show that the generated models yield state-of-the-art results compared to the modern popular networks on the open dataset Ninapro DB5. Moreover, compared to other networks, the generated models have fewer parameters and can be deployed in practical applications with less resource consumption.
Machine analysis of visual signals has been providing fundamental infrastructure support for artificial intelligence (AI) applications. As such, the compact representation of visual signal towards machine analysis is critical in current AI systems. A straightforward and general pipeline of visual signal compression for machines is to compress the visual signal, then perform the machine analysis on the reconstructed visual signal for ultimate results. This pipeline naturally supports multiple machine vision tasks and human perception, and is denoted as visual signal compression towards machines (VSCM). This paper provides an overview to summarize the VSCM related technologies in terms of various modules, and envision the future development of VSCM.
To reduce the decoder burden of performing the video analysis tasks and increase the accuracy of the task results, an effective solution is to perform the video analysis task in the encoder side and send the results to the decoder. In this paper, the object masks which are obtained by invoking a video analysis model in the encoder are represented by object mask pictures and encoded as auxiliary pictures. To enable this idea, a new auxiliary picture type was specified and a new supplemental enhancement information (SEI) message, the object mask information (OMI) SEI message, is introduced. The experiments are conducted by using versatile video coding test model (VTM) layered coding, and according to the experimental results, it is concluded that coding the object mask pictures as auxiliary pictures is feasible and efficient. And thus, the proposed OMI SEI message was adopted to working draft of versatile supplemental enhancement information version 4 and included in the technical report of optimization of encoders and receiving systems for machine analysis of coded video content in January 2024.
BACKGROUND AND OBJECTIVE:The accurate diagnosis of schizophrenia spectrum disorder plays an important role in improving patient outcomes, enabling timely interventions, and optimizing treatment plans. Functional connectivity analysis, utilizing functional magnetic resonance imaging data, has been demonstrated to offer invaluable biomarkers conducive to clinical diagnosis. However, previous studies mainly focus on traditional machine learning methods or hand-crafted neural networks, which may not fully capture the spatial topological relationship between brain regions. METHODS:This paper proposes an evolutionary algorithm (EA) based graph neural architecture search (GNAS) method. EA-GNAS has the ability to search for high-performance graph neural networks for schizophrenia spectrum disorder prediction. Moreover, we adopt GNNExplainer to investigate the explainability of the acquired architectures, ensuring that the model's predictions are both accurate and comprehensible. RESULTS:The results suggest that the graph neural network model, derived using genetic algorithm search, outperforms under five-fold cross-validation, achieving a fitness of 0.1850. Relative to conventional machine learning and other deep learning approaches, the proposed method yields superior accuracy, F1 score, and AUC values of 0.8246, 0.8438, and 0.8258, respectively. CONCLUSION:Based on a multi-site dataset from schizophrenia spectrum disorder patients, the findings reveal an enhancement over prior methods, advancing our comprehension of brain function and potentially offering a biomarker for diagnosing schizophrenia spectrum disorder.
Traditional end-to-end video coding is typically featured with sparsely distributed operational rate-distortion (R-D) points. This creates daunting challenges to rate control which is typically regarded as the indispensable coding optimization module. To tackle this problem, this paper proposes high efficiency rate control for end-to-end scale-adaptive video coding which enables the conversion from sparsely to densely distributed R-D points. The proposed scheme does not increase the number of models in the sparse-to-dense conversion and provides more flexibility in end-to-end video coding thereby leading to better coding performance. More specifically, R-D analyses for scale-adaptive coding are first conducted, shedding light on the design of the rate control algorithm. Subsequently, generalized R-D models are presented, based on which high efficiency rate control is achieved. Extensive experimental results provide evidence of the efficiency of the proposed method in terms of R-D performance, control accuracy and computational complexity..
How to compress face video is a crucial problem for a series of online applications, such as video chat/conference, live broadcasting and remote education. Compared to other natural videos, these face-centric videos owning abundant structural information can be compactly represented and high-quality reconstructed via deep generative models, such that the promising compression performance can be achieved. However, the existing generative face video compression schemes are faced with the inconsistency between the 3D facial motion in the physical world and the face content evolution in the 2D view. To solve this drawback, we propose a 3D-Keypoint-and-2D-Motion based generative method for Face Video Compression, namely FVC-3K2M, which can well ensure perceptual compensation and visual consistency between motion description and face reconstruction. In particular, the temporal evolution of face video can be characterized into separate 3D keypoints from the global and local perspectives, entailing great coding flexibility and accurate motion representation. Moreover, a cascade motion conversion mechanism is further proposed to internally convert 3D keypoints to 2D dense motion, enforcing the face video reconstruction to be perceptually realistic. Finally, an adaptive reference frame selection scheme is developed to enhance the adaptation of various temporal movements. Experimental results show that the proposed scheme can realize reliable video communication in the extremely limited bandwidth, e.g., 2 kbps. Compared to the state-of-the-art video coding standards and the latest face video compression methods, extensive comparisons demonstrate that our proposed scheme achieves superior compression performance in terms of multiple quality evaluations.
There has been an increasing consensus that the machine vision is gradually replacing human vision in numerous tasks, with the demonstrated success of artificial intelligence. In this paper, we propose a deep image compression scheme towards machine vision, with the principle of “begin with the end in mind”. In particular, a unified optimization scheme for end-to-end image compression towards machine vision is proposed, accompanied with the dedicated variable bitrate coding and generalized rate-accuracy optimization. The presented framework, which jointly optimizes the compression and the machine vision networks, exploits the utmost potential of robust machine vision for compressed images. The variable bitrate modules towards machine vision, which effectively shrink the storage space for model parameters, are further developed to accommodate to the real-world applications. Moreover, an iterative algorithm is presented to achieve the optimality in terms of the generalized rate-accuracy towards machine vision. Experimental results show that the proposed framework achieves the state-of-the-art object detection performance among the end-to-end image compression methods: in the exploration of Video Coding for Machines (VCM) in Moving Picture Experts Group (MPEG), and the proposed framework achieves 31.69% and 23.96% BD-rate gains compared with the VCM official test datasets, the Open Images dataset and the TVD dataset respectively, which are generated using the state-of-the-art standard Versatile Video Coding (VVC) standard. The generalization capability of the proposed framework is also verified with instance segmentation under various scenarios.
Muscle fatigue detection is of great significance to human physiological activities, but many complex factors increase the difficulty of this task. In this article, we integrate several effective techniques to distinguish muscle states under fatigue and nonfatigue conditions via surface electromyography (sEMG) signals. First, we perform an isometric contraction experiment of biceps brachii to collect sEMG signals. Second, we propose a neural architecture search (NAS) framework based on reinforcement learning to autogenerate neural networks. Finally, we present an effective two-step training strategy to improve the performance by combining CNN with three types of commonly used statistical algorithms. Meanwhile, we propose a data enhancement algorithm based on empirical mode decomposition (EMD) to generate time-series data for expanding the dataset. The results show that this search algorithm can hunt for high-performing networks, and the accuracy of the best-selected model combined with support vector machine (SVM) for the group is 96.5%. With the same architecture, the average accuracy in individual models is 97.8%. The proposed data enhancement technique can effectively improve the fatigue detection performance, which allows further implementations in the human-exoskeleton interaction systems.
High-quality face images are required to guarantee the stability and reliability of automatic face recognition (FR) systems in surveillance and security scenarios. However, a massive amount of face data is usually compressed before being analyzed due to limitations on transmission or storage. The compressed images may lose the powerful identity information, resulting in the performance degradation of the FR system. Herein, we make the first attempt to study just noticeable difference (JND) for the FR system, which can be defined as the maximum distortion that the FR system cannot notice. More specifically, we establish a JND dataset including 3530 original images and 137,670 compressed images generated by advanced reference encoding/decoding software based on the Versatile Video Coding (VVC) standard (VTM-15.0). Subsequently, we develop a novel JND prediction model to directly infer JND images for the FR system. In particular, in order to maximum redundancy removal without impairment of robust identity information, we apply the encoder with multiple feature extraction and attention-based feature decomposition modules to progressively decompose face features into two uncorrelated components, i.e., identity and residual features, via self-supervised learning. Then, the residual feature is fed into the decoder to generate the residual map. Finally, the predicted JND map is obtained by subtracting the residual map from the original image. Experimental results have demonstrated that the proposed model achieves higher accuracy of JND map prediction compared with the state-of-the-art JND models, and is capable of saving more bits while maintaining the performance of the FR system compared with VTM-15.0.
Intention recognition based on surface electromyography (sEMG) signals is pivotal in human-machine interaction (HMI), where continuous motion estimation with high accuracy has been the challenge. The convolutional neural network (CNN) possesses excellent feature extraction capability. Still, it is difficult for ordinary CNN to explore the dependencies of time-series data, so most researchers adopt the recurrent neural network or its variants (e.g., LSTM) for motion estimation tasks. This paper proposes a multi-feature temporal convolutional attention-based network (MFTCAN) to recognize joint angles continuously. First, we recruited ten subjects to accomplish the signal acquisition experiments in different motion patterns. Then, we developed a joint training mechanism that integrates MFTCAN with commonly used statistical algorithms, and the integrated architectures were named MFTCAN-KNR, MFTCAN-SVR and MFTCAN-LR. Last, we utilized two performance indicators (RMSE and [Formula: see text]) to evaluate the effect of different methods. Moreover, we further validated the performance of the proposed method on the open dataset (Ninapro DB2). When evaluating on the original dataset, the average RMSE of the estimations obtained by MFTCAN-KNR is 0.14, which is significantly less than the results obtained by LSTM (0.20) and BP (0.21). The average [Formula: see text] of the estimations obtained by MFTCAN-KNR is 0.87, indicating the anti-disturbance ability of the architecture. Moreover, MFTCAN-KNR also achieves high performance when evaluating on the open dataset. The proposed methods can effectively accomplish the task of motion estimation, allowing further implementations in the human-exoskeleton interaction systems.
The detection of muscle activation intervals is of great significance to the application of gait, gesture, and some other biomedical movements. Surface electromyographic (sEMG) signal can record the electrical activity of muscles effectively. This kind of signal is sometimes, however, corrupted by background noises. Traditionally, the amplitude analysis of the sEMG signal is an alternative approach for the identification of onsets and offsets. In this work, the sEMG signal was analyzed using global-based rapid composite multiscale sample entropy and compared to the other popular algorithms. A double threshold method with an interlocking structure was applied to complete the onsets and offsets detection task. The proposed algorithm was tested in semisynthetic and recorded signals, and it can take advantage of the nonlinear properties of entropy to distinguish the sEMG signal from motion artifacts and tonic spikes. By using the proposed algorithm, the median values of absolute error time for onsets and offsets estimation were 32 and 60 ms, respectively. Meanwhile, the accuracy, false-alarm rate, and missing-alarm rate were 86.1%, 5.7%, and 9.7%, respectively. Our findings suggest that the proposed method can effectively detect muscle activation intervals.
Compactly representing visual information plays a fundamental role in optimizing the ultimate utility of myriad visual data-centered applications. Numerous approaches have been proposed to efficiently compress the texture and visual features for human visual perception and machine intelligence, respectively; however, much less work has been dedicated to studying the interactions between them. Here, we investigate the integration of feature and texture compression and show that a universal and collaborative visual information representation can be achieved in a hierarchical way. In particular, we study feature and texture compression in a scalable coding framework, where the base layer serves as the deep learning feature and the enhancement layer targets to perfectly reconstruct the texture. Based on the strong generative capability of deep neural networks, the gap between the base feature layer and enhancement layer is further filled with feature-level texture reconstruction, with the goal of further constructing texture representations from features. As such, the residuals between the original and reconstructed texture could be further conveyed in the enhancement layer. To improve the efficiency of the proposed framework, the base layer neural network is trained in a multitask manner such that the learned features enjoy both high-quality reconstruction and high-accuracy analysis. The framework and optimization strategies are further applied in face image compression, and promising coding performance has been achieved in terms of both rate-fidelity and rate-accuracy evaluations.
The analysis of human muscle fatigue is of great significance to human physiological activities. Surface electromyography (sEMG) is the widely used technique to analyze muscle fatigue due to its non-invasiveness. However, sEMG signals are complex, non-linear, and multiple time scales. Hence, multiscale entropy is an effective method to quantify the process of muscle fatigue. In this study, we proposed rapid refined composite multiscale sample entropy (R2CMSE) and used it to analyze the process of muscle fatigue. Firstly, R2CMSE was utilized to characterize the complexity and validity of noise signals on different scales and lengths. Secondly, isometric contraction activities of the biceps brachii were recorded by using sEMG from ten subjects. Then, combined with the sEMG signals of all subjects, a three-dimensional map of scale-length-entropy was constructed to determine an appropriate time scale and data length. Meanwhile, a two-way repeated-measures ANOVA was operated, and the results showed that non-fatigue and fatigue conditions exist significant differences under all algorithms. Finally, R2CMSE was used to quantify the fatigue process to analyze the reliability of its application to different subjects. It was shown that compared with other multiscale entropy algorithms, R2CMSE was faster in the calculation, less dependent on different data lengths, and more robust at different time scales. The proposed algorithm can also extract the hidden information of sEMG signals and investigate the process of muscle fatigue more effectively.
In this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded in an end-to-end manner by optimizing the rate- distortion cost to achieve feature-in-feature representation. The multi-granularity constraint is further imposed, serving as the optimization objective to make the feature compression more "healthier" from the perspective of ultimate utility. More specifically, the analysis accuracy is considered in the coarse granularity level constraint, ensuring the capability of facial analysis with the reconstructed feature. Furthermore, at the fine granularity level the feature fidelity is involved to preserve the original feature quality. Moreover, a latent code level teacher-student enhancement model is proposed to efficiently transfer the low bit-rate representation into a high bit- rate one. Such a strategy further allows us to adaptively shift the representation cost to decoding computations, leading to more flexible feature compression with enhanced decoding capability. We verify the effectiveness of the proposed model with the facial feature, and experimental results reveal better compression performance in terms of rate-accuracy compared with existing models.
There has been an increasing consensus that precise understanding of the rate-distortion (RD) characteristics plays a critical role in image and video coding. In this paper, we explore the RD behaviors of end-to-end image compression in the real-world application scenario that the images could be corrupted by noise at different levels. With the RD behaviors that all images share, we develop a deep learning driven pre-analytical model which fully exploits the properties of RD functions and allows us to improve the quality with economized coding bits. The proposed approach does not require any prior knowledge of the noise level, and could effectively defend against the noise through the end-to-end compression. Extensive experimental results show that the proposed scheme offers the best promise in predicting RD behaviors, and naturally avoids the unnecessary bits consumption.
The visual signal compression is a long-standing problem. Fueled by the recent advances of deep learning, exciting progress has been made. Despite better compression performance, existing end-to-end compression algorithms are still designed towards better signal quality in terms of rate-distortion optimization. In this paper, we show that the design and optimization of network architecture could be further improved for compression towards machine vision. We propose an inverted bottleneck structure for the encoder of the end-to-end compression towards machine vision, which specifically accounts for efficient representation of the semantic information. Moreover, we quest the capability of optimization by incorporating the analytics accuracy into the optimization process, and the optimality is further explored with generalized rate-accuracy optimization in an iterative manner. We use object detection as a showcase for end-to-end compression towards machine vision, and extensive experiments show that the proposed scheme achieves significant BD-rate savings in terms of analysis performance. Moreover, the promise of the scheme is also demonstrated with strong generalization capability towards other machine vision tasks, due to the enabling of signal-level reconstruction.
In this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded in an end-to-end manner by optimizing the rate-distortion cost to achieve feature-in-feature representation. In order to further improve the compression performance, we present a latent code level teacher-student enhancement model, which could efficiently transfer the low bit-rate representation into a high bit rate one. Such a strategy further allows us to adaptively shift the representation cost to decoding computations, leading to more flexible feature compression with enhanced decoding capability. We verify the effectiveness of the proposed model with the facial feature, and experimental results reveal better compression performance in terms of rate-accuracy compared with existing models.