In multimodal emotion recognition, modality missing significantly degrades the recognition performance. Cross-Modal missing modality imagination is one of the mainstream solutions. However, with limited available modalities, it still faces two challenges: (i) effectively addressing the "modality gap" in cross-modal imagination and (ii) solving the "information scarcity" in cross-modal interaction. To address the first challenge, we propose a Cross Encoding Temporal Invariant Feature learning method which can effectively bridge modality gaps by learning temporal invariant features between paired modalities. To tackle the second challenge, we develop a novel Arbitrary Modality Enhancement strategy, which can fully exploit the knowledge of available modalities by enabling flexible mutual information enhancement between any modalities to overcome information scarcity. Finally, based on the two approaches, we design TIFAE, an architecture for multimodal emotion recognition with missing modalities. Experiments on three popular datasets show that TIFAE outperforms existing methods in recognition accuracy and robustness.
Anomaly detection is a critical task in industrial manufacturing, and leveraging artificial intelligence to identify product anomalies is essential for significantly enhancing production efficiency. However, most existing approaches follow the one-to-one paradigm, where a customized model is trained for each category, incurring substantial computational and memory costs. Although some methods have emerged for universal anomaly detection in recent years, they usually require carefully designed text prompts or have slow inference speeds. Moreover, most anomaly detection methods lack the capability of fine-grained anomaly classification, necessitating additional training of classification models for practical applications with different categories. To address these challenges, we propose YOLOSAM, a unified and efficient anomaly detection model based on auto mask prompt. YOLOSAM is a dual-branch architecture that can handle multi-class few-shot anomaly detection with a unified model, including both segmentation and classification branches. In the segmentation branch, we design an auto mask prompt generator that generates mask prompts directly from visual information, eliminating the need for complex prompt engineering. In the detection branch, we design a defect detection head that utilizes the visual information to achieve fine-grained anomaly classification. Additionally, we employed knowledge distillation techniques to compress the image encoder, and both branches share this distilled encoder, effectively preserving SAM’s general knowledge while significantly enhancing the inference speed. YOLOSAM achieved anomaly classification and segmentation results of 95.6
Gradient Orthogonal Projection (GOP) is an efficient strategy in continual learning to mitigate catastrophic forgetting. Despite its success so far, GOP-based methods often suffer from the learning capacity degradation problem with an increasing number of tasks. To address this problem, we propose a novel and plug-and-play method to learn new tasks in low-coherence subspaces rather than orthogonal subspaces. Specifically, we construct a unified cost function with the DNN parameters lying on the Oblique manifold. A corresponding gradient descent algorithm is developed to jointly minimize the cost function that involves both inter-task and intra-task coherence. We then provide a theoretical analysis to show the advantages of the proposed in the stability and plasticity. Experimental results show that the proposed method has prominent advantages in maintaining the learning capacity, when the number of tasks increases, especially on a large number of tasks, compared with baselines.
Time series forecasting plays an important role in various fields, such as energy, finance, transport, and weather.Temporal convolutional networks (TCNs) based on dilated causal convolution have been widely used in time series forecasting.However, two problems weaken the performance of TCNs.One is that in dilated casual convolution, causal convolution leads to the receptive fields of outputs being concentrated in the earlier part of the input sequence, whereas the recent input information will be severely lost.The other is that the distribution shift problem in time series has not been adequately solved.To address the first problem, we propose a subsequence-based dilated convolution method (SDC).By using multiple convolutional filters to convolve elements of neighboring subsequences, the method extracts temporal features from a growing receptive field via a growing subsequence rather than a single element.Ultimately, the receptive field of each output element can cover the whole input sequence.To address the second problem, we propose a difference and compensation method (DCM).The method reduces the discrepancies between and within the input sequences by difference operations and then compensates the outputs for the information lost due to difference operations.Based on SDC and DCM, we further construct a temporal subsequence-based convolutional network with difference (TSCND) for time series forecasting.The experimental results show that TSCND can reduce prediction mean squared error by 7.3% and save runtime, compared with state-of-the-art models and vanilla TCN.
Due to its complexity and privacy concerns, medical data is often difficult to collect fully at once. New data emerges with the discovery of new diseases and advances in medical technology, but privacy concerns limit its storage. When deep learning models require training on datasets of new diseases, they often suffer from catastrophic forgetting, where the models’ performance on previously learned data degrades significantly. Recently, pre-trained continual learning methods have been proposed. Unlike earlier continual learning methods, these new approaches do not require learning excessive amounts of new knowledge, furthermore, they can still achieve good performance even when the pre-trained feature extractor is completely frozen. For these approaches, the performance degradation is not primarily due to catastrophic forgetting, but rather confusion. To combat confusion, we propose a method using a buffer with feature tokens (BWFT) which is cheaper and safer than directly storing the original data. Additionally, we propose a feature merge method to further reduce storage and computation costs. In the class-incremental scenario, our method outperforms state-of-the-art methods on multiple medical datasets as well as natural image datasets. Code is available at https://github.com/1345475677/BWFT.
Crowd counting and localization have become increasingly important in computer vision due to their wide-ranging applications. While point-based strategies have been widely used in crowd counting methods, they face a significant challenge, i.e., the lack of an effective learning strategy to guide the matching process. This deficiency leads to instability in matching point proposals to target points, adversely affecting overall performance. To address this issue, we introduce an effective approach to stabilize the proposal-target matching in point-based methods. We propose Auxiliary Point Guidance (APG) to provide clear and effective guidance for proposal selection and optimization, addressing the core issue of matching uncertainty. Additionally, we develop Implicit Feature Interpolation (IFI) to enable adaptive feature extraction in diverse crowd scenarios, further enhancing the model's robustness and accuracy. Extensive experiments demonstrate the effectiveness of our approach, showing significant improvements in crowd counting and localization performance, particularly under challenging conditions. The source codes and trained models will be made publicly available.
Understanding an opposing player's behaviours and weaknesses is often the key to winning a badminton game. This study presents a system to extract game data from broadcast badminton videos, and visualize the extracted data to help coaches and players develop effective tactics. Specifically, we apply state‐of‐the‐art machine learning methods to partition a broadcast video into segments, in which each video segment shows a badminton rally. Next, we detect players' feet in each video frame and transform the player positions into the court coordinate system. Finally, we detect hit frames in each rally, in which the shuttle moves towards the opposite directions. By visualizing the extracted data, our system conveys when and where players hit the shuttle in historical games. Since players tend to smash or drop shuttles under a specific location, we provide users with interactive tools to filter data and focus on the distributions conditioned by player positions. This strategy also reduces visual clutter. Besides, our system plots the shuttle hitting distributions side‐by‐side, enabling visual comparison and analysis of player behaviours under different conditions. The results and the use cases demonstrate the feasibility of our system.
Amplitude-integrated electroencephalography (aEEG) is widely adopted for recognizing neonatal neurological disorders in clinics. Previous work has mainly analyzed aEEGs from a time series perspective, while clinicians are more concerned about images. This paper studies the efficacy of image representations in aEEG interpretation. To this end, we employ Amplitude-frequency contour map (AFCM) to express information on local amplitude distributions and overall temporal trend. During its generation, Multi-scale Box-Cox transformation is introduced to flexibly control the degree of amplitude compression and capture details in different amplitude regions. Moreover, one- and two-dimensional features are extracted and combined to enhance model predictions. The experimental results show the potential of image representation for aEEG recognition. This method can provide new ideas to analyze time series signals like aEEGs from the perspective of combining time series and image.
This paper reviews the NTIRE 2023 challenge on efficient single-image super-resolution with a focus on the proposed solutions and results. The aim of this challenge is to devise a network that reduces one or several aspects such as runtime, parameters, FLOPs, activations, memory footprint, and depth of RFDN while at least maintaining the PSNR of 29.00dB on DIV2K validation datasets. The challenge had 272 registered participants, and 35 teams made valid submissions. They gauge the state-of-the-art for efficient single-image super-resolution.
Crowd counting has recently attracted significant attention in the field of computer vision due to its wide applications to image understanding. Numerous methods have been proposed and achieved state-of-the-art performance for real-world tasks. However, existing approaches do not perform well under adverse weather such as haze, rain, and snow since the visual appearances of crowds in such scenes are drastically different from those images in clear weather of typical datasets. In this paper, we propose a method for robust crowd counting in adverse weather scenarios. Instead of using a two-stage approach that involves image restoration and crowd counting modules, our model learns effective features and adaptive queries to account for large appearance variations. With these weather queries, the proposed model can learn the weather information according to the degradation of the input image and optimize with the crowd counting module simultaneously. Experimental results show that the proposed algorithm is effective in counting crowds under different weather types on benchmark datasets. The source code and trained models will be made available to the public.
In recent years, segmentation of the multimodal brain tumor image puts forward high requirements for performance. To meet the accuracy requirements, we propose a multimodal brain tumor image segmentation method based on UNet and deep residual learning, which can make full use of multimodal information in order to obtain details. The model achieves the information extraction and integration of multimodal images at different scales by introducing a multiscale feature extraction module, a coordinate attention module, and a dense atrous spatial pyramid pooling module. The results show that the proposed method can fully exploit the multi-scale context information of multimodal images, alleviate the problem of blurred edges in segmentation regions, and achieve good performance on popular evaluation metrics.
Human pose estimation is viewed as a crucial step for understanding human behaviour. Although significant progress has been made in this area in recent years, most studies have focused on feature-level information fusion, while decision-level information fusion has rarely been explored. Compared with feature-level information, decision-level information contains more semantic and interpretable information and can help improve the performance of pose estimation in occluded and crowded scenes. In this paper, we focus on the fusion of decision-level information. We propose a View Fusion module for aggregating decision-level information from different stages to generate a more comprehensive estimation. An Auxiliary Task module is introduced to bridge the gap between the feature extractor and the View Fusion module and to provide prior information about the form of the decision-level information. Considering that the precision of predictions from different stages varies, we use different strategies to guide the learning process. Experiments show that our models outperform previous methods and achieve competitive results on the CrowdPose test set. Further experiments indicate that our method is flexible and can improve the performance of various backbones.
Ultraviolet photodetectors (UVPDs) based on Si-Zn-SnO (SZTO) thin-film transistors (TFTs) with a stacked dual-channel layer (DCL) structure with different carrier concentration and NiO capping layer (CL) to alleviate the trade-off between dark current ( I dark ) and photocurrent ( I ph ) are reported. Experimental results show that under 275 nm irradiation, the proposed SZTO TFT UVPD with a 30 nm thick upper layer stacked on a 50 nm thick channel layer and a patterned NiO CL exhibit excellent photoresponsivity and photosensitivity up to 1672 A W −1 and 1.03 × 10 7 A A −1 , which is about 272 and 137 times higher than conventional 30 nm thick single-channel layer SZTO TFT. These improvements are due to the use of a DCL which forms a high-low junction to reduce the effective channel thickness and increasing the space for UV illumination and the use of NiO CL lowers the I dark and causes a considerable negative threshold voltage shift under UV irradiation to significantly boost the I ph .
Single-image shadow removal aims to remove undesired shadow information from captured images. With the development of deep convolutional neural networks, several methods have been proposed to achieve promising performance in shadow removal. However, they still struggle with limited performance due to the non-homogeneous intensity distribution of the shadow. To address this issue, we propose a two-stage shadow removal architecture based on the transformer called TSRFormer. The proposed architecture is divided into shadow removal and content refinement networks. These two stages adopt different transformer architectures and remove the shadow based on different information to achieve effective shadow removal. Experiments performed on challenging benchmark show that the proposed model achieves the 2 nd highest SSIM in the NTIRE 2023 Image Shadow Removal Challenge. The source code will be public after the acceptance of this paper.
High-resolution non-homogeneous dehazing aims to generate a clear image from a 4000 × 6000 image with non-homogeneous haze. To the best of our knowledge, this task is a new challenge that was not addressed in the previous literature. To address this issue, we propose semantic-guided loss functions for high-resolution non-homogeneous dehazing. We find semantic information contains strong texture and color prior. Thus, we proposed to adopt the pre-trained model to generate the semantic mask to guide the neural network during the training phase. On the other hand, to handle the non-homogeneous dehazing process in the high-resolution scenario, we adjust the kernel size of the model to increase the receptive field. Furthermore, to deal with the different image sizes during the training and the testing phase, several post-processing methods are applied to improve the high-resolution non-homogeneous dehazing. Several experiments performed on challenging benchmark show that the proposed model achieves competitive performance in the NTIRE 2023 HR NonHomogeneous Dehazing Challenge.
Automatic video seizure detection can assist in the diagnosis of seizures. Existing works usually rely on the information of single granularity for learning seizure patterns and the performance of these works is often constrained by the limited annotated data. To overcome these problems, we introduce a multi-granularity information flow method to enhance skeleton-based infant seizure detection from different perspectives. A forward-backward consistency constraint is proposed to use the reverse temporal sequences as augmented data, helping make better predictions. Besides, an importance-aware module is proposed to adaptively pay more attention to important frames. Experimental results indicate that our method can improve the performance of seizure detection and achieve high performance with a relatively low amount of parameters.
The purpose of product entity matching tasks is to identify whether different product entity data entries in fact refer to the same real-world product entity. Such tasks are crucial to the collation of entity data. Existing models have only considered the one-to-one matching of goods, even though both one-to-one and many-to-one cases exist in offline entity matching. In this regard, this study proposed the combination of a supervised matching model based on ALBERT and an unsupervised model based on Fast-text to solve the matching tasks in both the above-mentioned situations. Through such a joint approach, the advantages of the multi-head self-attention mechanism of ALBERT in selecting candidates with consistent key information and then the merits of the Fast-text Matcher in capturing the details of the data with high accuracy were utilized to discover the matching candidates with the highest possibilities. In addition, lowering the cost was also one of the objectives of this study. The experimental results showed that the proposed ALFA Matcher neural network model outperformed other baseline models in performing the two tasks, effectively solving the two major problems of offline entity matching. Furthermore, both ALBERT and Fast-text have the characteristics of fewer training parameters and faster training speed, which make them convenient to deploy and use in practical applications.
This paper reviews the third biennial challenge on spectral reconstruction from RGB images, i.e., the recovery of whole-scene hyperspectral (HS) information from a 3-channel RGB image. This challenge presents the "ARAD_1K" data set: a new, larger-than-ever natural hyperspectral image data set containing 1,000 images. Challenge participants were required to recover hyper-spectral information from synthetically generated JPEG-compressed RGB images simulating capture by a known calibrated camera, operating under partially known parameters, in a setting which includes acquisition noise. The challenge was attended by 241 teams, with 60 teams com-peting in the final testing phase, 12 of which provided de-tailed descriptions of their methodology which are included in this report. The performance of these submissions is re-viewed and provided here as a gauge for the current state-of-the-art in spectral reconstruction from natural RGB images.
Depth estimation using stereo images can be achieved by calculating the disparity values between the left and the right images captured by two parallel cameras. Reconstructing depth information from 2D images is crucial in many applications, such as self-driving vehicles and robot navigation. Furthermore, most of these applications are employed with resource-constrained devices and have real-time requirements. In this paper, a high-speed, low-cost hardware implementation for disparity estimation is proposed. We adopted the novel disparity fusion method in our architecture, which can significantly reduce the number of calculations in the overall process. A refinement method is also designed to reduce the error rate of the resulting depth map and improve the tolerance to light noise. The proposed algorithm was implemented with the Kintex-7 field-programmable gate array. Its performance was tested by using the Middlebury-Version 2 and -Version 3 datasets. The proposed algorithm provides an operating speed of 118 fps with an error rate of only 6.36%. Compared with other state-of-the-art designs used for similar applications, the proposed method can achieve a 34.6% reduction in the error rate while providing the highest speed with competitive hardware cost.
Ming-Hsuan Yang合作论文数Vision and Learning Lab, University of California, Merced;Google DeepMind2