Under individual smoothness, the optimal incremental first-order oracle (IFO) complexity of nonconvex finite-sum optimization has remained open. Known algorithms use O(n+√(n) ΔL_max/ε^2) calls, while prior lower bounds miss a factor of √(n). We prove the matching lower bound for randomized IFO algorithms whose component indices and query points may depend on the complete preceding transcript and private randomness. This determines the minimax IFO complexity up to universal constants under both individual and mean-squared smoothness. Under the global Polyak-Lojasiewicz (PL) condition, the standard PAGE guarantee is not tight when κ_ms<√(n). Restarted PAGE attains O(n+nlog(Δ/ε)/(1+log(√(n)/κ_ms))) for 1_ms≤√(n), and O(n+κ_ms√(n)log(Δ/ε)) for κ_ms≥√(n). We prove matching lower bounds under individual smoothness for every κ_max≥ 3; the same hard instances also give the mean-squared lower bounds. In the small-κ_max range, their average objective is globally strongly convex. Our lower bounds use dense weak hiding. A fixed sign table spreads each hidden direction across the components. Each queried row carries little information, while the exact row average preserves the full signal after rescaling. A bounded radial map handles arbitrary query points, and a smooth gate makes unopened links invisible to both function values and gradients. Balancing the rows needed to reveal one stage with the number of stages allowed by individual smoothness yields the missing √(n) factor.
We study the problem of outlier correspondence pruning for non-rigid point cloud registration. In rigid registration, spatial consistency has been a commonly used criterion to discriminate outliers from inliers. It measures the compatibility of two correspondences by the discrepancy between the respective distances in two point clouds. However, spatial consistency no longer holds in non-rigid cases and outlier rejection for non-rigid registration has not been well studied. In this work, we propose Graph-based Spatial Consistency Network (GraphSCNet) to filter outliers for non-rigid registration. Our method is based on the fact that non-rigid deformations are usually locally rigid, or local shape preserving. We first design a local spatial consistency measure over the deformation graph of the point cloud, which evaluates the spatial compatibility only between the correspondences in the vicinity of a graph node. An attention-based non-rigid correspondence embedding module is then devised to learn a robust representation of non-rigid correspondences from local spatial consistency. Despite its simplicity, GraphSCNet effectively improves the quality of the putative correspondences and attains state-of-the-art performance on three challenging benchmarks. Our code and models are available at https://github.com/qinzheng93/GraphSCNet.
We study the problem of outlier correspondence pruning for non-rigid point cloud registration. In rigid registration, spatial consistency has been a commonly used criterion to discriminate outliers from inliers. It measures the compatibility of two correspondences by the discrepancy between the respective distances in two point clouds. However, spatial consistency no longer holds in non-rigid cases and outlier rejection for non-rigid registration has not been well studied. In this work, we propose Graph-based Spatial Consistency Network (GraphSCNet) to filter outliers for non-rigid registration. Our method is based on the fact that non-rigid deformations are usually locally rigid, or local shape preserving. We first design a local spatial consistency measure over the deformation graph of the point cloud, which evaluates the spatial compatibility only between the correspondences in the vicinity of a graph node. An attention-based non-rigid correspondence embedding module is then devised to learn a robust representation of non-rigid correspondences from local spatial consistency. Despite its simplicity, GraphSCNet effectively improves the quality of the putative correspondences and attains state-of-the-art performance on three challenging benchmarks. Our code and models are available at https://github.com/qinzheng93/GraphSCNet.
We study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods have shown great potential through bypassing the detection of repeatable keypoints which is difficult to do especially in low-overlap scenarios. They seek correspondences over downsampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer, or GeoTransformer for short, to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it invariant to rigid transformation and robust in low-overlap cases. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to $100$ times acceleration. Extensive experiments on rich benchmarks encompassing indoor, outdoor, synthetic, multiway and non-rigid demonstrate the efficacy of GeoTransformer. Notably, our method improves the inlier ratio by $18{\sim}31$ percentage points and the registration recall by over $7$ points on the challenging 3DLoMatch benchmark. Our code and models are available at \url{https://github.com/qinzheng93/GeoTransformer}.
In this work, we present Cascaded Visual-Geometric Encoding (CasViGE), a novel inter-modality feature learning method to improve point cloud registration via the visual information from RGB images. Registering point clouds based on solely the 3D geometric structure suffers from severe matching ambiguity caused by repeated geometric patterns and geometrically less discriminative regions. For this reason, recent methods attempt to inject the visual information from RGB images to obtain more accurate correspondences. They usually first extract the visual and the geometric features independently and then fuse them by point-wise concatenation. However, as 2D and 3D convolutions have different inductive biases, this simplistic method ignores the intrinsic correlation between the two modalities, which harms the distinctiveness of the point descriptors. To address this issue, CasViGE iteratively fuses the inter-modality features by leveraging the inductive biases of both 2D and 3D convolutions, which better considers the correlation between the two modalities. A geometry-centric encoding module first promotes the visual features with 3D convolutions in the geometric space to explicitly embed the geometric information. Next, a vision-centric encoding module casts the fused features back to the image space and enforces the local feature saliency and correlation with 2D convolutions. Extensive experiments on 3DMatch and 3DLoMatch benchmarks have demonstrated the efficacy of our method. Our CasViGE is plug-and-play and attains significant improvements consistently on various point cloud registration methods.
We study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods have shown great potential through bypassing the detection of repeatable keypoints which is difficult to do especially in low-overlap scenarios. They seek correspondences over downsampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer, or GeoTransformer for short, to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it invariant to rigid transformation and robust in low-overlap cases. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Extensive experiments on rich benchmarks encompassing indoor, outdoor, synthetic, multiway and non-rigid demonstrate the efficacy of GeoTransformer. Notably, our method improves the inlier ratio by $18{\sim }31$18∼31 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark.
The longitudinal vibration of the multi-layer winding hoisting wire rope is inevitable due to the time-varying dead weight and the fluctuation of hoisting acceleration. Longitudinal vibration will lead to relative sliding between the hoisting wire rope and the rope groove, resulting in friction and wear, and thus threaten the service life of the wire rope. In order to investigate the influence of longitudinal vibration on the tribological characteristics of the wire rope and its wear mechanism, the friction and wear tests between the sliding wire rope and the fixed wire rope groove under different longitudinal vibration amplitudes and frequencies were carried out with the help of a self-developed test rig. Results show that the coefficient of friction (CoF) of the wire rope has experienced three stages of rapid growth, "concave" transition and relative stability with the increase of sliding distance, which increases first and then decreases with increasing the amplitude in the relatively stable stage and has no obvious change with an increase in the frequency. Additionally, the maximum temperature rises at the wind-out region of the sliding wire rope and the middle of the contact wire ropes increase first and then decrease with the increase of the amplitude in the relatively stable stage, but increase gradually with an increase in the frequency. The vibration intensities in the wear regions of the wire ropes have a great increase in the rapid wear stage. Furthermore, the main vibration wear mechanisms of the wire rope are abrasive wear and adhesive wear, the content of O element in the furrows and the abrasive adhesion regions are higher, and the frequency has less effect on the oxidation degree of the worn surface of the wire rope than the amplitude.
We study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods bypass the detection of repeatable keypoints which is difficult in low-overlap scenarios, showing great potential in registration. They seek correspondences over down-sampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it robust in low-overlap cases and invariant to rigid transformation. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Our method improves the inlier ratio by 17∼30 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark. Our code and models are available at https://github.com/qinzheng93/GeoTransformer.
Emotion recognition in conversation (ERC) aims to detect the emotion in a conversation, which has drawn increasing interests due to its widely applications. Current methodologies mainly endeavor to capture a good representation of conversation context. However, we argue that the conversation context are not always consistent with the emotion evolution. This incongruity can greatly restrict the recognition performance. To address aforementioned challenges, in this paper, we propose an emotion evolution network for emotion recognition in conversation (E2Net). Specifically, a speaker-aware modeling methodology is firstly constructed to fuse the utterance from conversations. We employ the gated recurrent unit (GRU) encodes the utterance sequentially. For encoding the interaction between speakers, a listener state is introduced to aid in analyzing conversation context. Then, a Transformer-based method is proposed to capture the emotion evolution accompanying with the emotion transformation matrix. To demonstrate the superior performance of our proposed method, extensive experiments are conducted on four REC datasets and the experimental results suggest that our method is effective and outperforms the current state-of-the-art methods on multiple datasets.
Wound image segmentation is a critical component for the clinical diagnosis and in-time treatment of wounds. Recently, deep learning has become the mainstream methodology for wound image segmentation. However, the pre-processing of the wound image, such as the illumination correction, is required before the training phase as the performance can be greatly improved. The correction procedure and the training of deep models are independent of each other, which leads to sub-optimal segmentation performance as the fixed illumination correction may not be suitable for all images. To address aforementioned issues, an end-to-end dual-view segmentation approach was proposed in this paper, by incorporating a learn-able illumination correction module into the deep segmentation models. The parameters of the module can be learned and updated during the training stage automatically, while the dual-view fusion can fully employ the features from both the raw images and the enhanced ones. To demonstrate the effectiveness and robustness of the proposed framework, the extensive experiments are conducted on the benchmark datasets. The encouraging results suggest that our framework can significantly improve the segmentation performance, compared to the state-of-the-art methods.
We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice and face. These methods neglect the semantic content of different videos, introducing false-negative pairs as training noise. Furthermore, the positive pairs are constructed based on the natural correlation between audio clips and visual frames. However, this correlation might be weak or inaccurate in a large amount of real-world data, which leads to deviating positives into the contrastive paradigm. To address these issues, we propose the cross-modal prototype contrastive learning (CMPC), which takes advantage of contrastive methods and resists adverse effects of false negatives and deviate positives. On one hand, CMPC could learn the intra-class invariance by constructing semantic-wise positives via unsupervised clustering in different modalities. On the other hand, by comparing the similarities of cross-modal instances from that of cross-modal prototypes, we dynamically recalibrate the unlearnable instances' contribution to overall loss. Experiments show that the proposed approach outperforms state-of-the-art unsupervised methods on various voice-face association evaluation protocols. Additionally, in the low-shot supervision setting, our method also has a significant improvement compared to previous instance-wise contrastive learning.
The main shaft device (MSD) is a key component of the mine hoisting system. During lifting, vibration of the system will cause dynamic tension of wire rope at both ends, resulting in the complex stress state at the critical location of the MSD and thus considerably decreasing its service life. This study is aimed at investigating time-dependent reliability of the MSD based on fatigue damage accumulation. Firstly, the actual service-loading history of wire rope during lifting was simulated. Then, the residual strength degradation models of the MSD following different shifts were established based on Palmgren-Miner hypothesis. Furthermore, the performance function of time-dependent reliability considering residual strength degradation was established. Finally, time-dependent reliability of the MSD was evaluated by adopting moment-based saddlepoint approximation. The results show that when the MSD follows a heavy shift, the corresponding reliability drops rapidly with the passage of service time. In addition, a reliable shift of the MSD after gridding was made. According to this shift, the remaining reliable service life of the MSD for multi-shift can be predicted.
Audio classification aims to discriminate between different audio signal types, and it has received intensive attention due to its wide applications. In deep learning-based audio classification methods, researchers usually transform the raw signal of audios into different feature representations (such as Short Time Fourier Transform and Mel Frequency Cepstral Coefficients) as the inputs of networks. However, selecting the feature representation requires expert knowledge and extensive experimental verification. Besides, using a single type of feature representation may cause suboptimal results as the information implied in different kinds of feature representations may be complementary. Previous works show that ensembling the networks trained on different representations can greatly boost classification performance. However, making inferences using multiple networks is cumbersome and computation expensive. In this paper, we propose a novel end-to-end collaborative training framework for the audio classification task. The framework takes multiple representations as inputs to train the networks jointly with a knowledge distillation method. Consequently, our framework significantly promotes the performance of networks without increasing the computational overhead in the inference stage. Extensive experimental results demonstrate that the proposed approach improves classification performance and achieves competitive results on both acoustic scene classification tasks and general audio tagging tasks.
Multi-turn dialogue is challenging because semantic information is not only contained in the current utterance, but also in the dialogue context. In fact, understanding multiturn dialogue is a dynamic process. With the increase of dialogue turn, users' understanding is also changing. In this case, we propose a network based on utterance hidden state transfer for task-oriented dialogue (USET). In our model, we first extract the hidden state of previous utterance as the previous comprehension. Then, this comprehension is passed to the next turn. We take the comprehension as a prior knowledge to understand the semantic information in the dialogue context. Finally, we put the previous comprehension and current utterance together to understand current utterance. In order to realize the transfer of comprehension in dialogue, we propose a continuous sample training method: CST, which takes a multi-turn dialogue as a whole to understand. All the sentences in a dialogue are put into a batch for training. Our method makes use of previous comprehension and achieves the information exchange among dialogue. Experimental results on Stanford Multi-Domain dataset demonstrate that our model is superior to existing models. Code is available at https://github.com/season1blue/USET
Despite the fact that GPUs and accelerators are more efficient in deep learning (DL), commercial clouds like Facebook and Amazon now heavily use CPUs in DL computation because there are large numbers of CPUs which would otherwise sit idle during off-peak periods. Following the trend, CPU vendors have not only released high-performance many-core CPUs but also developed efficient math kernel libraries. However, current DL platforms cannot scale well to a large number of CPU cores, making many-core CPUs inefficient in DL computation. We analyze the memory access patterns of various layers and identify the root cause of the low scalability, i.e., the per-layer barriers that are implicitly imposed by current platforms which assign one single instance (i.e., one batch of input data) to a CPU. The barriers cause severe memory bandwidth contention and CPU starvation in the access-intensive layers (like activation and BN). This paper presents a novel approach called ParaX, which boosts the performance of DL on many-core CPUs by effectively alleviating bandwidth contention and CPU starvation. Our key idea is to assign one instance to each CPU core instead of to the entire CPU, so as to remove the per-layer barriers on the executions of the many cores. ParaX designs an ultralight scheduling policy which sufficiently overlaps the access-intensive layers with the computeintensive ones to avoid contention, and proposes a NUMA-aware gradient server mechanism for training which leverages shared memory to substantially reduce the overhead of per-iteration parameter synchronization. We have implemented ParaX on MXNet. Extensive evaluation on a two-NUMA Intel 8280 CPU shows that ParaX significantly improves the training/inference throughput for all tested models (for image recognition and natural language processing) by 1.73x similar to 2.93x.
Recent dialogue state tracking (DST) usually treats utterance, system action and ontology equally to estimate the slot types and values. In this way, the expression of slot in utterance is restricted. As the main way to directly express user semantics, utterance should receive further attention and its proportion in semantic expression should be dynamic according to the content. It’s common to recognize the different importance of information in all the DST models. However, most of them pay little attention to the position of slot in utterance. In fact, position and semantics are related due to human grammatical habits and expression habits. Therefore, we propose T-Mask 1 , a model to actively and accurately learn the token mask position of slot, and we further utilize the learned position information to influence the semantic expression of utterance. We verify the effectiveness of our model on DSTC2 and WoZ2.0. On WoZ2.0, we achieve 90.84 joint goal accuracy and 97.6 turn request accuracy, which is better than most existing models.
GPUs and CPUs have been widely used for model training of deep learning (DL) in the cloud, where both DL workloads and resource usage might heavily change over time. Traditional training methods require beforehand specification on the type (either GPUs or CPUs) and amount of computing devices, and thus cannot elastically schedule the dynamic DL workloads onto available GPUs/CPUs. In this paper, we propose Elastic Scheduler ( ES ), a novel approach that efficiently supports both heterogeneous training (with different device types) and dynamic training (with varying device numbers). ES (i) accumulates local gradients and simulates multiple virtual workers on one GPU to alleviate the performance gap between GPUs and CPUs for achieving similar accuracy in heterogeneous GPU‐CPU‐hybrid training as in homogeneous training and (ii) uses local gradients stabilizes batch sizes for high accuracy without long compensation. Experiments show that ES achieves significantly higher performance than existing methods for heterogeneous and dynamic training as well as inference.
Collecting a large amount of labeled data is crutial for training deep neural network, which is a limitation for medical image classification because it necessarily involves expert knowledge. To mitigate this problem of insufficient labeled medical data, in this work, we propose a novel semi-supervised framework for medical image classification. For unlabeled data, we apply the consistency-based strategy to produce high-quality pseudo label, which encourages model to output the same predictions under different perturbations. In addition, we present a novel mixed sample data augmentation CamMix to effectively exploit the relation between samples, mixing pairs of input data and labels according to the class activation map mask. We have evaluated our proposed method on two public medical image datasets, interstitial lung disease dataset and ISIC 2018 skin lesion analysis dataset. The results demonstrate superior performance of our method over other existing methods on the two datasets. Meanwhile, our proposed CamMix performs better than the current mixed sample data augmentation methods.
The mainstream SLU models, such as SDEN, take the joint training way of slot filling and intent detection because of their correlation and add contextual information to improve the model performance by the contextual vector. Although these models have proved effective, it also brings challenges for slot filling. The slot filling decoder is fed with the deep-layer semantic encoding without alignment information, which will affect the performance of slot filling. The alignment information of the history utterances is attenuated in the context vector because of the repeated fusion process, which is not conducive to the performance improvement of slot filling. In order to solve the above problems, we proposed a novel cross layer semantic enhanced SLU model with role context differentiated fusion, which contains two important improvements: 1) the word embedding information of the current utterance is introduced into the slot filling decoder to strengthen the alignment information based on the mutual attention mechanism; 2) the utterances of different roles are fused in different ways to reserve the alignment information of history utterances in the contextual vector. A large number of experiments were carried out on the standard dataset from SDEN, named KVRET*, and the results verify the effectiveness of our new model. Our model can increase the F1 score of slot filling by more than 7.5% than the existing models.
DM-cache is a component of the device mapper of Linux kernel, which has been widely used to map SSDs and HDDs onto higher-level virtual block devices that take fast SSDs as a cache for slow HDDs to achieve high I/O performance at low monetary cost. While enjoying the benefit of persistent caching where SSDs accelerate normal I/O without affecting durability, the current design of DM-cache suffers from long crash recovery times (at the scale of hours) and low availability. This is because its metadata of dirty bits has to be asynchronously persisted for high I/O performance, which consequently causes all cached data on SSDs to be assumed dirty and to be recovered after the system is restarted. This paper presents MapperX, a novel extension to DM-cache that uses an on-disk adaptive bit-tree (ABT) to synchronously maintain the metadata of dirty bits in a hierarchical manner. Leveraging spatial locality of block writes, MapperX achieves controlled metadata persistence overhead with fast crash recovery by adaptively adding/deleting leaves in the ABT where different levels represent the states of blocks with different granularity. We have implemented MapperX for Linux DM-cache module. Experimental results show that the MapperX based hybrid storage device outperforms the original DM-cache based hybrid device by orders of magnitude in crash recovery times while only introducing negligible metadata persistence overhead.