Mixed-precision quantization plays a pivotal role in deploying deep neural networks in resource-constrained environments. However, the task of finding the optimal bit-width configurations for different layers under deployable mixed-precision quantization has barely been explored and remains a challenge. In this work, we present Cobits, an efficient and effective deployable mixed-precision quantization framework based on the relationship between the range of real-valued input and the range of quantized real-valued. It assigns a higher bit-width to the quantizer with a narrower quantized real-valued range and a lower bit-width to the quantizer with a wider quantized real-valued range. Cobits employs a co-learning approach to entangle and learn quantization parameters across various bit-widths, distinguishing between shared and specific parts. The shared part collaborates, while the specific part isolates precision conflicts. Additionally, we upgrade the normal quantizer to dynamic quantizer to mitigate statistical issues in the deployable mixed-precision supernet. Over the trained mixed-precision supernet, we utilize the quantized real-valued ranges to derive quantized-bit-sensitivity, which can serve as importance indicators for efficiently determining bit-width configurations, eliminating the need for iterative validation dataset evaluations. Extensive experiments show that Cobits outperforms previous state-of-the-art quantization methods on the ImageNet and COCO datasets while retaining superior efficiency. We show this approach dynamically adapts to varying bit-width and can generalize to various deployable backends. The code will be made public in https://github.com/sunnyxiaohu/cobits.
Dataset distillation synthesizes a small dataset such that a model trained on this set approximates the performance of the original dataset. Recent studies on dataset distillation focused primarily on the design of the optimization process, with methods such as gradient matching, feature alignment, and training trajectory matching. However, little attention has been given to the issue of underutilized regions in synthetic images. In this paper, we propose UDD, a novel approach to identify and exploit the underutilized regions to make them informative and discriminate, and thus improve the utilization of the synthetic dataset. Technically, UDD involves two underutilized regions searching policies for different conditions, i.e., response-based policy and data jittering-based policy. Compared with previous works, such two policies are utilization-sensitive, equipping with the ability to dynamically adjust the underutilized regions during the training process. Additionally, we analyze the current model optimization problem and design a category-wise feature contrastive loss, which can enhance the distinguishability of different categories and alleviate the shortcomings of the existing multi-formation methods. Experimentally, our method improves the utilization of the synthetic dataset and outperforms the state-of-the-art methods on various datasets, such as MNIST, FashionMNIST, SVHN, CIFAR-10, and CIFAR-100. For example, the improvements on CIFAR-10 and CIFAR-100 are 4.0% and 3.7% over the next best method with IPC=1, by mining the underutilized regions.
Channel Pruning is one of the most widespread techniques used to compress deep neural networks while maintaining their performances. Currently, a typical pruning algorithm leverages neural architecture search to directly find networks with a configurable width, the key step of which is to identify representative subnet for various pruning ratios by training a supernet. However, current methods mainly follow a serial training strategy to optimize supernet, which is very time-consuming. In this work, we introduce PSE-Net, a novel parallel-subnets estimator for efficient channel pruning. Specifically, we propose a parallel-subnets training algorithm that simulate the forward-backward pass of multiple subnets by droping extraneous features on batch dimension, thus various subnets could be trained in one round. Our proposed algorithm facilitates the efficiency of supernet training and equips the network with the ability to interpolate the accuracy of unsampled subnets, enabling PSE-Net to effectively evaluate and rank the subnets. Over the trained supernet, we develop a prior-distributedbased sampling algorithm to boost the performance of classical evolutionary search. Such algorithm utilizes the prior information of supernet training phase to assist in the search of optimal subnets while tackling the challenge of discovering samples that satisfy resource constraints due to the longtail distribution of network configuration. Extensive experiments demonstrate PSE-Net outperforms previous state-of-the-art channel pruning methods on the ImageNet dataset while retaining superior supernet training efficiency. For example, under 300M FLOPs constraint, our pruned MobileNetV2 achieves 75.2% Top-1 accuracy on ImageNet dataset, exceeding the original MobileNetV2 by 2.6 units while only cost 30%/16% times than BCNet/AutoAlim.
Point-level weakly-supervised temporal action localization (P-TAL) aims to localize action instances in untrimmed videos through the use of single-point annotations in each instance. Existing methods predict the class activation sequences without any boundary information, and the unreliable sequences result in a significant misalignment between the quality of proposals and their corresponding confidence. In this paper, we surprisingly observe the most salient frame tend to appear in the central region of the each instance and is easily annotated by humans. Guided by the temporal saliency information, we present a novel proposal-level plug-in framework to relearn the aligned confidence of proposals generated by the base locators. The proposed approach consists of Center Score Learning (CSL) and Alignment-based Boundary Adaptation (ABA). In CSL, we design a novel center label generated by the point annotations for predicting aligned center scores. During inference, we first fuse the center scores with the predicted action probabilities to obtain the aligned confidence. ABA utilizes the both aligned confidence and IoU information to enhance localization completeness. Extensive experiments demonstrate the generalization and effectiveness of the proposed framework, showcasing state-of-the-art or competitive performances across three benchmarks. Our code is available at https://github.com/zyxia1009/CVPR2024-TSPNet.
In this work, we show that it is feasible to perform multiple tasks concurrently on point cloud with a straightforward yet effective multi-task network. Our framework, Poly-PC, tackles the inherent obstacles (e.g., different model architectures caused by task bias and conflicting gradients caused by multiple dataset domains, etc.) of multi-task learning on point cloud. Specifically, we propose a residual set abstraction (Res-SA) layer for efficient and effective scaling in both width and depth of the network, hence accommodating the needs of various tasks. We develop a weight-entanglement- based one-shot NAS technique to find optimal architectures for all tasks. Moreover, such technique entangles the weights of multiple tasks in each layer to offer task-shared parameters for efficient storage deployment while providing ancillary task-specific parameters for learning task-related features. Finally, to facilitate the training of Poly-PC, we introduce a task-prioritization-based gradient balance algorithm that leverages task prioritization to reconcile conflicting gradients, ensuring high performance for all tasks. Benefiting from the suggested techniques, models optimized by Poly-PC collectively for all tasks keep fewer total FLOPs and parameters and outperform previous methods. We also demonstrate that Poly-PC allows incremental learning and evades catastrophic forgetting when tuned to a new task.
In this part, we give the visualization of task prioritization when Poly-PC jointly optimizes three tasks in the main paper: 3D point shape classification, segmentation, and object detection. As shown in Fig. 1, the task prioritization of such three tasks adjusts dynamically and tends to converge with the progress of iteration. The task prioritization can enable our proposed gradient balance algorithm to prioritize the learning of difficult task at each epoch and ensures that all tasks converge to the optimal solution.
OBJECTIVE To observe the effect of conventional acupuncture combined with row-like puncture at sternocleidomastoid on peripheral facial palsy at recovery stage. METHODS A total of 60 patients with peripheral facial palsy at recovery stage were randomized into an observation group and a control group, 30 cases in each one. Acupuncture was applied at affected Cuanzhu (BL 2), Yangbai (GB 14), Sibai (ST 2), Quanliao (SI 18), Jiache (ST 6), Dicang (ST 4), Hegu (LI 4), Taichong (LR 3) and Zusanli (ST 36) in the control group. On the basis of the treatment in the control group, row-like puncture was applied at sternocleidomastoid (1 needle was punctured at muscle origin and insertion respectively, 3 to 4 needles were row-like punctured at the connection line of muscle origin and insertion). The treatment was given once a day, 5 times were as one course, with 2-day interval, totally 4 courses were required in the both groups. The house-brackmann (H-B) facial nerve function grade, facial nerve function rating system-dynamic view rating scale score and facial disability index (FDI) scale score [including scores of FDI physical function (FDIp) and FDI social life function (FDIs)] before and after treatment were observed, and the clinical efficacy was evaluated in the two groups. RESULTS After treatment, the H-B facial nerve function grades were improved compared before treatment in the both groups (P<0.05), and that in the observation group was superior to the control group (P<0.05). After treatment, the scores of dynamic view rating scale and FDIp were increased (P<0.05), FDIs scores were decreased compared before treatment in the both groups (P<0.05); compared with the control group, the scores of dynamic view rating scale and FDIp were increased (P<0.05), FDIs score was decreased in the observation group (P<0.05). The total effective rate was 93.3% (28/30) in the observation group, which was superior to 73.3% (22/30) in the control group (P<0.05). CONCLUSION Compared with conventional acupuncture, combination therapy with row-like puncture at sternocleidomastoid can improve the therapeutic effect of peripheral facial palsy at recovery stage.
We address the problem of action detection in continuous untrimmed video streams, based on the two-stage framework: one stage for action proposals generation and the other for proposals classification and refinement. The context features inside and outside a candidate region (proposal) are critical for classification in action detection. Therefore, effective integration of these features with different scales has become a fundamental problem. We contend that different action instances and candidate proposals may need different context features. To address this issue, we present a novel multiple scales based context-aware net (MSCA-Net) to effectively classify the action proposals for action detection in this paper. For each candidate action proposal, MSCA-Net takes its multiple regions with different temporal scales as input and then generates suitable context features. Based on the “candidate-control” mechanism of LSTM, the proposed MSCA-Net specially adopts the two-branch structure: Branch1 generates multi-scale context features for each candidate proposal, whereas Branch2 utilizes the context-aware gate function to control the message passing. Extensive experiments on THUMOS’14, Charades daily and ActivityNet action detection datasets, demonstrate the effectiveness of the designed structure and show how these context features influence the detection results.
In this paper, we focus on model generalization and adaptation for cross-domain person re-identification (Re-ID). Unlike existing cross-domain Re-ID methods, leveraging the auxiliary information of those unlabeled target-domain data, we aim at enhancing the model generalization and adaptation by discriminative feature learning, and directly exploiting a pre-trained model to new domains (datasets) without any utilization of the information from target domains. To address the discriminative feature learning problem, we surprisingly find that simply introducing the attention mechanism to adaptively extract the person features for every domain is of great effectiveness. We adopt two popular type of attention mechanisms, long-range dependency based attention and direct generation based attention. Both of them can perform the attention via spatial or channel dimensions alone, even the combination of spatial and channel dimensions. The outline of different attentions are well illustrated. Moreover, we also incorporate the attention results into the final output of model through skip-connection to improve the features with both high and middle level semantic visual information. In the manner of directly exploiting a pre-trained model to new domains, the attention incorporation method truly could enhance the model generalization and adaptation to perform the cross-domain person Re-ID. We conduct extensive experiments between three large datasets, Market-1501, DukeMTMC-reID and MSMT17. Surprisingly, introducing only attention can achieve state-of-the-art performance, even much better than those cross-domain Re-ID methods utilizing auxiliary information from the target domain.
Temporal action detection aims at not only recognizing action category but also detecting start time and end time for each action instance in an untrimmed video. The key challenge of this task is to accurately classify the actions and determine the temporal boundaries of each action instance. In temporal action detection benchmark: THUMOS 2014, large variations exist in the same action category while many similarities exist in different action categories, which always limit the performance of temporal action detection. To address this problem, we propose to use joint Identification-Verification network to reduce the intra-action variations and enlarge inter-action differences. The joint Identification-Verification network is a siamese network based on 3D ConvNets, which can simultaneously predict the action categories and the similarity scores for the input pairs of video proposal segments. Extensive experimental results on the challenging THUMOS 2014 dataset demonstrate the effectiveness of our proposed method compared to the existing state-of-art methods for temporal action detection in untrimmed videos. We further demonstrate that our model is a general framework by evaluating our approach on Charades dataset.
In this paper, we present GDA, a generalized decision aggregation framework that integrates information from distributed sensor nodes for decision making in a resource efficient manner. Different from traditional approaches, our proposed GDA framework is able to not only estimate the reliability of each sensor, but also take advantage of its confidence information, and thus achieves higher decision accuracy. Targeting generalized problem domains, our framework can naturally handle the scenarios where different sensor nodes observe different sets of events whose numbers of possible classes may also be different. GDA also makes no assumption about the availability level of ground truth label information, while being able to take advantage of any if present. For these reasons, our approach can be applied to a much broader spectrum of sensing scenarios. In this paper, we also propose two extensions of the GDA framework, i.e., incremental GDA (I-GDA) and parallel GDA (P-GDA) to deal with streaming and large-scale data. The advantages of our proposed methods are demonstrated through both theoretic analysis and extensive experiments.
Pedestrian detection has achieved great improvements in recent years, while complex occlusion handling is still one of the most important problems. To take advantage of the body parts and context information for pedestrian detection, we propose the part and context network (PCN) in this work. PCN specially utilizes two branches which detect the pedestrians through body parts semantic and context information, respectively. In the Part Branch, the semantic information of body parts can communicate with each other via recurrent neural networks. In the Context Branch, we adopt a local competition mechanism for adaptive context scale selection. By combining the outputs of all branches, we develop a strong complementary pedestrian detector with a lower miss rate and better localization accuracy, especially for occlusion pedestrian. Comprehensive evaluations on two challenging pedestrian detection datasets (i.e. Caltech and INRIA) well demonstrated the effectiveness of the proposed PCN.
Pedestrian detection has achieved great improve-ments in recent years, while complex occlusion handling and high-accurate localization are still the most important problems. To take advantage of the body part semantic information and the contextual information for pedestrian detection, we propose the part and context network (PCN) in this paper. A PCN is composed of three branches: the basic branch; the part branch; and the context branch. It specially utilizes two branches to detect the pedestrians through the body part semantic information and the contextual information, respectively. In the part branch, the semantic information of body parts can communicate with each other via long short-term memory (LSTM). In the context branch, we adopt a local competition mechanism (maxout) for adaptive context scale selection. By combining the outputs of all branches, we develop a strong complementary pedestrian detector with a lower miss rate and higher localization accuracy, especially for the occlusion pedestrian. The combination of the body part semantic information and the contextual information in pedestrian detection is fully explored in this paper. Comprehensive evaluations on three challenging pedestrian detection datasets (i.e., Caltech, INRIA and KITTI) well demonstrate the effectiveness of our proposed PCN. Code for PCN is publicly available on GitHub https://github.com/sunnyxiaohu/pcn_pedestrian.
This paper describes the implementation of a service to identify and geo-locate real world events that may be present as social activity signals in two different social networks. Specifically, we focus on content shared by users on Twitter and Instagram in order to design a system capable of fusing data across multiple networks. Past work has demonstrated that it is indeed possible to detect physical events using various social network platforms. However, many of these signals need corroboration in order to handle events that lack proper support within a single network. We leverage this insight to design an unsupervised approach that can correlate event signals across multiple social networks. Our algorithm can detect events and identify the location of the event occurrence. We evaluate our algorithm using both simulations and real world datasets collected using Twitter and Instagram. The results indicate that our algorithm significantly improves false positive elimination and attains high precision compared to baseline methods on real world datasets.
Some of the most widely deployed IoT devices in urban areas are smartphones in the possession of urban individuals. Their proliferation has led to the emergence of crowdsensing/crowdsourcing services, where humans collect data about their environment (using phones), and servers aggregate the data for various application purposes of interest. With the emergence of social media, a common alternative form of human data entry has become media posts (e.g., on Twitter). This leads to the prospect of building crowdsensing services on top of social media content, exploiting humans as "sensors". In this paper, we develop one such service, called StoryLine. The service detects and tracks physical urban events of interest to the user, such as car accidents, infrastructure damage (in the aftermath of a natural disaster), or instances of civil unrest. It offers an interface to client-side software that allows browsing such events in real time, as well as an interface for software applications to a structured representation of the events and their related statistics. The service embodies novel algorithms for real-time detection, demultiplexing, and tracking of physical events using social media data. In our evaluation with Twitter feeds, we show that our service outperforms two state-of-the-art baselines in event detection and demultiplexing. We also conduct two case-studies to show the effectiveness of the real-time event detection capability and event tracking performance of our system.
In this demo we present a tool that allows us to visualize the real world events on a map interface using the contents shared by users on Twiter and Instagram. We incorporate a novel algorithm that analyzes the data from both Twitter and Instagram for fusing the contents corresponding to the same event thereby enhancing the corroboration of the event detection techniques for the individual networks. Such a service enables tracking of events using explanation from both the social networks to give a rich experience to the users.
In this article, we describe a general methodology for enhancing sensing accuracy in cyber-physical systems that involve structured human interactions in noisy physical environment. We define structured human interactions as domain-specific workflow. A novel workflow-aware sensing model is proposed to jointly correct unreliable sensor data and keep track of states in a workflow. We also propose a new inference algorithm to handle cases with partially known states and objects as supervision. Our model is evaluated with extensive simulations. As a concrete application, we develop a novel log service called Emergency Transcriber , which can automatically document operational procedures followed by teams of first responders in emergency response scenarios. Evaluation shows that our system has significant improvement over commercial off-the-shelf (COTS) sensors and keeps track of workflow states with high accuracy in noisy physical environment.
This paper proposes an unsupervised framework for tracking real world events from their traces on Twitter and Instagram. Empirical data suggests that event detection from Instagram streams errs on the false-negative side due to the relative sparsity of Instagram data (compared to Twitter data), whereas event detection from Twitter can suffer from false-positives, at least if not paired with careful analysis of tweet content. To tackle both problems simultaneously, we design a unified unsupervised algorithm that fuses events detected originally on Instagram (called I-events) and events detected originally on Twitter (called T-events), that occur in adjacent periods, in an attempt to combine the benefits of both sources while eliminating their individual disadvantages. We evaluate the proposed framework with real data crawled from Twitter and Instagram. The results indicate that our algorithm significantly improves tracking accuracy compared to baselines.
This paper develops an algorithm that exploits picture-oriented social networks to localize urban events. We choose picture-oriented networks because taking a picture requires physical proximity, thereby revealing the location of the photographed event. Furthermore, most modern cell phones are equipped with GPS, making picture location, and time metadata commonly available. We consider Instagram as the social network of choice and limit ourselves to urban events (noting that the majority of the world population lives in cities). The paper introduces a new adaptive localization algorithm that does not require the user to specify manually tunable parameters. We evaluate the performance of our algorithm for various real-world datasets, comparing it against a few baseline methods. The results show that our method achieves the best recall, the fewest false positives, and the lowest average error in localizing urban events.
Suman Kumar Nath合作论文数Microsoft Research;Computer Scinece Department, University of Carnegie Mellon4