Point cloud annotation plays a pivotal role in computer vision and machine learning by facilitating the creation of volumetric annotations in 3D space. While prior research has explored point cloud annotation in VR environments, its practical implementation in space-constrained office settings, where data annotation is typically conducted, remains an open question. In this paper, we introduce Annorama, an interactive system that translates 3D point cloud scenes into miniature desk-scale dioramas, enabling annotation using a unique family of keyboard-assisted mid-air gestures inspired by direct manipulation. Through a within-subjects study with 16 participants, we demonstrate the feasibility of our system by assessing the efficacy of four types of mid-air gestures for drawing cuboid annotations. Our findings suggest that Annorama allows for rapid and accurate annotation of point cloud data, particularly with the Sizing and Two Point Gestures.
Object detection tasks are central to the development of datasets and algorithms in computer vision and machine learning. Despite its centrality, object detection remains tedious and time-consuming due to the inherent interactions that are often associated with drawing precise annotations. In this paper, we introduce Snapper, an interactive and intelligent annotation tool that intercepts bounding box annotations as they're drawn and "snaps" them to the nearby object edges in real-time. Through a mixed-design user study with 18 full-time annotators, we compare Snapper's annotation mode to alternative modes of annotation and find that Snapper enables participants to complete object detection tasks 39% more quickly without diminishing annotation quality. Further, we find that participants perceive Snapper as a tool that is interactively intuitive, trustworthy, and helpful. We conclude by discussing the implications of our findings as they relate to augmenting annotators' conventions for drawing annotations in practice.
Compositionality is a common property in many modalities including natural languages and images, but the compositional generalization of multi-modal models is not well-understood. In this paper, we identify two sources of visual-linguistic compositionality: linguistic priors and the interplay between images and texts. We show that current attempts to improve compositional generalization rely on linguistic priors rather than on information in the image. We also propose a new metric for compositionality without such linguistic priors.
In this paper, we present a method for automatic recognition of isolated Marathi handwritten numerals in which Zernike moments and Fourier Descriptors are used as features. After preprocessing the numeral image, Zernike moment and the Fourier Descriptor features of the numeral are extracted. These features are then fed in the k-NN classifier for classification. The proposed method is experimented on a database of 12690 samples of Marathi handwritten numeral. We have obtained recognition accuracy of 96. 58% using k-NN classifier.
In this paper, we present a novel method for automatic recognition of isolated Marathi handwritten numerals. Chain code and Fourier Descriptors that capture the information about the shape of the numeral are used as features. After preprocessing the numeral image, the normalized chain code and the Fourier descriptors of the contour of the numeral are extracted. These features are then fed in the Support Vector Machine (SVM) for classification. The proposed method is experimented on a database of 12690 samples of Marathi handwritten numeral using fivefold cross validation technique. We have obtained recognition accuracy of
We explore whether useful temporal neural generative models can be learned from sequential data without back-propagation through time. We investigate the viability of a more neurocognitively-grounded approach in the context of unsupervised generative modeling of sequences. Specifically, we build on the concept of predictive coding, which has gained influence in cognitive science, in a neural framework. To do so we develop a novel architecture, the Temporal Neural Coding Network, and its learning algorithm, Discrepancy Reduction. The underlying directed generative model is fully recurrent, meaning that it employs structural feedback connections and temporal feedback connections, yielding information propagation cycles that create local learning signals. This facilitates a unified bottom-up and top-down approach for information transfer inside the architecture. Our proposed algorithm shows promise on the bouncing balls generative modeling problem. Further experiments could be conducted to explore the strengths and weaknesses of our approach.
After a long period of decline, the European Otter is now in the process of recolonizing the French territory. In order to facilitate its expansion, one of the objectives of the National Action Plan is to define and locate the habitat suitability for this mustelid in France. For each river sub sector we gathered all available information about the factors (availability and quality of aquatic habitat, availability of food resources, human disturbances and general characteristics of the ecosystem) essential to the presence of the otter in order to create a Maxent model. According to this model, 30 % of the sub-sectors in metropolitan France are unlikely to offer favourable habitats for the Otter, 68 % should contain favourable habitats and 2 % could be considered as very favourable for the settling of European otters.
Determination of potential habitat suitability for the European Otter (Lutra lutra) in geographical sectors of metropolitan France. After a long period of decline, the European Otter is now in the process of recolonizing the French territory. In order to facilitate its expansion, one of the objectives of the National Action Plan is to define and locate the habitat suitability for this mustelid in France. For each river subsector we gathered all available information about the factors (availability and quality of aquatic habitat, availability of food resources, human disturbances and general characteristics of the ecosystem) essential to the presence of the otter in order to create a Maxent model. According to this model, 30 % of the sub-sectors in metropolitan France are unlikely to offer favourable habitats for the Otter, 68 % should contain favourable habitats and 2 % could be considered as very favourable for the settling of European otters.
Many real world problems can be formulated as optimization problems with various parameters to be optimized. Some problems only have one objective to be optimized, some may have multiple objectives to be optimized at the same time and some need to be optimized subjecting to one or more constraints. Thus numerous optimization algorithms have been proposed to solve these problems. Particle Swarm Optimizer (PSO) is a relatively new optimization algorithm which has shown its strength in the optimization world. This thesis presents two PSO variants, Comprehensive Learning Particle Swarm Optimizer (CLPSO) and Dynamic Multi-Swarm Particle Swarm Optimizer (DMS-PSO), which have good global search ability and can solve complex multi-modal problems for single objective optimization. The latter one is extended to solve constrained optimization and multi-objective optimization problems successfully with a novel constraint-handling mechanism and a novel updating criterion respectively. Subsequently, DMS-PSO is applied to determine the Bragg wavelengths of the sensors in an FBG sensor network and a tree search structure is designed to improve the accuracy and reduce the computation cost. Outstanding Chapter Award IEEE CIS UKRI Chapter, UK For promoting and supporting the dissemination of computational intelligence within the UKRI Section. IEEE Transactions on Neural Networks Outstanding Paper Award Long Cheng, Zeng-Guang Hou, Yingzi Lin, Min Tan, Wenjun Chris Zhang, Fang-Xiang Wu for their paper entitled “Recurrent Neural Network for Non-Smooth Convex Optimization Problems with Application to the Identification of Genetic Regulatory Networks”, vol. 22, no. 5, pp. 714–726, May 2011. Digital Object Identifier: 10.1109/ TNN.2011.2109735 Abstract—A recurrent neural network is proposed for solving the nonsmooth convex optimization problem with the convex inequality and linear equality constraints. Since the objective function and inequality constraints may not be smooth, the Clarke’s generalized gradients of the objective function and inequality constraints are employed to describe the dynamics of the proposed neural network. It is proved that the equilibrium point set of the proposed neural network is equivalent to the optimal solution of the original optimization problem by using the Lagrangian saddle-point theorem.
The growing amount of high dimensional data in different machine learning applications requires more efficient and scalable optimization algorithms. In this work, we consider combining two techniques, parallelism and Nesterov's acceleration, to design faster algorithms for L1-regularized loss. We first simplify BOOM, a variant of gradient descent, and study it in a unified framework, which allows us to not only propose a refined measurement of sparsity to improve BOOM, but also show that BOOM is provably slower than FISTA. Moving on to parallel coordinate descent methods, we then propose an efficient accelerated version of Shotgun, improving the convergence rate from $O(1/t)$ to $O(1/t^2)$. Our algorithm enjoys a concise form and analysis compared to previous work, and also allows one to study several connected work in a unified way.
With rapid growth in smart phones and mobile data, effectively managing cellular data networks is important in meeting user performance expectations. However, the scale, complexity and dynamics of a large 3G cellular network make it a challenging task to understand the diverse factors that affect its performance. In this paper we study the RNC (Radio Network Controller)-level performance in one of the largest cellular network carriers in US. Using large amount of datasets collected from various sources across the network and over time, we investigate the key factors that influence the network performance in terms of the round-trip times and loss rates (averaged over an hourly time scale). We start by performing the “first-order” property analysis to analyze the correlation and impact of each factor on the network performance. We then apply RuleFit - a powerful supervised machine learning tool that combines linear regression and decision trees - to develop models and analyze the relative importance of various factors in estimating and predicting the network performance. Our analysis culminates with the detection and diagnosis of both “transient” and “persistent” performance anomalies, with discussion on the complex interactions and differing effects of the various factors that may influence the 3G UMTS (Universal Mobile Telecommunications System) network performance.
Srinivas Bangalore合作论文数Interactions, LLC12