For the purpose of enhancing discriminability of convolutional neural networks (CNNs) and facilitating optimization, a multilayer structured variant of the maxout unit (named Multilayer Maxout Network, MMN) is proposed in this paper. CNNs with maxout units employ linear convolution filters followed by maxout units to abstract representations from less abstract ones. Our model instead applies MMNs as activation functions of CNNs to abstract representations, which inherits advantages of both maxout units and deep neural networks, and is a more general nonlinear function approximator as well. Experimental results show that our proposed model yields better performance on three image classification benchmark datasets (CIFAR-10, CIFAR-100 and MNIST) than some state-of-the-art methods. Furthermore, the influence of MMN in different hidden layers is analyzed, and a trade-off scheme between the accuracy and computing resources is given.
Image captioning with a natural language has been an emerging trend. However, the social image, associated with a set of user-contributed tags, has been rarely investigated for a similar task. The user-contributed tags, which could reflect the user attention, have been neglected in conventional image captioning. Most existing image captioning models cannot be applied directly to social image captioning. In this work, a dual attention model is proposed for social image captioning by combining the visual attention and user attention simultaneously.Visual attention is used to compress a large mount of salient visual information, while user attention is applied to adjust the description of the social images with user-contributed tags. Experiments conducted on the Microsoft (MS) COCO dataset demonstrate the superiority of the proposed method of dual attention.
Attention based encoder-decoder models have shown a great success on video captioning. Recent multi-modal video captioning mainly focused on applying the attention mechanism to all modalities and fusing them in the same level. However, the connections among specific modalities have not been investigated in the fusion process. In this paper, the expressivity of uni-modal is firstly investigated. Due to the characteristic of attention mechanism, an instance-level of visual content is exploited to refine the temporal features. Then, a semantic detection architecture based on CNN+RNN is also employed on the spatiotemporal content to exploit the correlations between semantic labels for better video semantic representation. Finally, a hierarchical attention-based multimodal fusion model for video captioning is proposed by jointly considering the intrinsic properties of multimodal features. Experimental results on the MSVD and MSR-VTT datasets show that the proposed method has achieved competitive performance compared with the related video captioning methods.
Cross modal (e.g., text-to-image or image-to-text) retrieval has received great attention with the flushed multi-modal social media data. It is of considerable challenge to stride across the heterogeneous gap between modalities. Existing methods project different modalities into a common space by minimizing the distance within the heterogeneous pairs (intra-pair) of the new latent space. However, the relationship among these multi-modal pairs (inter-pair) are neglected, which are beneficial to eliminate the heterogeneity. In this paper, we propose a novel algorithm based on canonical correlation analysis by considering the high-order relationship among pairs (HCCA) for cross-modal retrieval. Supervised with additional semantic labels and unsupervised without semantic labels are simultaneously considered by treating the intra- and inter-pair correlation discriminatively. Moreover, kernel tricks are also performed on HCCA to learn a non-linear projection, termed HKCCA. Extensive experiments conducted on three public datasets demonstrate the superiority of the proposed methods compared with the state-of-the-art approaches in cross modal retrieval.
A novel objective function of deep neuron networks with companion losses of both convolutional layers and non-linear activation functions is proposed, aiming to obtain more discriminative features. Conventional deep neuron networks were generally trained by the end-to-end supervised learning framework, whose performance is restricted by the training problems, such as the gradient vanishing problem, leading to less discriminative features, especially in lower layers. Instead, we build a novel objective function with two kinds of companion losses. The advantages of this framework are as follows: Firstly, it facilities the optimization by solving the gradient vanishing problem. Secondly, both kinds of companion supervised information contribute to obtain more discriminative features. Finally, a good initialization for fine-tuning could be obtained with the aid of the companion supervised training. Experimental results demonstrate the proposed model yielding better performances on the image classification benchmark dataset.
A novel objective function of deep neuron networks with companion losses of both convolutional layers and non-linear activation functions is proposed, aiming to obtain more discriminative features. Conventional deep neuron networks were generally trained by the end-to-end supervised learning framework, whose performance is restricted by the training problems, such as the gradient vanishing problem, leading to less discriminative features, especially in lower layers. Instead, we build a novel objective function with two kinds of companion losses. The advantages of this framework are as follows: Firstly, it facilities the optimization by solving the gradient vanishing problem. Secondly, both kinds of companion supervised information contribute to obtain more discriminative features. Finally, a good initialization for fine-tuning could be obtained with the aid of the companion supervised training. Experimental results demonstrate the proposed model yielding better performances on the image classification benchmark dataset.
The retrieval and recommendation of social media have provided an immense opportunity to exploit the collective behavior of community users through linked multi-modal data, such as images and tags, where tags provide context information, and images represent visual content. The stability of content information is more reliable than user contributed context information, which was ignored by many existing methods. In this paper, through discovering the latent feature space between visual features and context, we propose a novel approach for social image retrieval by imposing context regularization terms to constraint visual features. The method can effectively reflect the interior visual structure for social image representation. Experimental results on the NUS-WIDEOBJECT dataset demonstrate that the proposed approach obtains competitive performance compared with state-of-the-art methods.
In this paper, we discuss the existence and multiplicity of positive solutions for the singular fractional boundary value problem D0+αu(t)+f(t,u(t),D0+νu(t),D0+μu(t))=0,u(0)=u′(0)=u″(0)=u″(1)=0, where 3<α≤4, 0<ν≤1, 1<μ≤2, D0+α is the standard Riemann–Liouville fractional derivative, f is a Carathédory function and f(t,x,y,z) is singular at the value 0 of its arguments x,y,z. By means of a fixed point theorem, the existence and multiplicity of positive solutions are obtained.