Semantic hashing is an effective method for fast similarity search which maps high-dimensional data to a compact binary code that preserves the semantic information of the original data. Most existing text hashing approaches treat each document separately and only learn the hash codes from the content of the documents. However, in reality, documents are related to each other either explicitly through an observed linkage such as citations or implicitly through unobserved connections such as adjacency in the original space. The document relationships are pervasive in the real world while they are largely ignored in the prior semantic hashing work. In this paper, we propose node2hash, an unsupervised deep generative model for semantic text hashing by utilizing graph context. It is designed to incorporate both document content and connection information through a probabilistic formulation. Based on the deep generative modeling framework, node2hash employs deep neural networks to learn complex mappings from the original space to the hash space. Moreover, the probabilistic formulation enables a principled way to generate hash codes for unseen documents that do not have any connections with the existing documents. Besides, node2hash can go beyond one-hop connections about directed linked documents by considering more global graph information. We conduct comprehensive experiments on seven datasets with explicit and implicit connections. The results have demonstrated the effectiveness of node2hash over competitive baselines.
Ad-hoc retrieval models with implicit feedback often have problems, e.g., the imbalanced classes in the data set. Too few clicked documents may hurt generalization ability of the models, whereas too many non-clicked documents may harm effectiveness of the models and efficiency of training. In addition, recent neural network-based models are vulnerable to adversarial examples due to the linear nature in them. To solve the problems at the same time, we propose an adversarial sampling and training framework to learn ad-hoc retrieval models with implicit feedback. Our key idea is (i) to augment clicked examples by adversarial training for better generalization and (ii) to obtain very informational non-clicked examples by adversarial sampling and training. Experiments are performed on benchmark data sets for common ad-hoc retrieval tasks such as Web search, item recommendation, and question answering. Experimental results indicate that the proposed approaches significantly outperform strong baselines especially for high-ranked documents, and they outperform IRGAN in NDCG@5 using only 5% of labeled data for the Web search task.
An important linear algebra routine, GEneral Matrix Multiplication (GEMM), is a fundamental operator in deep learning. Compilers need to translate these routines into low-level code optimized for specific hardware. Compiler-level optimization of GEMM has significant performance impact on training and executing deep learning models. However, most deep learning frameworks rely on hardware-specific operator libraries in which GEMM optimization has been mostly achieved by manual tuning, which restricts the performance on different target hardware. In this paper, we propose two novel algorithms for GEMM optimization based on the TVM framework, a lightweight Greedy Best First Search (G-BFS) method based on heuristic search, and a Neighborhood Actor Advantage Critic (N-A2C) method based on reinforcement learning. Experimental results show significant performance improvement of the proposed methods, in both the optimality of the solution and the cost of search in terms of time and fraction of the search space explored. Specifically, the proposed methods achieve 24% and 40% savings in GEMM computation time over state-of-the-art XGBoost and RNN methods, respectively, while exploring only 0.1% of the search space. The proposed approaches have potential to be applied to other operator-level optimizations.
We propose sequenced-replacement sampling (SRS) for training deep neural networks. The basic idea is to assign a fixed sequence index to each sample in the dataset. Once a mini-batch is randomly drawn in each training iteration, we refill the original dataset by successively adding samples according to their sequence index. Thus we carry out replacement sampling but in a batched and sequenced way. In a sense, SRS could be viewed as a way of performing "mini-batch augmentation". It is particularly useful for a task where we have a relatively small images-per-class such as CIFAR-100. Together with a longer period of initial large learning rate, it significantly improves the classification accuracy in CIFAR-100 over the current state-of-the-art results. Our experiments indicate that training deeper networks with SRS is less prone to over-fitting. In the best case, we achieve an error rate as low as 10.10%.
L1 and L2 regularizers are critical tools in machine learning due to their ability to simplify solutions. However, imposing strong L1 or L2 regularization with gradient descent method easily fails, and this limits the generalization ability of the underlying neural networks. To understand this phenomenon, we investigate how and why training fails for strong regularization. Specifically, we examine how gradients change over time for different regularization strengths and provide an analysis why the gradients diminish so fast. We find that there exists a tolerance level of regularization strength, where the learning completely fails if the regularization strength goes beyond it. We propose a simple but novel method, Delayed Strong Regularization, in order to moderate the tolerance level. Experiment results show that our proposed approach indeed achieves strong regularization for both L1 and L2 regularizers and improves both accuracy and sparsity on public data sets. Our source code is published.
Copper selenide (of the type Cu2-xSe) film electrodes, prepared by combined electrochemical (ECD) followed by chemical bath deposition (CBD), may yield high photo-electrochemical (PEC) conversion efficiency (∼14.6%) with no further treatment. The new ECD/CBD-copper selenide film electrodes show enhanced PEC characteristics and exhibit high stability under PEC conditions, compared to the ECD or the CBD films deposited separately. The electrodes combine the advantages of both ECD-copper selenide electrodes (in terms of good adherence to FTO surface and high surface uniformity) and CBD-copper selenide electrodes (suitable film thickness). Effect of annealing temperature, on the ECD/CBD film electrode composition and efficiency, is discussed.
Interpreting black box classifiers, such as deep networks, allows an analyst to validate a classifier before it is deployed in a high-stakes setting. A natural idea is to visualize the deep network's representations, so as to "see what the network sees". In this paper, we demonstrate that standard dimension reduction methods in this setting can yield uninformative or even misleading visualizations. Instead, we present DarkSight, which visually summarizes the predictions of a classifier in a way inspired by notion of dark knowledge. DarkSight embeds the data points into a low-dimensional space such that it is easy to compress the deep classifier into a simpler one, essentially combining model compression and dimension reduction. We compare DarkSight against t-SNE both qualitatively and quantitatively, demonstrating that DarkSight visualizations are more informative. Our method additionally yields a new confidence measure based on dark knowledge by quantifying how unusual a given vector of predictions is.
Regularization plays an important role in generalization of deep neural networks, which are often prone to overfitting with their numerous parameters. L1 and L2 regularizers are common regularization tools in machine learning with their simplicity and effectiveness. However, we observe that imposing strong L1 or L2 regularization with stochastic gradient descent on deep neural networks easily fails, which limits the generalization ability of the underlying neural networks. To understand this phenomenon, we first investigate how and why learning fails when strong regularization is imposed on deep neural networks. We then propose a novel method, gradient-coherent strong regularization, which imposes regularization only when the gradients are kept coherent in the presence of strong regularization. Experiments are performed with multiple deep architectures on three benchmark data sets for image recognition. Experimental results show that our proposed approach indeed endures strong regularization and significantly improves both accuracy and compression (up to 9.9x), which could not be achieved otherwise.
This book constitutes the refereed proceedings of the 14th Information Retrieval Societies Conference, AIRS 2018, held in Taipei, Taiwan, in November 2018. The 8 full papers presented together with 9 short papers and 3 session papers were carefully reviewed and selected from 41 submissions. The scope of the conference covers applications, systems, technologies and theory aspects of information retrieval in text, audio, image, video and multimedia data.
Waste cadmium sulfide (CdS) film electrodes, originally deposited onto glass/fluorine doped tin oxide (glass/FTO) substrates, were used to prepare recycled CdS film electrodes. The waste glass/FTO/CdS were processed in acidic media to recover the glass/FTO substrates, the Cd2+ ions (in the acidic solutions) and the gaseous H2S (recaptured in basic media). All components of the waste electrodes were thus recovered. The recovered glass/FTO and the Cd2+ ions were then reused to produce new recycled glass/FTO/CdS electrodes by chemical bath deposition. The produced films were then characterized by X-ray diffractometry, scanning electron microscopy, electronic absorption spectroscopy and other techniques. The Cd2+ ions were recovered with efficiency higher than 90% from the waste films, as observed from atomic absorption spectrometry. The recycled films were assessed in photo-electrochemical conversion of light to electricity, and exhibited comparable efficiency to those freshly prepared from authentic starting materials and other literature values. Photoelectrochemical characteristics for the recovered films were further enhanced by avoiding stirring of the chemical deposition bath during preparation. The results manifest the feasibility of recycling CdS electrodes and enhancing their photoelectrochemical characteristics by simple low cost methods. Both environmental protection and economic goals can thus be potentially achieved.
A number of online marketplaces enable customers to buy or sell used products, which raises the need for ranking tools to help them find desirable items among a huge pool of choices. To the best of our knowledge, no prior work in the literature has investigated the task of used product ranking which has its unique characteristics compared with regular product ranking. While there exist a few ranking metrics (e.g., price, conversion probability) that measure the “goodness” of a product, they do not consider the time factor, which is crucial in used product trading due to the fact that each used product is often unique while new products are usually abundant in supply or quantity. In this paper, we introduce a novel time-aware metric—“sellability”, which is defined as the time duration for a used item to be traded, to quantify the value of it. In order to estimate the “sellability” values for newly generated used products and to present users with a ranked list of the most relevant results, we propose a combined Poisson regression and listwise ranking model. The model has a good property in fitting the distribution of “sellability”. In addition, the model is designed to optimize loss functions for regression and ranking simultaneously, which is different from previous approaches that are conventionally learned with a single cost function, i.e., regression or ranking. We evaluate our approach in the domain of used vehicles. Experimental results show that the proposed model can improve both regression and ranking performance compared with non-machine learning and machine learning baselines.
Understanding how users' search behavior is influenced by real world events is important both for social science research and for designing better search engines for users. In this paper, we study how to model the influence of events on user queries by framing it as a novel data mining problem. Specifically, given a text description of an event, we mine the search log data to identify queries that are triggered by it and further characterize the temporal trend of influence created by the same event on user queries. We solve this data mining problem by proposing computational measures that quantify the influence of an event on a query to identify triggered queries and then, proposing a novel extension of Hawkes process to model the evolutionary trend of the influence of an event on search queries. Evaluation results using news articles and search log data show that the proposed approach is effective for identification of queries triggered by events reported in news articles and characterization of the influence trend over time, opening up many interesting opportunities of applications such as comparative analysis of influential events and prediction of event-triggered queries by users.
Query auto-completion (QAC) systems suggest queries that complete a user's text as the user types each character. Such queries are typically selected among previously stored queries, based on specific attributes such as popularity. However, queries cannot be suggested if a user's text does not match any queries in the storage. In order to suggest queries for previously unseen text, we propose a neural language model that learns how to generate a query from a starting text, a prefix. Specifically, we employ a recurrent neural network to handle prefixes in variable length. We perform the first neural language model experiments for QAC, and we evaluate the proposed methods with a public data set. Empirical results show that the proposed methods are as effective as traditional methods for previously seen queries and are superior to the state-of-the-art QAC method for previously unseen queries.
Product reviews have become an important resource for customers before they make purchase decisions. However, the abundance of reviews makes it difficult for customers to digest them and make informed choices. In our study, we aim to help customers who want to quickly capture the main idea of a lengthy product review before they read the details. In contrast with existing work on review analysis and document summarization, we aim to retrieve a set of real-world user questions to summarize a review. In this way, users would know what questions a given review can address and they may further read the review only if they have similar questions about the product. Specifically, we design a two-stage approach which consists of question selection and question diversification. For question selection phase, we first employ probabilistic retrieval models to locate candidate questions that are relevant to a given review. A Recurrent Neural Network Encoder–Decoder is utilized to measure the “answerability” of questions to a review. We then design a set function to re-rank the questions with the goal of rewarding diversity in the final question set. The set function satisfies submodularity and monotonicity, which results in an efficient greedy algorithm of submodular optimization. Evaluation on product reviews from two categories shows that the proposed approach is effective for discovering meaningful questions that are representative of individual reviews.
Polycrystalline CdSe films have been deposited onto fluorine doped tin oxide (FTO/glass) substrates by three different techniques, electrochemical deposition (ECD), chemical bath deposition (CBD) and, for the first time, combined ECD and CBD (ECD/CBD). The films were comparatively characterized by photoluminescence spectra (PL), electronic absorption spectra, scanning electron microscopy (SEM) and X-ray diffraction (XRD). The SEM micrographs show that the films involved rod shaped agglomerates with various lengths and widths. XRD patterns show that the three systems involved nano-sized CdSe particles with cubic type crystals. Based on Scherrer's equation, the ECD film showed larger particle size than the CBD film, while the ECD/CBD film showed largest particles among the series. Similarly, the band gap values varied for different films as CBD>ECD>ECD/CBD. Photo-electrochemical (PEC) characteristics, including photo-current density vs. voltage (J-V) plots, conversion efficiency (ƞ), fill factor (FF) and stability were all studied for different film electrodes. The films exhibited n-type behaviors with direct band gaps. The new ECD/CBD-CdSe electrode exhibited higher conversion efficiency (ƞ% ~4.40) than other counterparts. The results show the added value of combining ECD and CBD methods in enhancing PEC characteristics of CdSe film electrodes, even with no additional treatment.
Product reviews have become an important resource for customers before they make purchase decisions. However, the abundance of reviews makes it difficult for customers to digest them and make informed choices. In our study, we aim to help customers who want to quickly capture the main idea of a lengthy product review before they read the details. In contrast with existing work on review analysis and document summarization, we aim to retrieve a set of real-world user questions to summarize a review. In this way, users would know what questions a given review can address and they may further read the review only if they have similar questions about the product. Specifically, we design a two-stage approach which consists of question retrieval and question diversification. We first propose probabilistic retrieval models to locate candidate questions that are relevant to a review. We then design a set function to re-rank the questions with the goal of rewarding diversity in the final question set. The set function satisfies submodularity and monotonicity, which results in an efficient greedy algorithm of submodular optimization. Evaluation on product reviews from two categories shows that the proposed approach is effective for discovering meaningful questions that are representative for individual reviews.
This communication describes for the first time how nano-size particles, sensitized with natural dye molecules of anthocyanin, can be used as catalysts in photo-degradation of gram negative Escherichia coli bacteria in water. The naked ZnO nano-particles degraded up to 83% of the bacteria under solar simulator light, while the dye-sensitized particles increased the bacterial loss by similar to 10%. Solar simulator light includes about 5% of UV tail (shorter than 400 nm) which means that both UV and visible light (longer than 400 nm) radiations could be involved. When a cut-off filter was used, the naked ZnO caused only 40% bacterial loss, in accordance with earlier literature that described killing of bacteria with ZnO particles both in the dark and under light. With the cut-off filter, the sensitized ZnO particles caused higher than 90% bacterial loss, which confirms sensitization of the ZnO particles to visible light. Moreover, the results show that the catalyzed photo-degradation process causes mineralization of the bacteria and their organic internal components which leach out by killing. The catalyst can be recovered and reused losing similar to 10% of its activity each time due to mineralization of the dye molecules. However, catalyst activity can be totally regained by re-sensitizing it with the anthocyanin dye. The effects of different experimental conditions, such as reaction temperature, pH, bacterial concentration and catalyst amount together with nutrient broth and saline media, will be discussed together with the role of the sensitizer. (C) 2016 Elsevier B.V. All rights reserved.
People often implicitly or explicitly express their needs in social media in the form of "user status text". Such text can be very useful for service providers and product manufacturers to proactively provide relevant services or products that satisfy people's immediate needs. In this paper, we study how to infer a user's intent based on the user's "status text" and retrieve relevant mobile apps that may satisfy the user's needs. We address this problem by framing it as a new entity retrieval task where the query is a user's status text and the entities to be retrieved are mobile apps. We first propose a novel approach that generates a new representation for each query. Our key idea is to leverage social media to build parallel corpora that contain implicit intention text and the corresponding explicit intention text. Specifically, we model various user intentions in social media text using topic models, and we predict user intention in a query that contains implicit intention. Then, we retrieve relevant mobile apps with the predicted user intention. We evaluate the mobile app retrieval task using a new data set we create. Experiment results indicate that the proposed model is effective and outperforms the state-of-the-art retrieval models.
Xiaohua Tony Hu合作论文数College of Computing & Informatics, Drexel University3