In the field of Natural Language Processing (NLP), sentence pair classification is important in various real-world applications. Bi-encoders are commonly used to address these problems due to their low-latency requirements, and their ability to act as effective retrievers. However, bi-encoders often under-perform compared to cross-encoders by a significant margin. To address this gap, many Knowledge Distillation (KD) techniques have been proposed. Most existing KD methods focus solely on utilizing the prediction scores of cross-encoder models and overlook the fact that cross-encoders and bi-encoders have fundamentally different input structures. In this work, we introduce a novel knowledge distillation approach called DISKCO, which DISentangles the Knowledge learned in Cross-encoder models especially from multi-head cross-attention models and transfers it to bi-encoder models. DISKCO leverages the information encoded in the cross-attention weights of the trained cross-encoder model, and provide it as contextual cues for the student bi-encoder model during training and inference. DISKCO combines the benefits of independent encoding for low-latency applications with the knowledge acquired from cross-encoders, resulting in improved performance. Empirically, we demonstrate the effectiveness of DISKCO on proprietary and on various publicly available datasets. Our experiments show that DISKCO outperforms traditional knowledge distillation methods by upto 2%.
Arindam Bhattacharya, Ankith Ms, Ankit Gandhi, Vijay Huddar, Atul Saroop, Rahul Bhagat. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track). 2023.
Semantic matching is an important component of a product search pipeline. Its goal is to capture the semantic intent of the search query as opposed to the syntactic matching performed by a lexical matching system. A semantic matching model captures relationships like synonyms, and also captures common behavioral patterns to retrieve relevant results by generalizing from purchase data. They however suffer from lack of availability of informative negative examples for model training. Various methods have been proposed in the past to address this issue based upon hard-negative mining and contrastive learning. In this work, we propose a novel method for semantic matching based on one-class classification called SMOCC. Given a query and a relevant product, SMOCC generates the representation of an informative negative which is then used to train the model. Our method is based on the idea of generating negatives by using adversarial search in the neighborhood of the positive examples. We also propose a novel approach for selecting the radius to generate adversarial negative products around queries based on the model's understanding of the query. Depending on how we select the radius, we propose two variants of our method: SMOCC-QS, that quantizes the queries using their specificity, and SMOCC-EM, that uses expectation-maximization paradigm to iteratively learn the best radius. We show that our method outperforms the state-of-the-art hard negative mining approaches by increasing the purchase recall by 3 percentage points, and improving the percentage of exacts retrieved by up to 5 percentage points while reducing irrelevant results by 1.8 percentage points.
In this paper, we focus on subspace learning problems on the Grassmann manifold. Interesting applications in this setting include low-rank matrix completion and low-dimensional multivariate regression, among others. Motivated by privacy concerns, we aim to solve such problems in a decentralized setting where multiple agents have access to (and solve) only a part of the whole optimization problem. The agents communicate with each other to arrive at a consensus, i.e., agree on a common quantity, via the gossip protocol. We propose a novel cost function for subspace learning on the Grassmann manifold, which is a weighted sum of several sub-problems (each solved by an agent) and the communication cost among the agents. The cost function has a finite-sum structure. In the proposed modeling approach, different agents learn individual local subspaces but they achieve asymptotic consensus on the global learned subspace. The approach is scalable and parallelizable. Numerical experiments show the efficacy of the proposed decentralized algorithms on various matrix completion and multivariate regression benchmarks.
Lack of calibrated product sizing in popular categories such as apparel and shoes leads to customers purchasing incorrect sizes, which in turn results in high return rates due to fit issues. We address the problem of product size recommendations based on customer purchase and return data. We propose a novel approach based on Bayesian logit and probit regression models with ordinal categories Small, Fit, Largeto model size fits as a function of the difference between latent sizes of customers and products. We propose posterior computation based on mean-field variational inference, leveraging the Polya-Gamma augmentation for the logit prior, that results in simple updates, enabling our technique to efficiently handle large datasets. Our Bayesian approach effectively deals with issues arising from noise and sparsity in the data providing robust recommendations. Offline experiments with real-life shoe datasets show that our model outperforms the state-of-the-art in 5 of 6 datasets. and leads to an improvement of 17-26% in AUC over baselines when predicting size fit outcomes.
We propose a novel latent factor model for recommending product size fits {Small, Fit, Large} to customers. Latent factors for customers and products in our model correspond to their physical true size, and are learnt from past product purchase and returns data. The outcome for a customer, product pair is predicted based on the difference between customer and product true sizes, and efficient algorithms are proposed for computing customer and product true size values that minimize two loss function variants. In experiments with Amazon shoe datasets, we show that our latent factor models incorporating personas, and leveraging return codes show a 17-21% AUC improvement compared to baselines. In an online A/B test, our algorithms show an improvement of 0.49% in percentage of Fit transactions over control.
In this paper, we propose novel gossip algorithms for decentralized subspace learning problems that are modeled as finite sum problems on the Grassmann manifold. Interesting applications in this setting include low-rank matrix completion and multi-task feature learning, both of which are naturally reformulated in the considered setup. To exploit the finite sum structure, the problem is distributed among different agents and a novel cost function is proposed that is a weighted sum of the tasks handled by the agents and the communication cost among the agents. The proposed modeling approach allows local subspace learning by different agents while achieving asymptotic consensus on the global learned subspace. The resulting approach is scalable and parallelizable. Our numerical experiments show the good performance of the proposed algorithms on various benchmarks, e.g., the Netflix dataset.
In this paper, we propose novel gossip algorithms for the low-rank decentralized matrix completion problem. The proposed approach is on the Riemannian Grassmann manifold that allows local matrix completion by different agents while achieving asymptotic consensus on the global low-rank factors. The resulting approach is scalable and parallelizable. Our numerical experiments show the good performance of the proposed algorithms on various benchmarks.
In this paper we investigate the message diffusion process in on-line social networks (OSNs) with the aim to understand how and why some messages become viral. We model peculiarities of messaging in OSNs, in particular, information aging and competing message streams. We present a mean-field analysis that gives an approximation to the diffusion dynamics in the limit of large N (the number of participants in an OSN). This approach allows us to precisely define the outbreak of a message and derive conditions for it. Our main results are threshold theorems, which imply that a message becomes viral if a certain threshold is crossed. The results show that owing to competing message streams, a message is required to cross a higher threshold in order to become viral. This, we believe, may be one of the reasons for the low incidence of viral messages in these networks. We provide simulation and numerical results to support our analyses. We also investigate the role of various factors which come into play and derive some insights for launching successful information campaigns on OSNs.
Microblogging websites such as Twitter are increasingly being used by businesses/campaigners for timely dissemination of information to their followers. The diffusion of a tweet depends on several factors: the activity of the follower nodes, the responsiveness of follower nodes to tweets from the source node, the out-degree of the follower nodes, the content of recent related tweets seen by the follower node, etc. Using such factors, in this paper, we propose a framework to measure the effectiveness of an information campaign over Twitter. We consider a positive as well as a negative metric to measure the impact of a tweet: while retweets are used to measure the positive impact, the lack of a timely response from an active follower node is taken as a potential negative impact. We investigate the scheduling of tweets to increase the net positive impact while keeping the net negative impact below a desired level. We propose and study several scheduling algorithms by casting the problem in a Markov Decision Process (MDP) framework. In order to compare our algorithms, we estimate the model parameters from tweet data collected using the Twitter API from an arbitrarily selected node and its 6837 followers over several months. For this dataset, we find that if successive tweets in the campaign are novel, then substantial gains over user activity based scheduling can be obtained by scheduling tweets in time slots where the ratio of the expected positive and negative metrics is high. We call this the MaxRatio policy and we show that it is optimal under certain conditions. In cases where we are not certain about the response of users to successive related tweets, we identify another algorithm (which we call MaxReach) as a robust alternative.
Online social networks are growing at a rapid pace, both in terms of addition of new links between existing nodes and addition of new nodes to the network. Due to this continuous evolution of such networks, it is important to constantly crawl for information the overall network in general, and specific subnetworks in times of need. Precise information about social networks is important for devising strategies for improved dispersion of targeted information through the masses, for fine tuning messaging in marketing campaigns and for measuring the effectiveness of such marketing efforts. With the objective of gathering precise up-to-date information, we explore designs of fast crawlers for online social networks. Our experiments, carried on data downloaded from Twitter, show that node discovery strategies of random walk with backtrack and random search show promise as fast network crawlers. We implement the random search crawler for purposes of crawling Twitter for large amounts of information on network structure, user profile information and Tweet-level data. We present a summary of the data thus collected from Twitter. We also try to design generative models for Twitter-like networks that can be used in our simulations going forward, rather than having to depend upon downloading of large amounts of network information related data from online social networks.
The objective of this work is to develop mathematical models and tools to aid decision makers in devising capacity plans for a manufacturing network in face of uncertainty in future demand. A manufacturing network is constituted by a set of plants serving a set of markets with a set of products. It is specified in terms of a product portfolio for each market and a product portfolio, resource capacities and a set of markets to serve, for each plant. Based on this definition, we address the following question: given a manufacturing network, determine when to change the manufacturing capacity, where to change it and by how much to change it. Our approach is to devise a plan that is robust with respect to certain perturbations in demand forecasts, rather than one that is optimal with respect to the inherent uncertainty in demand. Thus the idea we follow is to superimpose perturbations on demand forecasts used by decision-makers, and devise a plan that can withstand those variations. Representing manufacturing capacity through the notion of operating configurations, we formulate dynamic optimization models to address the problem.
The car customer in the U.S. market has come to ex- pect many incentives during her purchase process. Many of these incentives are based on the private profile infor- mation of the customer, which she would not want to dis- close to the set of all possible dealers. On the other hand, the dealers want to keep the exact information on various incentives on offer private as well due to various business reasons. In this paper, we present design of a privacy pre- serving marketplace that executes control on information disclosure while enabling the purchasing process. Market- place designs pertaining to both single dealer and multiple dealer participation are dealt with.
In this paper, we present a hybrid, two-round procurement auction that can be used when a buyer wants to procure a single unit of a multi-attribute item. In such cases, bids are measured on many attributes like price, quality, reliability, past history of the bidder, geographical distance between the locations of the bidder and the buyer. The problem is even more acute for Global Enterprises where additional attributes like tax and tariff structures of the country of the supplier become important as well. While such multi-attribute bids are commonplace in sealed bid tenders where the analysis of the bids can be carried out after all of them have been placed to determine the winner; it is difficult to handle such multi-attribute bids in other auction formats like English and Dutch auctions. The difficulty arises because in holding multi-attribute forms of English and Dutch auctions, the buyer needs to communicate information about his true preference amongst attributes to the participating suppliers. But by passing the information on preference between various attributes ( termed as the preference structure), the buyer risks revealing sensitive strategic information to the suppliers. In this paper, we present a two-phase auction mechanism that guides the multi-attribute bidding of the participating suppliers, but ensures that only limited information about the buyer's preference structure can be reverse interpreted by the buyers. We also provide results relating to proper choice of the amount of information that should be disclosed in such manner.
GM's research project investigates ways to quickly sense customer demand and ways to respond to it ... new analytical methodologies, IT framework, and collaborative decision-making processes are being developed to better match demand with supply ... the CPFR (collaborative planning, forecasting and replenishment) program is being adopted by the auto industry. General Motors Corporation (GM) is the world's largest automaker and has been the global industry sales leader since 1931. Founded in 1908, GM employs about 317,000 people around the world. We have manufacturing operations in 32 countries, and our vehicles are sold in 200 countries. In 2004, GM sold nearly 9 million cars and trucks globally, up 4 percent from the previous year, which is the second-highest total in the company's history. The General Motors extended enterprise, which includes suppliers, dealers, and logistics providers, is a large, complex network. The automotive industry in general has unique and significant challenges in sensing and responding to customer demand. Customer preferences change rapidly and products are complex with respect to numerous option configurations and long lead-times. Customers have many touch points (Internet, dealerships, etc.), and the associated databases are large. As a part of GM's central RD all of this leads to high volatility in demand. It is difficult to predict customer preferences and trends, and it is even more challenging to identify and acquire relevant demand signals that are both accurate and timely. …
In this paper, we identify various models from the optimization and econometrics literature that can potentially help sense customer demand in the e-business era. While modelling reality is a difficult task, many of these models come close to modelling the customer's decision-making process. We provide a brief overview of these techniques, interspersing the discussion occasionally with a tutorial introduction of the underlying concepts.
Recent advances in client-server, web-server and networking technology have made it possible to hold auctions over the Internet. The popularity of such auctions has grown very rapidly in the last few years, giving a new impetus to the study and analysis of auctions. Apart from technological problems, there are a variety of implementational, economic, and behavioral aspects that need to be studied. In this paper we describe a few representative issues, and outline promising quantitative and computer-based approaches that can lead to solutions. The three issues we examine are timed bids, forecasting of final bid values in incomplete auctions, and last-minute bidding. These and other related features of Internet auctions could have a major impact in the near future on the volume of business transactions over the Internet.
Amitava Bagchi合作论文数Management Information Systems1