Pricing is a key lever used by e-commerce companies to achieve “growth with profitability”. Given the huge catalog size in e-commerce, most products have very close substitutes and complements. These complementary/substitute products result in influencing the demand for one another. Moreover, in the context of fashion, the utility of a product is mostly subjective. In categories like electronics, it’s relatively easy to define the utility of a product based on its attributes, but the same is not directly applicable to fashion. Products with similar attributes can have different utilities for the customer and therefore can be priced differently. Taking these things into consideration we base our pricing strategy on the following 3-stage decision-making process: 1) identifying the items which influence each other 2) building demand models that include effects of demand transference 3) joint optimization of the prices to achieve revenue or profit margin targets. We discuss our contributions to building a real-world system that implements these 3 stages in the specific context of fashion e-commerce. Fashion e-commerce has its nuances when it comes to pricing compared to general e-commerce and we explain how we dealt with these difficulties. Moreover, in addition to the formulations, we also describe challenges faced in building working systems that scale to millions of products and hundreds of categories. In addition, we describe a unique approach to quantifying the dollar benefit under scenarios where true A/B testing is not possible for legal reasons. Lastly, we explain how this work has resulted in significant incremental revenue for a large fashion e-commerce company.
Community detection is a well-studied problem in machine learning and recommendation systems literature. In this paper, we study a novel variant of this problem where we assign predefined fashion communities to users in an Ecommerce ecosystem for downstream tasks. We model our problem as a link prediction task in knowledge graphs with multiple types of edges and multiple types of nodes depicting the intricate Ecommerce ecosystems. We employ Relational Graph Convolutional Networks (R-GCN) on top of this knowledge graph to determine whether a user should be assigned to a given community or not. We conduct empirical experiments on two real-world datasets from a leading fashion retailer. Our experiments demonstrate that the proposed graph-based approach performs significantly better than the non-graph-based baseline, indicating that higher order methods like GCN can improve the task of community assignment for fashion and Ecommerce users.
Unambiguous customer addresses are important for the e-Commerce companies in timely and accurate delivery of shipments. In many developing countries a prescribed structure is not usually followed in practice. It is observed that the customer addresses contain additional text such as instructions to the delivery team. Further not every address has an associated geolocation information in many countries. Thus address understanding, address classification, and clustering of similar but noisy addresses are critical. In addition to last mile delivery solutions, the models help reduce buyer fraud and understanding returns. In view of noisy nature of customer addresses and non-availability of associated geolocation, the above problems are effectively solved with the help of NLP. The proposed talk traces the challenges in the Indian addresses in the context of e-commerce, solution approaches, and their extensions during the last 8 years. The talk is based on the author's own experience, his publications as well as developments in this topic across the industry over these years.
Myntra is an online fashion e-commerce company based in India. At Myntra, a market leader in fashion e-commerce in India, customer experience is paramount and a significant portion of our resources are dedicated to it. Here we describe an algorithm that identifies eligible customers to enable preferential product return processing for them by Myntra. We declare the group of aforementioned eligible customers on the platform as elite customers. Our algorithm to identify eligible/elite customers is based on sound principles of game theory. It is simple, easy to implement and scalable.
In the e-commerce space, accurate prediction of delivery dates plays a major role in customer experience as well as in optimizing the supply chain operations. Predicting a date later than the actual delivery date might sometimes result in the customer not placing the order (lost sales) while promising a date earlier than the actual delivery date would lead to a bad customer experience and consequent customer churn. In this paper, we present a machine learning-based approach for penalizing incorrect predictions differently using non-conventional loss functions, while working under various uncertainties involved in making successful deliveries such as traffic disruptions, weather conditions, supply chain, and logistics. We examine statistical, deep learning, and conventional machine learning approaches, and we propose an approach that outperformed the pre-existing rule-based models. The proposed model is deployed internally for Fashion e-Commerce and is operational.
E-commerce customers in developing nations like India tend to follow no fixed format while entering shipping addresses. Parsing such addresses is challenging because of a lack of inherent structure or hierarchy. It is imperative to understand the language of addresses, so that shipments can be routed without delays. In this paper, we propose a novel approach towards understanding customer addresses by deriving motivation from recent advances in Natural Language Processing (NLP). We also formulate different pre-processing steps for addresses using a combination of edit distance and phonetic algorithms. Then we approach the task of creating vector representations for addresses using Word2Vec with TF-IDF, Bi-LSTM and BERT based approaches. We compare these approaches with respect to sub-region classification task for North and South Indian cities. Through experiments, we demonstrate the effectiveness of generalized RoBERTa model, pre-trained over a large address corpus for language modelling task. Our proposed RoBERTa model achieves a classification accuracy of around 90% with minimal text preprocessing for sub-region classification task outperforming all other approaches. Once pre-trained, the RoBERTa model can be fine-tuned for various downstream tasks in supply chain like pincode suggestion and geo-coding. The model generalizes well for such tasks even with limited labelled data. To the best of our knowledge, this is the first of its kind research proposing a novel approach of understanding customer addresses in e-commerce domain by pre-training language models and fine-tuning them for different purposes.
What a piece of work is a (hu)man!How noble in reason, how infinite in faculty!In form and moving how express and admirable!In action how like an angel, in apprehension how like a God!The beauty of the world.The paragon of animals.
E-Commerce companies face a number of challenges in return requests. Claims of missing-items is one such challenge, where customer claims that main product is missing from shipment through return comments. It is observed that dominant part of such claims are inadvertent given the limited literacy of customers. Some of them have fraud intent. At Flipkart, such claims are evaluated manually to examine whether the comment relates to missing item. Classification of the claim intent automatically saves human bandwidth and provides good customer experience by reducing the turn around time to customers. However, this is challenging as comments are replete with spell variations, non-English vernacular words, and are often incomplete and short. This is compounded by noisy labeling of such comments due to human bias and manual errors. To classify the claim intent, we apply conventional as well as deep learning methods. To handle label noise, we employed stateof-the-art noise-aware techniques, which fail to perform due to pattern specific label noise. Motivated by the wide pattern specific label noise, we encode domain heuristics as labeling functions (LFs) which label subsets of the data. However, LFs may conflict and prone to noise. We address the conflict by defining a conflict-score to rank the LFs. Proposed method of noise handling with LFs out performs all the state-of-the-art noise-aware baselines.
The customer addresses are important in e-Commerce for efficient shipment delivery. A predefined structure in addresses is not usually followed in developing countries. Further, some customer addresses are found to be noisy as they contain avoidable additional details such as directions to reach the place. In the presence of such challenges, understanding and equivalence mapping of the addresses becomes necessary for efficient shipment delivery as well as customer linking. We discuss the challenges with actual address data in Indian context. We propose effective methods for efficient large scale address clustering using conventional as well as deep learning approaches. We demonstrate effectiveness of these approaches through elaborate experimentation with real address dataset of an Indian e-commerce company. We further discuss effectiveness of such solution in fraud prediction models.
E-Commerce companies face a wide variety of fraudulent activities. Machine Learning models are continuously built to detect them for their effective mitigation. Randomly typed alphanumeric characters as customer addresses is one peculiar fraud. The orders are placed with such addresses. We refer to them as monkey-typed addresses. Accurate methods of identifying such addresses is important in reducing operational cost. The current work presents machine learning based approaches that classify a given address as normal or monkey-typed with a high accuracy. The approach integrates address preprocessing, novel feature generation and classification.
This book addresses the challenges of data abstraction generation using a least number of database scans, compressing data through novel lossy and non-lossy schemes, and carrying out clustering and classification directly in the compressed domain. Schemes are presented which are shown to be efficient both in terms of space and time, while simultaneously providing the same or better classification accuracy. Features:describes a non-lossy compression scheme based on run-length encoding of patterns with binary valued features; proposes a lossy compression scheme that recognizes a pattern as a sequence of features and identifying subsequences; examines whether the identification of prototypes and features can be achieved simultaneously through lossy compression and efficient clustering; discusses ways to make use of domain knowledge in generating abstraction; reviews optimal prototype selection using genetic algorithms; suggests possible ways of dealing with big data problems using multiagent systems.
Online retail focuses on optimal delivery system of ordered shipments. In the Last Mile context of a Supply Chain, automatic categorization of addresses is an important problem. An automated solution to this problem reduces manual effort significantly from physical reading of shipment addresses to automatic identification of the corresponding route. In general, addresses help to relate to a geolocation. In the absence of geolocation information in terms of latitude and longitude of individual houses and a definitive structure in the addresses, classifying a given address as belonging to a particular locality is a challenging task. In the current work we devised an accurate method to classify the addresses belonging to a region as belonging to predefined subregions in the background of the above challenges. The activity involves text processing, address preprocessing, clustering, classification using ensemble of classifiers, efficient ways to deal with large dataset with high dimensionality and increasing labeled dataset using semi-supervised classification. We discuss each of these stages that culminates in classification of addresses into sub-localities with a high classification accuracy. The solution is demonstrated in an operational setting in a major e-commerce organization. The solution is applicable to developing countries where geolocation information is not completely available.
Structural Support Vector Machines (SSVMs) have recently gained wide prominence in classifying structured and complex objects like parse-trees, image segments and Part-of-Speech (POS) tags. Typical learning algorithms used in training SSVMs result in model parameters which are vectors residing in a large-dimensional feature space. Such a high-dimensional model parameter vector contains many non-zero components which often lead to slow prediction and storage issues. Hence there is a need for sparse parameter vectors which contain a very small number of non-zero components. L1-regularizer and elastic net regularizer have been traditionally used to get sparse model parameters. Though L1-regularized structural SVMs have been studied in the past, the use of elastic net regularizer for structural SVMs has not been explored yet. In this work, we formulate the elastic net SSVM and propose a sequential alternating proximal algorithm to solve the dual formulation. We compare the proposed method with existing methods for L1-regularized Structural SVMs. Experiments on large-scale benchmark datasets show that the proposed dual elastic net SSVM trained using the sequential alternating proximal algorithm scales well and results in highly sparse model parameters while achieving a comparable generalization performance. Hence the proposed sequential alternating proximal algorithm is a competitive method to achieve sparse model parameters and a comparable generalization performance when elastic net regularized Structural SVMs are used on very large datasets.
Domain knowledge about the problem on hand always leads to an effective solution. In this chapter, we discuss ways to make use of domain knowledge in generating abstraction. We consider binary classifiers such as support vector machine (SVM) and adaptive boosting (AdaBoost) to classify 10-class handwritten digit data. We carry out statistical analysis on the data to derive inferences on domain knowledge. We combine it with human expert’s domain knowledge to arrive at a decision tree of depth 4 to classify 10-class data accurately. In this process, we provide an overview of multiclass classification approaches, decision trees, SVM, and AdaBoost. We combine prototype selection with both these methods to obtain high classification accuracy. In essence, the approach emphasizes exploitation of domain knowledge in mining large datasets, which in the present case results in significant compaction in the data and multiclass classification. We provide a discussion on relevant literature and a list of references at the end of the chapter.
Mining data in nonlossy compressed form provide an important direction in data mining. It is interesting to examine whether by preparing to avoid some features in the given representation of patterns, we would still be able to generate an abstraction that is as accurate in classification as the one with original feature set. In this chapter, we propose a lossy compression scheme. We demonstrate its efficiency and accuracy on practical datasets. The scheme consists of recognizing a pattern as a sequence of features and identifying subsequences. We identify blocks of features and find a subsequence of their values. We identify unique subsequences. The number of unique subsequences can be pruned by choosing only frequent items that exceed a chosen support value. The number of unique subsequences can further be pruned by choosing to replace less frequent subsequences by their more frequent nearest neighbors. This results in lossy compression in two levels. Generating compressed testing data forms an interesting scheme too. We demonstrate significant reduction in data and its working on large handwritten digit data. We provide bibliographic notes and references at the end of the chapter.
Mining large datasets in a compressed domain are an interesting direction in data mining. In this chapter, we propose a nonlossy compression scheme. It is based on run-length encoding of binary-valued features or floating-point-valued features that are appropriately quantized into binary-valued data. The proposed algorithm compresses a given dataset in terms of runs and computes the dissimilarity in the compressed domain directly. This results in significant gains in computation time and storage. We provide a detailed discussion on relevant terms, algorithms, and its performance. We demonstrate efficiency of its working on classification of unseen compressed patterns. We discuss applicability of the scheme to genetic algorithms where classification happens to be a fitness function. We provide a few application scenarios in data mining. We provide theoretical discussions on the scheme. Bibliographic notes provide a brief discussion on important relevant references. A list of references is provided in the end.
In the process of finding novel patterns, algorithms for mining large datasets face a number of issues. We discuss the issues related to efficiency in data mining. We elaborate some important data mining tasks such as clustering, classification, and association rule mining that are relevant to the content of the book. We discuss popular and representative algorithms of partitional and hierarchical data clustering. In classification, we discuss the nearest-neighbor classifier and the support vector machine. We use both these algorithms extensively in the book. We provide an elaborate discussion on issues in mining large datasets and possible solutions. We discuss each possible direction in detail. The discussion on clustering includes topics such as incremental clustering with focus on leader and BIRCH clustering algorithms, divide-and-conquer clustering algorithms, and clustering based on intermediate representation. The discussion on classification includes topics such as incremental classification and classification based on intermediate abstraction. We further discuss frequent-itemset mining with two directions such as divide-and-conquer itemset mining and intermediate abstraction for frequent-itemset mining. Bibliographic notes contain a brief discussion on the significant research contribution in each of the directions discussed in the chapter and literature for further study.
Agent mining interaction has attracted a lot of attention among researchers. It is possible to solve large data mining problems through multiagent systems. Big data is characterized by huge volumes of data that are not easily amenable for generating abstraction; variety of data formats, data frequency, types of data, and their integration; real or near-real time data processing for generating business or scientific value depending on nature of data. Data mining algorithms and machine learning have a large role to play in big data abstraction. We propose to deal with big data with multiagent systems. In this process of suggesting possible ways of dealing big data problems using multiagent systems, we provide discussion on big data and algorithms associated with massive data systems such as MapReduce and PageRank. We discuss agents, multiagent systems, issues with big data analytics, and how the divide-and-conquer approach of multiagent systems improves handling huge datasets. We propose four multiagent systems that can help generating abstraction with big data. We provide suggested reading and bibliographic notes. A list of references is provided in the end.
In mining large, high-dimensional sparse featured datasets, it is important to reduce the dimensionality for efficient processing. Some methods of reducing the features include conventional feature selection and extraction methods, frequent item support-based methods, and optimal feature selection approaches. In earlier chapters, we discussed feature selection based on frequent items. In the present chapter, we combine a nonlossy compression scheme with genetic algorithm-based feature selection in arriving at a scheme that results in efficient feature selection. In the process, we provide an overview of methods of feature selection, feature extraction, genetic algorithms, etc. We implement the proposed scheme of efficient optimal prototype selection using genetic algorithms that combines compressed data classification performance as a fitness function. We demonstrate working of the scheme by implementing it on a large dataset bringing out insights, and sensitivity of genetic operators is shown as a movement of cloud of solution space as the parameters vary. We provide notes on relevant literature and a list of references at the end of the chapter.