Bid Shading has become increasingly important in Online Advertising, with a large amount of commercial [4,12,13,29] and research work [11,20,28] recently published. Most approaches for solving the bid shading problem involve estimating the probability of win distribution, and then maximizing surplus [28]. These generally use parametric assumptions for the distribution, and there has been some discussion as to whether Log-Normal, Gamma, Beta, or other distributions are most effective [8,38,41,44]. In this paper, we show evidence that online auctions generally diverge in interesting ways from classic distributions. In particular, real auctions generally exhibit significant structure, due to the way that humans set up campaigns and inventory floor prices [16,26]. Using these insights, we present a nonparametric method for Bid Shading which enables the exploitation of this deep structure. The algorithm has low time and space complexity, and is designed to operate within the challenging millisecond Service Level Agreements of Real-Time Bid Servers. We deploy it in one of the largest Demand Side Platforms in the United States, and show that it reliably out-performs best in class Parametric benchmarks. We conclude by suggesting some ways that the best aspects of parametric and nonparametric approaches could be combined.
Since 2019, most ad exchanges and sell-side platforms (SSPs), in the online advertising industry, shifted from second to first price auctions. Due to the fundamental difference between these auctions, demand-side platforms (DSPs) have had to update their bidding strategies to avoid bidding unnecessarily high and hence overpaying. Bid shading was proposed to adjust the bid price intended for second-price auctions, in order to balance cost and winning probability in a first-price auction setup. In this study, we introduce a novel deep distribution network for optimal bidding in both open (non-censored) and closed (censored) online first-price auctions. Offline and online A/B testing results show that our algorithm outperforms previous state-of-art algorithms in terms of both surplus and effective cost per action (eCPX) metrics. Furthermore, the algorithm is optimized in run-time and has been deployed into VerizonMedia DSP as production algorithm, serving hundreds of billions of bid requests per day. Online A/B test shows that advertiser's ROI are improved by +2.4%, +2.4%, and +8.6% for impression based (CPM), click based (CPC), and conversion based (CPA) campaigns respectively.
Online auctions play a central role in online advertising, and are one of the main reasons for the industry's scalability and growth. With great changes in how auctions are being organized, such as changing the second- to first-price auction type, advertisers and demand platforms are compelled to adapt to a new volatile environment. Bid shading is a known technique for preventing overpaying in auction systems that can help maintain the strategy equilibrium in first-price auctions, tackling one of its greatest drawbacks. In this study, we propose a machine learning approach of modeling optimal bid shading for non-censored online first-price ad auctions. We clearly motivate the approach and extensively evaluate it in both offline and online settings on a major demand side platform. The results demonstrate the superiority and robustness of the new approach as compared to the existing approaches across a range of performance metrics.
This paper describes a new win-rate based bid shading algorithm (WR) that does not rely on the minimum-bid-to-win feedback from a Sell-Side Platform (SSP). The method uses a modified logistic regression to predict the profit from each possible shaded bid price. The function form allows fast maximization at run-time, a key requirement for Real-Time Bidding (RTB) systems. We report production results from this method along with several other algorithms. We found that bid shading, in general, can deliver significant value to advertisers, reducing price per impression to about 55% of the unshaded cost. Further, the particular approach described in this paper captures 7% more profit for advertisers, than do benchmark methods of just bidding the most probable winning price. We also report 4.3% higher surplus than an industry Sell-Side Platform shading service. Furthermore, we observed 3% - 7% lower eCPM, eCPC and eCPA when the algorithm was integrated with budget controllers. We attribute the gains above as being mainly due to the explicit maximization of the surplus function, and note that other algorithms can take advantage of this same approach.
Click-through rate (CTR) prediction is a critical task in online display advertising. The data involved in CTR prediction are typically multi-field categorical data, i.e., every feature is categorical and belongs to one and only one field. One of the interesting characteristics of such data is that features from one field often interact differently with features from different other fields. Recently, Field-aware Factorization Machines (FFMs) have been among the best performing models for CTR prediction by explicitly modeling such difference. However, the number of parameters in FFMs is in the order of feature number times field number, which is unacceptable in the real-world production systems. In this paper, we propose Field-weighted Factorization Machines (FwFMs) to model the different feature interactions between different fields in a much more memory-efficient way. Our experimental evaluations show that FwFMs can achieve competitive prediction performance with only as few as 4% parameters of FFMs. When using the same number of parameters, FwFMs can bring 0.92% and 0.47% AUC lift over FFMs on two real CTR prediction data sets.
Cost-per-action (CPA), or cost-per-acquisition, has become the primary campaign performance objective in online advertising industry. As a result, accurate conversion rate (CVR) prediction is crucial for any real-time bidding (RTB) platform. However, CVR prediction is quite challenging due to several factors, including extremely sparse conversions, delayed feedback, attribution gaps between the platform and the third party, etc. In order to tackle these challenges, we proposed a practical framework that has been successfully deployed on Yahoo! BrightRoll, one of the largest RTB ad buying platforms. In this paper, we first show that over-prediction and the resulted over-bidding are fundamental challenges for CPA campaigns in a real RTB environment. We then propose a safe prediction framework with conversion attribution adjustment to handle over-predictions and to further alleviate over-bidding at different levels. At last, we illustrate both offline and online experimental results to demonstrate the effectiveness of the framework.
Motivated by mass-spectrometry protein sequencing, we consider the problem of reconstructing a string from the multisets of its substring composition. We show that all strings of length 7, one less than a prime and one less than twice a prime, can be reconstructed uniquely up to reversal. For all other lengths, we show that unique reconstruction is not always possible and provide sometimes-tight bounds on the largest number of strings with given substring compositions. The lower bounds are derived by combinatorial arguments, while the upper bounds follow from algebraic approaches that lead to precise characterizations of the sets of strings with the same substring compositions in terms of the factorization properties of bivariate polynomials. Using results on the transience of multidimensional random walks, we also provide a reconstruction algorithm that recovers random strings over alphabets of size ≥ 4 from their substring compositions in optimal near-quadratic time. The problem considered is related to the well-known turnpike problem, and its solution may hence shed light on this longstanding open problem as well.
Motivated by the problem of deducing the structure of proteins using mass-spectrometry, we study the reconstruction of a string from the multiset of its substring compositions. We specialize the backtracking algorithm used for the more general turnpike problem for string reconstruction. Employing well known results about transience of random walks in ≥ 3 dimensions, we show that the algorithm reconstructs random strings over alphabet size ≥ 4 with high probability in near-optimal quadratic time.
We present an adaptive non-local means (NLM) denoising method for a sequence of images captured by a multiview imaging system, where direct extensions of existing single image NLM methods are incapable of producing good results. Our proposed method consists of three major components: (1) a robust joint-view distance metric to measure the similarity of patches; (2) an adaptive procedure derived from statistical properties of the estimates to determine the optimal number of patches to be used; (3) a new NLM algorithm to denoise using only a set of similar patches. Experimental results show that the proposed method is robust to disparity estimation error, out-performs existing algorithms in multiview settings, and performs competitively in video settings.
We consider two related problems of estimating properties of a collection of point processes: estimating the multiset of parameters of continuous-time Poisson processes based on their activities over a period of time t, and estimating the multiset of activity probabilities of discrete-time Bernoulli processes based on their activities over n time instants. For both problems, it is sufficient to consider the observations' profile - the multiset of activity counts, regardless of their process identities. We consider the profile maximum likelihood (PML) estimator that finds the parameter multiset maximizing the profile's likelihood, and establish some of its competitive performance guarantees. For Poisson processes, if any estimator approximates the parameter multiset to within distance ε with error probability δ, then PML approximates the multiset to within distance 2ε with error probability at most δ · e 4√t·S , where S is the sum of the Poisson parameters, and the same result holds for Bernoulli processes. In particular, for the L 1 distance metric, we relate the problems to the long-studied distribution-estimation problem and apply recent results to show that the PML estimator has error probability e -(t·S)0.9 for Poisson processes whenever the number of processes is k = O(tS log(tS)), and show a similar result for Bernoulli processes. We also show experimental results where the EM algorithm is used to compute the PML.
Pattern Maximum Likelihood (PML) is a method of probability estimation that works well for large alphabets. It does not assume that all elements from the unknown alphabet have been observed. PML outperforms the traditional Maximum Likelihood for sequences, and it is particularly useful when the sample size is small. In this dissertation we study both the theory and application of PML. For the theory part, we extend the the previous results on the properties of PML, and also show how to find the PML distributions analytically for patterns of simple forms. For general patterns, PML probabilities can be approximated using a previously developed EM algorithm, which we will prove to be equivalent to a generalized Gradient Ascend Method. We also use the algorithm to conduct experiments on different distributions and evaluate the performance of PML. In addition, we investigate the calculation of pattern probability. We show that the pattern probability is closely related to symmetric polynomials, and it can be written as a summation over graphs using power sums. Along the way we reveal a relation between pattern probability and the enumeration of certain connected graphs as well as inversion-free trees. For applications, we show how PML can be used to predict the number of new symbols that would appear in a future sample. We conduct experiments on various distributions and compare PML to the method of Good & Toulmin and the method of Efron & Thisted. We demonstrate that PML outperforms the other methods even if the future sample size is large. Finally we apply PML to authenticating the authorship of the Taylor poem, attributed to Shakespeare, and conclude that it is consistent with Efron and Thisted's models. PML deals with samples from a single distribution. In the last part of this dissertation we extend PML to set-patterns where multiple samples are observed from concurrent Bernoulli processes. Analogous to the single-process patterns, we show that for certain forms of set-patterns we can find the exact Set-pattern Maximum Likelihood (SPML) probabilities analytically. Furthermore, for general set-patterns we extend the previous EM algorithm to approximate the SPML probabilities. We also show that for samples taken from Poisson distributions the set-pattern is reduced to the single-process pattern problem.
Non-local means (NL-means) filter removes independent and identically distributed (i.i.d.) image noises using self-similarity. In this paper, we derive a generalized NL-means (GNL-means), which is specifically used to deal with non-i.i.d. noises in the NL-means filtered images. Inspired by BM3D and LPG-PCA, which perform denoising iteratively, our idea is also to iteratively apply NL-means. However, NL-means can't be applied directly due to the correlated noises in the image filtered by NL-means. We modify the original NL-means to incorporate noise dependence into the weight function, and show how the new weight can be calculated and give a reasonable estimator. We evaluate GNL-means on several benchmark images, and compare it to NL-means and other state-of-the-art non-local methods including BM3D and LPG-PCA. Our experimental results demonstrate that, while it is not surprising that BM3D essentially achieves the best denoising effect, GNL-means always performs better than NL-means, and better than LPG-PCA on average.
We test whether two sequences are generated by the same distribution or by two dierent ones. Unlike previous work, we make no assumptions on the distributions’ support size. Additionally, we compare our performance to that of the best possible test. We describe an eciently-computa ble algorithm based on pattern maximum likelihood that is near optimal whenever the best possible error probability is exp( 14n 2=3 ) using length-n sequences.
Image forgery detection is an important issue in digital forensics. We propose in this manuscript an algorithm for image composite tampering detection. In the algorithm, we first divide a possibly composite image into overlapped blocks, then define and extract a block measure factor which contains both re-sampling characteristics and JPEG compression characteristics for each block, and apply the block measure factor to discriminate the tampered regions from non-tampered regions. Experimental results show that with the proposed algorithm, composite images are correctly recognized, and at the same time, the tampered regions are determined effectively.
We consider the problem of classification, where the data of the classes are generated i.i.d. according to unknown probability distributions. The goal is to classify test data with minimum error probability, based on the training data available for the classes. The Likelihood Ratio Test (LRT) is the optimal decision rule when the distributions are known. Hence, a popular approach for classification is to estimate the likelihoods using well known probability estimators, e.g., the Laplace and Good-Turing estimators, and use them in a LRT. We are primarily interested in situations where the alphabet of the underlying distributions is large compared to the training data available, which is indeed the case in most practical applications. We motivate and propose LRT's based on pattern probability estimators that are known to achieve low redundancy for universal compression of large alphabet sources. While a complete proof for optimality of these decision rules is warranted, we demonstrate their performance and compare it with other well-known classifiers by various experiments on synthetic data and real data for text classification.
Motivated by protein sequencing, we consider the problem of reconstructing a string from the compositions of its substrings. We provide several results, including the following. General classes of strings that cannot be distinguished from their substring compositions. An almost complete characterization of the lengths for which reconstruction is possible. Bounds on the number of strings with the same substring compositions in terms of the number of divisors of the string length plus one. A relation to the turnpike problem and a bivariate polynomial formulation of string reconstruction.
A pattern is skewed if, as in 11123, one of its symbols repeats and the others appear once. We show that the pattern-maximum-likelihood distribution of essentially all skewed patterns consists of one discrete element whose probability is the fraction of times the repeated symbol appears in the pattern.
In computer graphics, ray tracing is a popular technique for rendering images with computers. In this project we implemented a serial version of ray tracing in C, and three parallelized versions in CUDA. We conducted experiments to demonstrate speedup with CUDA, as well as the importance of balancing workload among threads.
We derive several pattern maximum likelihood (PML) results, among them showing that if a pattern has only one symbol appearing once, its PML support size is at most twice the number of distinct symbols, and that if the pattern is ternary with at most one symbol appearing once, its PML support size is three. We apply these results to extend the set of patterns whose PML distribution is known to all ternary patterns, and to all but one pattern of length up to seven.