Large-scale industrial recommendation systems (RS) usually confront computational problems due to the enormous corpus size. Hence, an efficient indexing structure is a practical solution to retrieve and recommend the most relevant items within a limited response time. The existing approaches that adopted embedding or tree-based index structures cannot handle the long-tail phenomenon. To address this issue, we propose a HI erarchical T ree-based model with variable-length layers (HIT) for recommendation systems. HIT consists of a hierarchical tree index structure and a user preference prediction model. It can fully exploit all the training data by dynamically adjusting the lengths of layers in its tree index structure, which can effectively alleviate the long-tail problem. To assess the models’ resistance against the long-tail problem, we further define two types of equilibrium under our index structure. To satisfy the equilibrium, we propose a corresponding hierarchical tree learning algorithm. Furthermore, for those items with a rare appearance in the training data, on which the learning algorithm would fail, we design a dedicated bandit layer to solve them. Extensive experiments on three large-scale real-world datasets show that HIT can significantly outperform the existing methods in terms of efficient recommendations on items with different frequencies.
Mobile crowdsensing has become increasingly popular due to its ability to collect a massive amount of data with the help of many individual smartphone users. A crowdsensing platform can utilize the collected data to extract effective information and provide diverse services. Designing an incentive mechanism to compensate the participants for their resources consumption is critical in attracting more participation. Offline incentive mechanism design has been widely studied in various crowdsensing applications, whereas the online scenario, is much more challenging due to the unavailability of future information when the platform makes user selection decisions. In this paper, we investigate the problem of online crowdsensing by considering a critical property that the values of users' contributions decrease as time goes by. The time-discounting property is common in inter-temporal choice scenarios but has not been carefully addressed from the perspective of mechanism design. To handle this problem, we propose a new method to select users based on a time-dependent threshold, and present a strategy-proof framework where participants prefer to submit their true types, instead of manipulating the market by misreporting their private information. We consider two cases, one is that the total value is the summation of each participant's contributing value, the other is more general that the total value function is submodular. We call these two mechanisms TDM and TDMS, respectively. We prove that our two mechanisms can achieve computational efficiency, budget feasibility, strategy-proofness, and a constant competitive ratio, in the context of time-discounting values. By comparing our mechanisms with the state-of-the-art methods, we show that our design achieves better performance in terms of the total value.
移动群智感知网络已经成为一种新型的感知模式,被广泛应用在环境数据收集等多种不同场景.限制群智感知系统效能的一个关键问题在于:如何在一定预算范围内选取最合适的用户来进行感知任务,从而最大化收集到数据的信息量.这其中的关键挑战包括:(1)如何定义量化评价指标衡量数据的信息量,(2)如何在无先验知识的情况下有效地学习选择每个用户的成本,(3)如何设计有效的用户选择算法,最大程度地降低算法的累计遗憾.在本文中,我们采用高斯过程建模空间环境,并且提出一个基于互信息的信息量衡量指标.为了解决第二、第三个挑战,我们提出有预算限制的多臂老虎机用户选择问题模型,并为静态和动态场景分别设计了理论可证的低累计遗憾的多轮用户选择算法.我们的理论分析和仿真实验均证实我们提出的算法能够在预算限制情况下有效地选择最有信息量的用户,与基准方法相比提升约20%.
In recent years, mobile devices have gained increasingly development with stronger computation capability and larger storage. Some of the computation-intensive machine learning and deep learning tasks can now be run on mobile devices. To take advantage of the resources available on mobile devices and preserve users' privacy, the idea of mobile distributed machine learning is proposed. It uses local hardware resources and local data to solve machine learning sub-problems on mobile devices, and only uploads computation results instead of original data to contribute to the optimization of the global model. This architecture can not only relieve computation and storage burden on servers, but also protect the users' sensitive information. Another benefit is the bandwidth reduction, as various kinds of local data can now participate in the training process without being uploaded to the server. In this paper, we provide a comprehensive survey on recent studies of mobile distributed machine learning. We survey a number of widely-used mobile distributed machine learning methods. We also present an in-depth discussion on the challenges and future directions in this area. We believe that this survey can demonstrate a clear overview of mobile distributed machine learning and provide guidelines on applying mobile distributed machine learning to real applications.
Auctions are believed to be effective methods to solve the problem of wireless spectrum allocation. Existing spectrum auction mechanisms are all centralized and suffer from several critical drawbacks of the centralized systems, which motivates the design of distributed spectrum auction mechanisms. However, extending a centralized spectrum auction to a distributed one broadens the strategy space of agents from one dimension (bid) to three dimensions (bid, communication, and computation), and thus cannot be solved by traditional approaches from mechanism design. In this paper, we propose two distributed spectrum auction mechanisms, namely distributed VCG and FAITH. Distributed VCG implements the celebrated Vickrey-Clarke-Groves mechanism in a distributed fashion to achieve optimal social welfare, at the cost of exponential communication overhead. In contrast, FAITH achieves sub-optimal social welfare with tractable computation and communication overhead. We prove that both of the two proposed mechanisms achieve faithfulness, i.e., the agents' individual utilities are maximized, if they follow the intended strategies. Besides, we extend FAITH to adapt to dynamic scenarios where agents can arrive or depart at any time, without violating the property of faithfulness. We implement distributed VCG and FAITH, and evaluate their performance in various setups. Evaluation results show that distributed VCG results in optimal allocation, while FAITH is more efficient in computation and communication.
Mobile crowdsensing has become a novel and effective way to collect sensing data of people's surrounding environment. Among the data collected from multiple contributors, inconsistency often occurs due to noise, different sensor precision, or contributors' heterogeneous sensing behaviors. To tackle the data inconsistency, the problem of truth discovery has been widely studied to jointly infer the underlying ground truths and the contributors' data qualities. Existing truth discovery algorithms are based on the aggregation of large amounts of data so as to generate accurate estimations. However, in mobile crowdsensing, the collected data are usually sparsely distributed among a large sensing area, where each point of interest (PoI) may receive only a few sensing reports. In this case, traditional truth discovery algorithms may not provide an accurate truth estimation for each PoI. To tackle this challenge, in this paper, we propose an effective truth discovery method, namely Holmes, which takes advantage of the spatial correlations of the monitored phenomena by reusing each contributor's data for multiple nearby PoIs. We also take the issue of long-tail data phenomenon into the estimation of contributors' data quality levels, and proposed Holmes-LT. We further propose Holmes-OL to address the online streaming data scenarios. We evaluate the performance of our proposed algorithms on both real and synthetic datasets. The evaluation results demonstrate that our algorithms achieve significant performance improvements in terms of estimation accuracy over the existing truth discovery algorithms.
Mobile crowdsensing (MCS), as a novel and promising sensing paradigm, can utilize people's mobile devices to gather large amounts of data, such as environment information, traffic conditions, and human movements. The users of mobile crowdsensing are usually more capable than traditional sensors, and can reach locations that cannot be easily covered by static sensors, achieving more comprehensive coverage than traditional sensor networks. However, the uncertainty of the users' behaviors, as well as their uneven levels of qualities of contributed data, may also bring challenges to the coordination and supervision of mobile crowdsensing, causing the effectiveness of crowdsensing platform to significantly deviate from the theoretical optimum. In this paper, we address the users' uncertain behaviors by considering a quality- aware user steering problem, and propose to design user coordination algorithms so as to improve the mobile crowdsensing system's overall effectiveness. We jointly take two issues into account, i.e., data quality and coverage of sensing area, and propose a characterization of the system's effectiveness based on the two factors. Next, we consider optimizing the system's effectiveness in three different practical crowdsensing scenarios, and prove the NP-hardness of each of them. Given the infeasibility of calculating the global optimum in polynomial time, we propose three efficient algorithms to achieve suboptimal solutions to the three problems respectively. We extensively evaluate our proposed algorithms based on both real and synthetic datasets. The evaluation results show that our proposed algorithms can dramatically improve the crowdsensing system's effectiveness.
In the past decade, with the rapid development of wireless communication and sensor technology, ubiquitous smartphones equipped with increasingly rich sensors have more powerful computing and sensing abilities. Thus, mobile crowdsensing has received extensive attentions from both industry and academia. Recently, plenty of mobile crowdsensing applications come forth, such as indoor positioning, environment monitoring, transportation, and so on. However, most existing mobile crowdsensing systems lack of vast user bases, and thus urgently need appropriate incentive mechanisms to attract mobile users to guarantee the service quality. In this paper, we propose to incorporate sensing platform and social network applications, which already have large user bases to build a three-layer network model. Thus, we can publicize the sensing platform promptly in large scale, and provide long-term guarantee of data sources. Based on a three-layer network model, we design incentive mechanisms for both intermediaries and the crowdsensing platform, and provide a solution to cope with the problem of user overlapping among intermediaries. We indicate the properties of our proposed incentive mechanisms, including incentive compatibility, individual rationality, and efficiency.
Mobile crowdsensing has become a novel and promising paradigm in collecting environmental data. A critical problem in improving the QoS of crowdsensing is to decide which users to select to perform sensing tasks, in order to obtain the most informative data, while maintaining the total sensing costs below a given budget. The key challenges lie in (i) finding an effective measure of the informativeness of users' data, (ii) learning users' sensing costs which are unknown a priori, and (iii) designing efficient user selection algorithms that achieve low-regret guarantees. In this paper, we build Gaussian Processes (GPs) to model spatial locations, and provide a mutual information-based criteria to characterize users' informativeness. To tackle the second and third challenges, we model the problem as a budgeted multi-armed bandit (MAB) problem based on stochastic assumptions, and propose an algorithm with theoretically proven low-regret guarantee. Our theoretical analysis and evaluation results both demonstrate that our algorithm can efficiently select most informative users under stringent constraints.
Online advertising has become a tremendous business, and continues to grow rapidly. There are mainly two ways to distribute ads in online advertising: contracts and auctions. Previous researches usually focus on ad allocation through either contracts or auctions, while the problem of allocating ads simultaneously through both methods has not been carefully addressed. Jointly considering these two methods enables the publisher to globally allocates ads in a more efficient way, conforming to the actual needs of online advertising, but rises several critical challenges that cannot be solved by existing approaches. Our work aims to fill this gap and takes a deep investigation into this problem. In this paper, we consider the revenue maximization problem in online advertising, where the publisher sells ads to advertisers and also needs to satisfy the demands of the contracts. We discuss and tackle the emerged challenges based on a more practical contract model, and propose a bidding strategy for the publisher to satisfy the contract and maximize her total revenue. Our evaluation results demonstrate the good performance of our algorithm in fulfilling the contract and maximizing the revenue.
Crowdsensing has become increasingly popular due to its ability to collect a massive amount of real-time data with the help of many individual smartphone users. A crowdsensing platform can utilize the collected data to extract effective information and provide services to service requesters. Due to the rationality of smartphone users, designing an incentive mechanism to compensate the participants for their resources consumption is critical in attracting more participation. Offline incentive mechanism design has been widely studied in various crowdsensing applications, whereas the online scenario, is much more challenging due to the unavailability of future information when the platform has to make user selection decisions. In this paper, we investigate the problem of online crowdsensing by considering a critical property that the values of users' contributions decrease as time goes by. The time- discounting property is common in inter-temporal choice scenarios but has not been carefully addressed in mechanism design perspective. To handle this problem, we propose a new method to select users based on a time-related threshold, and present a strategyproof framework where participants prefer to submit their true types, instead of manipulating the market by misreporting their private information. We prove that our mechanism can achieve computational efficiency, budget feasibility, strategy-proofness, and a constant competitive ratio. By comparing our mechanism with two heuristic benchmarks, we show that our design achieves great performance in terms of the total obtained value.
Auctions are believed to be effective methods to solve the problem of wireless spectrum allocation. Existing spectrum auction mechanisms are all centralized and suffer from several critical drawbacks of the centralized systems, which motivates the design of distributed spectrum auction mechanisms. However, extending a centralized spectrum auction to a distributed one broadens the strategy space of agents from one dimension (bid) to three dimensions (bid, communication, and computation), and thus cannot be solved by traditional approaches from mechanism design. In this paper, we propose two distributed spectrum auction mechanisms, namely distributed VCG and FAITH. Distributed VCG implements the celebrated Vickrey-Clarke-Groves mechanism in a distributed fashion to achieve optimal social welfare, at the cost of exponential communication overhead. In contrast, FAITH achieves sub-optimal social welfare with tractable computation and communication overhead. We prove that both of the two proposed mechanisms achieve faithfulness, i.e., the agents' individual utilities are maximized, if they follow the intended strategies. We also implement FAITH and evaluate its performance in various setups. Evaluation results show that FAITH achieves superior performance compared with the Nash equilibrium based approach.
Participatory sensing has become a novel and promising paradigm in environmental data collection. However, the issue of data quality has not been carefully addressed. Low quality data contributions may undermine the effectiveness and prospects of participatory sensing, and thus motivates the need for approaches to guarantee the high quality of the contributed data. In this paper, we integrate quality estimation and monetary incentive, and propose a quality-based surplus sharing method for participatory sensing. Specifically, we design an unsupervised learning approach to quantify the users' data qualities and long-term reputations, and exploit an outlier detection technique to filter out anomalous data items. Furthermore, we model the process of surplus sharing as a cooperative game, and propose a Shapley value-based method to determine each user's payment. We have conducted a participatory sensing experiment, and the experiment results show that our approach achieves good performance in terms of both quality estimation and surplus sharing.