In present-day data markets, sellers increasingly monetize data whose economic value erodes as the underlying information becomes stale. In this paper, we study how a monopolist data seller should jointly manage update frequency, update effort, and the price of access to a dataset over a finite selling horizon. The model captures three salient features of data monetization: data quality decays between updates, update effort stochastically improves quality, and buyers' willingness to pay depends on realized quality. To reflect heterogeneity across data markets and update technologies, we study three canonical regimes in which returns to effort are increasing, decreasing, or S-shaped – that is, initially increasing and then decreasing. Our analysis characterizes the optimal update policy across these regimes. We show that the structure of optimal effort depends critically on how effort translates into quality improvement. When returns to effort are increasing, optimal effort is bang-bang (all-or-nothing); when returns are decreasing, optimal effort is state-dependent; and when returns are S-shaped, the seller updates only in diminishing-returns region. We then contrast the dynamic update policy with a static benchmark that fixes effort across periods. We show that dynamic adjustment increases profit by up to 20%, especially when setup costs are low, marginal effort costs are high, update outcomes are noisy, and the selling horizon is short. By mapping the model primitives to practical data-selling settings, our study shows that profitable data monetization requires coordinating prices and update decisions with the cost, uncertainty, and returns of restoring data quality.
Problem definition: In this paper, we examine how firms offering AI services can effectively acquire large volumes of training data from their consumers to improve the accuracy of the machine learning (ML) models that drive these services. Since consumers often incur privacy costs when sharing sensitive information, it is essential to design data-sharing mechanisms that balance data acquisition needs with consumers’ privacy. Methodology/results: Inspired by practice, we examine two fundamentally distinct data-sharing mechanisms: manual data-sharing, where consumers control the amount of data they share, and algorithmic data-sharing, where the firm’s algorithm redacts sensitive segments of data before using it to train the ML model. We obtain revenue-maximizing mechanisms for each approach and compare their impact on firm revenue, consumer surplus, and the volume of data collected. Our analysis highlights the conditions under which each mechanism yields superior outcomes in terms of revenue and consumer surplus. Managerial implications: Based on the comparative performance of the two mechanisms, we provide managerial guidelines that help firms choose the preferred data-sharing mechanism for different types of AI services and consumer characteristics.
Digital content platforms such as Netflix, Spotify, and YouTube increasingly fund content through a mix of subscription fees and advertising. Most of these platforms offer users only coarse options: a single ad-supported tier, a single paid tier, or a two-tier “free-with-ads versus paid-ad-free” menu. We ask how platforms should jointly deploy these two revenue levers. The answer has implications for managers, consumers, and regulators. We introduce the generalized ad-supported subscription (GAS) mechanism: a menu offering users a continuum of ad intensity and subscription fee combinations that nests today’s ad-only, subscription-only, and two-tier models as special cases and that is implementable in practice via a simple slider or a short finite menu. Although GAS is revenue-optimal within this broad class, we find the familiar two-tier menu is often near optimal. We pinpoint the conditions under which GAS’s added flexibility yields a material revenue gain. We calibrate the model to a video-on-demand setting through an incentive-compatible willingness-to-pay survey and find that GAS raises platform revenue by 148%, 20%, and 1.8% over the ad-only, subscription-only, and two-tier models, respectively. We further show how emerging regulation that caps advertising intensity (e.g., the EU Digital Services Act and Digital Fairness Act) and competitive pressure reshape the optimal menu. Both forces can shrink coverage of low-value users and push platforms toward less ad-intensive offerings. Finally, we distill our findings into a practical guide for when platforms should adopt each mechanism.
Rapid advances in Machine Learning (ML) have led to a proliferation of Artificial Intelligence (AI) services offered by firms. To develop a valuable AI service, a firm must build an accurate ML model which, in turn, requires a large amount of training data. Present-day firms usually obtain this data by offering incentives to consumers to share their data during the initial development phase of the AI service, and then use that data to re-train the ML models to improve the quality of the service. Consumers, on the other hand, incur privacy costs for sharing their data. Inspired by AI services such as speech-to-text conversion offered by Google, and ChatGPT and DALL-E offered by OpenAI, we analyze two popular data-sharing mechanisms that firms employ in practice: manual data-sharing and algorithmic data-sharing. In the former approach, consumers decide the amount of data to share with the firm, whereas in the latter, the firm uses algorithmic data-redaction – an established approach used by technology firms such as Amazon, IBM, and Oracle – to identify and censor sensitive segments of data, and determine the amount of data collected from consumers. For both the data-sharing approaches, we obtain revenue-maximizing mechanisms for the firm and analyze the fundamental differences between the two approaches in terms of the revenue accrued by the firm, the consumer surplus, and the volume of data collected. Our analysis uncovers several interesting economic effects: For instance, we show that the firm can obtain a higher revenue with an inferior data-redaction algorithm and highlight two nuanced effects – namely, a weak privacy-cost-compensation effect and a strong data-collection effect – underlying this behavior.
The promise of consumer data along with advances in information technology has spurred innovation not only in the way firms conduct their business operations but also in the manner in which data are collected. A prominent institutional structure that has recently emerged is a data cooperative—an organization that collects data from its members, and processes and monetizes the pooled data. A characteristic of consumer data is the externality it generates: Data shared by an individual reveal information about other similar individuals; thus, the marginal value of pooled data increases in both the quantity and quality of the data. A key challenge faced by a data cooperative is the design of a revenue‐allocation scheme for sharing revenue with its members. An effective scheme generates a beneficial cycle: It incentivizes members to share high‐quality data, which in turn results in high‐quality pooled data—this increases the attractiveness of the data for buyers and hence the cooperative's revenue, ultimately resulting in improved compensation for the members. While the cooperative naturally wishes to maximize its total surplus, two other important desirable properties of an allocation scheme are individual rationality and coalitional stability. We first examine a natural proportional allocation scheme—which pays members based on their individual contribution—and show that it simultaneously achieves individual rationality, the first‐best outcome, and coalitional stability, when members' privacy costs are homogeneous. Under heterogeneity in privacy costs, we analyze a novel hybrid allocation scheme and show that it achieves both individual rationality and the first‐best outcome, but may not satisfy coalitional stability. Finally, our RobinHood allocation scheme—which uses a fraction of the revenue to ensure coalitional stability and allocates the remaining based on the hybrid scheme—achieves all the desirable properties.
Network effects, economies of scale, and lock-in-effects increasingly lead to a concentration of digital resources and capabilities, hindering the free and equitable development of digital entrepreneurship (SDG9), new skills, and jobs (SDG8), especially in small communities (SDG11) and their small and medium-sized enterprises (“SMEs”). To ensure the affordability and accessibility of technologies, promote digital entrepreneurship and community well-being (SDG3), and protect digital rights, we propose data cooperatives [1,2] as a vehicle for secure, trusted, and sovereign data exchange [3,4]. In post-pandemic times, community/SME-led cooperatives can play a vital role by ensuring that supply chains to support digital commons are uninterrupted, resilient, and decentralized [5]. Digital commons and data sovereignty provide communities with affordable and easy access to information and the ability to collectively negotiate data-related decisions. Moreover, cooperative commons (a) provide access to the infrastructure that underpins the modern economy, (b) preserve property rights, and (c) ensure that privatization and monopolization do not further erode self-determination, especially in a world increasingly mediated by AI. Thus, governance plays a significant role in accelerating communities’/SMEs’ digital transformation and addressing their challenges. Cooperatives thrive on digital governance and standards such as open trusted Application Programming Interfaces (APIs) that increase the efficiency, technological capabilities, and capacities of participants and, most importantly, integrate, enable, and accelerate the digital transformation of SMEs in the overall process. This policy paper presents and discusses several transformative use cases for cooperative data governance. The use cases demonstrate how platform/data-cooperatives, and their novel value creation can be leveraged to take digital commons and value chains to a new level of collaboration while addressing the most pressing community issues. The proposed framework for a digital federated and sovereign reference architecture will create a blueprint for sustainable development both in the Global South and North.
Network effects, economies of scale, and lock-in-effects increasingly lead to a concentration of digital resources and capabilities, hindering the free and equitable development of digital entrepreneurship, new skills, and jobs, especially in small communities and their small and medium-sized enterprises (“SMEs”). To ensure the affordability and accessibility of technologies, promote digital entrepreneurship and community well-being, and protect digital rights, we propose data cooperatives as a vehicle for secure, trusted, and sovereign data exchange. In post-pandemic times, community/SME-led cooperatives can play a vital role by ensuring that supply chains to support digital commons are uninterrupted, resilient, and decentralized. Digital commons and data sovereignty provide communities with affordable and easy access to information and the ability to collectively negotiate data-related decisions. Moreover, cooperative commons (a) provide access to the infrastructure that underpins the modern economy, (b) preserve property rights, and (c) ensure that privatization and monopolization do not further erode self-determination, especially in a world increasingly mediated by AI. Thus, governance plays a significant role in accelerating communities’/SMEs’ digital transformation and addressing their challenges. Cooperatives thrive on digital governance and standards such as open trusted application programming interfaces (“APIs”) that increase the efficiency, technological capabilities, and capacities of participants and, most importantly, integrate, enable, and accelerate the digital transformation of SMEs in the overall process. This review article analyses an array of transformative use cases that underline the potential of cooperative data governance. These case studies exemplify how data and platform cooperatives, through their innovative value creation mechanisms, can elevate digital commons and value chains to a new dimension of collaboration, thereby addressing pressing societal issues. Guided by our research aim, we propose a policy framework that supports the practical implementation of digital federation platforms and data cooperatives. This policy blueprint intends to facilitate sustainable development in both the Global South and North, fostering equitable and inclusive data governance strategies.
Digital resources and capabilities are concentrated due to network effects and economies of scale, limiting equitable digital entrepreneurship and job creation, particularly in local communities through small and medium enterprises (SMEs). In response, this policy brief advocates for the establishment of data cooperatives to ensure sovereign data exchange, promote digital entrepreneurship, protect digital rights, and contribute to community well-being. In the post-pandemic era, community and/or SME-led cooperatives can secure decentralised supply chains, supporting digital commons and data sovereignty. Through such cooperatives, communities can access affordable information and collectively negotiate data-related decisions, supporting self-determination in an AI-driven world. This brief presents diverse case studies demonstrating the transformative potential of data cooperatives across sectors worldwide. It proposes a framework for a digitally federated and sovereign architecture, providing a blueprint for achieving the Sustainable Development Goals. It concludes with six key policy recommendations for the G20 to foster inclusive and sustainable digital ecosystems.
The unprecedented rate at which data are being generated has led to the growth of data markets where valuable data sets are bought and sold. A salient feature of this market is that a data‐buyer (agent) is endowed with multidimensional private information, namely, her “ideal” record that she values the most and how her valuation for a given record changes as its distance from her ideal record changes. Consequently, the revenue‐maximization problem faced by a data‐seller (principal), who serves multiple buyers, is a multidimensional mechanism‐design problem, which is well recognized as being difficult to solve. Our main result in this paper is an approximation scheme that guarantees a revenue within as close a positive amount from the optimal revenue as desired. The scheme generates a posted‐price menu consisting of a set of item–price pairs—each entry in the menu consists of an item, that is, a set of records from the data set, and the price corresponding to that item. As a trade‐off, the length of the menu resulting from the scheme increases as the desired guarantee gets closer to zero. For convenience in practice, data‐sellers may want the ability to limit the length of the menu used by the scheme. To facilitate this, we extend our analysis to obtain a general approximation guarantee corresponding to a menu of any given length. We also demonstrate how the seller can exploit buyers' preferences to generate intuitive and useful rules of thumb for an effective practical implementation of the scheme.
This study is motivated by the challenges faced by clinics in sub‐Saharan Africa in allocating scarce and unreliable supply of antiretroviral drugs (ARVs) among a large pool of eligible patients. Existing discussion of ARV allocation is focused on qualitative rules for prioritizing certain socioeconomic and demographic patient segments over others at the national level. However, such prioritization rules are of limited utility in providing quantitative guidance on scaling up of treatment programs at individual clinics. In this study, we take the perspective of a clinic administrator whose objective is to maximize the quality‐adjusted survival of the entire patient population in its service area by allocating scarce and unreliable supply of drugs among two activities: initiating treatment for untreated patients and continuing treatment for previously treated patients. The key trade‐off underlying this allocation decision is between the marginal health benefit obtained by initiating an untreated patient on treatment and that obtained by avoiding treatment interruption of a treated patient. This trade‐off has not been explicitly studied in the clinical literature, which focuses either on the incremental value obtained from initiating treatment (over no treatment) or on the value of providing continuous treatment (over interrupted treatment) but not on the difference of the two. We cast the clinic's problem as a stochastic dynamic program and provide a partial characterization of the optimal policy, which consists of dynamic prioritization of patient segments and is characterized by state‐dependent thresholds. We use this structure of the optimal policy to design a simpler Two‐Period heuristic and show that it substantially outperforms the Safety‐Stock heuristic, which is commonly used in practice. In our numerical experiments based on realistic parameter values, the performance of the Two‐Period heuristic is within 4% of the optimal policy whereas that of the Safety‐Stock heuristic can be as much as 20% lower than that of the optimal policy. Our model can serve as a basis for developing a decision support tool for clinics to design their ARV treatment program scale‐up plans.
The wide variety of pricing policies used in practice by data sellers suggests that there are significant challenges in pricing data sets. In this paper, we develop a utility framework that is appropriate for data buyers and the corresponding pricing of the data by the data seller. Buyers interested in purchasing a data set have private valuations in two aspects—their ideal record that they value the most, and the rate at which their valuation for the records in the data set decays as they differ from the buyers’ ideal record. The seller allows individual buyers to filter the data set and select the records that are of interest to them. The multidimensional private information of the buyers coupled with the endogenous selection of records makes the seller’s problem of optimally pricing the data set a challenging one. We formulate a tractable model and successfully exploit its special structure to obtain optimal and near-optimal data-selling mechanisms. Specifically, we provide insights into the conditions under which a commonly used mechanism—namely, a price-quantity schedule—is optimal for the data seller. When the conditions leading to the optimality of a price-quantity schedule do not hold, we show that the optimal price-quantity schedule offers an attractive worst-case guarantee relative to an optimal mechanism. Further, we numerically solve for the optimal mechanism and show that the actual performance of two simple and well-known price-quantity schedules—namely, two-part tariff and two-block tariff—is near optimal. We also quantify the value to the seller from allowing buyers to filter the data set.
At an ad exchange, impressions are sold to advertisers in real time through an auction mechanism. The traditional mechanism at these exchanges selects a single advertiser whose ad is displayed over the entire duration of an impression, that is, throughout the user's visit. We argue that such a mechanism leads to an allocative inefficiency. Our goal in this paper is to address this efficiency loss by offering mechanisms in which multiple ads can be displayed sequentially over the lifetime of the impression. We consider two plausible settings—one where each auction is individually rational for the advertisers and another where advertisers are better off relative to the traditional mechanism over the long run—and derive an optimal mechanism for each setting. Under the former mechanism, whereas the ad exchange always benefits relative to the traditional mechanism, the advertisers can either gain or lose—we demonstrate both these possibilities. The optimal mechanism for the latter setting is a mutually beneficial mechanism in that it guarantees a win–win for both the parties relative to the traditional mechanism over the long run. Happily, for both the mechanisms, the allocation of ads and the payments from the advertisers are efficiently computable, thereby making them amenable to real-time bidding.