Machine learning methods are being increasingly applied in sensitive societal contexts, where decisions impact human lives. Hence it has become necessary to build capabilities for providing easily-interpretable explanations of models' predictions. Recently in academic literature, a vast number of explanations methods have been proposed. Unfortunately, to our knowledge, little has been documented about the challenges machine learning practitioners most often face when applying them in real-world scenarios. For example, a typical procedure such as feature engineering can make some methodologies no longer applicable. The present case study has two main objectives. First, to expose these challenges and how they affect the use of relevant and novel explanations methods. And second, to present a set of strategies that mitigate such challenges, as faced when implementing explanation methods in a relevant application domain -- poverty estimation and its use for prioritizing access to social policies.
In recent decades the world has seen a simultaneous trend towards becoming more peaceful overall, but also towards higher homicide rates surging in focal regions in the developing world. Although abundant research exists on the nature and sociology of crime, few studies look into the damaging impact of crime and violence on the daily lives of affected communities. The present study proposes the use of societal-scale behavioral data—card transactions’ metadata—to elicit such impact. On the crime side, we use detailed homicide records for an entire middle-income country to identify salient crime shocks at the local level. On the behavioral side, we use debit card transaction volumes throughout the country to extract behavioral indices. We show that crime shocks have a substantial effect on communities’ consumption patterns. Moreover, we show that the effects of crime shocks distribute differently across population subgroups defined by gender and socioeconomic status— e.g., with reductions of up to 7% in females’ average volume of transactions—potentially exacerbating social inequalities. We conclude this work with policy recommendations on the use of ‘big data’ sources to monitor and help.
Tourism has been an increasingly significant contributor to the economy, society, and environment. Policy-making and research on tourism traditionally rely on surveys and economic datasets, which are based on small samples and depict tourism dynamics at a low granularity. Anonymous call detail record (CDR) is a novel source of data with enormous potential in areas of high societal value: epidemics, poverty, and urban development. This study demonstrates the added value of CDR in event tourism, especially for the analysis and evaluation of marketing strategies, event operations, and the externalities at the local and national levels. To achieve this aim, we formalize 14 indicators in high spatial and temporal resolutions to measure both the positive and the negative impacts of the touristic events. We exemplify the use of these indicators in a tourism country, Andorra, on 22 high-impact events including sports competitions, cultural performances, and music festivals. We analyze these touristic events using the large-scale CDR data across 2 years. Our approach serves as a prescriptive and a diagnostic tool with mobile phone data and opens up future directions for tourism analytics.
BACKGROUND Diabetes and hypertension are among top public health priorities, particularly in low and middle-income countries where their health and socioeconomic impact is exacerbated by the quality and accessibility of health care. Moreover, their connection with severe or deadly COVID-19 illness has further increased their societal relevance. Tools for early detection of these chronic diseases enable interventions to prevent high-impact complications, such as loss of sight and kidney failure. Similarly, prognostic tools for COVID-19 help stratify the population to prioritize protection and vaccination of high-risk groups, optimize medical resources and tests, and raise public awareness. METHODS We developed and validated state-of-the-art risk models for the presence of undiagnosed diabetes, hypertension, visual complications associated with diabetes and hypertension, and the risk of severe COVID-19 illness (if infected). The models were estimated using modern methods from the field of statistical learning (e.g., gradient boosting trees), and were trained on publicly available data containing health and socioeconomic information representative of the Mexican population. Lastly, we assembled a short integrated questionnaire and deployed a free online tool for massifying access to risk assessment. RESULTS Our results show substantial improvements in accuracy and algorithmic equity (balance of accuracy across population subgroups), compared to established benchmarks. In particular, the models: i) reached state-of-the-art sensitivity and specificity rates of 90% and 56% (0.83 AUC) for diabetes, 80% and 64% (0.79 AUC) for hypertension, 90% and 56% (0.84 AUC) for visual diminution as a complication, and 90% and 60% (0.84 AUC) for development of severe COVID disease; and ii) achieved substantially higher equity in sensitivity across gender, indigenous/non-indigenous, and regional populations. In addition, the most relevant features used by the models were in line with risk factors commonly identified by previous studies. Finally, the online platform was deployed and made accessible to the public on a massive scale. CONCLUSIONS The use of large databases representative of the Mexican population, coupled with modern statistical learning methods, allowed the development of risk models with state-of-the-art accuracy and equity for two of the most relevant chronic diseases, their eye complications, and COVID-19 severity. These tools can have a meaningful impact on democratizing early detection, enabling large-scale preventive strategies in low-resource health systems, increasing public awareness, and ultimately raising social well-being.
Background The automated screening of patients at risk of developing diabetic retinopathy represents an opportunity to improve their midterm outcome and lower the public expenditure associated with direct and indirect costs of common sight-threatening complications of diabetes. Objective This study aimed to develop and evaluate the performance of an automated deep learning–based system to classify retinal fundus images as referable and nonreferable diabetic retinopathy cases, from international and Mexican patients. In particular, we aimed to evaluate the performance of the automated retina image analysis (ARIA) system under an independent scheme (ie, only ARIA screening) and 2 assistive schemes (ie, hybrid ARIA plus ophthalmologist screening), using a web-based platform for remote image analysis to determine and compare the sensibility and specificity of the 3 schemes. Methods A randomized controlled experiment was performed where 17 ophthalmologists were asked to classify a series of retinal fundus images under 3 different conditions. The conditions were to (1) screen the fundus image by themselves (solo); (2) screen the fundus image after exposure to the retina image classification of the ARIA system (ARIA answer); and (3) screen the fundus image after exposure to the classification of the ARIA system, as well as its level of confidence and an attention map highlighting the most important areas of interest in the image according to the ARIA system (ARIA explanation). The ophthalmologists’ classification in each condition and the result from the ARIA system were compared against a gold standard generated by consulting and aggregating the opinion of 3 retina specialists for each fundus image. Results The ARIA system was able to classify referable vs nonreferable cases with an area under the receiver operating characteristic curve of 98%, a sensitivity of 95.1%, and a specificity of 91.5% for international patient cases. There was an area under the receiver operating characteristic curve of 98.3%, a sensitivity of 95.2%, and a specificity of 90% for Mexican patient cases. The ARIA system performance was more successful than the average performance of the 17 ophthalmologists enrolled in the study. Additionally, the results suggest that the ARIA system can be useful as an assistive tool, as sensitivity was significantly higher in the experimental condition where ophthalmologists were exposed to the ARIA system’s answer prior to their own classification (93.3%), compared with the sensitivity of the condition where participants assessed the images independently (87.3%; P=.05). Conclusions These results demonstrate that both independent and assistive use cases of the ARIA system present, for Latin American countries such as Mexico, a substantial opportunity toward expanding the monitoring capacity for the early detection of diabetes-related blindness.
In recent years, the use of sophisticated statistical models that influence decisions in domains of high societal relevance is on the rise. Although these models can often bring substantial improvements in the accuracy and efficiency of organizations, many governments, institutions, and companies are reluctant to their adoption as their output is often difficult to explain in human-interpretable ways. Hence, these models are often regarded as black-boxes, in the sense that their internal mechanisms can be opaque to human audit. In real-world applications, particularly in domains where decisions can have a sensitive impact--e.g., criminal justice, estimating credit scores, insurance risk, health risks, etc.--model interpretability is desired. Recently, the academic literature has proposed a substantial amount of methods for providing interpretable explanations to machine learning models. This survey reviews the most relevant and novel methods that form the state-of-the-art for addressing the particular problem of explaining individual instances in machine learning. It seeks to provide a succinct review that can guide data science and machine learning practitioners in the search for appropriate methods to their problem domain.
Social networks continuously change as new ties are created and existing ones fade. It is widely acknowledged that our social embedding has a substantial impact on what information we receive and how we form beliefs and make decisions. However, most empirical studies on the role of social networks in collective intelligence have overlooked the dynamic nature of social networks and its role in fostering adaptive collective intelligence. Therefore, little is known about how groups of individuals dynamically modify their local connections and, accordingly, the topology of the network of interactions to respond to changing environmental conditions. In this paper, we address this question through a series of behavioral experiments and supporting simulations. Our results reveal that, in the presence of plasticity and feedback, social networks can adapt to biased and changing information environments and produce collective estimates that are more accurate than their best-performing member. To explain these results, we explore two mechanisms: 1) a global-adaptation mechanism where the structural connectivity of the network itself changes such that it amplifies the estimates of high-performing members within the group (i.e., the network "edges" encode the computation); and 2) a local-adaptation mechanism where accurate individuals are more resistant to social influence (i.e., adjustments to the attributes of the "node" in the network); therefore, their initial belief is dispro-portionately weighted in the collective estimate. Our findings substantiate the role of social-network plasticity and feedback as key adaptive mechanisms for refining individual and collective judgments.
Targeted social policies are the main strategy for poverty alleviation across the developing world. These include targeted cash transfers (CTs), as well as targeted subsidies in health, education, housing, energy, childcare, and others. Due to the scale, diversity, and widespread relevance of targeted social policies like CTs, the algorithmic rules that decide who is eligible to benefit from them---and who is not---are among the most important algorithms operating in the world today. Here we report on a year-long engagement towards improving social targeting systems in a couple of developing countries. We demonstrate that a shift towards the use of AI methods in poverty-based targeting can substantially increase accuracy, extending the coverage of the poor by nearly a million people in two countries, without increasing expenditure. However, we also show that, absent explicit parity constraints, both status quo and AI-based systems induce disparities across population subgroups. Moreover, based on qualitative interviews with local social institutions, we find a lack of consensus on normative standards for prioritization and fairness criteria. Hence, we close by proposing a decision-support platform for distributed governance, which enables a diversity of institutions to customize the use of AI-based insights into their targeting decisions.
Background: The automated screening of patients at risk of developing diabetic retinopathy (DR), represents an opportunity to improve their mid-term outcome and lower the public expenditure associated with direct and indirect costs of a common sight-threatening complication of diabetes. Objective: In the present study, we aim at developing and evaluating the performance of an automated deep learning-based system to classify retinal fundus images from international and Mexican patients, as referable and non-referable DR cases. In particular, we study the performance of the automated retina image analysis (ARIA) system under an independent scheme (i.e. only ARIA screening) and two assistive schemes (i.e., hybrid ARIA + ophthalmologist screening), using a web-based platform for remote image analysis. Methods: We ran a randomized controlled experiment where 17 ophthalmologists were asked to classify a series of retinal fundus images under three different conditions: 1) screening the fundus image by themselves (solo), 2) screening the fundus image after being exposed to the opinion of the ARIA system (ARIA answer), and 3) screening the fundus image after being exposed to the opinion of the ARIA system, as well as its level of confidence and an attention map highlighting the most important areas of interest in the image according to the ARIA system (ARIA explanation). The ophthalmologists9 opinion in each condition and the opinion of the ARIA system were compared against a gold standard generated by consulting and aggregating the opinion of three retina specialists for each fundus image. Results: The ARIA system was able to classify referable vs. non-referable cases with an area under the Receiver Operating Characteristic curve (AUROC), sensitivity, and specificity of 98%, 95.1% and 91.5% respectively, for international patient-cases; and an AUROC, sensitivity, and specificity of 98.3%, 95.2%, 90% respectively for Mexican patient-cases. The results achieved on Mexican patient-cases outperformed the average performance of the 17 ophthalmologist participants of the study. We also find that the ARIA system can be useful as an assistive tool, as significant specificity improvements were observed in the experimental condition where participants were exposed to the answer of the ARIA system as a second opinion (93.3%), compared to the specificity of the condition where participants assessed the images independently (87.3%). Conclusions: These results demonstrate that both use cases of ARIA systems, independent and assistive, present a substantial opportunity for Latin American countries like Mexico towards an efficient expansion of monitoring capacity for the early detection of diabetes-related blindness.
Modern availability of rich geospatial datasets and analysis tools can provide insight germane to the design of field experiments. Design of field experiments, and in particular the choice of sampling strategy, requires careful consideration of its consequences on the external representativity and interference (SUTVA violations) of the experimental sample. This paper presents a methodology for a) modeling the geospatial and social interaction factors that drive interference in rural field experiments; and b) eliciting a set of nondominated sample options that approximate the Pareto-optimal tradeoff between interference and external representativity, as functions of sample choice. The study develops and tests the methodology in the context of a large-scale health experiment in rural Mexico, involving more than 3,000 pregnant women and 600 health clinics across 5 states. Relevant for the practitioner, the methodology is computationally tractable and can be implemented leveraging open sourced geo-spatial data and software tools.
Society increasingly relies on machine learning models for automated decision making. Yet, efficiency gains from automation have come paired with concern for algorithmic discrimination that can systematize inequality. Recent work has proposed optimal post-processing methods that randomize classification decisions for a fraction of individuals, in order to achieve fairness measures related to parity in errors and calibration. These methods, however, have raised concern due to the information inefficiency, intra-group unfairness, and Pareto sub-optimality they entail. The present work proposes an alternative active framework for fair classification, where, in deployment, a decision-maker adaptively acquires information according to the needs of different groups or individuals, towards balancing disparities in classification performance. We propose two such methods, where information collection is adapted to group- and individual-level needs respectively. We show on real-world datasets that these can achieve: 1) calibration and single error parity (e.g., equal opportunity); and 2) parity in both false positive and false negative rates (i.e., equal odds). Moreover, we show that by leveraging their additional degree of freedom, active approaches can substantially outperform randomization-based classifiers previously considered optimal, while avoiding limitations such as intra-group unfairness.
Targeted social programs, such as conditional cash transfers (CCTs), are a major vehicle for poverty alleviation throughout the developing world. Only in Mexico and Brazil, these reach nearly 80 million people (25% of population), distributing +8 billion USD yearly. We study the potential efficiency and fairness gains of targeting CCTs by means of artificial intelligence algorithms. In particular, we analyze the targeting decision rules and underlying poverty prediction models used by national-wide CCTs in three middleincome countries (Mexico, Ecuador, and Costa Rica). Our contribution is three-fold: 1) We show that, absent explicit measures aimed at limiting algorithmic bias, targeting rules can systematically disadvantage population subgroups, such as incurring exclusion errors 2.3 times higher on poor urban households compared to their rural counterparts, or exclusion errors 2.2 times higher on poor elderly households compared with poor traditional nuclear families. 2) We constrain the targeting algorithms towards achieving fairness, and show that, for example, mitigating urban/rural unfairness in Ecuador can imply substantial costs in overall accuracy, yet, we also show that in the case of Mexico mitigating unfairness across four different types of family structures can be achieved at no significant accuracy costs. 3) Finally, we provide an interactive decision-support platform that allows even non-expert stakeholders to explore the space of possible AI-based decision rules, visualize their implications in terms of efficiency, fairness, and their trade-offs; and ultimately choose designs that best fit their preferences and context.
Social networks continuously change as new ties are created and existing ones fade. It is widely noted that our social embedding exerts a strong influence on what information we receive and how we form beliefs and make decisions. However, most empirical studies on the role of social networks in collective intelligence have overlooked the dynamic nature of social networks and its role in fostering adaptive collective intelligence. It remains unknown (1) how network structures adapt to the attributes of individuals, and (2) whether this adaptation promotes the accuracy of individual and collective decisions. Here, we answer these questions through a series of behavioral experiments and supporting simulations. Our results reveal that social network plasticity in the presence of feedback, can adapt to biased and changing information environments, and produce collective estimates that are more accurate than their best-performing member. We explore two mechanisms that explain these results: (1) a global adaptation mechanism where the structural connectivity of the network itself changes such that it amplifies the estimates of high-performing members within the group; (2) a local adaptation mechanism where accurate individuals are more resistant to social influence, and therefore their initial belief is weighted in the collective estimate disproportionately. Thereby, our findings substantiate the role of social network plasticity and feedback as adaptive mechanisms for refining individual and collective judgments.
Today's age of data holds high potential to enhance the way we pursue and monitor progress in the fields of development and humanitarian action. We study the relation between data utility and privacy risk in large-scale behavioral data, focusing on mobile phone metadata as paradigmatic domain. To measure utility, we survey experts about the value of mobile phone metadata at various spatial and temporal granularity levels. To measure privacy, we propose a formal and intuitive measure of reidentification risk$\unicode{x2014}$the information ratio$\unicode{x2014}$and compute it at each granularity level. Our results confirm the existence of a stark tradeoff between data utility and reidentifiability, where the most valuable datasets are also most prone to reidentification. When data is specified at ZIP-code and hourly levels, outside knowledge of only 7% of a person's data suffices for reidentification and retrieval of the remaining 93%. In contrast, in the least valuable dataset, specified at municipality and daily levels, reidentification requires on average outside knowledge of 51%, or 31 data points, of a person's data to retrieve the remaining 49%. Overall, our findings show that coarsening data directly erodes its value, and highlight the need for using data-coarsening, not as stand-alone mechanism, but in combination with data-sharing models that provide adjustable degrees of accountability and security.
In many domains of life, business and management, numerous problems are addressed by small groups of individuals engaged in face-to-face discussions. While research in social psychology has a long history of studying the determinants of small group performances, the internal dynamics that govern a group discussion are not yet well understood. Here, we rely on computational methods based on network analyses and opinion dynamics to describe how individuals influence each other during a group discussion. We consider the situation in which a small group of three individuals engages in a discussion to solve an estimation task. We propose a model describing how group members gradually influence each other and revise their judgments over the course of the discussion. The main component of the model is an influence network—a weighted, directed graph that determines the extent to which individuals influence each other during the discussion. In simulations, we first study the optimal structure of the influence network that yields the best group performances. Then, we implement a social learning process by which individuals adapt to the past performance of their peers, thereby affecting the structure of the influence network in the long run. We explore the mechanisms underlying the emergence of efficient or maladaptive networks and show that the influence network can converge towards the optimal one, but only when individuals exhibit a social discounting bias by downgrading the relative performances of their peers. Finally, we find a late-speaker effect, whereby individuals who speak later in the discussion are perceived more positively in the long run and are thus more influential. The numerous predictions of the model can serve as a basis for future experiments, and this work opens research on small group discussion to computational social sciences.
Energy efficiency is a key challenge for building modern sustainable societies. World’s energy consumption is expected to grow annually by 1.6
Large-scale datasets of human behavior have the potential to fundamentally transform the way we develop cities, fight disease and crime, and respond to natural disasters. However, understanding the privacy of these data sets is key to their broad use and potential impact, for these consist of sensitive information such as citizens' geo-location. Moreover, recent research has shown adversarial methods that successfully associate sensitive information in the datasets to individuals, even under pseudonymization of all personal identifiers. This thesis conceptualizes, relates, and generalizes salient methodologies for disclosure analysis of pseudonymized data that have been developed in the last two decades, such as: k-anonymity, t-closeness, and unicity. Data at the core of the so-called "big data" revolution is fundamentally high-dimensional. We show implications of high-dimensionality as paradigmatic to modern disclosure analysis. Consequently, we propose and analyze a methodological framework that couples information-theoretic concepts from t-closeness and J-disclosure with the partial adversarial knowledge model introduced by unicity [1] [2], as well as its possible extensions. The various methodologies were applied and compared on a large dataset of mobile phone records (CDRs), where results empirically showed ordinal equivalence among unicity measures and information distance measures EM-disclosure and KL-disclosure. Advantages of the proposed framework are highlighted, and future research avenues identified. We also investigate the tradeoff between data privacy and data usefulness related to mobile phone metadata (CDRs) and its real-world applications. On the disclosure side, four spatio-temporal points were enough to identify uniquely +95% of individuals, at a [ZIP code, 1 hour] spatiotemporal granularity consistent with main results in the literature. As the dataset was coarsened in space and time, the ratio (unicity) decreased to values below 0.2% for data specified at [District, 1 week] granularity or lower. We confirmed the existence of a utility-privacy tradeoff for the
Tourism has been an increasingly important factor in global economy, society and environment, accounting for a significant share of GDP and labor force. Policy and research on tourism traditionally rely on surveys and economic datasets, which are based on small samples and depict tourism dynamics at low spatial and temporal granularity. Anonymous call detail records (CDRs) are a novel source of data, showing enormous potential in areas of high societal value: such as epidemics, poverty, and urban development. This study demonstrates the added value of using CDRs for the formulation, analysis and evaluation of tourism strategies, at the national and local levels. In the context of the European country of Andorra, we use CDRs to evaluate marketing strategies in tourism, understand tourists' experiences, and evaluate revenues and externalities generated by touristic events. We do this by extracting novel indicators in high spatial and temporal resolutions, such as tourist flows per country of origin, flows of new tourists, tourist revisits, tourist externalities on transportation congestion, spatial distribution, economic impact, and profiling of tourist interests. We exemplify the use of these indicators for the planning and evaluation of high impact touristic events, such as cultural festivals and sports competitions.
This paper presents a methodology for a) modeling the geo-spatial and social interaction factors that drive interference (SUTVA violations) in randomized field experiments; and b) eliciting a set of non-dominated sample options that approximate the Pareto-optimal tradeoff between interference and external representativity as functions of sample choice. We develop and test the methodology in the context of a large-scale health experiment in rural Mexico, involving more than 5,000 pregnant women and 600 health clinics across five states. Relevant for the practitioner, we show the methodology is computationally tractable and can be implemented leveraging novel open sourced geo-spatial data and software tools.