Although much literature has established the presence of demographic bias in natural language processing (NLP) models, most work relies on curated bias metrics that may not be reflective of real-world applications. At the same time, practitioners are increasingly using algorithmic tools in high-stakes settings, with particular recent interest in NLP. In this work, we focus on one such setting: child protective services (CPS). CPS workers often write copious free-form text notes about families they are working with, and CPS agencies are actively seeking to deploy NLP models to leverage these data. Given well-established racial bias in this setting, we investigate possible ways deployed NLP is liable to increase racial disparities. We specifically examine word statistics within notes and algorithmic fairness in risk prediction, coreference resolution, and named entity recognition (NER). We document consistent algorithmic unfairness in NER models, possible algorithmic unfairness in coreference resolution models, and little evidence of exacerbated racial bias in risk prediction. While there is existing pronounced criticism of risk prediction, our results expose previously undocumented risks of racial bias in realistic information extraction systems, highlighting potential concerns in deploying them, even though they may appear more benign. Our work serves as a rare realistic examination of NLP algorithmic fairness in a potential deployed setting and a timely investigation of a specific risk associated with deploying NLP in CPS settings.
Introduction: Loss aversion when using gamification is incompletely understood. The aim of this study was therefore to examine how participants alter their behavior vis-a-vis meeting a daily step goal based on the prospect of losing or gaining a gamification level. Methods: We enrolled 602 participants across four arms who were given pedometers. In the three experimental arms, participants began at the medium level and were allocated 70 points each week, losing 10 points each day they did not meet their step goal. Having at least 40 points at the end of the week resulted in a level increase, otherwise they lost a level. We fit a generalized estimating equation, clustered on participants, modeling step goal attainment on day 7. Our primary predictor was a categorical variable simultaneously indicating what level the participants began the week in and whether they had more than, less than, or exactly 40 points after 6 days. Results: Participants at risk of losing the highest level were 18.40% (confidence interval [95% CI]: 18.26-19.90) more likely to meet their step goal than those who had secured the highest level. Participants who could potentially move from the low to the medium level were 10.61% (95% CI: 9.98-11.24) more likely to meet their step goal than those in the Control group. Those in the Medium group were similarly more likely to achieve their step goal on day 7 (10.00%, 95% CI: 9.15-10.85) than those who had already secured an increase to the high level. Discussion: We find that participants in this trial generally exhibit loss aversion so long as the loss relates to something that was earned rather than endowed. This knowledge can be incorporated in future interventions using gamification by requiring participants to earn all levels as they progress. ClinicalTrials.gov identifier: NCT03311230.
Participants often vary in their response to behavioral interventions, but methods to identify groups of participants that are more likely to respond are lacking. In this secondary analysis of a randomized clinical trial, we used baseline characteristics to group participants into distinct behavioral phenotypes and evaluated differential responses to a physical activity intervention. Latent class analysis was used to segment participants based on baseline participant data including demographics, validated measures of psychosocial variables, and physical activity behavior. The trial included 602 adults from 40 U.S. states with body mass index ≥25 who were randomized to control or one of three gamification interventions (supportive, collaborative, or competitive) to increase physical activity. Daily step counts were monitored using a wearable device for a 24-week intervention with 12 weeks of follow-up. The model segmented participants into three classes named for key defining traits: Class 1, extroverted and motivated; Class 2, less active and less social; Class 3, less motivated and at-risk. Adjusted regression models were used to test for differences in intervention response relative to control within each behavioral phenotype. In Class 1, only participants in the competitive arm increased their mean daily steps during the intervention (adjusted difference, 945; 95% CI, 352-1537; P = .002), but it was not sustained during follow-up. In Class 2, participants in all three gamification arms significantly increased their mean daily steps compared to control during the intervention (supportive arm adjusted difference 1172; 95% CI, 363-1980; P = .005; collaborative arm adjusted difference 1119; 95% CI, 319-1919; P = .006; competitive arm adjusted difference 1179; 95% CI, 400-1957; P = .003) and all three had sustained impact during follow-up. In Class 3, none of the interventions had a significant effect on physical activity. Three behavioral phenotypes were identified, each with a different response to the interventions. This approach could be used to better target behavioral interventions to participants that are more likely to respond to them.
Importance Gamification, the use of game design elements in nongame contexts, is increasingly being used in workplace wellness programs and digital health applications. However, the best way to design social incentives in gamification interventions has not been well examined. Objective To assess the effectiveness of support, collaboration, and competition within a behaviorally designed gamification intervention to increase physical activity among overweight and obese adults. Design, Setting, and Participants This 36-week randomized clinical trial with a 24-week intervention and 12-week follow-up assessed 602 adults from 40 states with body mass indexes (calculated as weight in kilograms divided by height in meters squared) of 25 or higher from February 12, 2018, to March 17, 2019. Interventions Participants used a wearable device to track daily steps, established a baseline, selected a step goal increase, were randomly assigned to a control (n = 151) or to 1 of 3 gamification interventions (support [n = 151], collaboration [n = 150], and competition [n = 150]), and were remotely monitored. The control group received feedback from the wearable device but no other interventions for 36 weeks. The gamification arms were entered into a 24-week game designed using insights from behavioral economics with points and levels for achieving step goals. No gamification interventions occurred during follow-up. Main Outcomes and Measures The primary outcome was change in mean daily steps from baseline through the 24-week intervention period. Results A total of 602 participants (mean [SD] age, 39 [10] years; mean [SD] body mass index, 30 [5]; 427 [70.9%] male) were included in the study. Compared with controls, participants had a significantly greater increase in mean daily steps from baseline during the intervention in the competition arm (adjusted difference, 920; 95% CI, 513-1328; P < .001), support arm (adjusted difference, 689; 95% CI, 267-977; P < .001), and collaboration arm (adjusted difference, 637; 95% CI, 258-1017; P = .001). During follow-up, physical activity remained significantly greater in the competition arm than in the control arm (adjusted difference, 569; 95% CI, 142-996; P = .009) but was not significantly greater in the support (adjusted difference, 428; 95% CI, 19-837; P = .04) and collaboration (adjusted difference, 126; 95% CI, -248 to 468; P = .49) arms than in the control arm. Conclusions and Relevance All 3 gamification interventions significantly increased physical activity during the 24-week intervention, and competition was the most effective. Physical activity was lower in all arms during follow-up and only remained significantly greater in the competition arm than in the control arm. Trial Registration ClinicalTrials.gov identifier: NCT03311230.
BackgroundLess than half of adults in the United States (US) obtain the recommended level of physical activity. Social incentives, the influences that impact individuals to adjust their behaviors based on social ties or connections, are ubiquitous and could be leveraged within gamification interventions to provide a scalable, low-cost approach to increase engagement. Gamification, or the use of game design in non-game situations, is commonly used in the real world, but in most cases has not appropriately leveraged principles from theories of health behavior.MethodsWe are conducting a four-arm, randomized, controlled trial of 602 overweight and obese adults to evaluate the effectiveness of gamification interventions that leverage insights from behavioral economics to enhance either supportive, competitive, or collaborative social incentives. Daily step counts are monitored using wearable devices that transmit data to the study platform. Participants established a baseline step count, selected a step goal increase, and then were randomly assigned to control or one of three interventions for a 24-week intervention and 12-week follow-up period. To understand predictors of strong or poor performance, we had participants complete validated questionnaires on a range of areas including their personality, risk preferences, social network, and habits relating to physical activity, eating, and sleep. Trial enrollment was conducted in partnership with Deloitte Consulting and included employees from 40 states across the US.ConclusionThe STEP UP Trial represents a scalable model and interventions found to be effective could be deployed more broadly to increase physical activity.Trial registrationClinicaltrials.gov Identifier: NCT03311230
Business analytics, occupying the intersection of the worlds of management science, computer science and statistical science, is a potent force for innovation in both the private and public sectors. The successes of business analytics in strategy, process optimization and competitive advantage has led to data being increasingly recognized as a valuable asset in many organizations. In recent years, thanks to a dramatic increase in the volume, variety and velocity of data, the loosely defined concept of "Big Data" has emerged as a topic of discussion in its own right -- with different viewpoints in both the business and technical worlds. From our perspective, it is important for discussions of "Big Data" to start from a well-defined business goal, and remain moored to fundamental principles of both cost/benefit analysis as well as core statistical science. This note discusses some business case considerations for analytics projects involving "Big Data", and proposes key questions that businesses should ask. With practical lessons from Big Data deployments in business, we also pose a number of research challenges that may be addressed to enable the business analytics community bring best data analytic practices when confronted with massive data sets.
( Headache 2010;50:219‐223) Objective.— To evaluate the effectiveness of nonpharmacologic treatment for migraine in children younger than age 6 years. Background.— The mean age of onset of migraine in children is 7.2 years for boys and 10.9 years for girls. Treatment consists of individually tailored pharmacologic and nonpharmacologic interventions. However, data on migraine management in preschoolers are very sparse. Methods.— Demographic, clinical, and outcome data were collected from the files of patients with migraine who attended a pediatric headache clinic. Only those treated by nonpharmacologic measures, namely, good sleep hygiene, diet free of food additives, and limited sun exposure, were included. Clinical factors and response to treatment were compared between children younger than 6 years and older children. Results.— Of the 92 children identified, 32 were younger than 6 years and 60 were older. There was no difference between the age groups in most of the demographic and clinical parameters. The younger group was characterized by a significantly lower frequency of migraine attacks and shorter disease duration (in months). Mean age of the patients with no response to treatment (grade 1) was 10.588 ± 3.254 years; partial response (grade 2), 9.11 ± 4.6 years; and complete response (grade 3), 8.11 ± 3.93 years ( P = .02). The percentage of patients with complete to partial response as opposed to no response was significantly higher in the younger group ( P = .00075). Conclusion.— As the primary option, conservative therapy for migraine appears to be more effective in children younger than 6 years than in older children, perhaps because of their shorter duration of disease until treatment and lower frequency of attacks.
The characteristics of nontuberculous mycobacteria cheek lesions in 7 children were reviewed. The lesions usually presented as nontender erythematous nodules and were associated with a positive purified protein derivate tuberculin skin test. Mycobacterium haemophilum was isolated in 4 cases (57%) and Mycobacterium avium complex in 3 (43%). Cytology and imaging were noncontributory. Resolution was prolonged.
Classifying nodes in networks is a task with a wide range of applications. It can be particularly useful in anomaly and fraud detection. Many resources are invested in the task of fraud detection due to the high cost of fraud, and being able to automatically detect potential fraud quickly and precisely allows human investigators to work more efficiently. Many data analytic schemes have been put into use; however, schemes that bolster link analysis prove promising. This work builds upon the belief propagation algorithm for use in detecting collusion and other fraud schemes. We propose an algorithm called SNARE (Social Network Analysis for Risk Evaluation). By allowing one to use domain knowledge as well as link knowledge, the method was very successful for pinpointing misstated accounts in our sample of general ledger data, with a significant improvement over the default heuristic in true positive rates, and a lift factor of up to 6.5 (more than twice that of the default heuristic). We also apply SNARE to the task of graph labeling in general on publicly-available datasets. We show that with only some information about the nodes themselves in a network, we get surprisingly high accuracy of labels. Not only is SNARE applicable in a wide variety of domains, but it is also robust to the choice of parameters and highly scalable-linearly with the number of edges in a graph.
In recent years, there have been several large accounting frauds where a company's financial results have been intentionally misrepresented by billions of dollars. In response, regulatory bodies have mandated that auditors perform analytics on detailed financial data with the intent of discovering such misstatements. For a large auditing firm, this may mean analyzing millions of records from thousands of clients. This paper proposes techniques for automatic analysis of company general ledgers on such a large scale, identifying irregularities - which may indicate fraud or just honest errors - for additional review by auditors. These techniques have been implemented in a prototype system, called Sherlock, which combines aspects of both outlier detection and classification. In developing Sherlock, we faced three major challenges: developing an efficient process for obtaining data from many heterogeneous sources, training classifiers with only positive and unlabeled examples, and presenting information to auditors in an easily interpretable manner. In this paper, we describe how we addressed these challenges over the past two years and report on experiments evaluating Sherlock.
Retrieving information from heterogeneous database systems involves a complex process and remains a challenging research area. We propose a cognitively guided approach for developing an information-retrieval agent that takes the user's information request, identifies relevant information sources, and generates a multidatabase access plan. Our work is distinctive in that the agent design is based on an empirical study of how human experts retrieve information from multiple, heterogeneous database systems. To improve on empirically observed information-retrieval capabilities, the design incorporates mathematical models and algorithmic components. These components optimize the set of information sources that need to be considered to respond to a user query and are used to develop efficient multidatabase-access plans. This agent design, which integrates cognitive and mathematical models, has been implemented using Soar, a knowledge-based architecture.
From the mounds of raw information available electronically today, what professionals really need are targeted, timely nuggets of knowledge that can guide the solution to business problems. Today's common information tools -Web full-text search engines and the like - do not fully support this conversion of raw information into knowledge. In examining the common knowledge management problems faced by Price Waterhouse professionals, we have found that converting information to knowledge requires not only finding raw information, but also filtering through it for relevance, formatting it appropriately for the knowledge task at hand, and forwarding it to the right people. A fifth stage, feedback from the users, can allow the effectiveness of each stage to increase with time. In this paper, we describe each stage of this knowledge cycle and discuss the potential role that AI- based technology can play in its automation. We illustrate the possibilities through case studies of deployed knowledge management tools we have built at Price Waterhouse. These tools demonstrate that for targeted business tasks, AI-based technology can potentially facilitate much of the knowledge cycle, providing users with useful business knowledge that provides competitive advantage.
We describe an intelligent assistant that is being designed to help information specialists select and combine data sources to produce new information services. Rather than fully planning out and an- swering queries autonomously, the assistant will provide interactive knowledge-based support for a specialist designing a query plan. The assistant uses heuristics and constraints to help its user navigate through a potentially large set of data sources to locate those needed to produce a de- sired information service.
We explore the potential to include automatic learning in a design agent by implementing a simple distillation sequencing system, CPD-Soar, within Soar. Soar is an integrated software architecture with a build-in set of mechanisms for exhibiting intelligent behavior, including problem-solving, learning and interaction with the environment. Soar has a number of scientific uses: computer scientists build artificially intelligent agents with Soar as a foundation, and cognitive psychologists use Soar to model human cognition. CPD-Soar illustrates how design-related tasks can be cast within Soar's framework, hence demonstrating the functioning and potential of its problem-solving and learning mechanisms. This simple example system, which involves computations with real numbers, automatically learns things which are too specific, leading to the hypothesis that the generalization an agent infers from specific examples is strongly dependent upon the model the agent brings to the learning process.