Large Language Models are traditionally finetuned on large instruction datasets. However recent studies suggest that small, high-quality datasets can suffice for general purpose instruction following. This lack of consensus surrounding finetuning best practices is in part due to rapidly diverging approaches to LLM evaluation. In this study, we ask whether a small amount of diverse finetuning samples can improve performance on both traditional perplexity-based NLP benchmarks, and on open-ended, model-based evaluation. We finetune open-source MPT-7B and MPT-30B models on instruction finetuning datasets of various sizes ranging from 1k to 60k samples. We find that subsets of 1k-6k instruction finetuning samples are sufficient to achieve good performance on both (1) traditional NLP benchmarks and (2) model-based evaluation. Finally, we show that mixing textbook-style and open-ended QA finetuning datasets optimizes performance on both evaluation paradigms.
We study the behavior of an economic platform (e.g., Amazon, Uber Eats, Instacart) under shocks, such as COVID-19 lockdowns, and the effect of different regulation considerations. To this end, we develop a multi-agent simulation environment of a platform economy in a multi-period setting where shocks may occur and disrupt the economy. Buyers and sellers are heterogeneous and modeled as economically-motivated agents, choosing whether or not to pay fees to access the platform. We use deep reinforcement learning to model the fee-setting and matching behavior of the platform, and consider two major types of regulation frameworks: (1) taxation policies and (2) platform fee restrictions. We offer a number of simulated experiments that cover different market settings and shed light on regulatory tradeoffs. Our results show that while many interventions are ineffective with a sophisticated platform actor, we identify a particular kind of regulation—fixing fees to the optimal, no-shock fees while still allowing a platform to choose how to match buyers and sellers—as holding promise for promoting the efficiency and resilience of the economic system.
Real-world economies can be modeled as a network with many heterogeneous and strategic agents. In this setting, it is very challenging to find optimal mechanisms, e.g., taxes, 1) when taking strategic best responses into account and 2) even when using restrictive assumptions, e.g., that supply always meets demand. Deep multi-agent reinforcement learning (MARL) is a natural framework to learn mechanisms and model strategic best responses, but independent MARL often collapses to trivial solutions (e.g., where nobody works) as joint exploration severely distorts rewards and constraints. Here, we show how to use structured learning curricula and GPU-accelerated simulations to find non-trivial solutions in networks with many heterogeneous agents. We validate our approach in models with 100 worker-consumers, 10 firms, and a social planner who taxes and redistributes. We use empirical best-response analyses across agent types to show that it is difficult for agents to benefit by deviating from the learned solutions. In particular, we find income and corporate taxes that achieve 15% higher social welfare compared to baselines.
Real economies can be modeled as a sequential imperfect-information game with many heterogeneous agents, such as consumers, firms, and governments. Dynamic general equilibrium (DGE) models are often used for macroeconomic analysis in this setting. However, finding general equilibria is challenging using existing theoretical or computational methods, especially when using microfoundations to model individual agents. Here, we show how to use deep multi-agent reinforcement learning (MARL) to find $\epsilon$-meta-equilibria over agent types in microfounded DGE models. Whereas standard MARL fails to learn non-trivial solutions, our structured learning curricula enable stable convergence to meaningful solutions. Conceptually, our approach is more flexible and does not need unrealistic assumptions, e.g., continuous market clearing, that are commonly used for analytical tractability. Furthermore, our end-to-end GPU implementation enables fast real-time convergence with a large number of RL economic agents. We showcase our approach in open and closed real-business-cycle (RBC) models with 100 worker-consumers, 10 firms, and a social planner who taxes and redistributes. We validate the learned solutions are $\epsilon$-meta-equilibria through best-response analyses, show that they align with economic intuitions, and show our approach can learn a spectrum of qualitatively distinct $\epsilon$-meta-equilibria in open RBC models. As such, we show that hardware-accelerated MARL is a promising framework for modeling the complexity of economies based on microfoundations.
Optimizing economic and public policy is critical to address socioeconomic issues and trade-offs, e.g., improving equality, productivity, or wellness, and poses a complex mechanism design problem. A policy designer needs to consider multiple objectives, policy levers, and behavioral responses from strategic actors who optimize for their individual objectives. Moreover, real-world policies should be explainable and robust to simulation-to-reality gaps, e.g., due to calibration issues. Existing approaches are often limited to a narrow set of policy levers or objectives that are hard to measure, do not yield explicit optimal policies, or do not consider strategic behavior, for example. Hence, it remains challenging to optimize policy in real-world scenarios. Here we show that the AI Economist framework enables effective, flexible, and interpretable policy design using two-level reinforcement learning (RL) and data-driven simulations. We validate our framework on optimizing the stringency of US state policies and Federal subsidies during a pandemic, e.g., COVID-19, using a simulation fitted to real data. We find that log-linear policies trained using RL significantly improve social welfare, based on both public health and economic outcomes, compared to past outcomes. Their behavior can be explained, e.g., well-performing policies respond strongly to changes in recovery and vaccination rates. They are also robust to calibration errors, e.g., infection rates that are over or underestimated. As of yet, real-world policymaking has not seen adoption of machine learning methods at large, including RL and AI-driven simulations. Our results show the potential of AI to guide policy design and improve social welfare amidst the complexity of the real world.
AI and reinforcement learning (RL) have improved many areas, but are not yet widely adopted in economic policy design, mechanism design, or economics at large. At the same time, current economic methodology is limited by a lack of counterfactual data, simplistic behavioral models, and limited opportunities to experiment with policies and evaluate behavioral responses. Here we show that machine-learning-based economic simulation is a powerful policy and mechanism design framework to overcome these limitations. The AI Economist is a two-level, deep RL framework that trains both agents and a social planner who co-adapt, providing a tractable solution to the highly unstable and novel two-level RL challenge. From a simple specification of an economy, we learn rational agent behaviors that adapt to learned planner policies and vice versa. We demonstrate the efficacy of the AI Economist on the problem of optimal taxation. In simple one-step economies, the AI Economist recovers the optimal tax policy of economic theory. In complex, dynamic economies, the AI Economist substantially improves both utilitarian social welfare and the trade-off between equality and productivity over baselines. It does so despite emergent tax-gaming strategies, while accounting for agent interactions and behavioral change more accurately than economic theory. These results demonstrate for the first time that two-level, deep RL can be used for understanding and as a complement to theory for economic design, unlocking a new computational learning-based approach to understanding economic policy.
Tackling real-world socio-economic challenges requires designing and testing economic policies. However, this is hard in practice, due to a lack of appropriate (micro-level) economic data and limited opportunity to experiment. In this work, we train social planners that discover tax policies in dynamic economies that can effectively trade-off economic equality and productivity. We propose a two-level deep reinforcement learning approach to learn dynamic tax policies, based on economic simulations in which both agents and a government learn and adapt. Our data-driven approach does not make use of economic modeling assumptions, and learns from observational data alone. We make four main contributions. First, we present an economic simulation environment that features competitive pressures and market dynamics. We validate the simulation by showing that baseline tax systems perform in a way that is consistent with economic theory, including in regard to learned agent behaviors and specializations. Second, we show that AI-driven tax policies improve the trade-off between equality and productivity by 16% over baseline policies, including the prominent Saez tax framework. Third, we showcase several emergent features: AI-driven tax policies are qualitatively different from baselines, setting a higher top tax rate and higher net subsidies for low incomes. Moreover, AI-driven tax policies perform strongly in the face of emergent tax-gaming strategies learned by AI agents. Lastly, AI-driven tax policies are also effective when used in experiments with human participants. In experiments conducted on MTurk, an AI tax policy provides an equality-productivity trade-off that is similar to that provided by the Saez framework along with higher inverse-income weighted social welfare.
While using shaped rewards can be beneficial when solving sparse reward tasks, their successful application often requires careful engineering and is problem specific. For instance, in tasks where the agent must achieve some goal state, simple distance-to-goal reward shaping often fails, as it renders learning vulnerable to local optima. We introduce a simple and effective model-free method to learn from shaped distance-to-goal rewards on tasks where success depends on reaching a goal state. Our method introduces an auxiliary distance-based reward based on pairs of rollouts to encourage diverse exploration. This approach effectively prevents learning dynamics from stabilizing around local optima induced by the naive distance-to-goal reward shaping and enables policies to efficiently solve sparse reward tasks. Our augmented objective does not require any additional reward engineering or domain expertise to implement and converges to the original sparse objective as the agent learns to solve the task. We demonstrate that our method successfully solves a variety of hard-exploration tasks (including maze navigation and 3D construction in a Minecraft environment), where naive distance-based reward shaping otherwise fails, and intrinsic curiosity and reward relabeling strategies exhibit poor performance.
Acquiring abilities in the absence of a task-oriented reward function is at the frontier of reinforcement learning research. This problem has been studied through the lens of empowerment, which draws a connection between option discovery and information theory. Information-theoretic skill discovery methods have garnered much interest from the community, but little research has been conducted in understanding their limitations. Through theoretical analysis and empirical evidence, we show that existing algorithms suffer from a common limitation -- they discover options that provide a poor coverage of the state space. In light of this, we propose 'Explore, Discover and Learn' (EDL), an alternative approach to information-theoretic skill discovery. Crucially, EDL optimizes the same information-theoretic objective derived from the empowerment literature, but addresses the optimization problem using different machinery. We perform an extensive evaluation of skill discovery methods on controlled environments and show that EDL offers significant advantages, such as overcoming the coverage problem, reducing the dependence of learned skills on the initial state, and allowing the user to define a prior over which behaviors should be learned. Code is publicly available at this https URL.
In many real-world scenarios, an autonomous agent often encounters various tasks within a single complex environment. We propose to build a graph abstraction over the environment structure to accelerate the learning of these tasks. Here, nodes are important points of interest (pivotal states) and edges represent feasible traversals between them. Our approach has two stages. First, we jointly train a latent pivotal state model and a curiosity-driven goal-conditioned policy in a task-agnostic manner. Second, provided with the information from the world graph, a high-level Manager quickly finds solution to new tasks and expresses subgoals in reference to pivotal states to a low-level Worker. The Worker can then also leverage the graph to easily traverse to the pivotal states of interest, even across long distance, and explore non-locally. We perform a thorough ablation study to evaluate our approach on a suite of challenging maze tasks, demonstrating significant advantages from the proposed framework over baselines that lack world graph knowledge in terms of performance and efficiency.
Deep learning has achieved remarkable successes in solving challenging reinforcement learning (RL) problems. However, it still often suffers from the need to engineer a reward function that not only reflects the task but is also carefully shaped. This limits the applicability of RL in the real world. It is therefore of great practical importance to develop algorithms which can learn from unshaped, sparse reward signals, e.g. a binary signal indicating successful task completion. We propose a novel method called competitive experience replay, which efficiently supplements a sparse reward by placing learning in the context of an exploration competition between a pair of agents. Our method complements the recently proposed hindsight experience replay (HER) by inducing an automatic exploratory curriculum. We evaluate our approach on the tasks of reaching various goal locations in an ant maze and manipulating objects with a robotic arm. Each task provides only binary rewards indicating whether or not the goal is completed. Our method asymmetrically augments these sparse rewards for a pair of agents each learning the same task, creating a competitive game designed to drive exploration. Extensive experiments demonstrate that this method leads to faster converge and improved task performance.
Questions that require counting a variety of objects in images remain a major challenge in visual question answering (VQA). The most common approaches to VQA involve either classifying answers based on fixed length representations of both the image and question or summing fractional counts estimated from each section of the image. In contrast, we treat counting as a sequential decision process and force our model to make discrete choices of what to count. Specifically, the model sequentially selects from detected objects and learns interactions between objects that influence subsequent selections. A distinction of our approach is its intuitive and interpretable output, as discrete counts are automatically grounded in the image. Furthermore, our method outperforms the state of the art architecture for VQA on multiple metrics that evaluate counting.
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . iii List of figures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . viii Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix
OBJECTIVES We estimated the seroprevalence of both acute and chronic HIV infection by using a random sample of emergency department (ED) patients from a region of the United States with low-to-moderate HIV prevalence. METHODS This cross-sectional seroprevalence study consecutively enrolled patients aged 18 to 64 years within randomly selected sampling blocks in a Midwestern urban ED in a region of lower HIV prevalence in 2008 to 2009. Participants were compensated for providing a blood sample and health information. After de-identification, we assayed samples for HIV antibody and nucleic acid. RESULTS There were 926 participants who consented and enrolled. Overall, prevalence of undiagnosed HIV was 0.76% (95% confidence interval [CI] = 0.30%, 1.56%). Three participants (0.32%; 95% CI = 0.09%, 0.86%) were nucleic acid-positive but antibody-negative and 4 (0.43%; 95% CI = 0.15%, 1.02%) were antibody-positive. CONCLUSIONS Even when the absolute prevalence is low, a considerable proportion of undetected HIV cases in an ED population are acute. Identification of acute HIV in ED settings should receive increased priority.
Objective:Universal HIV screening is recommended but challenging to implement. Selectively targeting those at risk is thought to miss cases, but previous studies are limited by narrow risk criteria, incomplete implementation, and absence of direct comparisons. We hypothesized that targeted HIV screening, when fully implemented and using maximally broad risk criteria, could detect nearly as many cases as universal screening with many fewer tests. Methods:This single-center cluster-randomized trial compared universal and targeted patient selection for HIV screening in a lower prevalence urban emergency department. Patients were excluded for age (<18 and >64 years), known HIV infection, or previous approach for HIV testing that day. Targeted screening was offered for any risk indicator identified from charts, staff referral, or self-disclosure. Universal screening was offered regardless of risk. Baseline seroprevalence was estimated from consecutive deidentified blood samples. Results:There were 9572 eligible visits during which the patient was approached. For universal screening, 40.8% (1915/4692) consented with 6 being newly diagnosed [0.31%, 95% confidence interval (CI): 0.13% to 0.65%]. For targeted screening, 37% (1813/4880) had no testing indication. Of the 3067 remaining, 47.4% (1454) consented with 3 being newly diagnosed (0.22%, 95% CI: 0.06% to 0.55%). Estimated seroprevalence was 0.36% (95% CI: 0.16% to 0.70%). Targeted screening had a higher proportion consenting (47.4% vs. 40.8%, P < 0.002), but a lower proportion of ED encounters with testing (29.7% vs. 40.7%, P < 0.002). Conclusions:Targeted screening, even when fully implemented with maximally permissive selection, offered no important increase in positivity rate or decrease in tests performed. Universal screening diagnosed more cases, because more were tested, despite a modestly lower consent rate.
Approximately 41,000 central line-associated blood stream infections (CLABSI) occur in U.S. hospitals each year. These infections typically cause prolongation of hospital stay, increased cost, risk of mortality and are a National Patient Safety Imperative. CLABSI can be minimized by implementation of published quality improvement, infection control initiatives and has been driven as low as 1.17 infections per 1000 patient days in the 350 hospital On the CUSP: Stop BSI AHRQ funded project. Central venous catheters (CVCs) are commonly placed in the emergency department (ED) setting and represent a core procedure for the emergency physician. Reporting is not specifically required for CLABSI rates of lines placed in the ED's outpatient setting, however, examination of these lines represents an important and rarely reported area of possible patient morbidity and safety. Our study objective was to examine the CLABSI rates for lines inserted in the ED setting at three large urban EDs. This IRB approved, cohort study includes all patients with a CVC inserted in the EDs of three hospitals in Cincinnati, Ohio between January, 2007 and June, 2012. The hospitals include a large, urban, academic medical center, and two smaller, community hospitals representing 160,000 patient visits per year. Patients were identified through billing and ICD-9 data followed by chart review utilizing a priori definitions and standardized case report forms by a trained research RN. A representative sample was examined by the investigators to determine average number of line days. Blood stream infections were identified through standard, daily, hospital-based quality/infection control RN rounds. During the study period, 1932 lines were placed. Average total complication rate was 6.45%. Average line duration was 3 days and 3 BSIs were identified, for a CLABSI rate of 0.52/1000 patient days. Reporting of CLABSI rates in the outpatient setting of the ED is not yet required by government agencies or payers in general. However, this study shows that it is possible, with implementation of recommended infection control practices, to improve upon published rates of inpatient unit CLABSI, in the unique environment of the ED.
In the United States, physicians insert more than 5 million central venous catheters (CVCs) every year. CVC insertion is a core procedure for emergency physicians. Unfortunately, the use of CVCs is associated with adverse events that increase cost, and patient morbidity. Reported mechanical and thrombotic complication rates vary from 1.5 to 19% and 2 to 26%, respectively. The complication rates of CVC inserted in the ED by emergency physicians have not been reported and could be an under-recognized area of patient safety and morbidity. Our study objective is to estimate the mechanical and thrombotic complication rates for CVCs inserted by emergency physicians. This cohort study includes all patients with a CVC inserted in the EDs of three hospitals in Cincinnati, Ohio between January 2007 and June 2012. The hospitals include an urban, academic medical center, and 2 smaller, community hospitals. Combined volume is ∼160,000 visits per year. Patients with CVC placement were identified through billing and ICD-9 data, and evaluated by chart review using a priori definitions and standardized case report forms by a trained research nurse. During the study period, 1932 lines were placed. Complication rates are reported in the Table. Both mechanical and thrombotic complications are low in this large cohort of emergency medicine patients relative to reported rates in other settings. The majority of CVCs in this study were placed by PGY-2 emergency medicine residents and of the 121 events, only 20 (1.0%) required intervention. More study is required to further characterize this important area of EM patient safety in order to establish best practice parameters and to minimize future morbidity.TableCVC Insertion Complication RatesClassMechanicalThromboticTypeArterial PuncturePneumothoraxHemothoraxDysrythmiaAir EmbolismDVTHematomaTotalN8113162414121%4.20.670.050.310.100.210.726.3 Open table in a new tab
Background: Lumbar puncture is a common emergency department (ED) procedure with widely variable complication rates reported between 0.98-19% in various experimental protocol driven settings. In the current emergency medicine environment of ABEM required performance improvement, Flexner report recommendations for quality improvement and CMS non-pay for complications, it is vital that robust performance improvement data be reported that reflect actual practice experience, without the artificiality of experimentally controlled conditions. These performance improvement results must be disseminated for comparison, benchmarking and hypothesis generation. This study reports the results of an ongoing performance improvement initiative to reduce lumbar puncture complication rates and to improve patient outcomes. Study Objectives: We hypothesize that multiple, knowledge translation-driven interventions improve the rate of post-dural puncture headache. Methods: This is an institutional review board approved, retrospective analysis of 3 hospitals, including a large urban trauma center, and 2 community hospitals. The operators were emergency medicine residents and attending physicians. Complications were identified from equipment charge data, followed by chart review utilizing standardized definitions and a CRF by a trained quality RN who was blinded to the hypothesis. Data were analyzed with SPSS 18.0 for Windows (SPSS Inc., Chicago, IL). Results: (See Chart and Table) A total of 1,049 lumbar puncture's were performed over 4 years. 41 (3.8%) patients experienced post-dural puncture headaches, 17 (5.9%) in 2007 falling to 6 (2.4%)in 2010. The probability of complication fell from 0.06 to 0.02 (p=0.0602). No other complications were found. Five interventions relating to the performance of lumbar puncture were instituted during the study period.Tabled 1 Conclusion: While the results do not prove causation, and just missed arbitrary significance of 0.05, multiple, sequential, continuous performance improvement interventions trend toward improved post-dural puncture headache rates in our setting.
Mate selection is critical to ensuring the survival of a species. In the fruit fly, Drosophila melanogaster, genetic and anatomical studies have focused on mate recognition and courtship initiation for decades. This model system has proven to be highly amenable for the study of neural systems controlling the decision making process. However, much less is known about how courtship quality is regulated in a temporally dynamic manner in males and how a female assesses male performance as she makes her decision of whether to accept copulation. Here, we report that the courting male dynamically adjusts the relative proportions of the song components, pulse song or sine song, by assessing female locomotion. Male flies deficient for olfaction failed to perform the locomotion-dependent song modulation, indicating that olfactory cues provide essential information regarding proximity to the target female. Olfactory mutant males also showed lower copulation success when paired with wild-type females, suggesting that the male's ability to temporally control song significantly affects female mating receptivity. These results depict the consecutive inter-sex behavioral decisions, in which a male smells the close proximity of a female as an indication of her increased receptivity and accordingly coordinates his song choice, which then enhances the probability of his successful copulation.
OBJECTIVE:Controversy surrounds the linkage of prevention counseling with emergency department (ED)-based HIV testing. Further, the effectiveness and feasibility of prevention counseling in the ED setting is unknown. We investigate these issues by conducting a preliminarily exploration of several related aspects of our ED's HIV prevention counseling and testing program. METHODS:Our urban, academic ED provides formal client-centered prevention counseling in conjunction with HIV testing. Five descriptive, exploratory observations were conducted, involving surveys and analysis of electronic medical records and programmatic data focused on (1) patient perception and feasibility of prevention counseling in the ED, (2) patient perceptions of the need to link prevention counseling with testing, and (3) potential effectiveness of providing prevention counseling in conjunction with ED-based HIV testing. RESULTS:Of 110 ED patients surveyed after prevention counseling and testing, 98% believed privacy was adequate, and 97% reported that their questions were answered. Patients stated that counseling would lead to improved health (80%), behavioral changes (72%), follow-up testing (77%), and discussion with partners (74%). However, 89% would accept testing without counseling, 32% were willing to seek counseling elsewhere, and 26% preferred not to receive the counseling. Correct responses to a 16-question knowledge quiz increased by 1.6 after counseling (95% confidence interval 1.3 to 12.0). The program completed counseling for 97% of patients tested; however, 6% of patients had difficulty recalling the encounter and 13% denied received testing. Among patients undergoing repeated testing, there was no consistent change in self-reported risk behaviors. CONCLUSION:Participants in the ED prevention counseling and testing program considered counseling acceptable and useful, though not required. Given adequate resources, prevention counseling can be provided in the ED, but it is unlikely that all patients benefit.