Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can correct errors in these automated labels using a probability-sampled audit set, but they usually treat the audit labels as correct. In practice, human audit labels are often noisy, and only some audited items are reviewed by an expert or adjudicator. We propose Partially Adjudicated Design-Based Supervised Learning (PA-DSL), a method for this setting. It uses adjudicated cases to correct noisy human labels and then uses the corrected audit information to debias analyses based on the full set of automated labels. The estimator is valid for a broad class of downstream analyses when the audit and adjudication probabilities are known. In synthetic and Wikipedia Detox semi-synthetic experiments, PA-DSL maintains nominal coverage and reduces RMSE by 10-17
Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic variation arising from ambiguous items, divergent interpretations, and simple mistakes. Machine learning research commonly treats all disagreement as noise, which obscures these distinctions and limits our understanding of what models actually learn. This paper reframes annotation as a measurement process and introduces a statistical framework for decomposing labeling outcomes into interpretable sources of variation: instance difficulty, annotator bias, situational noise, and relational alignment. The framework extends classical measurement-error models to accommodate both shared and individualized notions of truth, reflecting traditional and human label variation interpretations of error, and provides a diagnostic for assessing which regime better characterizes a given task. Applying the proposed model to a multi-annotator natural language inference dataset, we find empirical evidence for all four theorized components and demonstrate the effectiveness of our approach. We conclude with implications for data-centric machine learning and outline how this approach can guide the development of a more systematic science of labeling.
Data extraction is a critical but error-prone and labor-intensive task in evidence synthesis. Unlike other artificial intelligence (AI) technologies, large language models (LLMs) do not require labeled training data for data extraction. To compare an AI-assisted to a traditional y data extraction process. Study within reviews (SWAR) utilizing a prospective, parallel group comparison with blinded data adjudicators. Workflow validation within six ongoing systematic reviews of interventions under real-world conditions. Initial data extraction using an LLM (Claude versions 2.1, 3.0 Opus, and 3.5 Sonnet) verified by a human reviewer. Concordance, time on task, accuracy, recall, precision, and error analysis. The six systematic reviews of the SWAR contributed 9,341 data elements, extracted from 63 studies. Concordance between the two methods was 77.2%. The accuracy of the AI-assisted approach compared with enhanced human data extraction was 91.0%, with a recall of 89.4% and a precision of 98.9%. The AI-assisted approach had fewer incorrect extractions (9.0% vs. 11.0%) and similar risks of major errors (2.5% vs. 2.7%) compared to the traditional human-only method, with a median time saving of 41 minutes per study. Missed data items were the most frequent errors in both approaches. Assessing the concordance of data extractions and classifying errors required subjective judgment. Tracking time on task consistently was challenging. The use of an LLM can improve accuracy of data extraction and save time in evidence synthesis. Results reinforce previous findings that human-only data extraction is prone to errors. US Agency for Healthcare Research and Quality, RTI International SWAR28 Gerald Gartlehner (2023 FEB 11 2102).pdf
Differential privacy (DP) is becoming increasingly important for deployed machine learning applications because it provides strong guarantees for protecting the privacy of individuals whose data is used to train models. However, DP mechanisms commonly used in machine learning tend to struggle on many real world distributions, including highly imbalanced or small labeled training sets. In this work, we propose a new scalable DP mechanism for deep learning models, SWAG-PPM, by using a pseudo posterior distribution that downweights by-record likelihood contributions proportionally to their disclosure risks as the randomized mechanism. As a motivating example from official statistics, we demonstrate SWAG-PPM on a workplace injury text classification task using a highly imbalanced public dataset published by the U.S. Occupational Safety and Health Administration (OSHA). We find that SWAG-PPM exhibits only modest utility degradation against a non-private comparator while greatly outperforming the industry standard DP-SGD for a similar privacy budget.
BACKGROUND:Data extraction is a critical but error-prone and labor-intensive task in evidence synthesis. Unlike other artificial intelligence (AI) technologies, large language models (LLMs) do not require labeled training data for data extraction. OBJECTIVE:To compare an AI-assisted versus a traditional, human-only data extraction process. DESIGN:Study within reviews (SWAR) using a prospective, parallel-group comparison with blinded data adjudicators. SETTING:Workflow validation within 6 ongoing systematic reviews of interventions under real-world conditions. INTERVENTION:Initial data extraction using an LLM (Claude, versions 2.1, 3.0 Opus, and 3.5 Sonnet) verified by a human reviewer. MEASUREMENTS:Concordance, time on task, accuracy, sensitivity, positive predictive value, and error analysis. RESULTS:The 6 systematic reviews in the SWAR yielded 9341 data elements from 63 studies. Concordance between the 2 methods was 77.2% (95% CI, 76.3% to 78.0%). Compared with the reference standard, the AI-assisted approach had an accuracy of 91.0% (CI, 90.4% to 91.6%) and the human-only approach an accuracy of 89.0% (CI, 88.3% to 89.6%). Sensitivities were 89.4% (CI, 88.6% to 90.1%) and 86.5% (CI, 85.7% to 87.3%), respectively, with positive predictive values of 99.2% (CI, 99.0% to 99.4%) and 98.9% (CI, 98.6% to 99.1%). Incorrect data were extracted in 9.0% (CI, 8.4% to 9.6%) of AI-assisted cases and 11.0% (CI, 10.4% to 11.7%) of human-only cases, with corresponding proportions of major errors of 2.5% (CI, 2.2% to 2.8%) versus 2.7% (CI, 2.4% to 3.1%). Missed data items were the most frequent error type in both approaches. The AI-assisted method reduced data extraction time by a median of 41 minutes per study. LIMITATIONS:Assessing concordance and classifying errors required subjective judgment. Consistently tracking time on task was challenging. CONCLUSION:Data extraction assisted by AI may offer a viable, more efficient alternative to human-only methods. PRIMARY FUNDING SOURCE:Agency for Healthcare Research and Quality and RTI International.
BACKGROUND:In 2021, we used the National COVID Cohort Collaborative (N3C) as part of the National Institutes of Health RECOVER Initiative to develop a machine learning pipeline to identify patients with a high probability of having post-acute sequelae of SARS-CoV-2 infection or long COVID. However, the increased home testing, missing documentation, and reinfections that characterise the pandemic beyond 2022 necessitated the re-engineering of our original model to account for these changes in the COVID-19 research landscape. METHODS:Trained on 72 745 patient records (36 238 with long COVID and 36 507 with no evidence of long COVID), our updated XGBoost model gathered data for each patient in overlapping 100-day periods that progressed through time and issued a probability of long COVID for each 100-day period. We ran the model on patients in N3C (n=5 875 065) who met at least one of the following criteria from Jan 1, 2020, to June 22, 2023: a U07·1 (COVID-19) diagnosis code; a positive SARS-CoV-2 test; a U09·9 (post-acute sequelae of SARS-CoV-2 infection) diagnosis code; a prescription for nirmatrelvir-ritonavir or remdesivir; or an M35·81 (multisystem inflammatory syndrome in children [MIS-C]) diagnosis code. Each patient was given a model score that predicted long COVID status for each 100-day window in which they were aged ≥18 years. If a patient had known acute COVID-19 during any 100-day window (including reinfections), we censored the data from 7 days before the diagnosis or positive test date to 28 days after. We ran the model on controls selected from pre-2020 data to assess the likelihood of false positives. FINDINGS:The updated model had an area under the receiver operating characteristic curve of 0·90. Precision and recall could be adjusted according to a given use case, depending on whether greater sensitivity or specificity was warranted. Using our model, we estimate the overall prevalence of long COVID among the COVID-19 positive cohort within N3C repository to be 10.4%. INTERPRETATION:By eschewing the COVID-19 index date as an anchor point for analysis, we can assess the probability of long COVID among patients who might have tested at home, or with suspected (but untested) cases of COVID-19, or multiple SARS-CoV-2 reinfections. We view this exercise as a model for maintaining and updating any machine learning pipeline used for clinical research and operations. FUNDING:National Institutes of Health RECOVER Initiative.
Accurate data extraction is a key component of evidence synthesis and critical to valid results. The advent of publicly available large language models (LLMs) has generated interest in these tools for evidence synthesis and created uncertainty about the choice of LLM. We compare the performance of two widely available LLMs (Claude 2 and GPT-4) for extracting pre-specified data elements from 10 published articles included in a previously completed systematic review. We use prompts and full study PDFs to compare the outputs from the browser versions of Claude 2 and GPT-4. GPT-4 required use of a third-party plugin to upload and parse PDFs. Accuracy was high for Claude 2 (96.3%). The accuracy of GPT-4 with the plug-in was lower (68.8%); however, most of the errors were due to the plug-in. Both LLMs correctly recognized when prespecified data elements were missing from the source PDF and generated correct information for data elements that were not reported explicitly in the articles. A secondary analysis demonstrated that, when provided selected text from the PDFs, Claude 2 and GPT-4 accurately extracted 98.7% and 100% of the data elements, respectively. Limitations include the narrow scope of the study PDFs used, that prompt development was completed using only Claude 2, and that we cannot guarantee the open-source articles were not used to train the LLMs. This study highlights the potential for LLMs to revolutionize data extraction but underscores the importance of accurate PDF parsing. For now, it remains essential for a human investigator to validate LLM extractions.
Data extraction is a crucial, yet labor-intensive and error-prone part of evidence synthesis. To date, efforts to harness machine learning for enhancing efficiency of the data extraction process have fallen short of achieving sufficient accuracy and usability. With the advent of Large Language Models (LLMs), new possibilities have emerged to increase efficiency and accuracy of data extraction for evidence synthesis. The objective of this proof-of-concept study was to assess the performance of an LLM (Claude 2) in extracting data elements from published studies, compared with human data extraction as employed in systematic reviews. Our analysis utilized a convenience sample of 10 English-language, open-access publications of randomized controlled trials included in a single systematic review. We selected 16 distinct types of data, posing varying degrees of difficulty (160 data elements across 10 studies). We used the browser version of Claude 2 to upload the portable document format of each publication and then prompted the model for each data element. Across 160 data elements, Claude 2 demonstrated an overall accuracy of 96.3% with a high test-retest reliability (replication 1: 96.9%; replication 2: 95.0% accuracy). Overall, Claude 2 made 6 errors on 160 data items. The most common errors (n=4) were missed data items. Importantly, Claude 2’s ease of use was high; it required no technical expertise or training data for effective operation. Based on findings of our proof-of-concept study, leveraging LLMs has the potential to substantially enhance the efficiency and accuracy of data extraction for evidence syntheses.
While pregnancy has been associated with an altered immune response and distinct clinical manifestations of COVID-19, the influence of pregnancy on the persistence and severity of post-acute sequelae of SARS-CoV-2 infection (PASC), or Long COVID, remains uncertain. This study investigated PASC risk in individuals with SARS-CoV-2 infection during pregnancy and compared it with that in reproductive-age females with SARS-CoV-2 infection outside of pregnancy. This retrospective analysis identified 72,151 individuals who contracted SARS-CoV-2 during pregnancy and 1,439,354 females who contracted SARS-CoV-2 outside of pregnancy, aged 18 to 50 years old, from March 2020 to June 2023 in the National Patient-Centered Clinical Research Network (PCORnet) and the National COVID Cohort Collaborative (N3C). A comprehensive list of PASC outcomes was investigated, including a PCORnet rule-based PASC definition, an N3C PASC machine learning (ML) Phenotype, unspecified PASC ICD-10 diagnoses (ICD10 codes U09.9 or B94.8), and a cluster of cognitive, fatigue, and respiratory conditions. Overall, the estimated risk of PASC at 180 days of follow-up for those infected during pregnancy was 16.47 events per 100 persons (95% CI, 16.00 to 16.95) in the PCORnet cohort, based on the PCORnet rule-based PASC definition, and 4.37 events per 100 persons (95% CI, 4.18 to 4.57) in the N3C cohort based on the ML model. The risks of unspecified PASC diagnoses were 0.19 events per 100 persons (95% CI, 0.14 to 0.25) in PCORnet, and 0.23 events per 100 persons (95% CI, 0.19 to 0.28) in N3C; and the risks of any post-acute cognitive, fatigue, and respiratory condition were 4.86 events per 100 persons (95% CI, 4.59 to 5.14) in PCORnet, and 6.83 events per 100 persons (95% CI, 6.59 to 7.08) in N3C. The PASC risk varied across different subpopulations within pregnant females. The observed risk factors for PASC included self-reported Black race, advanced maternal age, infection during the first two trimesters, obesity, and the presence of baseline comorbid conditions. While the findings suggest a high incidence of PASC in individuals following SARS-CoV-2 infection during pregnancy, the risk of PASC in pregnant females was lower than in matched non-pregnant females.
Deductive coding is a widely used qualitative research method for determining the prevalence of themes across documents. While useful, deductive coding is often burdensome and time consuming since it requires researchers to read, interpret, and reliably categorize a large body of unstructured text documents. Large language models (LLMs), like ChatGPT, are a class of quickly evolving AI tools that can perform a range of natural language processing and reasoning tasks. In this study, we explore the use of LLMs to reduce the time it takes for deductive coding while retaining the flexibility of a traditional content analysis. We outline the proposed approach, called LLM-assisted content analysis (LACA), along with an in-depth case study using GPT-3.5 for LACA on a publicly available deductive coding data set. Additionally, we conduct an empirical benchmark using LACA on 4 publicly available data sets to assess the broader question of how well GPT-3.5 performs across a range of deductive coding tasks. Overall, we find that GPT-3.5 can often perform deductive coding at levels of agreement comparable to human coders. Additionally, we demonstrate that LACA can help refine prompts for deductive coding, identify codes for which an LLM is randomly guessing, and help assess when to use LLMs vs. human coders for deductive coding. We conclude with several implications for future practice of deductive coding and related research methods.
MOTIVATION:As the number of public data resources continues to proliferate, identifying relevant datasets across heterogenous repositories is becoming critical to answering scientific questions. To help researchers navigate this data landscape, we developed Dug: a semantic search tool for biomedical datasets utilizing evidence-based relationships from curated knowledge graphs to find relevant datasets and explain why those results are returned. RESULTS:Developed through the National Heart, Lung and Blood Institute's (NHLBI) BioData Catalyst ecosystem, Dug has indexed more than 15 911 study variables from public datasets. On a manually curated search dataset, Dug's total recall (total relevant results/total results) of 0.79 outperformed default Elasticsearch's total recall of 0.76. When using synonyms or related concepts as search queries, Dug (0.36) far outperformed Elasticsearch (0.14) in terms of total recall with no significant loss in the precision of its top results. AVAILABILITY AND IMPLEMENTATION:Dug is freely available at https://github.com/helxplatform/dug. An example Dug deployment is also available for use at https://search.biodatacatalyst.renci.org/. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
BACKGROUND:Social media are important for monitoring perceptions of public health issues and for educating target audiences about health; however, limited information about the demographics of social media users makes it challenging to identify conversations among target audiences and limits how well social media can be used for public health surveillance and education outreach efforts. Certain social media platforms provide demographic information on followers of a user account, if given, but they are not always disclosed, and researchers have developed machine learning algorithms to predict social media users' demographic characteristics, mainly for Twitter. To date, there has been limited research on predicting the demographic characteristics of Reddit users.OBJECTIVE:We aimed to develop a machine learning algorithm that predicts the age segment of Reddit users, as either adolescents or adults, based on publicly available data.METHODS:This study was conducted between January and September 2020 using publicly available Reddit posts as input data. We manually labeled Reddit users' age by identifying and reviewing public posts in which Reddit users self-reported their age. We then collected sample posts, comments, and metadata for the labeled user accounts and created variables to capture linguistic patterns, posting behavior, and account details that would distinguish the adolescent age group (aged 13 to 20 years) from the adult age group (aged 21 to 54 years). We split the data into training (n=1660) and test sets (n=415) and performed 5-fold cross validation on the training set to select hyperparameters and perform feature selection. We ran multiple classification algorithms and tested the performance of the models (precision, recall, F1 score) in predicting the age segments of the users in the labeled data. To evaluate associations between each feature and the outcome, we calculated means and confidence intervals and compared the two age groups, with 2-sample t tests, for each transformed model feature.RESULTS:The gradient boosted trees classifier performed the best, with an F1 score of 0.78. The test set precision and recall scores were 0.79 and 0.89, respectively, for the adolescent group (n=254) and 0.78 and 0.63, respectively, for the adult group (n=161). The most important feature in the model was the number of sentences per comment (permutation score: mean 0.100, SD 0.004). Members of the adolescent age group tended to have created accounts more recently, have higher proportions of submissions and comments in the r/teenagers subreddit, and post more in subreddits with higher subscriber counts than those in the adult group.CONCLUSIONS:We created a Reddit age prediction algorithm with competitive accuracy using publicly available data, suggesting machine learning methods can help public health agencies identify age-related target audiences on Reddit. Our results also suggest that there are characteristics of Reddit users' posting behavior, linguistic patterns, and account features that distinguish adolescents from adults.
Previous qualitative studies and data science studies using Reddit for tobacco research are limited by the lack of available demographic information. Social media investigations are often limited to manual qualitative coding or machine learning classification in isolation. This study combines both machine learning methods and manual qualitative coding to provide contextual age nuance to social media analysis. By being able to predict a Redditor’s age using publicly available data, the most popular posts can be analyzed and qualitatively coded to provide nuanced comparisons on thematic topics by age group. The current study combines these two methods to 1) predict Reddit users’ age into two categories (13-20, 21-54) and 2) qualitatively code Electronic Nicotine Delivery System [ENDS] related Reddit posts within the two age groups. We identified Reddit posts on three topics: Vaping in General, Tobacco 21 Minimum Age Laws, and Flavor Restriction Policies. An age algorithm was used to predict Reddit users’ ages (13-20 or 21-54 year old users). The 25 posts with the highest karma score (number of upvotes minus number of downvotes) for each query and each predicted age group were qualitatively coded. The top three, two of which were part of the query, out of nine, topics that emerged were “Flavor Restriction Policies”, “Tobacco 21 Policies”, and “Use”. Tobacco 21 and Flavor Restriction Policy posts were prominent coding categories. Opposition to flavor restriction policies was a prominent sub-category for both groups, but more common in the 21-54 group. The 13-20 group was more likely to discuss opposition to minimum age laws as well as access to flavored ENDS products. The 21-54 group more commonly mentioned general vaping use behavior. Users predicted to be in the 13-20 age group posted about different ENDS-related topics on Reddit than users predicted to be in the 21-54 age group. Future studies could use these complementary methods with social media data to gain insights from target audiences.
Social Network Analysis (SNA) is a promising yet underutilized tool in the international development field. SNA entails collecting and analyzing data to characterize and visualize social networks, where nodes represent network members and edges connecting nodes represent relationships or exchanges among them. SNA can help both researchers and practitioners understand the social, political, and economic relational dynamics at the heart of international development programming. It can inform program design, monitoring, and evaluation to answer questions related to where people get information; with whom goods and services are exchanged; who people value, trust, or respect; who has power and influence and who is excluded; and how these dynamics change over time. This brief advances the case for use of SNA in international development, outlines general approaches, and discusses two recently conducted case studies that illustrate its potential. It concludes with recommendations for how to increase SNA use in international development.
Accurate projections of seasonal agricultural output are essential for improving food security. However, the collection of agricultural information through seasonal agricultural surveys is often not timely enough to inform public and private stakeholders about crop status during the growing season. Acquiring timely and accurate crop estimates can be particularly challenging in countries with predominately smallholder farms because of the large number of small plots, intense intercropping, and high diversity of crop types. In this study, we used RGB images collected from unmanned aerial vehicles (UAVs) flown in Rwanda to develop a deep learning algorithm for identifying crop types, specifically bananas, maize, and legumes, which are key strategic food crops in Rwandan agriculture. The model leverages advances in deep convolutional neural networks and transfer learning, employing the VGG16 architecture and the publicly accessible ImageNet dataset for pretraining. The developed model performs with an overall test set F1 of 0.86, with individual classes ranging from 0.49 (legumes) to 0.96 (bananas). Our findings suggest that although certain staple crops such as bananas and maize can be classified at this scale with high accuracy, crops involved in intercropping (legumes) can be difficult to identify consistently. We discuss the potential use cases for the developed model and recommend directions for future research in this area.
Non-typhoidal Salmonella is a significant foodborne pathogen causing over a million illnesses each year in the United States. Poultry is one of the food commodities most frequently associated with Salmonella infections. While government, research, and industry efforts have reduced Salmonella contamination in poultry to some extent, the incidence of salmonellosis has not changed significantly and is still above the public health goals of Healthy People 2020, and novel and more comprehensive approaches are needed. In this paper, the public health impact of implementing different microbiological criteria (MC) for Salmonella in chicken parts was evaluated using a quantitative risk assessment approach. Four hypothetical scenarios, including a no-action baseline and three alternative scenarios, were considered. Scenario 1 modeled a prevalence-based microbiological criterion based on the proportion of positive samples in an establishment, Scenario 2 modeled a microbiological criterion based on the concentration of Salmonella in samples, and Scenario 3 modeled a combination of the two. With exception of the baseline, all three scenarios assumed that different interventions would be adopted for non-compliant establishments (Scenario 1) or lots (Scenario 2), with Scenario 3 combining establishment-level and lot-level interventions. The product was assumed to be sold to consumers as raw, and contamination via undercooked product as well as cross contamination in consumer kitchens were considered as potential exposure routes. Risk was characterized by the probability of illness and the preventable fraction of risk, which was calculated for each scenario in comparison with the baseline. Simulation results show that, depending on the parameters of specific sampling strategies, both prevalence-based and concentration-based MC coupled with interventions could significantly lower risk (range of 60-88% in mean preventable fraction of risk). Overall, while the model is preliminary and subject to the stated limitations, it is likely that a combination approach including establishment-level and lot-level interventions would be highly effective in reducing risk and, therefore, benefit public health. The effectiveness of all MC was impacted by several assumptions and model parameters. In particular, the prevalence MC threshold and the concentration reduction associated with the establishment-level intervention impacted the preventable fraction of risk for Scenario 1, and the concentration MC threshold and the variability across lots impacted the risk outcomes for Scenario 2. Overall, high variance in risk outputs was observed, mainly associated with a high variance in concentration inputs. This model provides a risk-based approach to test different MC approaches for chicken parts at both lot and establishment levels, and over a wide range of scenarios of input contamination distributions, interventions, and consumer behaviors. Model estimates, as well as the ability to distinguish between variability and uncertainty, could be improved by additional data on the distribution of Salmonella concentrations across and within establishments.
The results of many large-scale federal or multi-site evaluations are typically compiled into long reports which end up sitting on policymaker's shelves. Moreover, the information policymakers need from these reports is often buried in the report, may not be remembered, understood, or readily accessible to the policymaker when it is needed. This is not a new challenge for evaluators, and advances in statistical methodology, while they have created greater opportunities for insight, may compound the challenge by creating multiple lenses through which evidence can be viewed. The descriptive evidence from traditional frequentist models, while familiar, are frequently misunderstood, while newer Bayesian methods provide evidence which is intuitive, but less familiar. These methods are complementary but presenting both increases the amount of evidence stakeholders and policymakers may find useful. In response to these challenges, we developed an interactive dashboard that synthesizes quantitative and qualitative data and allows users to access the evidence they want, when they want it, allowing each user a customized, and customizable view into the data collected for one large-scale federal evaluation. This offers the opportunity for policymakers to select the specifics that are most relevant to them at any moment, and also apply their own risk tolerance to the probabilities of various outcomes.