
In the evolving field of survey research, leveraging machine learning to predict response behavior has transformative potential for the efficiency of survey operations. Integrating multiple data sources may improve response prediction by providing more nuanced insights for household outreach. This study presents a model-driven approach to enhancing respondent cooperation in the Medical Expenditure Panel Survey (MEPS) by combining features from disparate data sources. MEPS is a longitudinal household survey with 5 rounds of interviewing over 2.5 years. Its sample is derived prior National Health Interview Survey (NHIS) participants. MEPS Round 1 response rates are critical for sustaining representativeness throughout each panel. In this study, we constructed a multimodal machine learning model to predict (1) the likelihood of a positive response for an upcoming contact attempt and (2) the likelihood that new panel households complete a Round 1 interview. The model integrates tract-level data from the American Community Survey (ACS), outcomes from the Advance Call Records (ACR) made prior to MEPS Round 1, and paradata from the early contact period. We also explored the relative contributions of these sources to model performance. Our model aims to help manage field labor by identifying complex cases needing specialized support. It can also assist in determining the optimal mode for the next contact to increase the chance of a completed interview. Beyond improving MEPS operations, this study offers a roadmap for incorporating additional data sources to support fieldwork.
The mean-variance portfolio model, based on the risk-return trade-off for optimal asset allocation, remains fundamental in portfolio optimization. However, its reliance on restrictive assumptions about asset return distributions limits its applicability to real-world data. Parametric copula structures provide a novel way to overcome these limitations by accounting for asymmetry, heavy tails, and time-varying dependencies. Existing methods have been shown to rely on fixed or static dependence structures, thus overlooking the dynamic nature of the financial market. In this study, a semiparametric model is proposed that combines nonparametrically estimated copulas with parametrically estimated marginals to allow all parameters to dynamically evolve over time. A novel framework was developed that integrates time-varying dependence modeling with flexible empirical beta-copula structures. Marginal distributions were modeled using the skewed generalized t-family. This effectively captures asymmetry and heavy tails and makes the model suitable for predictive inferences in real-world scenarios. Furthermore, the model was applied to rolling windows of financial returns from the USA, India, and Hong Kong economies to understand the influence of dynamic market conditions. The approach addresses the limitations of models that rely on parametric assumptions. By accounting for asymmetry, heavy tails, and cross-correlated asset prices, the proposed method offers a robust solution to optimize diverse portfolios in an interconnected financial market. Through adaptive modeling, it allows for better management of risk and return across varying economic conditions, leading to more efficient asset allocation and improved portfolio performance.
Privacy-preserving machine learning methods seek to train useful models that do not disclose information about the data on which they were trained. Such methods are vital when organizations train neural networks on sensitive individual-level data and seek to release the models publicly. Their goal poses a trade-off between predictive performance (utility) and privacy protection. That trade-off makes privacy-preserving machine learning methods difficult to apply in practice, usually requiring extensive iteration and hyperparameter tuning. Yet, practitioners often have little guidance for navigating competing statistical, computational, and privacy demands. We present an implementation algorithm for the Stochastic Weight Averaging–Gaussian Pseudo Posterior Mechanism (SWAG-PPM), a Bayesian differentially private deep learning method. The implementation algorithm focuses on the joint tuning of two key hyperparameters whose interaction governs model convergence and the privacy–utility trade-off. We introduce novel diagnostic tools to evaluate convergence and guide hyperparameter adjustments. Using a transformer model for occupational injury classification, we demonstrate that diagnostic-guided tuning with SWAG-PPM can achieve strong privacy protection and utility. While our case study uses a specific dataset and model architecture, all methodological steps can apply to other settings where privacy risk is heterogeneously distributed.
We propose a fine-grained attention-based multiple instance classification (FAMIC) model for interpretable word-level sentiment analysis (SA) using only document-level sentiment labels. By operating at the word level, FAMIC enhances interpretability while maintaining competitive performance in document-level classification. The model generates interpretable outputs such as contextual weighting, word neutrality, and negation cues, offering insights into how context shapes sentiment and how the model arrives at its predictions. FAMIC is built on a straightforward yet effective architecture that combines a multiple instance classification framework with self-attention and positionally encoded self-attention blocks. This design enables the model to capture both local and global contextual dependencies, supporting nuanced sentiment interpretation. We evaluate FAMIC on two sentiment classification datasets and provide an extensive analysis of its interpretability and performance.
Survey researchers are increasingly adopting hybrid sampling designs to address the limitations of traditional probability sampling, especially when studying rare or hard-to-reach populations. Challenges such as high screening costs, low statistical efficiency, and operational constraints make purely probability-based approaches impractical in many contexts. This article uses public data from the National Health and Nutrition Examination Survey to demonstrate how one can make population estimates from a hybrid sampling strategy that combines data from a stratified, multistage probability sample with data from a non-probability sample within the same primary sampling units as the probability sample. We outline a framework and discuss methods for analyzing data from a hybrid sample such as this, where covariates and survey outcomes are observed in both the probability and non-probability samples. We present a case study to illustrate the framework. We provide the case study R code in the supplementary material.
The self-organizing map (SOM) is an unsupervised, competitive learning neural network that projects high-dimensional data onto a low-dimensional grid, effectively showcasing the topological relationships within the original dataset. However, the conventional SOM training algorithm is restricted to numeric data. Categorical data typically needs to be converted into binary format before SOM training, which can lead to the loss of crucial similarity information between categorical values. As a result, the trained SOM may not accurately reflect the true topological order. While a training data splitting method (TDSM) can help identify perfect representative neurons and enhance clustering outcomes, the training data itself often lacks sufficient information, such as data distribution, and can be uncertain and ambiguous. Even when perfect neurons are identified, further improvements in clustering results become challenging. This paper investigates the possibility of improving the performance of supervised TDSM SOM clustering by utilizing unsupervised self-organization granule encoding for discrete data. This approach to unsupervised learning is advantageous for uncovering uncertain and ambiguous information within discrete data, leading to a more effective topological representation of the training data.
Hierarchical linear mixed models are commonly used in many scientific fields. However, without a strong statistical background, it can be hard to understand the relationships between the random effect variables and the inferences that can be made when a model has nested random effects. Visualizing relationships makes it easier for the practitioner to understand what relationships the model is capable of estimating and testing. We present an R package modeldiagramR that seamlessly creates a visualization of the model based on the data and the model object created when fitting a linear mixed model using either lme4 or nlme.
Satellite precipitation products have the potential to be employed for the purpose of better understanding extreme precipitation events in remote mountainous terrain, where weather stations and radar data tend to be sparse. For this reason, it is crucial to assess how closely satellite estimates agree with ground observations during extreme events, and how that agreement varies across such regions. We use asymptotic dependence from multivariate extreme value theory as the primary tool in this study. After presenting two measures of asymptotic dependence and their associated estimators, we illustrate these ideas using simulated data. We then model the level of asymptotic dependence between PERSIANN-CDR and SNOTEL station data over the US Northern Rocky Mountains. We consider both asymptotic dependence estimators, and based on hypothesis tests and visual diagnostics, both estimates of asymptotic dependence indicate positive spatial dependence. We also investigate whether geographical factors influence the levels of asymptotic dependence over this region. Using a spatial correlation analysis, we find that elevation is negatively correlated with both asymptotic dependence estimators and average summer temperature is positively correlated with both asymptotic dependence estimators. However, we did not find any geographical covariates to be statistically significant in the model.
This study investigates how user ability to manipulate plot features affects graphical perception, by extending a previous graphical study (Vanderplas and Hofmann, 2017) with an interactive framework. Similar to the original study, statistical lineups included two target patterns (a linear trend and a clustering pattern), as well as eighteen null plots generated from three different mixture proportions of the combined cluster and trend models. Participants were asked to select two plots that they perceived as ‘most different’, and were able to interact with the graphics by toggling aesthetic features such as cluster coloring, cluster ellipses, linear trendlines, and regression error bands. We found that toggle workflow varied across participants, revealing a divide between “maximalists,” who enabled all features, and “minimalists,” who used few or none, with most toggling occurring before the first selection. Starting features aesthetics did not have a significant effect on target choice. A generalized linear mixed model identified mixture proportion as the strongest predictor of target selection, with additional interactions involving the enabled ending features. These findings contribute to understanding how users engage with interactive graphical tools and how such tools support data interpretation in exploratory data analysis.
Statistical survey metadata contains essential contextual information that underpins the accurate interpretation, discovery, and reuse of statistical data. However, traditional metadata formats are not optimized for consumption by large language models (LLMs), which increasingly function as interfaces for data exploration, question-answering, and decision support. This work introduces a knowledge graph-based approach to modeling survey metadata using semantic web standards and linked data principles, specifically designed to make metadata machine-understandable and LLM-compatible. The core metadata entities, including surveys, datasets, variables, concepts, populations, and provenance, are modeled as rich interlinked nodes that allow reasoning, contextual enrichment, and structured prompting. The graph integrates established ontologies such as the Resource Description Framework (RDF) to promote interoperability and alignment with global standards. We demonstrate how this structure allows LLMs to surface relevant metadata, ground their outputs in authoritative sources, and generate semantically precise responses. This approach enhances transparency, facilitates metadata reuse, and supports the development of artificial intelligence (AI) applications powered by statistical products.
The United States Department of Agriculture’s (USDA’s) National Agricultural Statistics Service (NASS) conducted a pilot study in 2024 to obtain data collected onboard farm machinery and explore their uses for statistical purposes. NASS has recognized high value in these machine-logged data (MLD) systems as they can potentially augment, or even replace, traditional survey efforts while providing additional benefits of reducing respondent burden and improving crop-related estimates. This pilot study ultimately addressed four topics: 1) understanding the obstacles in obtaining MLD from farmers; 2) creating geographic workflows to manage inherent geospatial MLD; 3) developing the linkages to NASS’s tabular list frame information; and 4) assessing the use of MLD to replace survey data for time-sensitive estimates. To study each topic, field-level information was gathered from the MLD systems of dozens of producers over hundreds of fields across the central United States (US) for the 2023 growing season. Results showed that 90% of the fields could be linked to a producer on the NASS list frame. Of those producers, the consistency of MLD versus traditional survey reporting was highly variable for those who were selected for a survey in 2023. Comparisons showed median MLD values were larger than historical NASS survey values. Approximately 48% of survey comparisons showed a difference of 25% or less between MLD and historical NASS survey values. MLD shows promise for use in official statistics; however, further analyses with additional producers’ data and enhancements to MLD collection processes are needed before supplementing traditional survey methods.
Traditionally z-scores specified from the WHO population growth curves have been used to describe a child’s growth in relation to his age- and sex-matched population distribution. We propose a new regression approach that offers a straightforward interpretation of the relative growth in terms of the original anthropometric variable. We create a hybrid data set consisting of the observations from the study of interest and counterpart pseudo-population observations imputed from the WHO population growth curves matched to each study participant. We then fit linear and quantile regression models to the hybrid data incorporating demographic variables (usually age and biological sex) corresponding to the growth curves of demographically-similar individuals, a study versus population indicator, and its interactions with demographic variables. We further control for confounding variables from the study by adding their interactions with the study indicator variable. The interaction terms between the study indicator and the demographic variables age and biologic sex can be interpreted as relative growth parameters that depict the differences in means (or quantiles) between the study participants and their pseudo-population counterparts of the original anthropometric variables, rather than the associated z-scores. We use anthropometric growth data from a prospective birth cohort study conducted in Uganda for illustration.
The U.S. Rehabilitation Services Administration (RSA) has partnered with state vocational rehabilitation (VR) agencies since 1973 to improve employment outcomes for individuals with disabilities. A critical resource in this effort is the RSA-911 dataset, a quarterly collection of standardized participant data. However, its complex structure, including high rates of missing or ambiguous values, poses significant challenges for effective analysis. We address these challenges by developing an R package designed to streamline the cleaning and analysis of RSA-911 data, as well as the newly introduced Transition Readiness Toolkit (TRT) scores data (R Core Team, 2021). The TRT assesses participants’ improvement across services and offers a critical measure of VR program effectiveness. Using this R package, our work offers the first analysis of the relationship between TRT pre-post scores and RSA-911 demographic data, providing insights into program outcomes. Additionally, we deliver a user-friendly online dashboard, built with the shiny framework, to allow VR counselors and researchers to independently analyze RSA-911 and TRT data (Chang et al., 2024). This dashboard features intuitive visualizations and workflows, making it easier to generate reproducible analyses without requiring extensive technical expertise. By automating data preparation and providing accessible analysis tools, this project contributes to the field of vocational rehabilitation by facilitating more efficient research and empowering VR professionals with data-driven insights. The tools presented offer a framework for future studies, enhancing the consistency, flexibility, and reproducibility of VR data analysis.
Artificial intelligence systems deployed in real-world environments often experience performance degradation due to dynamic and non-stationary data distributions. Existing approaches predominantly adopt model-centric optimization strategies, assuming static data conditions and relying on frequent retraining to address performance decay. However, such strategies are computationally expensive and operationally impractical in continuous deployment settings. This study addresses the research gap in adaptive data-centric artificial intelligence by proposing an automated framework that prioritizes continuous data adaptation rather than repeated model modification. The proposed framework integrates automated data profiling, concept drift detection, and adaptive data refinement mechanisms to maintain decision-making robustness under evolving data conditions. The methodology evaluates the framework across multiple real-world datasets characterized by temporal variation, noise, and class imbalance, simulating realistic deployment scenarios. Performance is compared against conventional static data pipelines using identical model architectures to isolate the impact of data-centric adaptation. Experimental results demonstrate that the adaptive data-centric framework consistently outperforms static pipelines in terms of predictive accuracy, decision stability, and generalization consistency. In particular, the framework achieves sustained accuracy improvements following detected drift events and significantly reduces performance volatility over time. Moreover, these gains are obtained with substantially lower computational overhead compared to retraining-based strategies. The goal of this research is to establish adaptive data-centric optimization as a scalable and practical paradigm for long-term AI system reliability. The findings provide empirical evidence that intelligent data adaptation can effectively mitigate concept drift and enhance operational resilience in dynamic decision-making environments.
Advances in AI and automation are reshaping qualitative research workflows, making processes more efficient, accurate, consistent, and scalable. This paper presents innovations developed for the Illinois Needs Assessment project, a statewide initiative led by the Illinois State Board of Education and the American Institutes for Research to conduct comprehensive needs assessments for schools that need intensive or comprehensive support. To address the scale and tight timeline requirements of the project, the team designed three interconnected pipelines that work together to produce a finalized report. The first, an Audio Pipeline, uses Whisper and generative AI to automate transcription, text-based speaker role attribution, thematic coding, and insight generation from focus groups and interviews. The second, a Report Generation Pipeline, integrates Airtable automations with AWS infrastructure to produce customized school reports that merge AI-generated findings with survey data, school performance metrics, and contextual comparisons. Third, the Needs Assessment Summary Report automates the assembly of all quantitative and qualitative inputs into a polished, customizable deliverable that combines efficiency with expert review. Together, these pipelines replace ad hoc manual workflows with reproducible, consistent systems that enhance data quality, reduce error, and broaden access for non-technical users. The integrated design demonstrates how automation and generative AI can reduce manual burdens, shorten delivery timelines, and support timely, data-informed, and human-centered decision-making in education.
Feature engineering remains one of the most time-intensive and expertise-dependent stages in machine learning pipelines, often limiting scalability and reproducibility. Despite advances in automated machine learning, existing systems largely emphasize model and hyperparameter optimization while leaving feature construction partially manual and task-specific. This reveals a critical research gap: the absence of a transferable, experience-driven mechanism capable of generalizing feature engineering knowledge across heterogeneous datasets. To address this limitation, this study proposes a meta-learning–based automated feature engineering framework that models transformation selection as a learnable mapping between dataset meta-characteristics and transformation utility. The framework constructs a reusable meta-knowledge layer trained on historical task–transformation–performance relationships and applies ranked transformation strategies to unseen datasets under computational constraints. Experiments conducted on diverse classification and regression datasets demonstrate that the proposed approach achieves up to 4.2% improvement in F1-score and 8.3% reduction in RMSE compared to raw-feature baselines, while maintaining performance comparable to or exceeding manually engineered pipelines. In addition, development time is reduced by up to 55%, and search complexity decreases by approximately 60% through ranking-based pruning. These findings confirm that feature engineering can be formalized as a transferable meta-learning problem, enabling scalable, efficient, and generalizable data science workflows. The study advances the automation of representation construction and supports the integration of intelligent meta-knowledge reuse in next-generation AutoML systems.