The linear varying coefficient models posits a linear relationship between an outcome and covariates in which the covariate effects are modeled as functions of additional effect modifiers. Despite a long history of study and use in statistics and econometrics, state-of-the-art varying coefficient modeling methods cannot accommodate multivariate effect modifiers without imposing restrictive functional form assumptions or involving computationally intensive hyperparameter tuning. In response, we introduce VCBART, which flexibly estimates the covariate effect in a varying coefficient model using Bayesian Additive Regression Trees. With simple default settings, VCBART outperforms existing varying coefficient methods in terms of covariate effect estimation, uncertainty quantification, and outcome prediction. We illustrate the utility of VCBART with two case studies: one examining how the association between later-life cognition and measures of socioeconomic position vary with respect to age and socio-demographics and another estimating how temporal trends in urban crime vary at the neighborhood level. An R package implementing VCBART is available at https://github.com/skdeshpande91/VCBART
Background In response to the Companies Act of 2013, Corporate Social Responsibility (CSR) initiatives have exploded in India. However, little research exists on Corporate Social Marketing (CSM) in India. Focus of the Article The current study addresses this gap and synthesizes insights from 31 CSR interventions in India across diverse industries, organized according to Campbell et al.'s (2022) corporate social marketing benchmarks to assess program development, research methods, audience targeting, partnerships, and market mix components. Research Question To what extent does the presence of corporate social marketing benchmarks predict the effectiveness of corporate social marketing interventions? Method The study employs qualitative content analysis to examine CSM programs across diverse industries in India, including FMCG, financial services, and manufacturing. The companies were selected from a list of nominated companies (including some who won or were included as honorable mentions) for the National CSR Awards in India in 2019 and 2020. The study evaluates the alignment of these CSM interventions with Sustainable Development Goals, adherence to CSM benchmarks, and the effectiveness of communication strategies. Findings The analysis reveals a significant link between the higher presence of corporate social marketing benchmarks and the success of the interventions (chi(2) = 9.64, p < .01). Strategic (such as 4Ps) and planning benchmarks (outcome setting objectives and partners) were observed in most interventions, while research and audience understanding benchmarks were the least frequent. The study also highlights the strengths and weaknesses of the interventions and proposes strategies to improve corporate involvement in social change. Recommendations to Research/Practice The current study validates the benchmarks presented by Campbell et al. (2022) as a practical tool to evaluate the effectiveness of CSM interventions in India. Additionally, this study underscores the expanding role of corporate social marketing (CSM) in driving behaviour change and advancing the broader social marketing agenda. Importance to the Social Marketing Field The current study advances the understanding of CSM in India by analyzing CSR interventions through a corporate social marketing benchmark lens, examining their strategic aspects, gaps, and key strengths to enhance corporate engagement in social change.
The multivariate regression interpretation of the Gaussian chain graph model simultaneously parametrizes (i) the direct effects of $p$ predictors on $q$ outcomes and (ii) the residual partial covariances between pairs of outcomes. We introduce a new method for fitting sparse Gaussian chain graph models with spike-and-slab LASSO (SSL) priors. We develop an Expectation Conditional Maximization algorithm to obtain sparse estimates of the $p \times q$ matrix of direct effects and the $q \times q$ residual precision matrix. Our algorithm iteratively solves a sequence of penalized maximum likelihood problems with self-adaptive penalties that gradually filter out negligible regression coefficients and partial covariances. Because it adaptively penalizes individual model parameters, our method is seen to outperform fixed-penalty competitors on simulated data. We establish the posterior contraction rate for our model, buttressing our method's excellent empirical performance with strong theoretical guarantees. Using our method, we estimated the direct effects of diet and residence type on the composition of the gut microbiome of elderly adults.
Musculoskeletal pain is a global public health problem. Social marketing aims to increase adoption of desired behaviours in target audiences and may uncover new strategies to improve uptake of helpful pain-related behaviours at the population-level. We systematically evaluated effects of contemporary mass media campaigns targeting musculoskeletal pain and used social marketing benchmarking to explore strategies associated with campaign success. Published evaluations of campaigns involving an online/digital component and a comparator/control condition were eligible. The primary outcome was population beliefs; secondary outcomes were healthcare provider beliefs, behavioural (e.g., healthcare-related, work-related), clinical (e.g., pain), and economic outcomes. Decision-rules and meta-analyses (random-effects models) were used to synthesise findings. Eight databases and grey literature were searched from inception to May-2024. Thirteen eligible publications evaluated eight campaigns (N=5 back pain, N=2 rheumatic pain; N=1 work-related pain) from eight Western/high-income countries. All evaluations reported historical control data (interrupted time-series/before-and-after designs); three also compared selected outcomes to an unexposed geographical region (quasi-experimental designs). Risk of bias was weak-moderate for all evaluations. Population beliefs improved from baseline vs. final follow-up (1.5-10yrs) for items related to ‘staying active’ [RR=1.38 (95%CI: 1.14-1.67), N=4 campaigns, n=12,568 participants] and ‘rest’ [RR=1.35 (95%CI: 1.14-1.60), N=5 campaigns, n=14,571 participants] for pain management, however, certainty of evidence was very low. Other outcomes were not pooled due to heterogeneity, and evidence was mixed. Greater numbers of social marketing benchmarks were associated with successful campaign outcomes. Future campaigns should implement social marketing strategies beyond education alone, including behaviour change support, to facilitate adoption of desired pain-related behaviours. Registration PROSPERO registration number: CRD42023400456; Open Science Framework (detailed Social Marketing Benchmarking analysis plan): https://osf.io/npyck/ Perspective We systematically evaluated contemporary mass media campaigns targeting musculoskeletal pain. Promising improvements in population beliefs about pain supports continued investment into campaigns. Our review provides critical new information including social marketing strategies to ensure future campaign efforts shift population-level pain-related behaviours, towards reducing the societal burden of pain.
We study the impact of teenage sports participation on early-adulthood health using data from the National Study of Youth and Religion. We focus on two primary outcomes measured at ages 23 to 28 (self-rated health and PHQ9 Patient Depression Questionnaire score) and control for several demographic and socioeconomic confounders. To probe the possibility that certain types of sports participation may have larger effects on health than others, we conduct matched observational studies at each level within a hierarchy of exposures. Our hierarchy ranges from broadly defined exposures (e.g., participation in any organized after-school activity) to narrow (e.g., participation in collision sports). We maintained a fixed family-wise error rate using an ordered testing approach that exploits the hierarchical relationships between our exposure definitions. Compared to teenagers who did not participate in any after-school activities, those who participated in sports had statistically significantly better self-rated and mental health outcomes in early adulthood.
We study the asymptotic properties of Deshpande et al. (2019)'s multivariate spike-and-slab LASSO (mSSL) procedure for simultaneous variable and covariance selection in the sparse multivariate linear regression problem. In that problem, q correlated responses are regressed onto p covariates and the mSSL works by placing separate spike-and-slab priors on the entries in the matrix of marginal covariate effects and off-diagonal elements in the upper triangle of the residual precision matrix. Under mild assumptions about these matrices, we establish the posterior contraction rate for the mSSL posterior in the asymptotic regime where both p and q diverge with n. By "de-biasing" the corresponding MAP estimates, we obtain confidence intervals for each covariate effect and residual partial correlation. In extensive simulation studies, these intervals displayed close-to-nominal frequentist coverage in finite sample settings but tended to be substantially longer than those obtained using a version of the Bayesian bootstrap that randomly re-weights the prior. We further show that the de-biased intervals for individual covariate effects are asymptotically valid.
Commercial marketing literature highlights benefits from brand advocates who recruit and promote in the interest of the commercial entity. However, a similar focus is lacking on how advocacy can extend the effectiveness of social change initiatives. We utilise a case study to demonstrate the benefit of social advocacy and its impact on behaviour change, and thereby propose an advocacy model. To develop this conceptual model, we discuss several key areas; behaviour change and advocacy, advocate identification, and how to influence advocacy within communities and individuals. This research provides a guiding framework for practitioners to develop programs and interventions with advocacy triggers and strategies to enhance the longevity and effectiveness of social change programs through participant-based advocacy. Thus, giving intervention programs in a variety of organisational structures e.g. non-profit, corporate, government etc. a specific model to increase the effectiveness of social programs. Our paper extends behaviour change literature by leveraging social marketing concepts to modify and extend the transtheoretical model.
Most implementations of Bayesian additive regression trees (BART) one-hot encode categorical predictors, replacing each one with several binary indicators, one for every level or category. Regression trees built with these indicators partition the discrete set of categorical levels by repeatedly removing one level at a time. Unfortunately, the vast majority of partitions cannot be built with this strategy, severely limiting BART's ability to partially pool data across groups of levels. Motivated by analyses of baseball data and neighborhood-level crime dynamics, we overcame this limitation by re-implementing BART with regression trees that can assign multiple levels to both branches of a decision tree node. To model spatial data aggregated into small regions, we further proposed a new decision rule prior that creates spatially contiguous regions by deleting a random edge from a random spanning tree of a suitably defined network. Our re-implementation, which is available in the flexBART package, often yields improved out-of-sample predictive performance and scales better to larger datasets than existing implementations of BART.
Although it is an extremely effective, easy-to-use, and increasingly popular tool for nonparametric regression, the Bayesian Additive Regression Trees (BART) model is limited by the fact that it can only produce discontinuous output. Initial attempts to overcome this limitation were based on regression trees that output Gaussian Processes instead of constants. Unfortunately, implementations of these extensions cannot scale to large datasets. We propose ridgeBART, an extension of BART built with trees that output linear combinations of ridge functions (i.e., a composition of an affine transformation of the inputs and non-linearity); that is, we build a Bayesian ensemble of localized neural networks with a single hidden layer. We develop a new MCMC sampler that updates trees in linear time and establish posterior contraction rates for estimating piecewise anisotropic Hölder functions and nearly minimax-optimal rates for estimating isotropic Hölder functions. We demonstrate ridgeBART's effectiveness on synthetic data and use it to estimate the probability that a professional basketball player makes a shot from any location on the court in a spatially smooth fashion.
Modern statistics provides an ever-expanding toolkit for estimating unknown parameters. Consequently, applied statisticians frequently face a difficult decision: retain a parameter estimate from a familiar method or replace it with an estimate from a newer or more complex one. While it is traditional to compare estimates using risk, such comparisons are rarely conclusive in realistic settings. In response, we propose the "c-value" as a measure of confidence that a new estimate achieves smaller loss than an old estimate on a given dataset. We show that it is unlikely that a large c-value coincides with a larger loss for the new estimate. Therefore, just as a small p-value supports rejecting a null hypothesis, a large c-value supports using a new estimate in place of the old. For a wide class of problems and estimates, we show how to compute a c-value by first constructing a data-dependent high-probability lower bound on the difference in loss. The c-value is frequentist in nature, but we show that it can provide validation of shrinkage estimates derived from Bayesian models in real data applications involving hierarchical models and Gaussian processes. Supplementary materials for this article are available online.
Abstract: We will study the impact of adolescent sports participation on early-adulthood health using longitudinal data from the National Study of Youth and Religion. We focus on two primary outcomes measured at ages 23–28 — self-rated health and total score on the PHQ9 Patient Depression Questionnaire — and control for several potential confounders related to demographics and family socioeconomic status. Comparing outcomes between sports participants and matched non-sports participants with similar confounders is straightforward. Unfortunately, an analysis based on such a broad exposure cannot probe the possibility that participation in certain types of sports (e.g., collision sports like football or soccer) may have larger effects on health than others. In this study, we introduce a hierarchy of exposure definitions, ranging from broad (participation in any after-school organized activity) to narrow (e.g., participation in limited-contact sports). We will perform separate matched observational studies, one for each definition, to estimate the health effects of several levels of sports participation. In order to conduct these studies while maintaining a fixed family-wise error rate, we deployed an ordered testing approach that exploits the logical relationships between exposure definitions. Our study will also consider several secondary outcomes including body mass index, life satisfaction, and problematic drinking behavior.
Expected points is a value function fundamental to player evaluation and strategic in-game decision-making across sports analytics, particularly in American football. To estimate expected points, football analysts use machine learning tools, which are not equipped to handle certain challenges. They suffer from selection bias, display counter-intuitive artifacts of overfitting, do not quantify uncertainty in point estimates, and do not account for the strong dependence structure of observational football data. These issues are not unique to American football or even sports analytics; they are general problems analysts encounter across various statistical applications, particularly when using machine learning in lieu of traditional statistical models. We explore these issues in detail and devise expected points models that account for them. We also introduce a widely applicable novel methodological approach to mitigate overfitting, using a catalytic prior to smooth our machine learning models.
We introduce a three-step framework to determine at which pitches Major League batters should swing. Unlike traditional plate discipline metrics, which implicitly assume that all batters should always swing at (resp. take) pitches inside (resp. outside) the strike zone, our approach explicitly accounts not only for the players and umpires involved in the pitch but also in-game contextual information like the number of outs, the count, baserunners, and score. We first fit flexible Bayesian nonparametric models to estimate (i) the probability that the pitch is called a strike if the batter takes the pitch; (ii) the probability that the batter makes contact if he swings; and (iii) the number of runs the batting team is expected to score following each pitch outcome (e.g. swing and miss, take a called strike, etc.). We then combine these intermediate estimates to determine whether swinging increases the batting team's run expectancy. Our approach enables natural uncertainty propagation so that we can not only determine the optimal swing/take decision but also quantify our confidence in that decision. We illustrate our framework using a case study of pitches faced by Mike Trout in 2019.
In the last quarter of a century, algebraic statistics has established itself as an expanding field which uses multilinear algebra, commutative algebra, computational algebra, geometry, and combinatorics to tackle problems in mathematical statistics. These developments have found applications in a growing number of areas, including biology, neuroscience, economics, and social sciences. Naturally, new connections continue to be made with other areas of mathematics and statistics. This paper outlines three such connections: to statistical models used in educational testing, to a classification problem for a family of nonparametric regression models, and to phase transition phenomena under uniform sampling of contingency tables. We illustrate the motivating problems, each of which is for algebraic statistics a new direction, and demonstrate an enhancement of related methodologies.
Current implementations of Bayesian Additive Regression Trees (BART) are based on axis-aligned decision rules that recursively partition the feature space using a single feature at a time. Several authors have demonstrated that oblique trees, whose decision rules are based on linear combinations of features, can sometimes yield better predictions than axis-aligned trees and exhibit excellent theoretical properties. We develop an oblique version of BART that leverages a data-adaptive decision rule prior that recursively partitions the feature space along random hyperplanes. Using several synthetic and real-world benchmark datasets, we systematically compared our oblique BART implementation to axis-aligned BART and other tree ensemble methods, finding that oblique BART was competitive with – and sometimes much better than – those methods.
Test log-likelihood is commonly used to compare different models of the same data or different approximate inference algorithms for fitting the same probabilistic model. We present simple examples demonstrating how comparisons based on test log-likelihood can contradict comparisons according to other objectives. Specifically, our examples show that (i) approximate Bayesian inference algorithms that attain higher test log-likelihoods need not also yield more accurate posterior approximations and (ii) conclusions about forecast accuracy based on test log-likelihood comparisons may not agree with conclusions based on root mean squared error.
We demonstrate how Hahn et al.'s Bayesian Causal Forests model (BCF) can be used to estimate conditional average treatment effects for the longitudinal dataset in the 2022 American Causal Inference Conference Data Challenge. Unfortunately, existing implementations of BCF do not scale to the size of the challenge data. Therefore, we developed flexBCF -- a more scalable and flexible implementation of BCF -- and used it in our challenge submission. We investigate the sensitivity of our results to the choice of propensity score estimation method and the use of sparsity-inducing regression tree priors. While we found that our overall point predictions were not especially sensitive to these modeling choices, we did observe that running BCF with flexibly estimated propensity scores often yielded better-calibrated uncertainty intervals.
Purpose: While domestic and family violence against people with disabilities is an ongoing and crucial public health concern, and awareness of the extent of violence against people with disabilities is growing, research on the field is still limited. Thus, the present review aims to systematically identify and synthesize evidence and effectiveness from intervention strategies to increase the awareness and skills of those with disabilities to reduce and prevent domestic and family violence against them. Method: PRISMA guidelines were followed to perform a systematic search of seven scientific databases to identify the peer-reviewed literature. Results: A total of 17 eligible studies were identified (14 evaluations and 3 descriptive studies), with most taking place in developed countries. Children and women are the most frequent victims, and they were therefore the most common target audience of the included studies. Sexual, physical, and verbal abuse were the most reported types of abuse, while financial abuse and neglect were studied less often. Interventions also focused on a diversity of disabilities, including learning, intellectual, mental, and physical impairments. Overall, the intervention strategies reflected a substantial homogeneity: focus on training and education as well as setting up channels and facilities for victims to seek help. Nine studies yielded significant positive outcomes using various strategies and techniques, while five studies had mixed results, and three studies only reported on the intervention strategies but did not evaluate the results. Conclusions: This review confirms a significant gap in the literature on domestic and family violence against people with disabilities and how to prevent and address the violence through evidence-based interventions. Several recommendations to improve future research and practice are proposed.
This paper performs a comprehensive analysis of academic research on impulse buying following a systematic literature review approach. Drawing on the TCCM framework suggested by Paul and Rosado-Serrano, we synthesize the impulse buying literature and develop a future research agenda. Accordingly, this review synthesizes impulse buying research in terms of theory development, context, characteristics, and methodologies to examine the development of the literature over time. This systematic review shows that impulse buying research is fragmented and still developing due to its transition from a traditional retail environment into different online channels. Furthermore, this paper proposes a conceptual framework based on the literature synthesis, presenting antecedents and mediators of impulse buying behaviour. Finally, this review identifies overlooked areas in impulse buying literature and provides insightful directions to advance research in the domain. Overall, this research effort makes a significant contribution to consumer behaviour literature, specifically to impulse buying literature.
Background Segmentation use in social marketing especially in improving the health of young adults is limited, and theory use within segmentation remains infrequent. A generalisable segmentation structure that can be reliably applied across different young adult's samples may assist social marketers to move beyond one size fits all healthy eating programs. Focus of the Article Segmentation is an essential marketing principle which allows customising marketing activities to the needs of specific segments. Evidence shows that behaviour change is more likely when more principles are used, yet segmentation remains underutilised and a cross-sample validation of segments across different populations remains to be demonstrated. Importance to the Social Marketing Field Delivery of healthy eating programs targeted to group differences and accommodating a broader theory-based socio-ecological viewpoint is needed to engage with a cross section of young adults more effectively along with a cross-sample validation of segments across different populations to identify a valid segmentation structure that can be reliably applied across the Australian young adult population. Methods A replication study was conducted using the same constructs, items and analytical procedures as in the original study. Data was collected online and in person using a paper survey in two military bases to ensure a mix of Australian Defence Force (ADF) trainee types. Psychographic variables informed by the MOA framework were collected and used to segment the sample with two-step cluster analysis along with a demographic measure (education) and behavioural measure (eating behaviour) to repeat the segmentation analysis. Results The ability of the MOA framework to explain eating behaviour was confirmed in the ADF trainee sample, and two-step cluster analysis produced a similar segment structure to the original study with education, opportunity and motivation to eat healthy being the most important variables in segment formation. Recommendations for Research or Practice Segmentation is important for developing understanding that enables social marketers to design social change programs to meet the needs of young adults. This empirical replication study confirmed a similar theory-driven healthy eating segment solution across two young adult populations illustrating the value of using behavioural theories to draw segments and utilising the same theory to cross-validate the constructs in a comparable sample. Future research could use this approach to identify a valid segmentation structure that can be reliably applied across different populations and behavioural contexts.