
Service-learning programs play an important role in the recruitment and development of the public health workforce. Such programs serve as necessary pathways for trainees to enter public health and related fields (McClamroch & Montgomery, 2009; Horney, et al., 2014; Yeager, Beitsch, & Hasbrouch, 2016; Leider, Resnick, & Erwin, 2022; Leider et al., 2023), providing participants with hands-on career experience and supplying organizations access to a pool of early career applicants (Furco, 1996; Cashman & Seifer, 2008; Thacker et al., 2008; Meritt & Murphy, 2019; Markaki, et al., 2021). Service-learning participants offer valuable insight into program quality and effectiveness, and gathering this input through surveys is among the most widely used approaches to evaluate training and professional development programs (Gelmon, et al., 2001; Brown, 2005; Kirkpatrick & Kirkpatrick, 2006). Although certain scales designed to evaluate different components of service-learning have been examined previously (e.g., Eyler, et al., 1997; Shiarella, et al., 2000: Moely, et al., 2002; Snell & Lau, 2020; Lee et al, 2021), the overall body of evidence derived from psychometric evaluation is limited (Gelmon et al., 2001; Toncar, et al., 2006; Ma et al., 2019; Snell & Lau, 2020). This is particularly true for service-learning programs in public health and related fields and in programs sponsored by non-academic institutions. The Public Health Associate Program (PHAP) Service-Learning Scale (PSLS) (Appendix) was first developed in 2016. It was designed to evaluate participant experience and satisfaction with PHAP, a service-learning fellowship program managed by the Centers for Disease Control and Prevention (CDC). Using an exploratory factor analysis (EFA), the initial pilot of assessment of PSLS provided evidence of validity and reliability and as an underlying factor structure for the scale (Colman et al., 2018). For the pilot study, EFA was more appropriate methodology because the scale was still in development and hypothesized factors had not been generated (Kelloway, 1995). As explained by Hurley et al., (1997), psychometric research on a particular scale can be phased, beginning with the EFA study and succeeded by a CFA study to see what can be confirmed. The current study purpose is to reexamine and confirm previous findings of the factor structure of subscales and provide evidence of its validity using a confirmatory factor analysis (CFA). While this sample for these psychometric evaluations has been limited to PHAP participants, if the instrument is validated, this scale has utility for a plethora of service-learning programs.
The Rehabilitation Services Administration (RSA)-911 is a potent dataset. The purpose of this study is to improve the veracity of research using the RSA-911 dataset. To achieve this, 43 research studies that were published between 2018 and 2022 that used Rehabilitation Services Administration (RSA)-911 data were evaluated in a scoping review. The articles were assessed on several characteristics–reporting of data cleaning strategies utilized, discussion of variable definitions, and the methodological rigor of the statistical analyses that were reported. Opportunities for improvement in data cleaning strategies, reporting accurate definitions of the variables that were selected for the studies, improving the rigor of study methodologies, and recommendations for practice in State Vocational Rehabilitation (VR) agencies were found. Recommendations are provided that may improve research studies that are conducted with the RSA-911 data which may ultimately improve services for participants with disabilities served in State VR agencies.
Ideas are the primary currency of academic success. They can be elaborated upon to yield publications and grants, and high-quality ideas can result in tenure. Unfortunately, mechanisms for producing high quality ideas are underexplored, and ways of generating and maintaining good ideas are seldom a direct goal of university instruction.
Raven’s Progressive Matrices measure logical reasoning and are often included in large multi-topic surveys in low and middle-income countries. The matrices are image-based items that do not require formal knowledge of language or math to complete. As such, they are attractive items to measure logical reasoning in international development contexts. Many of these large field surveys include short item Raven’s sets because space is too limited to fit a full suite. However, short sets can result in restricted variation in terms of test scores. In this paper, we use a nominal response model (NRM) form of item response theory (IRT) to uncover hidden variation in right and wrong answers using short-item Raven’s tests from two large field surveys in Malawi and Zambia. We also analyze relationships between a set of other variables, comparing performance of different versions of the logical reasoning scores as both independent and dependent variables, checking the validity of the new scores. The new NRM-estimated logical reasoning scores follow a more normal distribution in both samples. Validity checks suggest that when relationships are less strong, NRM-estimated scores can capture more nuance than summed scores or even 2 parameter logistic IRT-estimated scores. NRM can uncover differences that are not apparent when using simple summed scores.
Variance in the Eysenck Personality Questionnaire Revised Short Form’s (EPQ-RS) Neuroticism scale is divisible into a general factor (Neuroticism) and two special factors (Anxious-Tense and Worried-Vulnerable), and although all three factors are associated with poorer mental health, their associations with physical health differ: the general Neuroticism factor was associated with poorer health, the association between the Anxious-Tense factor and health was mixed, and the Worried-Vulnerable factor was associated with better health. One unanswered question is how these factors map onto the domains of the Five-Factor Model of personality, and these domains’ lower-order facets? I addressed this question by collecting data from 230 first year psychology undergraduates. These participants completed the 12-item EPQ-RS Neuroticism scale and the 30-item short form version of the Big Five Inventory-2 (BFI-2-S). The general Neuroticism factor was associated positively with higher Neuroticism and its facets of Anxiety, Depression, and Emotional volatility. This factor was also associated negatively with Extraversion and its facet Energy level, Agreeableness and its facet Trust, and with Conscientiousness. The Anxious-Tense factor was associated positively with Neuroticism and its facet Anxiety, and negatively with Extraversion and its facet Assertiveness. The Worried-Vulnerable factor was associated positively only with Neuroticism and its facet Anxiety. Future epidemiological studies should be cautious when interpreting the effects of Neuroticism when it is measured using the EPQ-RS and should seek to replicate the present findings in larger, representative samples, and with comprehensive measures of the Five-Factor Model, such as the NEO Inventories.
Machine learning has become one of the important methods to process big data. It has made a breakthrough in the limitations of traditional statistical models dealing with high-dimensional data. The current study is to introduce and discuss about how machine learning method can be implemented in high-dimensional education data and help with increasing the model efficacy in dealing with high-dimensional education data. A demonstration of the implementation with an empirical data set is also provided.
The logic of the SMART (Sequential Multiple Assignment Randomization Trial) design was applied to assess the replicability of original-replicate study pairs for Open Science Collaboration (OSC) intervention studies. Within SMART, we utilized both subtests of the correspondence test (CT) to assess study pair comparability. First, we implemented a CT difference test to determine if an effect size difference between the original study and its replicate pair was close to zero; second, we implemented a CT equivalence test to determine if the effect size difference of that study pair was within a designated threshold. In Stage 1 of SMART, each study pair was randomly assigned to one of two alphas (.01 and .05), thereby creating two, probabilistically similar subsets of study pairs. Within each alpha subset, successful difference tests (test of significance was not significantly different than zero) and unsuccessful difference tests were then determined. In Stage 2 of SMART, study pairs in each combination of alpha level and successful or unsuccessful difference tests were randomly assigned to one of two thresholds (±.25 SD, ±.50 SD). Equivalence tests were then conducted for all study pairs in each of these four subsets. Successful equivalence occurred when the distance between an original and its replicate pair was statistically significantly less than a given threshold. Thus, initial randomization followed by a second randomization was used to gauge comparability of each OSC original study and its replicate, for two alpha levels and two thresholds. In the first set of results, to mirror the common replicability assessment case in which only difference tests are conducted, 16 of 96 difference tests (16.7%) conducted in Stage 1 were successful. In the second set of results, for initially successful difference tests, two thresholds were used to determine the percent of study pairs that also passed the equivalence test. Depending on α and threshold, 8.0%-13.8% of studies successfully passed both difference and equivalence CT subtests. In the third set of results, using SMART, after randomization to two α-values and contingent on success or lack of success of a difference test, study pairs were randomized to two thresholds and a statistical test of equivalence conducted. Using meta-analysis methods within SMART-based subsets of study pairs, original-replicate average effect size differences were compared to differences in the second set of results. We found a similarly-sized 10.3% of study pairs passed both CT subtests (nine of 87 study pairs successfully passed the difference test at either alpha and successfully passed the equivalence test at either threshold). Reflecting the importance of incorporating both CT subtests, of 16 study pairs that initially passed the difference test, nearly half (43.7%) failed the equivalence test. Thus, for CT success, we found that α choice had little impact, while threshold choice was an important determinant. In all three sets of results, the percent of successful replications was substantially smaller than the 36% of OSC replicates that were statistically significant. To confirm this study’s replicability, we found very similar patterns of CT success and lack of success for two, SMART-based tables, one for alpha = .01 and one for alpha = .05. The current research extends the utility of CT established by Steiner and Wong (2018) in which results were based on simulation data.
A new method for citing articles and books in scientific publications is proposed. The method all but eliminates the need to list references. In addition to identifying and illustrating the basic rules involved, this article uses the proposed method. Thus, while citations appear throughout, no references are presented. Instead, readers can locate each cited publication by simply copying the citation verbatim and inserting it into the dialogue box of Google Scholar. Two more recommendations for improving the transmission of scientific are also proposed.
Qualitative methods can enhance our understanding of constructs that have not been well portrayed and enable nuanced depiction of experience from study participants who have not been broadly studied. However, qualitative data require time and effort to train raters to achieve validity and reliability. This study compares recent advances in Natural Language Processing (NLP) models with human coding. This web-based study (N=1,253; 3,046 free-text entries, averaging 64 characters per entry) included people with Duchenne Muscular Dystrophy (DMD), their siblings, and a representative comparison group. Human raters (n=6) were trained over multiple sessions in content analysis as per a comprehensive codebook. Three prompts addressed distinct aspects of participants’ aspirations. Unsupervised NLP was implemented using Latent Dirichlet Allocation (LDA), which extracts latent topics across all the free-text entries. Supervised NLP was done using a Bidirectional Encoder Representations from Transformers (BERT) model, which requires training the algorithm to recognize relevant human-coded themes across free-text entries. We compared the human-, LDA-, and BERT-coded themes. Study sample contained 286 people with DMD, 355 DMD siblings, and 997 comparison participants, age 8-69. Human coders generated 95 codes across the three prompts and had an average inter-rater reliability (Fleiss’s kappa) of 0.77, with minimal rater-effect (pseudo R2=4%). Compared to human coders, LDA does not yield easily interpretable themes. BERT correctly classified only 61-70% of the validation set. LDA and BERT required technical expertise to program and took approximately 1.15 minutes per open-text entry, compared to 1.18 minutes for human raters including training time. LDA and BERT provide potentially viable approaches to analyzing large-scale qualitative data, but both have limitations. When text entries are short, LDA yields latent topics that are hard to interpret. BERT accurately identified only about two thirds of new statements. Humans provided reliable and cost-effective coding in the web-based context. The upfront training enables BERT to process enormous quantities of text data in future work, which should examine NLP’s predictive accuracy given different quantities of training data.
The Journal of Methods and Measurement in the Social Sciences (JMM) publishes articles related to methodology and research design, measurement, and data analysis.The journal is published twice yearly, and features theoretical, empirical and educational articles.JMM is meant to further our understanding of methodology and how to formulate the right questions.It is broadly concerned with improving the methods used to conduct research, the measurement of variables used in the social sciences, and improving the applications of data analysis.In addition to research articles, the journal welcomes instructional articles and brief reports or commentaries.We welcome sound, original contributions.
Using an example from animal cognition, I argue that the problems of bias—inherent in choosing null hypotheses or setting Bayesian priors—can sometimes be avoided altogether by collecting more and better observational data before setting up tests of any sort.
Petrinovich highlighted many salient issues in the behavioral and social sciences that are of concern to this day, such as insufficient attention to construct validity. Structural equation modeling, particularly with regard to latent variables, is introduced and discussed in this context. Though conceptual issues remain, analytic and statistical techniques have made immense strides in the past three decades since the article was written, and properly used, offer solutions to many problems Petrinovich identified.
A barrier that prevents many social scientists from pursuing big data research is the lack of technical training required to assemble and organize big data. In an effort to address this barrier, we provide an introductory tutorial into machine learning for social scientists by demonstrating the basic steps and fundamental concepts involved in binary classification. We first describe the data and libraries required for analysis. We then demonstrate data cleaning methods, feature engineering, the model-building process, model assessment, and feature importance. Last, we discuss the ways in which social scientists can use machine learning to complement inference-based approaches and how it can contribute to a richer understanding of social science.
There are multiple levels of backstory to this paper. I was among the last few doctoral students trained by Professor Lewis (“Lew”) Franklin Petrinovich, and collaborated with him on various research projects for decades after I received my PhD in Comparative Psychology, so I witnessed the whole origin story for this paper unfold over the years. While still a graduate student at the University of California at Riverside, I took the course that Lew regularly taught on research methodology. I still consider that one course to be the biggest eye-opener of my graduate training. It revolutionized my thinking on how to do science and has continued to influence my professional career as a researcher in psychology.
Petrinovich’s target article focused on how behavioral science is done, including how it is often done wrong, and how it should be done. I identify another malign influence on behavioral science, which, so far as I know, has, until now, been ignored (I would be happy to be shown that I am wrong on this). To wit, the way that Introductions to papers are written creates a niche that can be exploited for the purposes of promoting one’s work to obtain resources or status, or for self-aggrandizement. I offer a few, probably wrongheaded, suggestions for ending this practice.