IntroductionAs artificial intelligence (AI) technology becomes more widespread in the classroom environment, educators have relied on data-driven machine learning (ML) techniques and statistical frameworks to derive insights into student performance patterns. Bayesian methodologies have emerged as a more intuitive approach to frequentist methods of inference since they link prior assumptions and data together to provide a quantitative distribution of final model parameter estimates. Despite their alignment with four recent ML assessment criteria developed in the educational literature, Bayesian methodologies have received considerably less attention by academic stakeholders prompting the need to empirically discern how these techniques can be used to provide actionable insights into student performance.MethodsTo identify the factors most indicative of student retention and attrition, we apply a Bayesian framework to comparatively examine the differential impact that the amalgamation of traditional and AI-driven predictors has on student performance in an undergraduate in-person science, technology, engineering, and mathematics (STEM) course.ResultsInteraction with the course learning management system (LMS) and performance on diagnostic concept inventory (CI) assessments provided the greatest insights into final course performance. Establishing informative prior values using historical classroom data did not always appreciably enhance model fit.DiscussionWe discuss how Bayesian methodologies are a more pragmatic and interpretable way of assessing student performance and are a promising tool for use in science education research and assessment.
Prior studies of active learning (AL) efficacy have typically lacked dosage designs (e.g., varying intensities rather than simple presence or absence) or specification of whether misconceptions were part of the instructional treatments. In this study, we examine the extent to which different doses of AL (approximately 10%, 15%, 20%, 36% of unit time), doses of misconception-focused instruction (MFI; approximately 0%, 8%, 11%, 13%), and their intersections affect evolution learning. A quantitative, quasiexperimental study (N > 1500 undergraduates) was conducted using a pretest, posttest, delayed posttest design with multiple validated measures of evolution understanding. The student background variables (e.g., binary sex, race or ethnicity), evolution acceptance, and prior coursework were controlled. The results of hierarchical linear and logistic models indicated that higher doses of AL and MFI were associated with significantly larger knowledge and abstract reasoning gains and misconception declines. MFI produced significant learning above and beyond AL. Explicit misconception treatments, coupled with AL, should be explored in more areas of life science education.
Filter methods are a class of feature selection techniques used to identify a subset of informative features during data preprocessing. While the differential efficacy of these techniques has been extensively compared in data science pipelines for predictive outcome modeling, less work has examined how their stability is impacted by underlying corpora properties. A set of six stability metrics (Davis, Dice, Jaccard, Kappa, Lustgarten, and Novovičová) was compared during cross-validation in a Monte Carlo simulation study on synthetic data to examine variability in the stability of three filter methods in data pipelines for binary classification, considering five underlying data properties: (1) error of measurement in the independent covariates, (2) number of training observations, (3) number of features, (4) class imbalance magnitude, and (5) missing data pattern. Feature selection stability was platykurtic and was negatively impacted by measurement error and a smaller number of training observations included in the input corpora. The Novovičová stability metric yielded the highest mean stability values, while the Davis stability metric was the most unstable method. The distribution of all stability metrics was negatively skewed, and the Jaccard metric exhibited the largest amount of variability across all five data properties. A statistical analysis of the synergistic effects between filter feature selection techniques, filter cutoffs, data corpora properties, and machine learning (ML) algorithms on overall pipeline efficacy, quantified using the area the under curve (AUC) evaluation metric, is also presented and discussed.
Educators seek to develop accurate and timely prediction models to forecast student retention and attrition. Although prior studies have generated single point estimates to quantify predictive efficacy, much less education research has examined variability in student performance predictions using nonparametric bootstrap algorithms in data pipelines. In this study, bootstrapping was applied to examine performance variability among five data mining methods (DMMs) and four filter preprocessing feature selection techniques for forecasting course grades for 3225 students enrolled in an undergraduate biology class. While the median area under the curve (AUC) values obtained from bootstrapping were significantly lower than the AUC point estimates obtained without resampling, DMMs and feature selection techniques impacted variability in different ways. The ensemble technique elastic net regression (GLMNET) significantly outperformed all other DMMs and exhibited the least amount of variability in the AUC. However, all filter feature selection techniques significantly increased variability in student success predictions, compared to when this step was omitted from the data pipeline. We discuss the potential benefits and drawbacks of incorporating bootstrapping into prediction pipelines to track, monitor, and forecast classroom performance, as well as highlight the risks of only examining point estimates.
To analyse data, a computationally feasible pipeline must be developed for data modelling. Corpora properties affect performance variability of machine learning (ML) techniques in pipelines; however, this has not been thoroughly investigated using simulation methodologies. A Monte Carlo study is used to compare differences in the area under the curve (AUC) metric for large-n-small-p-corpora examining: 1) the choice of ML algorithm; 2) size of the training database; 3) measurement error; 4) class imbalance magnitude; 5) missing data pattern. Our simulations are consistent with established results under which these algorithms and corpora properties perform best, while providing insights into their synergistic effects. Measurement error negatively impacted pipeline performance across all corpora factors and ML algorithms. A larger training corpus ameliorated the decrease in predictive efficacy resulting from measurement error, class imbalance magnitudes, and missing data patterns. We discuss the implications of these findings for designing pipelines to enhance prediction performance.
High levels of attrition characterize undergraduate science courses in the USA. Predictive analytics research seeks to build models that identify at-risk students and suggest interventions that enhance student success. This study examines whether incorporating a novel assessment type (concept inventories [CI]) and using machine learning (ML) methods (1) improves prediction quality, (2) reduces the time point of successful prediction, and (3) suggests more actionable course-level interventions. A corpus of university and course-level assessment and non-assessment variables (53 variables in total) from 3225 students (over six semesters) was gathered. Five ML methods were employed (two individuals, three ensembles) at three time points (pre-course, week 3, week 6) to quantify predictive efficacy. Inclusion of course-specific CI data along with university-specific corpora significantly improved prediction performance. Ensemble ML methods, in particular the generalized linear model with elastic net (GLMNET), yielded significantly higher area under the curve (AUC) values compared with non-ensemble techniques. Logistic regression achieved the poorest prediction performance and consistently underperformed. Surprisingly, increasing corpus size (i.e., amount of historical data) did not meaningfully impact prediction success. We discuss the roles that novel assessment types and ML techniques may play in advancing predictive learning analytics and addressing attrition in undergraduate science education.
Educators seek to harness knowledge from educational corpora to improve student performance outcomes. Although prior studies have compared the efficacy of data mining methods (DMMs) in pipelines for forecasting student success, less work has focused on identifying a set of relevant features prior to model development and quantifying the stability of feature selection techniques. Pinpointing a subset of pertinent features can (1) reduce the number of variables that need to be managed by stakeholders, (2) make “black-box” algorithms more interpretable, and (3) provide greater guidance for faculty to implement targeted interventions. To that end, we introduce a methodology integrating feature selection with cross-validation and rank each feature on subsets of the training corpus. This modified pipeline was applied to forecast the performance of 3225 students in a baccalaureate science course using a set of 57 features, four DMMs, and four filter feature selection techniques. Correlation Attribute Evaluation (CAE) and Fisher’s Scoring Algorithm (FSA) achieved significantly higher Area Under the Curve (AUC) values for logistic regression (LR) and elastic net regression (GLMNET), compared to when this pipeline step was omitted. Relief Attribute Evaluation (RAE) was highly unstable and produced models with the poorest prediction performance. Borda’s method identified grade point average, number of credits taken, and performance on concept inventory assessments as the primary factors impacting predictions of student performance. We discuss the benefits of this approach when developing data pipelines for predictive modeling in undergraduate settings that are more interpretable and actionable for faculty and stakeholders.
This chapter consists almost entirely of genetic associationGenetic association tests that allow for heterogeneity. The underlying mixture differs depending upon the statistical test. This chapter consists almost entirely of genetic associationGenetic association tests that allow for heterogeneity. The underlying mixture differs depending upon the statistical test. Specifically, we consider phenotype and genotype misclassificationMisclassification, genotype, locusLocus heterogeneityHeterogeneity, locus, and phenotype heterogeneityHeterogeneity, phenotype. Our genomic dataGenomic data consist of genotype, multi-locusLocus genotypeGenotype, multi-locus (MLG), or next-generation sequence data. Our phenotype data are all categorical, with two or more phenotype categories. Virtually all of the tests are based on the EM algorithm. For each EM-based test, we provide definitions of terms, the null hypothesis for the test, closed-form solutions of the parameter estimatesParameter estimate for each iteration step, and a formula for how the test statistic is computed. For a subset of the statistics, we provide examples of how the EM algorithm is applied to a fictitious data set.
This chapter is critically important for researchers who work in statistical genetics. Here, we provide mathematical statistics methods that allow researchers to quantify the effects of different types of mixtures in their genetic linkage and association studies before they have collected any samples. As a result, they can modify their sample size calculations to insure that they have sufficient statistical power even in the presence of different forms of heterogeneity. We consider tests of linkage and association and qualitative and quantitative phenotypes. Also, our methods allow for the determination of statistical power for a fixed sample size and significance levelSignificance level.
We develop Pleiotropyan analytic approach toQuantitative Trait (QT) compute statistical power for a fixed sample size and given significance levelSignificance level, or MSSNMinimum Sample Size Necessary (MSSN) (in terms of casesCase/affectedAffected and controlsControl/unaffectedUnaffected) to achieve fixed power at a given significance levelSignificance level. Our method is threshold-selected, in the sense that we transform individuals with quantitative phenotype-vector-values into either casesCase or controlsControl. The vector may have dimension one (univariate) or greater (multivariate or pleiotropyPleiotropy). The test statistics we consider are the trend test and the chi-square test of independence on genotypes.
This book offers a unified resource on heterogeneity in statistical genetics. It provides an overview of past developments as well as new methodologies and applications. It should appeal to established investigators and advanced students active in statistics and genetics.
We start with the definition of heterogeneity and how this concept applies to genetic data. The remainder of the chapter focuses primarily on genetic modeling under homogeneity. We provide formulas for important genetic concepts like Hardy-Weinberg EquilibriumHardy Weinberg Equilibrium (HWE), haplotype and diplotypeDiplotype frequencies, genetic model-free and model-based methods for computing conditional genotype frequenciesFrequency (frequencies), conditional genotype, threshold-selected methods for computing conditional genotype frequencies, test statistics considered in this book, non-centrality parametersNon-Centrality Parameter (NCP), and an introduction to the EM algorithm.
We provide examples of how different forms of heterogeneity may arise. Also, we provide mathematical models for phenotype heterogeneityHeterogeneity, phenotype. We document the effects on population-based and family-based statistical tests of association. Finally, we discuss the relative effects of phenotype misclassificationMisclassification, phenotype versus genotype misclassificationMisclassification, genotype.
Background: The adverse consequences of major depressive disorder (MDD) and posttraumatic stress disorder (PTSD) affect a significant portion of the US population every year (i.e., 15 million for MDD; 8 million for PTSD) and are of public health concern. The current study examines tobacco, alcohol, and marijuana use as possible longitudinal predictors of MDD and/or PTSD. Methods: A community sample of 674 participants (53% African Americans and 47% Puerto Ricans; 405 females and 269 males) were recruited from the Harlem Longitudinal Development Study. We used Mplus software to obtain the triple trajectories of tobacco, alcohol, and marijuana use from mean age 14 to 36. Logistic regression analyses were then conducted to examine the associations between those triple trajectory groups and a single diagnosis of MDD or PTSD as well as a dual diagnosis of MDD with PTSD at age 36. Results: The observed percentages of MDD, PTSD, and the comorbidity of MDD and PTSD were 17%, 8%, and 5%, respectively. The heavy use of all 3 substances group was associated with an increased likelihood of having MDD (adjusted odds ratio [AOR] = 3.14, P < .01), PTSD (AOR = 3.91, P < .05), and MDD with PTSD (AOR = 6.64, P < .01), as compared with the tobacco and alcohol use group. Conclusions: Treatment programs to quit or reduce the use of tobacco, alcohol, and marijuana may help decrease the prevalence of MDD and PTSD. This could lead to improvements in individualized treatments for patients who use tobacco, alcohol, and marijuana and who have both MDD and PTSD.
Objectives Since the number of individuals who use substances in the United States has markedly increased every year, substance use is a significant public health concern. The current study examines the possible risk and protective factors associated with triple comorbid trajectories of longitudinal alcohol, tobacco, and cannabis use from age 14 to 36. Methods A community sample of 674 participants (53% African Americans and 47% Puerto Ricans; 60% females) were recruited from the Harlem Longitudinal Development Study. Multinomial logistic regression analyses were conducted to examine the associations between the risk (low self-control, peer drug use) and protective (parent-child attachment, family church attendance) factors at age 14 and membership in the triple trajectory groups derived from a multivariate growth mixture model. Results Low self-control and peer drug use were associated with an increased likelihood of being a member in the triple comorbid trajectory groups compared to the reference group (i.e., low alcohol, no tobacco, and no cannabis use). On the other hand, parent-child attachment and family church attendance were associated with a decreased likelihood of being a member in the triple comorbid trajectory groups compared to the reference group. Conclusions Treatment programs for adolescents who use substances may be more helpful if their parents and/or friends could also participate together with the adolescent, rather than only the adolescent participates in the treatment programs. Further research is needed to gain a greater understanding of the conceptual nature of the relationship between earlier risk and protective factors and later substance use patterns.
Background: Posttraumatic stress disorder (PTSD) symptoms are related to a number of adverse consequences such as substance use and general medical conditions. The present longitudinal study seeks to find the longitudinal patterns of cannabis use as precursors of PTSD symptoms. Such information will serve as a guide for intervention programs for PTSD. Methods: Growth mixture modeling was conducted to identify the cannabis use trajectory groups using a community sample of 674 participants (53% African Americans, 47% Hispanics of Puerto Rican decent; 60% females) from the Harlem Longitudinal Development Study. Logistic regression analyses were performed to examine the association between earlier trajectories of cannabis use (ages 14 to 36) and later symptoms of PTSD (at age 36) for the full model including the entire sample (N = 674) as well as the reduced model including only participants who had experienced a traumatic event (n = 205). Results: Five trajectory groups of cannabis use were obtained. The chronic use group (full model: adjusted odds ratio [AOR] = 4.68, P<.01; reduced model: AOR = 4.27, P<.05), the late quitting group (full model: AOR = 6.18, P<.01; reduced model: AOR = 6.67, P<.01), and the moderate use group (full model: AOR = 3.97, P<.01; reduced model: AOR = 3.32, P<.05) were all associated with an increased likelihood of having PTSD symptoms at age 36 compared with the no use group. Conclusions: The findings provide information that PTSD symptoms in the mid-30s can possibly be reduced by decreasing membership in the chronic cannabis use trajectory group, the late quitting trajectory group, and the moderate cannabis use trajectory group.
The current study examines longitudinal patterns of cigarette smoking and depressive symptoms as predictors of generalized anxiety disorder using data from the Harlem Longitudinal Development Study. There were 674 African American (53%) and Puerto Rican (47%) participants. Among the 674 participants, 60% were females. In the logistic regression analyses, the indicators of membership in each of the joint trajectories of cigarette smoking and depressive symptoms from the mid-20s to the mid-30s were used as the independent variables, and the diagnosis of generalized anxiety disorder in the mid-30s was used as the dependent variable. The high cigarette smoking with high depressive symptoms group and the low cigarette smoking with high depressive symptoms group were associated with an increased likelihood of having generalized anxiety disorder as compared to the no cigarette smoking with low depressive symptoms group. The findings shed light on the prevention and treatment of generalized anxiety disorder.
Adult maladaptive behaviors including antisocial personality disorder (ASPD) and marijuana use are major public health concerns. At the present time, there is a dearth of research showing the interrelationships among the possible predictors of adult maladaptive behaviors (i.e., ASPD and marijuana use). Therefore, the current study examines the pathways from adverse family environments in late adolescence to these maladaptive behaviors in adulthood. There were 674 participants (52 % African Americans, 48 % Puerto Ricans). Sixty percent of the sample was female. Structural equation modeling in the current study included 4 waves of data collection (mean ages 19, 24, 29, and 36). An adverse family environment in late adolescence was related to greater externalizing personality in late adolescence, which in turn, was related to greater marijuana use in emerging adulthood. This in turn was positively associated with partner marijuana use in young adulthood, which in turn, was ultimately related to maladaptive behaviors in adulthood. An adverse family environment in late adolescence was also related to greater marijuana use in emerging adulthood, which in turn, was associated with an adverse relationship with one's partner in young adulthood. Such a negative partner relationship was related to maladaptive behaviors in adulthood. The findings suggest that family-focused interventions (Kumpfer and Alvarado in Am Psychol 58(6-7): 457-465, 2003) for dysfunctional families may be most helpful when they include the entire family.