
ABSTRACT The rapid expansion of online education has increased flexibility and access while also creating instructional challenges for statistics courses that require students to engage with formulas, software output, graphical representations, and applied data analysis simultaneously. This paper examines an instructional approach for teaching introductory statistics online through the integration of real‐time annotation, statistical software demonstrations, simulations, recorded lectures, and multimedia feedback using SMART Podium technology. The approach was designed to reduce fragmentation between explanation, computation, visualization, and interpretation during online instruction. Examples involving regression, sampling distributions, and confidence intervals illustrate how annotation and integrated software demonstrations were incorporated into instruction. A mixed‐methods course‐based evaluation combining student survey responses and course‐level performance indicators suggested generally positive student perceptions of instructor annotations, recorded instructional materials, software‐based demonstrations, and personalized feedback within the online course context.
This paper examines the use of virtual populations and artificial intelligence (AI) tools in a course-based project within the undergraduate statistics course "Experimental Design." Consistent with our prior research, we found that most students supported the use of project-based learning and reported gains in conceptual understanding, report writing, and communication skills. We build on this work by explicitly allowing students to use AI tools and introducing the Island, a virtual population designed to simulate realistic human experiments, thereby deepening students' understanding of experimental design. Both quantitative and qualitative findings suggest that students view generative AI and the Island as valuable resources for learning and completing their projects. Additionally, our analysis indicates that international students are more likely than domestic students to regret participating in the project when given the option to opt out.
Women continue to be underrepresented in STEM fields, including data science. Although much research has focused on gender gaps at the undergraduate level, there has been less focus on women's experiences in postgraduate programs that combine statistics and applied data science. This study investigated the experiences of women enrolled in a postgraduate applied data science program in Ghana. Using a qualitative descriptive approach, data were gathered through semi-structured interviews with eight female postgraduate students. Thematic analysis showed that, despite a limited background in advanced research and statistical techniques, the women exhibited strong motivation to learn these skills through postgraduate studies. They also expressed clear expectations for hands-on use of statistical software, supportive interactions with lecturers, and experiential learning opportunities. The results emphasize the need for supportive teaching environments, early skills development, and practical learning in postgraduate education in statistics and data science.
This paper proposes a classroom-oriented approach to teaching the coefficient of variation () across school and university levels. The CV is introduced through its relationship with the mean and standard deviation, with attention to interpretation, common misunderstandings, and practical limitations (e.g., when the mean is zero). The main contribution is a set of adaptable learning activities: problem-based tasks for university students, collaborative comparisons for secondary classrooms, and hands-on activities for younger learners to build intuitive ideas of relative variability. Supporting materials include simulated datasets and Stata code to facilitate implementation. By positioning the CV alongside familiar measures such as the range and standard deviation, the approach aims to strengthen students' understanding of variability and support meaningful comparisons across different contexts and units.
Introductory statistics courses for first-year undergraduate students can convey a positive or negative attitude toward the subject. In a traditional teacher-centered classroom setting, students often learn statistical concepts without being able to relate what they learn to the real world. This makes statistics seem more daunting, confusing, and hard than useful. In this article, the author describes an active methodology that uses Socrates' instruction-through-questioning approach in teaching statistics. Students actively engage in dialogue about a real-life issue in a challenging and thought-provoking learning environment that allows students to experience statistics as a process of logical discovery. Using Socrates' dialogue approach as a core form of teaching statistics fosters interest in the subject, boosts retention of the acquired analytical skills, and promotes life-long learning. The author elucidates what inspired her to redesign her introductory statistics course, illustrates the course structure, and pinpoints some of the constraints on scaling this reform.
Data visualization is pivotal in agricultural research, yet tools like R, Python, and Excel pose significant challenges. R's advanced graphics are offset by a steep learning curve, Python often encounters runtime errors, and Excel lacks robust functionality. To overcome these obstacles, we introduce grapesDraw, an open-source web application, and R package tailored for agricultural research and education. Built on Shiny, grapesDraw integrates R's ggplot2, corrplot, packcircles, and factoextra packages, offering an intuitive interface that minimizes coding demands. This enables researchers and students to create high-quality visualizations effortlessly. A survey across state agricultural universities (SAUs) confirmed grapesDraw's usability and effectiveness, particularly among novices. Freely available on shinyapps.io and as an R package on GitHub, with detailed documentation, grapesDraw enhances accessibility and efficiency in agricultural data visualization.
This article presents a methodological approach that combines conversation analysis (CA) and interaction analysis (IA) to examine how students reason with data in collaborative settings, using food justice as an illustrative case. While traditional analytical approaches in data science education often rely on individual cognitive measures or final products, this approach captures the dynamic, interactional nature of data reasoning as it unfolds in real time. By analyzing video-recorded interactions, we demonstrate how CA/IA can reveal the micro-processes through which students negotiate meaning, challenge interpretations, and engage with social and ethical dimensions of data. Though demonstrated here in a food justice context, the approach is broadly applicable to any setting where understanding the social organization of collaborative data reasoning matters, including formal classrooms, workplace teams, and community-based learning environments, and contributes both methodological and practical insights for data science education researchers.
Traditional teaching often emphasizes correct methods, limiting opportunities to explore analytical errors and biases. We introduce a seminar framework that integrates peer-to-peer teaching with intentional exposure to statistical and machine learning mishaps through flawed, student-designed case studies. Students delivered two presentations: one teaching a chosen mishap and another presenting an original case study embedding errors. Personalized feedback before and after each presentation supported iterative improvement. This structure created a safe environment to intentionally produce, analyze, and discuss errors, helping students understand how mishaps arise, affect results, and can be addressed. Class discussions and peer analysis encouraged deeper engagement and critical thinking. A follow-up questionnaire administered 1 year later showed that seminar participants significantly outperformed peers in identifying and correcting data analysis errors. Our results suggest that peer-led, error-focused learning with repeated feedback enhances engagement, retention, and preparedness for real-world data challenges, offering a valuable complement to traditional, instructor-led teaching.
Statistical hypothesis testing (SHT) is widely employed across numerous scientific disciplines, and a clear understanding of its underlying logic is essential for the broader scientific community. Here, drawing upon both epistemological and statistical perspectives, we aim to clarify-primarily for educational purposes-the logical relationship between proof by contradiction (PBC) and SHT. We begin by outlining the logics of SHT and PBC, followed by a discussion of their key similarities and differences. We then explore the pedagogical value of these analogies and distinctions through illustrative examples. Our objective is to help prevent common misinterpretations in the application of SHT, particularly in the interpretation of its outcomes. By elucidating the conceptual parallels between SHT and PBC, we aim to make the logical structure of SHT more transparent. Finally, we highlight and discuss several critical aspects relevant to teaching, with the intention of enhancing the pedagogical effectiveness of instruction in this area.
When teaching how to describe and apply good practices for visualizing data, we need to define "good". Several sets of guidelines about good visualization practice exist in the literature and online, though each set focuses on different aspects of visualization and their level ranges from very general to very specific. We present five principles and associated guidance that is: (i) appropriate for an entry-level undergraduate data science course where students produce static visualizations using Python or R plotting libraries, (ii) actionable, meaning students and markers can assess visualizations against the guidance, and (iii) concise enough to fit on one page, provided as a resource. We describe how the resource helps our teaching and assessment, and the advice we give students to address the common problem of plots with inaccessibly small text. Informally, student responses to the principles are positive and are continuing to inform changes to the detailed guidance.
This study explores teachers' levels of statistical knowledge for teaching through the application of a novel model. The participants comprised ten experienced middle school mathematics teachers. Data were collected using two primary tools: an instrument designed to assess statistical knowledge for teaching and clinical interviews. Quantitative data were analyzed using the Rasch model, while qualitative data from the interviews provided deeper insights into teachers' responses. Results revealed that teachers' statistical knowledge for teaching was generally at an emerging level. An examination of components showed that many teachers demonstrated the competent level in key developmental understandings, while their knowledge of curriculum, understanding of students, and pedagogically powerful ideas remained at the emerging level. Their knowledge of teaching was found to be at the aware level. It is recommended to focus on the practices that enable preservice teachers to improve their statistical knowledge for teaching if they are to be well-informed teachers.
Concepts relating to sampling variation are known to be difficult for learners in introductory classes. There is some evidence that web-based visualization tools, or "applets," may aid the learning of these challenging ideas. Four freely available applets developed at the authors' institution were extensively tested through sessions involving 42 undergraduate students. Details are given on how students interacted with each applet, the student feedback, and insights into learning following exploration of each applet via a structured activity.
The academic experiences at colleges and universities were significantly disrupted by the coronavirus disease 2019 pandemic with notable impacts on student learning, engagement, and success. We explore its effect on academic performance and student interaction at a mid-sized public institution in the southeastern United States. Data collected from approximately 1500 introductory statistics students between 2019 and 2021 were used to compare student performance before and during the pandemic. During the pandemic, differing modes of instruction were offered, and overall performance on assessments in 2020 and 2021 differed little from performance prior to the pandemic in 2019. However, more absences and academic integrity issues were reported prior to the pandemic. Performance and student interactions varied based on modality, with students meeting face-to-face tending to have better outcomes. Finally, we observe that gender, prior academic performance, and prior course loads were the only academic and demographic descriptors distinguishable between which modalities students opted to attend.
The adjustment in the sample variance formulas is commonly justified on the grounds of obtaining an unbiased estimate of the population variance in inferential statistics. However, its use in descriptive statistics and nonrandom datasets is often applied without sufficient attention to the analytic purpose. This study examines the default use of the correction through real-world examples, classroom-based teaching scenarios, and instructionally motivated simulations implemented in the statistical software R. Our analysis shows that dividing by is not only acceptable in many practical contexts but may also be more appropriate when the goal is purely descriptive. The discussion emphasizes how variance calculations should align with data structure and analytic intent. We advocate for a context-driven approach to teaching variance and standard deviation, one that discourages formulaic rule-following and supports deeper conceptual understanding of variability.
Keeping pace with rapidly evolving technology is a key challenge in teaching statistics. To equip students with essential skills for the modern workplace, educators must integrate relevant technologies into the statistical curriculum where possible. University-level statistics education has experienced substantial technological change, particularly in the tools and practices that underpin teaching and learning. Statistical programming has become central to many courses, with R widely used and Python increasingly incorporated into statistics and data analytics programs. Additionally, coding practices, database management, and machine learning now feature within some statistics curricula. Looking ahead, we anticipate a growing emphasis on artificial intelligence (AI), particularly the pedagogical implications of generative AI tools such as ChatGPT. In this article, we discuss these technological developments and present our perspectives on their integration into contemporary statistics education.
Causal effect estimation has gained much attention in recent years, and it is not uncommon to see graduate students incorporate these analyses into their research. However, students may not be aware how model assumptions and causal reasoning can critically influence their results. We borrowed a simple dataset that was previously used to demonstrate the use of directed acyclic graphs and designed a set of accompanying data analysis multiple-choice questions to engage students in inquiry-based learning, guiding them to discover the importance of data visualization and subject-specific causal reasoning for assessing different methods of causal effect estimation.
Educators are being encouraged to teach with "big" datasets that have more cases and attributes than are typically used in the classroom. When introduced carefully, these types of datasets can allow students to engage in complex and self-directed reasoning, develop data management and inquiry skills, and experience data analysis in a way that is more authentic to professional practice. However, "big data" is also often unwieldy. It can overwhelm students, overload their software tools, and interfere with planful analysis. How can we make such datasets manageable? This paper presents pedagogical strategies and technical methods for educators, educational designers, and young investigators to reduce the number of cases or attributes in a large dataset in ways that do not unduly compromise their analyses. We illustrate these strategies with interactive demonstrations in two freely-available open source tools: the Choosy plugin in CODAP, and a Python Jupyter notebook.