The Coefficient of Dependence is introduced as an avenue for aiding student understanding of dependent, statistical events. The essential features of this coefficient are studied using different models and explored with multiple examples. Conditional probabilities are ultimately understood to be simple transformations of marginal probabilities via the Coefficient of Dependence, which is an idea well within the grasp of all undergraduate students. For completeness, proofs for the main results are relegated to the appendix for those interested.
In statistical inference, oftentimes the data are assumed to be normally distributed. Consequently, testing the validity of the normality assumption is an integral part of such statistical analyses. Here, we investigate twelve currently available tests for normality using Monte-Carlo simulation. Alternative distributions are used to calculate the empirical power of the tests studied here. The distributions considered arise from three different categories: symmetric short-tailed, symmetric long-tailed and asymmetric. In addition, power is calculated for several contaminated alternatives. As a direct consequence of this study, we recommend a two-tier approach: (i) observe the shape of the empirical data distribution using graphical methods, then (ii) select an appropriate test based on the likely distributional shape and the corresponding sample size. In general, with respect to power considerations, it is observed that for asymmetric distributions, the Shapiro-Wilk and Ryan-Joiner tests perform fairly well for all sample sizes studied here. Additionally, the Jarque-Bera, Modified Jarque-Bera, and Ryan-Joiner tests perform fairly well for contaminated normal distributions. The popular methods available in current software packages, such as the Shapiro-Wilk test, the Ryan-Joiner Normality test, and the Anderson-Darling goodness of test, work at least moderately well for most of the cases we considered.
The Pythagorean Win-Loss formula can be effectively used to estimate winning percentages for sporting events. This formula was initially developed by baseball statistician Bill James and later was extended by other researchers to sports such as football, basketball, and ice hockey. Although one can calculate actual winning percentages based on the outcomes of played games, that approach does not take into account the margin of victory. The key benefit of the Pythagorean formula is its utilization of actual average runs scored and actual average runs allowed. This article presents the application of the Pythagorean Win-Loss formula to two different types of limited-overs cricket formats, namely One Day International cricket (ODI) and Twenty20 cricket. The data for the application was used from the matches played by the top 10 International Cricket Council (ICC) members who participated in the 2019 ICC Cricket World Cup. For matches for which the second batting team won, runs scored were estimated by considering the remaining amount of resources, based on the Duckworth–Lewis method.
Bowling effectiveness is a key factor in winning cricket matches. The team captain should decide when to use the right bowler at the right moment so that the team can optimize the outcome of the game. In this study, we investigate the effectiveness of different types of bowlers at different stages of the game, based on the conceded percentage of runs from the innings total, for each over. Bowlers are generally categorized into three types: fast bowlers, medium-fast bowlers, and spinners. In this article, the authors divided the twenty over spell of a T20I match into four stages; namely, Stage 1: overs 1-6 (PowerPlay), Stage 2: overs 7-10, Stage 3: overs 11-15, and Stage 4: overs 16-20. To understand the broad spectrum of the behavior of game variables, a Quantile Regression methodology is used for statistical analysis. Following that, a Bayesian approach to Quantile Regression is undertaken, and it confirms the initial results.
In cricket, all-rounders play an important role. A good all-rounder should be able to contribute to the team by both bat and ball as needed. However, these players still have their dominant role by which we categorize them as batting all-rounders or bowling all-rounders. Current practice is to do so by mostly subjective methods. In this study, the authors have explored different machine learning techniques to classify all-rounders into bowling all-rounders or batting all-rounders based on their observed performance statistics. In particular, logistic regression, linear discriminant function, quadratic discriminant function, naïve Bayes, support vector machine, and random forest classification methods were explored. Evaluation of the performance of the classification methods was done using the metrics accuracy and area under the ROC curve. While all the six methods performed well, logistic regression, linear discriminant function, quadratic discriminant function, and support vector machine showed outstanding performance suggesting that these methods can be used to develop an automated classification rule to classify all-rounders in cricket. Given the rising popularity of cricket, and the increasing revenue generated by the sport, the use of such a prediction tool could be of tremendous benefit to decision-makers in cricket.
Formulation optimization and antidotal combination therapy are the two important tools to enhance the antidotal protection of the cyanide (CN) antidote dimethyl trisulfide (DMTS). The focus of this study is to demonstrate how the formulation with polysorbate 80 (Poly80), an excipient used in pharmaceutical technology, and the combinations with other CN antidotes having different mechanisms of action enhance the antidotal efficacy of the unformulated (neat) DMTS. The LD50 for CN was determined by the statistical Dixon up-and-down method on mice. Antidotal efficacy was expressed as antidotal potency ratio (APR). CN was injected subcutaneously one minute prior to the antidotes' injection intramuscularly. The APR values of 1.17 (dose: 25 mg/kg bodyweight) and 1.45 (dose: 50 mg/kg bodyweight) of the neat DMTS were significantly enhanced by the Poly80 formulation at both investigated doses to 2.03 and 2.33, respectively. The combination partners for the Poly80 formulated DMTS (DMTS-Poly80; 25 and 50 mg/kg bodyweight) were 4-nitrocobinamide (4NCbi) (20 mg/kg bodyweight) and aquohydroxocobinamide (AHCbi; 50, 100, and 250 mg/kg bodyweight). When DMTS-Poly80 (25 and 50 mg/kg bodyweight; APR = 2.03 and 2.33, respectively) was combined with 4NCbi (20 mg/kg bodyweight; APR = 1.35), significant increase in the APR values were noted at both DMTS doses (APR = 2.38 and 3.12, respectively). AHCbi enhanced the APR of DMTS-Poly80 (100 mg/kg bodyweight; APR = 3.29) significantly only at the dose of 250 mg/kg bodyweight (APR = 5.86). These studies provided evidence for the importance of the formulation with Poly80 and the combinations with cobinamide derivatives with different mechanisms of action for DMTS as a CN antidote candidate.
Increasing classroom attendance rates is important to improving success in developmental mathematics courses. Our results indicated that absence rates begin to increase after the first exam. Furthermore, the number of absences gradually increased throughout the semester and a higher proportion of class meeting time was missed in classes that met three days per week compared to two days per week. Thus, attempts to emphasize to students the importance of classroom attendance on course grades needs to begin before the first test. When absences were measured by proportion of class meeting time, effect of absences on students’ failing the course was higher for classes that met two days per week compared to classes that met three days per week.
While serving as critical tools against bacterial infections, antimicrobial therapies can also result in serious side effects, such as antibiotic-associated entercolitis. Recent studies utilizing next generation sequencing to generate community 16S gene profiles have shown that antibiotics can strongly alter community composition and deplete diversity. However, how these community changes in the microbiota are related to the host side effects is still unclear. We have used the freshwater Western mosquitofish (Gambusia affinis) as a tractable vertebrate model system to study host effects following exposure to a broad spectrum antibiotic, rifampicin. After 3days of exposure, the bacterial communities of the mucosal skin and gut microbiomes lost diversity and shifted composition. Compared to unexposed controls, treated fish were more susceptible to a specific pathogen, Edwardsiella ictaluri, yet displayed no survival differences when subjected to a polymicrobial water challenge of soil or feces. Treated fish were more susceptible to osmotic stress from NaCl, but not to the toxin nitrate. Treated fish failed to gain weight as well as controls over one month when fed a matched diet. Because of small sample sizes, pathogen susceptibility and weight gain differences were not statistically significant. This study provides supporting evidence in an experimental laboratory system that an antibiotic can have significant and persistent negative host effects, and provides for future study into the mechanisms of these effects.
This paper investigates the powerplay in one-day cricket. The rules concerning the powerplay have been tinkered with over the years, and therefore the primary motivation of the paper is the assessment of the impact of the powerplay with respect to scoring. The form of the analysis takes a "what if" approach where powerplay outcomes are substituted with what might have happened had there been no powerplay. This leads to a paired comparisons setting consisting of actual matches and hypothetical parallel matches where outcomes are imputed during the powerplay period. Some of our findings include (a) the various forms of the powerplay which have been adopted over the years have different effects, (b) recent versions of the powerplay provide an advantage to the batting side, (c) more wickets also occur during the powerplay than had there been no powerplay and (d) there is some effect in run production due to the over where the powerplay is initiated. We also investigate individual batsmen and bowlers and their performances during the powerplay. (C) 2015 Elsevier B.V. All rights reserved.
s Invited Talks (1) Friday, March 21, 1:30pm-4:50pm Low-Dimensional Approximations for Functional Data with Covariates Joan Staniswalis Department of Mathematical Sciences, University of Texas at El Paso http://faculty.utep.edu/Default.aspx?alias=faculty.utep.edu/jstaniswalis Ramsay (1996) first proposed the method of principal differential analysis (PDA) for fitting a differential equation to a collection of noisy data curves. Each data curve is modeled as a (possibly noisy observation of a) smooth curve in the null space of a linear differential operator of order m. Smooth estimates of the coefficient functions defining the linear differential operator are obtained by minimizing a penalized sum of the squared norm of the forcing functions: that part of the data curve that is not annihilated by the linear differential operator. Once the linear differential operator is estimated, a nonparametric basis of functions for the null space is computed using iterative methods. The nonparametric basis functions can be used to provide a smooth low dimensional approximation to the data curves. This paper extends PDA to allow for the coefficients in the linear differential equation to smoothly depend upon continuous subject covariates. This is implemented with local smoothing in the Rsoftware and used to explore how the cortical auditory evoked potentials (CAEP) of subjects with normal hearing change with age. Data-based selection of the smoothing parameters is addressed. The Role of Modern Social Media Data in Surveillance and Prediction of Infectious Diseases: from Time Series to Networks Yulia Gel Department of Mathematical Sciences, University of Texas at Dallas yxg142030@utdallas.edu The prompt detection and forecasting of infectious diseases with rapid transmission and high virulence are critical in the effective defense against these diseases. Despite many promising approaches in modern surveillance methodology, the lack of observations for near real-time forecasting is still the key challenge obstructing operational prediction and control of disease dynamics. For instance, even CDC data for well monitored areas in USA are two weeks behind, as it takes time to confirm influenza like illness (ILI) as flu, while two weeks is a substantial time in terms of flu transmission. These limitations have ignited the recent interest in searching for alternative near real-time data sources on the current epidemic state and, in particular, in the wealth of health-related information offered by modern social media. For example, Google Flu Trends uses flu-related searches to predict a future epidemiological state at a local level, and more recently, Twitter has also proven to be a very valuable resource for a wide spectrum of public health applications. In this talk we will review capabilities and limitations of such social media data as early warning indicators of influenza dynamics in conjunction with traditional time series epidemiological models and with more recent random network approaches accounting for heterogeneous social interaction patterns.
Metagenomics and bacterial culture were used to determine the normal skin microbiome of the Western mosquitofish (Gambusia affinis).This is the first study of G. affinis, and the most in-depth study of any fish skin, utilizing a combination of 16S profile pyrosequencing and culture analysis.Over 1800 sequences obtained from three individuals reveal that over half of all sequences come from five invariant genera, Acinetobacter, Sphingomonas, Acidovorax, Enhydrobacter, and Aquabacterium.The microbiome is diverse but has low equitability, with a total of 81 genera detected.Challenge studies suggest that non-native bacteria cannot colonize the skin.This definition of the normal skin microbiome lays the foundation for future studies with this model system.
The impact of the factors home field , winning the toss and team superiority on the outcome of Limited Overs One International (ODI) cricket matches is investigated. Our results show that the home-field-advantage is a significant factor but only for Day matches. Confirming previous studies, we also show that winning the toss does not give any statistically significant advantage towards the outcome of the game.
Performance analysis of cricket players is always an intricate task due to the correlated nature of the variables used to quantify contributions to the team. Lack of transparency of current methods, probably due to commercial confidentiality, creates a necessity for new and lucent evaluative methods. Here, we present a simple, yet straightforward, method for analyzing the performance of T20-World Cup Cricket 2012 players that can be easily adapted to other team sports.DOI: http://dx.doi.org/10.4038/sljastats.v14i1.5873
Principal Component Analysis is widely used in applied multivariate data analysis, and this article shows how to motivate student interest in this topic using cricket sports data. Here, principal component analysis is successfully used to rank the cricket batsmen and bowlers who played in the 2012 Indian Premier League (IPL) competition. In particular, the first principal component is seen to explain a substantial portion of the variation in a linear combination of some commonly used measures of cricket prowess. This application provides an excellent, elementary introduction to the topic of principal component analysis.
Underweight status in older adults is well documented. Causes include physiologic and biochemical changes as well as psychosocial changes associated with aging. The assisted living industry has grown in recent years in response to the need of older adults with such challenges. A number of studies have been conducted addressing the food service and dietary modifications in assisted living facilities. Most have been assessed in terms of quantity of food consumed. This study addresses changes in body weight of older adults following admission to an assisted living facility. Data relevant to weight changes from records of residents at assisted living facilities were retrospectively reviewed over the first ninety days of admission. Using admission weight as a baseline, weight changes were assessed at 30, 60 and 90 days. These changes were further analyzed based on initial nutritional status based on Body Mass Index (BMI) and gender of the residents. Surveys of the facilities and their food service were also completed. A significant number of residents of assisted living facilities studied were nutritionally at risk or underweight at the time of admission. A statistically significant increase in weight is seen over the first ninety days of admission to an assisted living facility in females. Furthermore, a significant weight gain is found in residents with normal BMI on admission. The results of this pilot study clearly justify the neediness of a large scale study to address the issue of weight changes of the residents after admission to assisted living facilities.
A receiver operating characteristic (ROC) curve visually demonstrates the tradeoff between sensitivity and specificity as a function of varying a classification threshold. It is a common practice to use ROC curves to measure the accuracy of predictions by different methods. Although this method has been used primarily in medical and engineering fields, it could be used effectively in sports as well. More precisely, an ROC plots the sensitivity versus (1 - specificity), and the area under the curve gives a measure of the prediction. So, the ideal best prediction should have one square unit of area under the ROC curve, where it achieves both 100% sensitivity and 100% specificity (which, in reality, rarely happens). Consequently, when we compare two methods, the one with the greater area under its ROC is judged best. This paper shows the effectiveness of using the ROC curves in analyzing cricket data. In particular, the quality of the decisions made by umpires is investigated. Also, the comparison of the accuracy of the methods for revising the target for matches shortening due to weather interruptions is the key interest in our investigation.
INTRODUCTION Even after a good first course, it is surprising to see that a large percentage of students do not have a clear understanding of basic concepts in probability and statistics. Several excellent articles discuss student misconceptions of basic probabilistic ideas. Gibbons [1] provides an analytical explanation contrasting differences between mutually exclusive events, independence and zero correlation. Rossman and Short [2] address the teaching aspects of conditional probability and its role in statistics education reform. Hirsch and O 'Donnei [3] also study some common student misconceptions concerning statistical reasoning, while Acker [4] discusses the sequential method of student thinking about conditional probabilities. Here, we consider some common student misconceptions concerning probabilistic independence and mutual exclusivity derived from survey data, and offer clear examples to use in the classroom. We focus discussion on college students enrolled in one or two introductory, undergraduate statistics courses as adjunct to their degree program as well as students taking advanced placement (AP) statistics courses. 1. PEDAGOGY Many first courses in probability and statistics begin with the graphical and numerical aspects of descriptive statistics: bar charts, histograms, pie charts, box plots, averages, standard deviations, and the like. In unison, students are usually exposed to the power of calculator (or software) technology to help them quickly describe data. Indeed, most inexpensive calculators available on the market today are capable of doing a host of descriptive statistical tasks quite efficiently. Typically, a full discussion of elementary probability concepts only comes later in the course, and on first brush it appears to students to be a digression. With the introduction of probability ideas, students begin to realize that the course is more challenging than first anticipated, and it entails more than summarizing columns of numbers and constructing tables and graphs. Once the elementary properties of probability are introduced, it is natural to consider the roles of independence and mutual exclusivity in the topical sequence. Additionally, it is common pedagogical practice to explore the concept of conditional probability and study its relationship to independence. Because many popular textbooks deemphasize probability concepts, especially the contrast between independence and mutually exclusivity, this is a critical juncture and instructors must be very careful at this point in the course development. Surprisingly, we have noticed that many students, even those in graduate courses, have a persistent misunderstanding of independence and mutual exclusivity. Perhaps we, their instructors, are to blame for not providing enough clarity. In this article, we show the surprising outcomes of a survey that indicates the magnitude of the problem, and we offer possible ways to present this material so that students actually understand. Throughout, we assume only that students: (i) have been exposed to the idea of a probabilistic (or statistical) experiment, (ii) have an elementary notion of the concept of an event and its occurrence or nonoccurrence, and (iii) have studied the basic requirements for a probability function. 2. DEFINITIONS AND NOTES We begin with an introduction to the concepts of independence and mutual exclusivity. If the information about the occurrence or nonoccurrence of a particular event does not give any probabilistic information about occurrence or nonoccurrence of another event, then those two events are said to be probabilistically independent. More formally, Definition 1: Independence Events A and B are called independent if and only if P(A and B) = P(A) × P(B) . Otherwise, they are said to be dependent (or not independent). If the simultaneous occurrence of two events A and B is impossible, then these events are said to be mutually exclusive or incompatible. …
Present studies have focused on nano-intercalated rhodanese in combination with sulfur donors to prevent cyanide lethality in a prophylactic mice model for future development of an effective cyanide antidotal system. Our approach is based on the idea of converting cyanide to the less toxic thiocyanate before it reaches the target organs by utilizing sulfurtransferases (e.g., rhodanese) and sulfur donors in a close proximity by injecting them directly into the blood stream. The inorganic thiosulfate (TS) and the garlic component diallydisulfide (DADS) were compared as sulfur donors with the nano-intercalated rhodanese in vitro and in vivo. The in vivo and in vitro experiments showed that DADS is not a more efficient sulfur donor than TS. However, the utilization of external rhodanese significantly enhanced the in vivo efficacy of both sulfur donor-nitrite combinations, indicating the potential usefulness of enzyme nano-delivery systems in developing antidotal therapeutic agents.