In 1935, Edgar Anderson collected size measurements for 150 flowers from three species of Iris on the Gaspe Peninsula in Quebec, Canada. Since then, Anderson's Iris observations have become a classic dataset in statistics, machine learning, and data science teaching materials. It is included in the base R datasets package as iris, making it easy for users to access without knowing much about it. However, the lack of data documentation, presence of non-intuitive variables (e.g. "sepal width "), and perfectly balanced groups with zero missing values make iris an inadequate and stale dataset for teaching and learning modern data science skills. Users would benefit from working with a more representative, real-world environmental dataset with a clear link to current scientific research. Importantly, Anderson's Iris data appeared in a 1936 publication by R. A. Fisher in the Annals of Eugenics (which is often the first-listed citation for the dataset), inextricably linking iris to eugenics research. Thus, a modern alternative to iris is needed. In this paper, we introduce the palmerpenguins R package (Horst et al., 2020), which includes body size measurements collected from 2007 - 2009 for three species of Pygoscelis penguins that breed on islands throughout the Palmer Archipelago, Antarctica. The penguins dataset in palmerpenguins provides an approachable, charismatic, and near drop-in replacement for iris with topical relevance for polar climate change and environmental impacts on marine predators. Since the release on CRAN in July 2020, the palmerpenguins package has been downloaded over 462,000 times, highlighting the demand and widespread adoption of this viable iris alternative. We directly compare the iris and penguins datasets for selected analyses to demonstrate that R users, in particular teachers and learners currently using iris, can switch to the Palmer Archipelago penguins for many use cases including data wrangling, visualization, linear modeling, multivariate analysis (e.g., PCA), cluster analysis and classification (e.g., by k-means).
Conversational impairments are well known among people with autism spectrum disorder (ASD), but their measurement requires time-consuming manual annotation of language samples. Natural language processing (NLP) has shown promise in identifying semantic difficulties when compared to clinician-annotated reference transcripts. Our goal was to develop a novel measure of lexico-semantic similarity – based on recent work in natural language processing (NLP) and recent applications of pseudo-value analysis – which could be applied to transcripts of children’s conversational language, without recourse to some ground-truth reference document. We hypothesized that: (a) semantic coherence, as measured by this method, would discriminate between children with and without ASD and (b) more variability would be found in the group with ASD. We used data from 70 4- to 8-year-old males with ASD ( N = 38) or typically developing (TD; N = 32) enrolled in a language study. Participants were administered a battery of standardized diagnostic tests, including the Autism Diagnostic Observation Schedule (ADOS). ADOS was recorded and transcribed, and we analyzed children’s language output during the conversation/interview ADOS tasks. Transcripts were converted to vectors via a word2vec model trained on the Google News Corpus. Pairwise similarity across all subjects and a sample grand mean were calculated. Using a leave-one-out algorithm, a pseudo-value, detailed below, representing each subject’s contribution to the grand mean was generated. Means of pseudo-values were compared between the two groups. Analyses were co-varied for nonverbal IQ, mean length of utterance, and number of distinct word roots (NDR). Statistically significant differences were observed in means of pseudo-values between TD and ASD groups ( p = 0.007). TD subjects had higher pseudo-value scores suggesting that similarity scores of TD subjects were more similar to the overall group mean. Variance of pseudo-values was greater in the ASD group. Nonverbal IQ, mean length of utterance, or NDR did not account for between group differences. The findings suggest that our pseudo-value-based method can be effectively used to identify specific semantic difficulties that characterize children with ASD without requiring a reference transcript.
Abstract Background Drug use-associated infective endocarditis (DUA-IE) is typically treated with 4-6 weeks of in hospital intravenous antibiotics (IVA). Outpatient parenteral antimicrobial therapy (OPAT) and partial oral antibiotics (PO) may be as effective as IVA, though long-term outcomes and costs remain unknown. We evaluated the clinical outcomes and cost-effectiveness of four antibiotic treatment strategies for DUA-IE. Methods We used a validated microsimulation model to compare: 1) 4-6 weeks of inpatient IVA along with opioid detoxification, status quo (SQ); 2) 4-6 weeks of inpatient IVA along with inpatient addiction care services (ACS) which offers medications for opioid use disorder (SQ with ACS); 3) 3 weeks of inpatient IVA with ACS followed by OPAT (OPAT); and 4) 3 weeks of IVA with ACS followed by PO antibiotics (PO). We derived model inputs from clinical trials and observational cohorts. All patients were eligible for either in-home or post-acute care OPAT. Outcomes included life years (LYs), discounted costs, incremental cost-effectiveness ratios (ICERs), proportion of DUA-IE cured, and mortality attributable to DUA-IE. Costs (&US) were annually discounted at 3%. We performed probabilistic sensitivity analyses (PSA) to address uncertainty. Results The SQ scenario resulted in 18.64 LY at a cost of &416,800/person with 77.4% hospitalized DUA-IE patients cured and 5% of deaths in the population were attributable to DUA-IE. Life expectancy was extended by each strategy: 0.017y in SQ with ACS, 0.011 in OPAT, and 0.024 in PO. The PO strategy provided the highest cure rate (80.2%), compared to 77.9% in SQ with ACS and 78.5% in OPAT and X in SQ. OPAT was the least expensive strategy at &412,300/person, Compared to OPAT, PO had an ICER of &141,500/LY. Both SQ strategies provided worse clinical outcomes for money invested than either OPAT or PO (dominated). All scenarios decreased deaths attributable to DUA-IE compared to SQ. Findings were robust in PSA. Table 1 Selected cost and clinical outcomes comparing treatment strategies for drug-use associated infective endocarditis including the status quo, status quo with addiction care services, outpatient parenteral antimicrobial therapy, and partial oral antibiotics. Conclusion Treating DUA-IE with OPAT along with ACS increases the number of people completing treatment, decreases DUA-IE mortality, and is cost-saving compared to the status quo. The PO strategy also improves clinical outcomes, but may not be cost-effective at the willingness-to-pay threshold of &100,000. Disclosures Simeon D. Kimmel, MD, MA, Abt Associates for a Massachusetts Department of Public Health project to improve access to medications for opioid use disorder in nursing facilities (Consultant)
Participatory live coding is a technique in which a teacher or instructor writes and narrates code out loud as they teach and invites learners to join them by writing and executing the same code. Learners watch as an instructor writes code live in real time, typically via 1 or more projector screens that show the same screen as the instructor sees. Instructors also read out loud what they type, explaining the different elements and principles that are relevant for learners to understand the code. At the same time, each learner is invited to copy and execute the exact code or commands that are being written on their own work station. Learners thus “code-along” with the instructor. There are frequent, often short, exercises, in which learners are asked to solve a small relevant problem on their own. This approach aims to be an improvement on teaching programming through lecturing showing static code or relying on learners reading a textbook or compendium. What is taught is immediately applied rather than just shown on a slide or on paper: It embodies the “I do, we do, you do” approach to knowledge transfer that is used both formally and informally to teach everything from laboratory bench skills to grant writing [1]. It also slows the instructor down, giving learners more time to actively engage with the material before moving on to the next concept. Importantly, the thought process behind coding can also be made explicit. Learner’s questions can immediately be answered and misconceptions corrected by coding them. Exercises enable immediate practice using the material. Crucially, the technique also allows for teaching handling of mistakes. Beyond deliberately introducing mistakes during the live coding, instructors will often make unplanned mistakes. Novice learners are likely to make many such mistakes themselves, and diagnosing and solving mistakes is an integral aspect of learning programming. The participatory aspect engages learners, which helps them become active practitioners rather than passive observers of the programming process. Participatory live coding is most beneficial for novices who are unfamiliar with the tools. More experienced learners may gain enough by listening passively or engaging with the material and classroom differently. Participatory live coding for teaching programming should not be confused with live coding used to demonstrate software (for example, at a conference, with an audience passively observing) or live streaming programming [2] or live coding used as a form of performing art (e.g., while creating computer music [3]). A video recording demonstrating the participatory live coding technique can be found here: https://vimeo.com/139316669. PLOS COMPUTATIONAL BIOLOGY
This article provides a review of substantial findings and methodological issues in epidemiological surveys of autism. Studies published since 2 000 are reviewed and indicate huge heterogeneity of methods across surveys. Prevalence estimates vary widely, with a range of prevalence from 0.7 % to 1.5 % being consistent with recent, well-designed studies. Factors that explain time trends in prevalence are examined, including changes in diagnostic concepts and criteria, diagnostic substitution and improved awareness and case ascertainment. Finally, we review how factors such as social class and ethnic minority status affect prevalence in subgroups.
This article provides a review of substantial findings and methodological issues in epidemiological studies on autism. Studies published since 2000 were reviewed and a large degree of heterogeneity of methods across surveys can be observed. Prevalence estimates vary widely, with a range of prevalence from 0.7 % to 1.5 % being consistent with recent, well-designed studies. Factors that explain time trends in prevalence are examined, including changes in diagnostic concepts and criteria and improved awareness and case ascertainment. Finally, we review how factors such as social class and ethnic minority status affect prevalence in subgroups.
Inconsistent findings regarding sex differences in cognition have been found in people with autism spectrum disorder (ASD). This study evaluated sex differences in cognitive-developmental functioning in a large clinical sample of young children diagnosed with ASD. The sample included children 18–68 months of age who received the Mullen Scales of Early Learning (MSEL) through Autism Treatment Network (ATN) sites from 2007 to 2013 (N = 1587, 16.7% female). In this large clinically referred sample of young children with ASD in the United States, no significant differences were found between the sexes for the MSEL Early Learning Composite (ELC) standard score, domain T Scores or age equivalents. These findings persisted when examining different age ranges, cognitive levels and domain profiles.
Cet article passe en revue les résultats importants et les problèmes méthodologiques rencontrés lors des enquêtes épidémiologiques sur l’autisme. Les études publiées depuis 2000 sont passées en revue et indiquent une énorme hétérogénéité des méthodes entre les enquêtes. Les estimations de la prévalence varient considérablement, la fourchette de prévalence allant de 0,7 % à 1,5 %, en cohérence avec les études récentes et bien conçues. Les facteurs expliquant les changements de prévalence au cours du temps sont examinés, notamment les changements de concepts et de critères diagnostiques et l’amélioration de la sensibilisation à l’autisme et à sa détermination. Enfin, sont examinés comment des facteurs tels que la classe sociale et le statut de minorité ethnique affectent la prévalence dans les sous-groupes.
Deficits in social communication, particularly pragmatic language, are characteristic of individuals with autism spectrum disorder (ASD). Speech disfluencies may serve pragmatic functions such as cueing speaking problems. Previous studies have found that speakers with ASD differ from typically developing (TD) speakers in the types and patterns of disfluencies they produce, but fail to provide sufficiently detailed characterizations of the methods used to categorize and quantify disfluency, making cross-study comparison difficult. In this study we propose a simple schema for classifying major disfluency types, and use this schema in an exploratory analysis of differences in disfluency rates and patterns among children with ASD compared to TD and language impaired (SLI) groups. 115 children ages 4-8 participated in the study (ASD = 51; SLI = 20; TD = 44), completing a battery of experimental tasks and assessments. Measures of morphological and syntactic complexity, as well as word and disfluency counts, were derived from transcripts of the Autism Diagnostic Observation Schedule (ADOS). High inter-annotator agreement was obtained with the use of the proposed schema. Analyses showed ASD children produced a higher ratio of content to filler disfluencies than TD children. Relative frequencies of repetitions, revisions, and false starts did not differ significantly between groups. TD children also produced more cued disfluencies than ASD children.
DSM-5 Autism Spectrum Disorder (ASD) comprises a set of neurodevelopmental disorders characterized by deficits in social communication and interaction and repetitive behaviors or restricted interests, and may both affect and be affected by multiple cognitive mechanisms. This study attempts to identify and characterize cognitive subtypes within the ASD population using our Functional Random Forest (FRF) machine learning classification model. This model trained a traditional random forest model on measures from seven tasks that reflect multiple levels of information processing. 47 ASD diagnosed and 58 typically developing (TD) children between the ages of 9 and 13 participated in this study. Our RF model was 72.7% accurate, with 80.7% specificity and 63.1% sensitivity. Using the random forest model, the FRF then measures the proximity of each subject to every other subject, generating a distance matrix between participants. This matrix is then used in a community detection algorithm to identify subgroups within the ASD and TD groups, and revealed 3 ASD and 4 TD putative subgroups with unique behavioral profiles. We then examined differences in functional brain systems between diagnostic groups and putative subgroups using resting-state functional connectivity magnetic resonance imaging (rsfcMRI). Chi-square tests revealed a significantly greater number of between group differences (p < .05) within the cingulo-opercular, visual, and default systems as well as differences in inter-system connections in the somato-motor, dorsal attention, and subcortical systems. Many of these differences were primarily driven by specific subgroups suggesting that our method could potentially parse the variation in brain mechanisms affected by ASD.
1Assistant Professor, Center for Spoken Language Understanding Institute on Development & Disability, Oregon Health & Science University 2Assistant Professor, Division of General Pediatrics Doernbecher Children’s Hospital, Oregon Health & Science University Affiliate Assistant Professor, Oregon Health & Science University – Portland State University School of Public Health 3Professor, Department of Psychiatry Director of Autism Research, Institute for Development & Disability Oregon Health & Science University
Atypical pragmatic language is often present in individuals with autism spectrum disorders (ASD), along with delays or deficits in structural language. This study investigated the use of the “fillers” uh and um by children ages 4–8 during the autism diagnostic observation schedule. Fillers reflect speakers' difficulties with planning and delivering speech, but they also serve communicative purposes, such as negotiating control of the floor or conveying uncertainty. We hypothesized that children with ASD would use different patterns of fillers compared to peers with typical development or with specific language impairment (SLI), reflecting differences in social ability and communicative intent. Regression analyses revealed that children in the ASD group were much less likely to use um than children in the other two groups. Filler use is an easy‐to‐quantify feature of behavior that, in concert with other observations, may help to distinguish ASD from SLI. Autism Res 2016, 9: 854–865 . © 2016 International Society for Autism Research, Wiley Periodicals, Inc.
Background: A subgroup of young children with autism spectrum disorders (ASD) have significant language impairments (phonology, grammar, vocabulary), although such impairments are not considered to be core symptoms of and are not unique to ASD. Children with specific language impairment (SLI) display similar impairments in language. Given evidence for phenotypic and possibly etiologic overlap between SLI and ASD, it has been suggested that language- impaired children with ASD (ASD + language impairment, ALI) may be characterized as having both ASD and SLI. However, the extent to which the language phenotypes in SLI and ALI can be viewed as similar or different depends in part upon the age of the individuals studied. The purpose of the current study is to examine differences in memory abilities, specifically those that are key "markers" of heritable SLI, among young school-age children with SLI, ALI, and ALN (ASD + language normal).Methods: In this cross-sectional study, three groups of children between ages 5 and 8 years participated: SLI (n = 18), ALI (n = 22), and ALN (n = 20). A battery of cognitive, language, and ASD assessments was administered as well as a nonword repetition (NWR) test and measures of verbal memory, visual memory, and processing speed.Results: NWR difficulties were more severe in SLI than in ALI, with the largest effect sizes in response to nonwords with the shortest syllable lengths. Among children with ASD, NWR difficulties were not associated with the presence of impairments in multiple ASD domains, as reported previously. Verbal memory difficulties were present in both SLI and ALI groups relative to children with ALN. Performance on measures related to verbal but not visual memory or processing speed were significantly associated with the relative degree of language impairment in children with ASD, supporting the role of verbal memory difficulties in language impairments among early school-age children with ASD.Conclusions: The primary difference between children with SLI and ALI was in NWR performance, particularly in repeating two- and three-syllable nonwords, suggesting that shared difficulties in early language learning found in previous studies do not necessarily reflect the same underlying mechanisms.
OBJECTIVE:Overweight and obesity are increasingly prevalent in the general pediatric population. Evidence suggests that children with autism spectrum disorders (ASDs) may be at elevated risk for unhealthy weight. We identify the prevalence of overweight and obesity in a multisite clinical sample of children with ASDs and explore concurrent associations with variables identified as risk factors for unhealthy weight in the general population.METHODS:Participants were 5053 children with confirmed diagnosis of ASD in the Autism Speaks Autism Treatment Network. Measured values for weight and height were used to calculate BMI percentiles; Centers for Disease Control and Prevention criteria for BMI for gender and age were used to define overweight and obesity (≥85th and ≥95th percentiles, respectively).RESULTS:In children age 2 to 17 years, 33.6% were overweight and 18% were obese. Compared with a general US population sample, rates of unhealthy weight were significantly higher among children with ASDs ages 2 to 5 years and among those of non-Hispanic white origin. Multivariate analyses revealed that older age, Hispanic or Latino ethnicity, lower parent education levels, and sleep and affective problems were all significant predictors of obesity.CONCLUSIONS:Our results indicate that the prevalence of unhealthy weight is significantly greater among children with ASD compared with the general population, with differences present as early as ages 2 to 5 years. Because obesity is more prevalent among older children in the general population, these findings raise the question of whether there are different trajectories of weight gain among children with ASDs, possibly beginning in early childhood.
Aggressive behavior problems (ABP) are frequent yet poorly understood in children with Autism Spectrum Disorders (ASD) and are likely to co-vary significantly with comorbid problems. We examined the prevalence and sociodemographic correlates of ABP in a clinical sample of children with ASD (N = 400; 2-16.9 years). We also investigated whether children with ABP experience more intensive medical interventions, greater impairments in behavioral functioning, and more severe comorbid problems than children with ASD who do not have ABP. One in four children with ASD had Child Behavior Checklist scores on the Aggressive Behavior scale in the clinical range (T-scores ≥ 70). Sociodemographic factors (age, gender, parent education, race, ethnicity) were unrelated to ABP status. The presence of ABP was significantly associated with increased use of psychotropic drugs and melatonin, lower cognitive functioning, lower ASD severity, and greater comorbid sleep, internalizing, and attention problems. In multivariate models, sleep, internalizing, and attention problems were most strongly associated with ABP. These comorbid problems may hold promise as targets for treatment to decrease aggressive behavior and proactively identify high-risk profiles for prevention.
Autism Spectrum Disorders (ASDs) and childhood obesity (OBY) are rising public health concerns. This study aimed to evaluate the prevalence of overweight (OWT) and OBY in a sample of 376 Oregon children with ASD, and to assess correlates of OWT and OBY in this sample. We used descriptive statistics, bivariate, and focused multivariate analyses to determine whether socio-demographic characteristics, ASD symptoms, ASD cognitive and adaptive functioning, behavioral problems, and treatments for ASD were associated with OWT and OBY in ASD. Overall 18.1% of children met criteria for OWT and 17.0% met criteria for OBY. OBY was associated with sleep difficulties, melatonin use, and affective problems. Interventions that consider unique needs of children with ASD may hold promise for improving weight status among children with ASD.
In this selective review of the literature, we present the most recent prevalence estimates for autism spectrum disorder (ASD) and discuss the limitations and challenges in interpreting changes in prevalence estimates over time. Increases in ASD prevalence estimates cannot currently be attributed to a true increase in the incidence of ASD due to multiple confounding factors. These include broader diagnostic criteria and a greater awareness of ASD. The current average prevalence of ASD is approximately 66/10,000, which translates to approximately 1 in 152 children affected, with males consistently outnumbering females by about 5:1. Several recent studies have reported higher estimates ranging from 147 (one in 68) to 264 (one in 38) per 10,000. This is in sharp contrast to the figures of about 1-5/10,000 quoted in earlier studies that used a narrow definition of autistic disorder and were not inclusive of all disorders falling onto the autism spectrum. It remains to be seen how changes to diagnostic criteria introduced in the DSM-5 will impact estimates of ASD prevalence.
We report on an automatic technique for quantifying two types of repetitive speech: repetitions of what the child says him/herself (self‐repeats) and of what is uttered by an interlocutor (echolalia). We apply this technique to a sample of 111 children between the ages of four and eight: 42 typically developing children ( TD ), 19 children with specific language impairment ( SLI ), 25 children with autism spectrum disorders ( ASD ) plus language impairment ( ALI ), and 25 children with ASD with normal, non‐impaired language ( ALN ). The results indicate robust differences in echolalia between the TD and ASD groups as a whole ( ALN + ALI ), and between TD and ALN children. There were no significant differences between ALI and SLI children for echolalia or self‐repetitions. The results confirm previous findings that children with ASD repeat the language of others more than other populations of children. On the other hand, self‐repetition does not appear to be significantly more frequent in ASD , nor does it matter whether the child's echolalia occurred within one (immediate) or two turns (near‐immediate) of the adult's original utterance. Furthermore, non‐significant differences between ALN and SLI , between TD and SLI , and between ALI and TD are suggestive that echolalia may not be specific to ALN or to ASD in general. One important innovation of this work is an objective fully automatic technique for assessing the amount of repetition in a transcript of a child's utterances. Autism Res 2013, ●●: ●●–●●. © 2013 International Society for Autism Research, Wiley Periodicals, Inc.