Modeling and simulations offer significant benefits for self-directed learning as learners can independently design experiments and investigate their own hypotheses at their discretion. However, many studies on the use of such tools focus on pedagogical contexts in K-12 education with well-defined problems, learning goals, assessments, and outcomes. This study investigates how online learners engage in self-directed modeling by analyzing the behaviors of 315 learners across 822 models within VERA, an ecological modeling tool. Through learning analytics techniques, including activity sequence analysis, hierarchical clustering, and Markov chain models, we identify three distinct behavioral patterns-Observation, Construction, and Exploration-that reflect varying levels of engagement and interaction. Our findings reveal that learners transition from construction-focused behaviors to more active, hypothesis-driven exploration, with observation consistently present across all learning phases. This research contributes to the fields of AI in education and self-directed learning by providing a framework for understanding how learners interact with modeling tools outside traditional classroom settings. Future work may build on these patterns to offer adaptive and personalized learning.
Developing a predictive science of the biosphere depends heavily on the rapidly expanding biodiversity data that are commonly stored in biodiversity databases. Despite the proliferation of biodiversity databases, their independent operation has limited data discovery, comparison, and synthesis. Therefore, the biodiversity informatics community has called for improved alignment among these efforts to better catalog Earth's biodiversity. The primary challenges are incomplete knowledge of existing databases and incongruent taxonomic systems and data schemas. Addressing these issues will require development of a database registry, means to compare database contents, taxonomic harmonization, and tools that enable users to merge disparate databases based on their needs, all within a community of practice that enables people of various skill levels and roles to participate. We believe that synthesis and integration, driven by a growing and thriving community, will be the next stage of biodiversity informatics and will help unlock the full potential of biodiversity information.
We described a study on the use of an online laboratory for self-directed learning by constructing and simulating conceptual models of ecological systems. In this study, we could observe only the modeling behaviors and outcomes; the learning goals and outcomes were unknown. We used machine learning techniques to analyze the modeling behaviors of 315 learners and 822 conceptual models they generated. We derive three main conclusions from the results. First, learners manifest three types of modeling behaviors: observation (simulation focused), construction (construction focused), and full exploration (model construction, evaluation and revision). Second, while observation was the most common behavior among all learners, construction without evaluation was more common for less engaged learners and full exploration occurred mostly for more engaged learners. Third, learners who explored the full cycle of model construction, evaluation and revision generated models of higher quality. These modeling behaviors provide insights into self-directed learning at large.
Inquiry-based modeling is essential to scientific practice. However, modeling is difficult for novice scientists in part due to limited domain-specific knowledge and quantitative skills. VERA is an interactive tool that helps users construct conceptual models of ecological phenomena, run them as simulations, and examine their predictions. VERA provides cognitive scaffolding for modeling by supplying access to large-scale domain knowledge. The VERA system was tested by college-level students in two different settings: a general ecology lecture course (N=91) at a large southeastern R1 university and a controlled experiment in a research laboratory (N=15). Both studies indicated that engaging students in ecological modeling through VERA helped them better understand basic biological concepts. The latter study additionally revealed that providing access to domain knowledge helped students build more complex models.
We describe a study on the use of an online laboratory for self-directed learning through the construction and simulation of conceptual models of ecological systems. We analyzed the modeling behaviors of 315 learners and 822 instances of learner-generated models using a sequential pattern mining technique. We found three types of learner behaviors: observation, construction, and exploration. We found that while the observation behavior was most common, exploration led to models of higher quality.
We present an interactive modeling tool, VERA, that scaffolds the acquisition of domain knowledge involved in conceptual modeling and agent-based simulations. We describe the knowledge engineering process of contextualizing large-scale domain knowledge. Specifically, we use the ontology of biotic interactions in Global Biotic Interactions, and the trait data of species in Encyclopedia of Life to facilitate the model construction. Learners can use VERA to construct qualitative conceptual models of ecological phenomena, run them as quantitative simulations, and review their predictions.
Traits have become a crucial part of ecological and evolutionary sciences, helping researchers understand the function of an organism's morphology, physiology, growth and life history, with effects on fitness, behaviour, interactions with the environment and ecosystem processes. However, measuring, compiling and analysing trait data comes with data‐scientific challenges. We offer 10 (mostly) simple rules, with some detailed extensions, as a guide in making critical decisions that consider the entire life cycle of trait data. This article is particularly motivated by its last rule, that is, to propagate good practice. It has the intention of bringing awareness of how data on the traits of organisms can be collected and managed for reuse by the research community. Trait observations are relevant to a broad interdisciplinary community of field biologists, synthesis ecologists, evolutionary biologists, computer scientists and database managers. We hope these basic guidelines can be useful as a starter for active communication in disseminating such integrative knowledge and in how to make trait data future‐proof. We invite the scientific community to participate in this effort at http://opentraits.org/best‐practices.html.
Modeling is an important aspect of scientific problem-solving. However, modeling is a difficult cognitive process for novice learners in part due to the high dimensionality of the parameter search space. This work investigates 50 college students’ parameter search behaviors in the context of ecological modeling. The study revealed important differences in behaviors of successful and unsuccessful students in navigating the parameter space. These differences suggest opportunities for future development of adaptive cognitive scaffolds to support different classes of learners.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
The Encyclopedia of Life currently hosts ~8M attribute records for ~400k taxa (March 2019, not including geographic categories, Fig. 1). Our aggregation priorities include Essential Biodiversity Variables (Kissling et al. 2018) and other global scale research data priorities. Our primary strategy remains partnership with specialist open data aggregators; we are also developing tools for the deployment of evolutionarily conserved attribute values that scale quickly for global taxonomic coverage, for instance: tissue mineralization type (aragonite, calcite, silica...); trophic guild in certain clades; sensory modalities. To support the aggregation and integration of trait information, data sets should be well structured, properly annotated and free of licensing or contractual restrictions so that they are ‘findable, accessible, interoperable, and reusable’ for both humans and machines (FAIR principles; Wilkinson et al. 2016). To this end, we are improving the documentation of protocols for the transformation, curation, and analysis of EOL data, and associated scripts and software are made available to ensure reproducibility. Proper acknowledgement of contributors and tracking of credit through derived data products promote both open data sharing and the use of aggregated resources. By exposing unique identifiers for data products, people, and institutions, data providers and aggregators can stimulate the development of automated solutions for the creation of contribution metrics. Since different aspects of provenance will be significant depending on the intended data use, better standardization of contributor roles (e.g., author, compiler, publisher, funder) is needed, as well as more detailed attribution guidance for data users. Global scale biodiversity data resources should resolve into a graph, linking taxa, specimens, occurrences, attributes, localities, and ecological interactions, as well as human agents, publications and institutions. Two key data categories for ensuring rich connectivity in the graph will be taxonomic and trait data. This graph can be supported by existing data hubs, if they share identifiers and/or create mappings between them, using standards and sharing practices developed by the biodiversity data community. Versioned archives of the combined graph could be published at intervals to appropriate open data repositories, and open source tools and training provided for researchers to access the combined graph of biodiversity knowledge from all sources. To achieve this, good communication among data hubs will be needed. We will need to share information about preferred vocabularies and identifier management practices, and collaborate on identifier mappings.
Synthesising trait observations and knowledge across the Tree of Life remains a grand challenge for biodiversity science. Despite the well-recognised importance of traits for addressing ecological and evolutionary questions, trait-based approaches still struggle with several basic data requirements to deliver openly accessible, reproducible, and transparent science. Here, we introduce the Open Traits Network (OTN) – a decentralised alliance of international researchers and institutions focused on collaborative integration and standardisation of the exponentially increasing availability of trait data across all organisms. The OTN embraces the use of Open Science principles in trait research, particularly open data, open source, and open methodology protocols and workflows, to accelerate the synthesis of trait data across the Tree of Life. Increased efforts at all levels – from individual scientists, research networks, scientific societies, funding agencies, to publishers – are necessary to fully exploit the opportunities offered by Open Science in trait research. Democratising access to data, tools and resources will facilitate rapid advances in the biological sciences and our ability to address pressing environmental and societal demands.
The practice of biologically inspired design requires access to general biological knowledge. In this paper, we describe AskEOL, a question-answering tool for accessing biological knowledge from Encyclopedia of Life (EOL), the world’s largest knowledgebase of biological taxa AskEOL operates in the context of a virtual research assistant, Vera, that provides an interactive environment for building conceptual models of ecological systems and runs experimental simulations on those models.
Citizen scientists have the potential to expand scientific research. The virtual research assistant called VERA empowers citizen scientists to engage in environmental science in two ways. First, it automatically generates simulations based on the conceptual models of ecological phenomena for repeated testing and feedback. Second, it leverages the Encyclopedia of Life biodiversity knowledgebase to support the process of model construction and revision.
Citizen inquiry combines the strengths of citizen science and inquiry-based learning, which has been applied mostly in informal learning settings. This chapter introduces the motivation theory and a rationale for exploring the particular factors in relation to students' motivation and contribution. It describes the procedure, materials, and results of the Tree and Bird Observation on Campus (TBOC) project. In citizen science, lack of feedback may discourage volunteers from continuing to contribute; feedback lets volunteers know that their efforts are appreciated, which keeps them from feeling peripheral to the scientific endeavour. The chapter then encompassed a field experiment based on a citizen science project followed by interviews to clarify experimental results. The study consisted of four phases: preparing students by introducing citizen science to them; engaging students in TBOC and completion of Situational Motivation Scale (SIMS); interviewing a subset of students; and analysing data associated with students' motivation, data quantity and quality, and the interview transcriptions.
The size of biodiversity data sets, and the size of people’s questions around them, are outgrowing the capabilities of desktop applications, single computers, and single developers. Numerous articles in the corporate sector (Delgado 2016) have been written on how much time professionals spend manipulating and formatting large data sets compared to the time they spend on the important work of doing analysis and modeling. To efficiently move large research questions forward, the biodiversity domain needs to transition towards shared infrastructure with the goal of providing a mise en place for researchers to do research with large data. The GUODA (Global Unified Open Data Access) collaboration was formed to explore tools and use cases for this type of collaborative work on entire biodiversity data sets. Three key parts of that exploration have been: the software and hardware infrastructure needed to be able to work with hundreds of millions of records and terabytes of data quickly, removing the impediment of data formatting and preparation, and workflows centered around GitHub for interacting with peers in an open and collaborative manner. We will describe our experiences building an infrastructure based on Apache Mesos, Apache Spark, HDFS, Jupyter Notebooks, Jenkins, and Github. We will also enumerate what resources are needed to do things like join millions of records, visualize patterns in whole data sets like iDigBio and the Biodiversity Heritage Library, build graph structures of billions of nodes, analyze terabytes of images, and use natural language processing to explore gigabytes of text. In addition to the hardware and software, we will describe the kinds of skills needed by staff to design, build, and use this sort of infrastructure and highlight some experiences we have with training students. Our infrastructure is one of many that are possible. We hope that by showing the amount and type of work we have done to the wider community, other organizations can understand what they would need to speed up their research programs by developing their own collaborative computation and development environments.
Biodiversity data are well-indexed by taxonomic names. While names reconciliation remains a challenge, there has been tremendous progress in recent years, and integration with available phylogenetic information can support sophisticated analyses for evolutionary questions. However, organisms are also linked to each other by relationships of ecology, geographic proximity, shared habitat, management categories, and other attributes, not yet recorded in a well-structured way. These data are best modeled as a graph, which makes these relationships explicit, and available for reasoning across - just like taxonomic relationships. This would support broad analyses of life on Earth not only from an evolutionary perspective but also across many other axes. This case study will describe how several categories of data are being modeled in the Encyclopedia of Life (EOL) v3 using ontology terms. It will focus on several areas where we anticipate sufficient taxonomic coverage to underlie significant search and analytical power: habitat, distribution, body size and metabolism, and provenance. Habitat and distribution terms are good examples of data terms in well structured hierarchies that could support powerful search. Habitat terms are available from and hierarchically organized in the Environment Ontology (ENVO). Geographic distribution knowledge can often be structured by geographic terms based on verbatim locality text when geocoordinates are not available. Geographic terms are available from several providers, notably Geonames (geonames.org), Marineregions.org and Wikidata. Both habitat and distribution terms can also be connected to simpler and less formalized but commonly used hierarchies like the World Wildlife Fund (WWF) Ecoregions. The hierarchy information made available for habitat and geography by the semantic structure of these ontologies supports searches like "wetland plants of South America," which requires the intersection of taxonomic, geographic, and habitat hierarchies. Body size and metabolism traits interact in a particular use case, illustrating the importance of precision of categorical data terms for informing calculations of quantitative traits. The use case EOL is currently working to support is the parameterization of food web interactions in ecological modelling software. Default or starting values are needed for the content of energy (or carbon) within an organism, and the rate of loss thereof through metabolism. This, plus assimilation efficiency, allows the modeling of carbon flow through the food web. Traits available for estimating carbon content and metabolic rate include various measures of body size, for which conversion factors and formulae are available. For phytoplankton, for instance, size may be reported as cell dimensions, cell volume, cell wet mass, cell dry mass, and/or carbon biomass. For an automated tool to derive parameters from these which are fit for use, the different types of data must all be findable, but the measurement types must be distinguished from one another so the correct conversions are performed for each - all in a machine readable way, so the process can be automated. The need for semantically structured data terms in this case is different, but just as critical to the success of the use case. Future work: Other important structured connections can be made through provenance metadata. These connect taxa and specimens to literature, authors, collectors, wildlife observers and other agents. The Social Media of biodiversity data, rendered explicit, could increase connectivity and communication in the global community - particularly benefitting young researchers in isolated regions without the benefit of professional travel or literature subscriptions. To accomplish this, we must leverage human identifiers such as those made available by Open Researcher And Contributor ID (ORCID) and Wikidata.
Spencer Rugaber合作论文数College of Computing;Georgia Institute of Technology8