Ecologists often use a hidden Markov model to decode a latent process, such as a sequence of an animal's behaviours, from an observed biologging time series. Modern technological devices such as video recorders and drones now allow researchers to directly observe an animal's behaviour. Using these observations as labels of the latent process can improve a hidden Markov model's accuracy when decoding the latent process. However, many wild animals are observed infrequently. Including such rare labels often has a negligible influence on parameter estimates, which in turn does not meaningfully improve the accuracy of the decoded latent process. We introduce a weighted likelihood approach that increases the relative influence of labelled observations. We use this approach to develop hidden Markov models to decode the foraging behaviour of killer whales (Orcinus orca) off the coast of British Columbia, Canada. Using cross-validated evaluation metrics and a detailed simulation study, we show that our weighted likelihood approach produces more accurate and understandable decoded latent processes compared to existing hidden Markov models and single-frame machine learning methods. Thus, our method effectively leverages sparse labels to enhance researchers' ability to accurately decode hidden processes across various fields.
Hidden Markov models (HMMs) are popular models to identify a finite number of latent states from sequential data. However, fitting them to large data sets can be computationally demanding because most likelihood maximization techniques require iterating through the entire underlying data set for every parameter update. We propose a novel optimization algorithm that updates the parameters of an HMM without iterating through the entire data set. Namely, we combine a partial E step with variance-reduced stochastic optimization within the M step. We prove the algorithm converges under certain regularity conditions. We test our algorithm empirically using a simulation study as well as a case study of kinematic data collected using suction-cup attached biologgers from eight northern resident killer whales (Orcinus orca) off the western coast of Canada. In both, our algorithm converges in fewer epochs and to regions of higher likelihood compared to standard numerical optimization techniques. Our algorithm allows practitioners to fit complicated HMMs to large time-series data sets more efficiently than existing baselines.
Measuring breathing rates is a means by which oxygen intake and metabolic rates can be estimated to determine food requirements and energy expenditure of killer whales (Orcinus orca) and other cetaceans. This relatively simple measure also allows the energetic consequences of environmental stressors to cetaceans to be understood but requires knowing respiration rates while they are engaged in different behaviours such as resting, travelling and foraging. We calculated respiration rates for different behavioural states of southern and northern resident killer whales using video from UAV drones and concurrent biologging data from animal-borne tags. Behavioural states of dive tracks were predicted using hierarchical hidden Markov models (HHMM) parameterized with time-depth data and with labeled tracks of drone-identified behavioural states (from drone footage that overlapped with the time-depth data). Dive tracks were sequences of dives and surface intervals lasting ≥ 10 minutes cumulative duration. We calculated respiration rates and estimated oxygen consumption rates for the predicted behavioural states of the tracks. We found that juvenile killer whales breathed at a higher rate when travelling (1.6 breaths min-1) compared to resting (1.2) and foraging (1.5)-and that adult males breathed at a higher rate when travelling (1.8) compared to both foraging (1.7) and resting (1.3). The juveniles in our study were estimated to consume 2.5-18.3 L O2 min-1 compared with 14.3-59.8 L O2 min-1 for adult males across all behaviours based on estimates of mass-specific tidal volume and oxygen extraction. Our findings confirm that killer whales take single breaths between dives and indicate that energy expenditure derived from respirations requires using sex, age, and behavioural-specific respiration rates. These findings can be applied to bioenergetics models on a behavioural-specific basis, and contribute towards obtaining better predictions of dive behaviours, energy expenditure and the food requirements of apex predators.
Educational resources, such as web apps and self-directed tutorials, have become popular tools for teaching and active learning. Ideally, students - the intended users of these resources - should be involved in the resource development stage. However, in practice students often only interact with fully developed resources, when it might be too late to incorporate changes. Previous work has addressed this by involving students in the development of new resources via in-person focus groups and interviews. In these, the resource developers observe students interacting with the resource. This allows developers to incorporate their observations and students' direct feedback into further development of the resource. However, as a result of the COVID-19 pandemic, carrying out in-person focus groups became infeasible due to social distancing restrictions. Instead, online meetings and classes became ubiquitous. In this work, we describe a fully-online methodology to evaluate new resources in development. Specifically, our methodology consists of carrying out student focus groups via online video conferencing software. We assessed two educational resources for introductory statistics using our methodology and found that the online setting allowed us to obtain rich, detailed information from the students. We also found online focus groups to be more efficient: students and researchers did not need to travel and scheduling was not restricted by the availability of physical space. Our findings suggest that online focus groups are an attractive alternative to in-person focus groups for student assessment of resources in development, even now that pandemic restrictions are being eased.
Data sets comprised of sequences of curves sampled at high frequencies in time are increasingly common in practice, but they can exhibit complicated dependence structures that cannot be modelled using common methods of Functional Data Analysis (FDA). We detail a hierarchical approach which treats the curves as observations from a hidden Markov model (HMM). The distribution of each curve is then defined by another fine-scale model which may involve auto-regression and require data transformations using moving-window summary statistics or Fourier analysis. This approach is broadly applicable to sequences of curves exhibiting intricate dependence structures. As a case study, we use this framework to model the fine-scale kinematic movement of a northern resident killer whale (Orcinus orca) off the coast of British Columbia, Canada. Through simulations, we show that our model produces more interpretable state estimation and more accurate parameter estimates compared to existing methods.
Most statistics undergraduate curricula provide only a brief introduction to Bayesian inference. Furthermore, there is evidence that learners can fail to appreciate core concepts of the Bayesian framework because of the focus on the mathematical formalism that is common in traditional instruction of Bayesian inference. Guided exercises and interactive simulations are promising alternatives for introducing Bayesian inference. However, these two types of resources have thus far been developed separately, which may make them less effective. In this work, we bridge this gap by developing interactive, structured resources in the form of web apps and accompanying activity sheets with guiding exercises. We also incorporate student feedback via think-aloud sessions. This allows us to pinpoint sources of confusion in students, e.g., problems constructing their priors.
Building on recent work in statistical science, the paper presents a theory for modelling natural phenomena that unifies physical and statistical paradigms based on the underlying principle that a model must be nondimensionalizable. After all, such phenomena cannot depend on how the experimenter chooses to assess them. Yet the model itself must be comprised of quantities that can be determined theoretically or empirically. Hence, the underlying principle requires that the model represents these natural processes correctly no matter what scales and units of measurement are selected. This goal was realized for physical modelling through the celebrated theories of Buckingham and Bridgman and for statistical modellers through the invariance principle of Hunt and Stein. Building on recent research in statistical science, the paper shows how the latter can embrace and extend the former. The invariance principle is extended to encompass the Bayesian paradigm, thereby enabling an assessment of model uncertainty. The paper covers topics not ordinarily seen in statistical science regarding dimensions, scales, and units of quantities in statistical modelling. It shows the special difficulties that can arise when models involve transcendental functions, such as the logarithm which is used e.g. in likelihood analysis and is a singularity in the family of Box-Cox family of transformations. Further, it demonstrates the importance of the scale of measurement, in particular how differently modellers must handle ratio- and interval-scales
Functional data often exhibit both amplitude and phase variation around a common base shape, with phase variation represented by a so called warping function. The process removing phase variation by curve alignment and inference of the warping functions is referred to as curve registration. When functional data are observed with substantial noise, model-based methods can be employed for simultaneous smoothing and curve registration. However, the nonlinearity of the model often renders the inference computationally challenging. In this paper, we propose an alternative method for model-based curve registration which is computationally more stable and efficient than existing approaches in the literature. We apply our method to the analysis of elephant seal dive profiles and show that more intuitive groupings can be obtained by clustering on phase variations via the predicted warping functions.
We develop and apply an approach for analyzing multi-curve data where each curve is driven by a latent state process. The state at any particular point determines a smooth function, forcing the individual curve to switch from one function to another. Thus each curve follows what we call a switching nonparametric regression model. We develop an EM algorithm to estimate the model parameters. We also obtain standard errors for the parameter estimates of the state process. We consider several types of state processes: independent and identically distributed, independent but depending on a covariate and Markov. Simulation studies show the frequentist properties of our estimates. We apply our methods to a data set of a building's power usage.
We consider estimation of the drift function of a stationary diffusion process when we observe high-frequency data with microstructure noise over a long time interval. We propose to estimate the drift function at a point by a Nadaraya–Watson estimator that uses observations that have been pre-averaged to reduce the noise. We give conditions under which our estimator is consistent and asympotically normal. Its rate and asymptotic bias and variance are the same as those without microstructure noise. To use our method in data analysis, we propose a data-based cross-validation method to determine the bandwidth in the Nadaraya–Watson estimator. Via simulation, we study several methods of bandwidth choices, and compare our estimator to several existing estimators. In terms of mean squared error, our new estimator outperforms existing estimators.
Understanding the energy consumption patterns of different types of consumers is essential in any planning of energy distribution. However, obtaining consumption information for single individuals is often either not possible or too expensive. Therefore, we consider data from aggregations of energy use, that is, from sums of individuals' energy use, where each individual falls into one of C consumer classes. Unfortunately, the exact number of individuals of each class may be unknown: consumers do not always report the appropriate class, due to various factors including differential energy rates for different consumer classes. We develop a methodology to estimate the expected energy use of each class as a function of time and the true number of consumers in each class. We also provide some measure of uncertainty of the resulting estimates. To accomplish this, we assume that the expected consumption is a function of time that can be well approximated by a linear combination of B-splines. Individual consumer perturbations from this baseline are modeled as B-splines with random coefficients. We treat the reported numbers of consumers in each category as random variables with distribution depending on the true number of consumers in each class and on the probabilities of a consumer in one class reporting as another class. We obtain maximum likelihood estimates of all parameters via a maximization algorithm. We introduce a special numerical trick for calculating the maximum likelihood estimates of the true number of consumers in each class. We apply our method to a data set and study our method via simulation.
Vision is a sensory modality of fundamental importance for many animals, aiding in foraging, detection of predators and mate choice. Adaptation to local ambient light conditions is thought to be commonplace, and a match between spectral sensitivity and light spectrum is predicted. We use opsin gene expression to test for local adaptation and matching of spectral sensitivity in multiple independent lake populations of threespine stickleback populations derived since the last ice age from an ancestral marine form. We show that sensitivity across the visual spectrum is shifted repeatedly towards longer wavelengths in freshwater compared with the ancestral marine form. Laboratory rearing suggests that this shift is largely genetically based. Using a new metric, we found that the magnitude of shift in spectral sensitivity in each population corresponds strongly to the transition in the availability of different wavelengths of light between the marine and lake environments. We also found evidence of local adaptation by sympatric benthic and limnetic ecotypes to different light environments within lakes. Our findings indicate rapid parallel evolution of the visual system to altered light conditions. The changes have not, however, yielded a close matching of spectrum-wide sensitivity to wavelength availability, for reasons we discuss.
Understanding the patterns of genetic variation and constraint for continuous reaction norms, growth trajectories, and other function-valued traits is challenging. We describe and illustrate a recent analytical method, simple basis analysis (SBA), that uses the genetic variance-covariance (G) matrix to identify "simple" directions of genetic variation and genetic constraints that have straightforward biological interpretations. We discuss the parallels between the eigenvectors (principal components) identified by principal components analysis (PCA) and the simple basis (SB) vectors identified by SBA. We apply these methods to estimated G matrices obtained from 10 studies of thermal performance curves and growth curves. Our results suggest that variation in overall size across all ages represented most of the genetic variance in growth curves. In contrast, variation in overall performance across all temperatures represented less than one-third of the genetic variance in thermal performance curves in all cases, and genetic trade-offs between performance at higher versus lower temperatures were often important. The analyses also identify potential genetic constraints on patterns of early and later growth in growth curves. We suggest that SBA can be a useful complement or alternative to PCA for identifying biologically interpretable directions of genetic variation and constraint in function-valued traits.
We develop and apply a new approach for analyzing a building's business day power usage. We treat each business day as a replicate and model power usage as arising from two smooth functions, one function giving power usage when the cooling system is off, the other function giving power usage when the cooling system is on. The condition chiller on/chiller off at any particular time cannot be observed directly, thus forming a latent process. In general, our method can be applied to multi-curve data where each curve is driven by a latent state process. The state at any particular point determines a smooth function. Thus each curve follows what we call a switching nonparametric regression model. We develop an EM algorithm to estimate the parameters of the latent process and the function corresponding to each state. We also obtain standard errors for the parameter estimates of the state process. Simulations studies show the frequentist properties of our estimates.
We propose a methodology to analyze data arising from a curve that, over its domain, switches among J states. We consider a sequence of response variables, where each response y depends on a covariate x according to an unobserved state z. The states form a stochastic process and their possible values are j=1,...,J. If z equals j the expected response of y is one of J unknown smooth functions evaluated at x. We call this model a switching nonparametric regression model. We develop an EM algorithm to estimate the parameters of the latent state process and the functions corresponding to the J states. We also obtain standard errors for the parameter estimates of the state process. We conduct simulation studies to analyze the frequentist properties of our estimates. We also apply the proposed methodology to the well-known motorcycle data set treating the data as coming from more than one simulated accident run with unobserved run labels.
Principal Components Analysis (PCA) is a common way to study the sources of variation in a high-dimensional data set. Typically, the leading principal components are used to understand the variation in the data or to reduce the dimension of the data for subsequent analysis. The remaining principal components are ignored since they explain little of the variation in the data. However, the space spanned by the low variation principal components may contain interesting structure, structure that PCA cannot find. Prinsimp is an R package that looks for interesting structure of low variability. "Interesting" is defined in terms of a simplicity measure. Looking for interpretable structure in a low variability space has particular importance in evolutionary biology, where such structure can signify the existence of a genetic constraint.
We propose a new version of functional data model for analyzing familial related individuals, where the within-subject correlation depends smoothly on a covariate such as age and the between-subject correlation follows family-wise genetic association. Our motivating example concerns measurements of weight as a function of age in sibling cows from independent families. Observations are sparsely sampled from trajectories of a phenotype contaminated with measurement error, where the phenotypic trajectory consists of a genetic component and an environmental component. By combining information across individuals, the genetic and environmental covariances are estimated via smoothing techniques. We study the genetic and environmental effects using principal component analysis, taking into account the genetic correlation to enhance the subject-level signal extraction. We show via the real data and simulations that incorporating the correlation structure improves predictions of individual phenotypic trajectories.
Teleconnections are quasi-periodic changes in atmospheric circulation that oscillate over long periods of time and impact climate over large regions. These patterns are often linked to long-term variations in climate and extreme weather events and may explain regional differences in climate vulnerability. We apply methods of functional data analysis to examine regional impacts of teleconnections on climate in British Columbia, Canada, between 1951 and 2000. We focus on monthly mean temperature as an overall determinant of crop growth and apply functional principal components analysis (FPCA) to study variations in the impacts of four major teleconnection indices affecting the Northern Hemisphere (the Southern Oscillation Index, the Pacific North American (PNA), Pacific Decadal Oscillation, and the North American Oscillation indices). Two challenges we consider are that the impacts of teleconnections cannot be observed directly and that fine scale data required to study regional variations may come from different sources with highly varied records. We first fit thin-plate regression splines to the raw data to construct complete series of pseudo-data at fixed grid points. Regression models incorporating Bayesian P-splines were then fit to the pseudo-data to estimate the impacts of the four teleconnections over time. Finally, FPCA was then applied to study regional variations in these effects. Our analysis identified strong variations in mean temperature associated with the PNA. The resulting spatial patterns also reveal areas of increased/decreased temperature variability that may have higher climate risk or be suitable for expansion of agricultural activity.
Linear mixed effects methods for the analysis of longitudinal data provide a convenient framework for modelling within-individual correlation across time. Using spline functions allows for flexible modelling of the response as a smooth function of time. A computational connection between linear mixed effects modelling and spline smoothing has resulted in use of spline functions in longitudinal data analysis and the use of mixed effects software in smoothing analyses. However, care must be taken in exploiting this connection, as resulting estimates of the underlying population mean might not track the data well and associated standard errors might not reflect the true variability in the data. We discuss these shortcomings and suggest some easy-to-compute methods to eliminate them.