The importance of linear habitat elements connecting core habitat patches for biodiversity conservation is still poorly understood. We surveyed reed strips along drainage ditches and reed marshes in an agricultural landscape to assess how both the density of linear habitat elements and the area of core habitat affect diversity and community composition of spiders, ground beetles, and long‐legged flies. For each taxonomic group, species composition of both ‘all’ and ‘typical wetland’ species, but not species richness was different between ditches and marshes. Overall local species richness and richness of species of conservation interest were affected at a landscape scale both by the density of ditches and by the area of core wetland. Strength and direction of these effects differed among groups. An increase in the density of reed ditches positively affected the total species richness of spiders and ground beetles and the species richness of typical wetland ground beetles, but not for long‐legged flies and typical wetland spiders. The positive effects were explained by improved network functionality, rather than by increase in available habitat area at landscape level. The number of red list spiders and long‐legged flies increased only with increasing core wetland area, while no significant effects were found for the number of red list ground beetles. Our study revealed that preserving or increasing the density of habitat corridors (more reed ditches) can be beneficial for the species richness of particular predatory arthropods, including species of conservation concern (especially ground beetles). Other groups react indifferently or are only positively impacted by an increase of core wetland area.
Although most models for incomplete longitudinal data are formulated within the selection model framework, pattern-mixture models have gained considerable interest in recent years [R.J.A. Little, Pattern-mixture models for multivariate incomplete data, J. Am. Stat. Assoc. 88 (1993), pp. 125–134; R.J.A. Lrittle, A class of pattern-mixture models for normal incomplete data, Biometrika 81 (1994), pp. 471–483], since it is often argued that selection models, although many are identifiable, should be approached with caution, especially in the context of MNAR models [R.J. Glynn, N.M. Laird, and D.B. Rubin, Selection modeling versus mixture modeling with nonignorable nonresponse, in Drawing Inferences from Self-selected Samples, H. Wainer, ed., Springer-Verlag, New York, 1986, pp. 115–142]. In this paper, the focus is on several strategies to fit pattern-mixture models for non-monotone categorical outcomes. The issue of under-identification in pattern-mixture models is addressed through identifying restrictions. Attention will be given to the derivation of the marginal covariate effect in pattern-mixture models for non-monotone categorical data, which is less straightforward than in the case of linear models for continuous data. The techniques developed will be used to analyse data from a clinical study in psychiatry.
Summary 1. Worldwide, the floristic composition of temperate forests bears the imprint of past land use for decades to centuries as forests regrow on agricultural land. Many species, however, display significant interregional variation in their ability to (re)colonize post‐agricultural forests. This variation in colonization across regions and the underlying factors remain largely unexplored. 2. We compiled data on 90 species and 812 species × study combinations from 18 studies across Europe that determined species’ distribution patterns in ancient (i.e. continuously forested since the first available land use maps) and post‐agricultural forests. The recovery rate (RR) of species in each landscape was quantified as the log‐response ratio of the percentage occurrence in post‐agricultural over ancient forest and related to the species‐specific life‐history traits and local (soil characteristics and light availability) and regional factors (landscape properties as habitat availability, time available for colonization, and climate). 3. For the herb species, we demonstrate a strong (interactive) effect of species’ life‐history traits and forest habitat availability on the RR of post‐agricultural forest. In graminoids, however, none of the investigated variables were significantly related to the RR. 4. The better colonizing species that mainly belonged to the short‐lived herbs group showed the largest interregional variability. Their recovery significantly increased with the amount of forest habitat within the landscape, whereas, surprisingly, the time available for colonization, climate, soil characteristics and light availability had no effect. 5. Synthesis. By analysing 18 independent studies across Europe, we clearly showed for the first time on a continental scale that the recovery of short‐lived forest herbs increased with the forest habitat availability in the landscape. Small perennial forest herbs, however, were generally unsuccessful in colonizing post‐agricultural forest – even in relatively densely forested landscapes. Hence, our results stress the need to avoid ancient forest clearance to preserve the typical woodland flora.
SummaryMuch research has been devoted to modelling strategies for longitudinal data with missingness, recently especially within the missingness not at random context. In this paper, the relatively unexplored but practically highly relevant domain of non-monotone missingness with multivariate ordinal responses is broached. For this, a dedicated version of the multivariate Dale model is formulated. Furthermore, we also assess the sensitivity of these models to their assumptions, by using the technique of global influence.
Many models to analyze incomplete data that allow the missingness to be non-random have been developed. Since such models necessarily rely on unverifiable assumptions, considerable research nowadays is devoted to assess the sensitivity of resulting inferences. A popular sensitivity route, next to local influence (Cook in J Roy Stat Soc Ser B 2:133–169, 1986; Jansen et al. in Biometrics 59:410–419, 2003) and so-called intervals of ignorance (Molenberghs et al. in Appl Stat 50:15–29, 2001), is based on contrasting more conventional selection models with members from the pattern-mixture model family. In the first family, the outcome of interest is modeled directly, while in the second family the natural parameter describes the measurement process, conditional on the missingness pattern. This implies that a direct comparison ought not to be done in terms of parameter estimates, but rather should pass by marginalizing the pattern-mixture model over the patterns. While this is relatively straightforward for linear models, the picture is less clear for the nevertheless important setting of categorical outcomes, since models ordinarily exhibit a certain amount of non-linearity. Following ideas laid out in Jansen and Molenberghs (Pattern-mixture models for categorical outcomes with non-monotone missingness. Submitted for publication, 2007), we offer ways to marginalize pattern-mixture-model-based parameter estimates, and supplement these with asymptotic variance formulas. The modeling context is provided by the multivariate Dale model. The performance of the method and its usefulness for sensitivity analysis is scrutinized using simulations.
The process of monoisotopic mass determination, i.e., nomination of the correct peak of an isotopically resolved group of peptide peaks as a monoisotopic peak, requires prior information about the isotopic distribution of the peptide. This points immediately to the difficulty of monoisotopic mass determination, whereas a single mass spectrum does not contain information about the atomic composition of a peptide and therefore the isotopic distribution of the peptide remains unknown. To solve this problem a technique is required, which is able to estimate the isotopic distribution given the information of a single mass spectrum. Senko et al. calculated the average isotopic distribution for any mass peptide via the multinomial expansion (Yergey 1983) [1], using a scaled version of the average amino acid Averagine (Senko et al. 1995) [2]. Another method, introduced by Breen et al., approximates the result of the multinomial expansion by a Poisson model (Breen et al. 2000) [3]. Although both methods perform well, they have their specific limitations. In this manuscript, we propose an alternative method for the prediction of the isotopic distribution based on a model for consecutive ratios of peaks from the isotopic distribution, similar in spirit to the approach introduced by Gay et al. (1999) [5]. The presented method is computationally simple and accurate in predicting the expected isotopic distribution. Further, we extend our method to estimate the isotopic distribution of sulphur-containing peptides. This is important because the naturally occurring isotopes of sulphur have an impact on the isotopic distribution of a peptide.
We present an approach to construct a classification rule based on the mass spectrometry data provided by the organizers of the "Classification Competition on Clinical Mass Spectrometry Proteomic Diagnosis Data." Before constructing a classification rule, we attempted to pre-process the data and to select features of the spectra that were likely due to true biological signals (i.e., peptides/proteins). As a result, we selected a set of 92 features. To construct the classification rule, we considered eight methods for selecting a subset of the features, combined with seven classification methods. The performance of the resulting 56 combinations was evaluated by using a cross-validation procedure with 1000 re-sampled data sets. The best result, as indicated by the lowest overall misclassification rate, was obtained by using the whole set of 92 features as the input for a support-vector machine (SVM) with a linear kernel. This method was therefore used to construct the classification rule. For the training data set, the total error rate for the classification rule, as estimated by using leave-one-out cross-validation, was equal to 0.16, with the sensitivity and specificity equal to 0.87 and 0.82, respectively.
In this paper, we propose a system for finding partial positive and negative coregulated gene clusters in microarray data. Genes are clustered together if they show the same pattern of changing tendencies in a user definied number of condition pairs. It is assumed that genes which show similar expression patterns under a number of conditions are under the control of the same transcription factor and are related to a similar function in the cell. Taking positive and negative coregulation of genes into account, we find two types of information:(1) clusters of genes showing the same changing tendency and (2) relationships between two such clusters whose respective members show opposite changing tendency. Because genes may be coregulated by different transcription factors under different environmental conditions, our algorithm allows the same gene to fall into different clusters. Overlapping gene clusters are allowed because coregulation normally takes place in only a fraction of the investigated condition pairs, and because the gene expression data is noisy so that the approach should be tolerant to errors. In a first step, the gene expression matrix is transformed to a binned matrix of changing tendencies between all condition pairs. For the binning of the gene expression levels, a statistical technique is used, for which no arbitrary threshold needs to be chosen, which automatically corrects for multiple testing, and which is able to handle replicates for the different conditions, immediately accounting for the random variability of gene expression data. To present the results of a clustering a new structure called coregulation graph is proposed.
The authors analyze data on marital satisfaction, obtained from couples at two distinct moments in time (1990, 1995). The data are of a bivariate longitudinal type. Moreover, some couples provide incomplete records only, usually because the 1995 follow-up interview has not taken place. The authors propose a hierarchical modeling strategy that takes all these features into account and is more generally valid than a classical complete case or single imputation-based strategy.
Models for incomplete longitudinal data under missingness not at random have gained some popularity. At the same time, cautionary remarks have been issued regarding their sensitivity to often unverifiable modeling assumptions. Consequently, there is evidence for a shift towards using ignorable methodology, supplemented with sensitivity analyses to explore the impact of potential deviations of this assumption in the direction of missingness at random. One such tool is local influence. It is shown that local influence tends to pick up a lot of different anomalies in the data at hand, not just deviations in the MNAR mechanism. This particular behavior is described and insight offered in terms of the non-standard behavior of the likelihood ratio test statistic for MAR missingness versus MNAR missingness within a model of the Diggle and Kenward type.
Commonly used methods to analyze incomplete longitudinal clinical trial data include complete case analysis (CC) and last observation carried forward (LOCF). However, such methods rest on strong assumptions, including missing completely at random (MCAR) for CC and unchanging profile after dropout for LOCF. Such assumptions are too strong to generally hold. Over the last decades, a number of full longitudinal data analysis methods have become available, such as the linear mixed model for Gaussian outcomes, that are valid under the much weaker missing at random (MAR) assumption. Such a method is useful, even if the scientific question is in terms of a single time point, for example, the last planned measurement occasion, and it is generally consistent with the intention-to-treat principle. The validity of such a method rests on the use of maximum likelihood, under which the missing data mechanism is ignorable as soon as it is MAR. In this paper, we will focus on non-Gaussian outcomes, such as binary, categorical or count data. This setting is less straightforward since there is no unambiguous counterpart to the linear mixed model. We first provide an overview of the various modeling frameworks for non-Gaussian longitudinal data, and subsequently focus on generalized linear mixed-effects models, on the one hand, of which the parameters can be estimated using full likelihood, and on generalized estimating equations, on the other hand, which is a nonlikelihood method and hence requires a modification to be valid under MAR. We briefly comment on the position of models that assume missingness not at random and argue they are most useful to perform sensitivity analysis. Our developments are underscored using data from two studies. While the case studies feature binary outcomes, the methodology applies equally well to other discrete-data settings, hence the qualifier "discrete" in the title.