We show that data censored by linear constraints can be fit within the generalized linear model (GLM) framework, recovering the expected latent counts together with their uncertainty in a single Fisher-scoring procedure. Our motivating application is origin-destination (OD) estimation in transportation studies, where the latent data are the OD trip demands and constraints are either origin and destination marginal counts per zone, as in OD matrix estimation, or observed link counts, as in network tomography. Casting the problem as a linearly constrained GLM lets us treat, within one model, three features that arise routinely in practice: unmatched constraints, trip predictors such as costs, and prior information such as diagonal dominance and seed counts. We demonstrate the methodology in synthetic and small scale studies and in a larger case study based on simulated data from the BO4Mob study.
We study a general factor analysis framework where the n-by-p data matrix is assumed to follow a general exponential family distribution entry-wise. While this model framework has been proposed before, we here further relax its distributional assumption by using a quasi-likelihood setup. By parameterizing the mean-variance relationship on data entries, we additionally introduce a dispersion parameter and entry-wise weights to model large variations and missing values. The resulting model is thus not only robust to distribution misspecification but also more flexible and able to capture mean-dependent covariance structures of the data matrix. Our main focus is on efficient computational approaches to perform the factor analysis. Previous modeling frameworks rely on simulated maximum likelihood (SML) to find the factorization solution, but this method was shown to lead to asymptotic bias when the simulated sample size grows slower than the square root of the sample size n, eliminating its practical application for data matrices with large n. Borrowing from expectation-maximization (EM) and stochastic gradient descent (SGD), we investigate three estimation procedures based on iterative factorization updates. Our proposed solution does not show asymptotic biases, and scales even better for large matrix factorizations with error O(1/p). To support our findings, we conduct simulation experiments and discuss its application in four case studies.
Principal component analyses are often applied to spatial data towards inference on latent modes of spatial variation. These analyses are widespread across domains including spatial transcriptomics and environmental sciences, where the modes of spatial variation are represented by corresponding factors of gene expression or remotely sensed time series measurements. Many methods have been proposed for incorporating spatial information into a probabilistic PCA framework; however, there are three main drawbacks to currently available approaches. First, the loadings matrices are not orthogonal, and subsequent orthogonalization of those loadings corrupts the original prior spatial information. Furthermore, currently proposed methods assume stationarity in their spatial prior. Finally, current methods typically do not achieve linear-time computational complexity with respect to the number of spatial locations. To resolve these problems, we first parameterize the model directly with orthogonal loadings. For the prior distribution, we derive the sampling distribution of an SVD transformation with k unique and m-k repeated singular values. We then show under this model that the maximum a posteriori estimator for the orthogonal loadings is the eigendecomposition of S + 1/nΣ, where S is the empirical covariance matrix and Σ is the prior spatial covariance. We develop a minorization-maximization-within-EM algorithm that is linear in computational complexity with respect to the number of spatial locations. We further extend our MM-EM algorithm to handle held-out locations and develop a validation strategy for optimizing the nonstationary prior covariance. Our methodology is used to infer the spatial distribution of direction-specific length scales in a human brain spatial transcriptomics case study, as well as a continental-scale phenology case study in sub-Saharan Africa.
Ixodes fuscipes is a tick species found in the Southern Cone of America and the only member of the Ixodes ricinus complex present in Uruguay. Members of this complex are particularly recognized as vectors of diseases affecting human health, such as babesiosis, caused by parasites of the genus Babesia (Apicomplexa: Piroplasmida). However, even though potential hosts of I. fuscipes in Uruguay (rodents, birds, and artiodactyls) are known carriers of Babesia species, the potential role of I. fuscipes as a vector of piroplasmids has not been studied. In this study, questing I. fuscipes ticks were collected from five locations in Uruguay, and the presence of piroplasmid DNA was assessed using polymerase chain reaction (PCR) to amplify fragments of the small subunit ribosomal RNA (18S rRNA) and cytochrome c oxidase subunit 1 (COI) genes. A total of 953 ticks (larvae, nymphs, and adults) were collected; 14 samples (two larval pools and 12 nymphs) tested positive. Genetic analyses using 18S rDNA and COI sequences revealed the presence of undescribed Babesia lineages, belonging to the Babesia odocoilei clade and others to the Babesia microti sensu stricto clade. This work represents the first association of Babesia spp. with I. fuscipes and highlights the importance of this type of study to detect and mitigate the emergence of diseases associated with these arthropods.
Ensuring food security in a framework of environmental sustainability is the greatest challenge of the 21st century. The rapid population growth together with changing consumption patterns associated with new lifestyles mean that total demand for food is increasing at a faster pace than that of the production capacity. To overcome this challenge we need to transform agriculture into a more efficient and less polluting activity. This may be achieved through harnessing the trillions of organisms inhabiting the soil. In this work we show how the use of microbial consortia can contribute to increased productivity and efficient use of nutrients in a maize field. The combined use of arbuscular mycorrhizal fungi and bacteria that promote plant growth, when applied in a field experiment were able to compensate for the reduction in fertilizer by 33%. The results show that the application of the microbial consortium increased nitrogen fixing and phosphorus solubilising bacteria in the soil, which may explain the increased uptake of these nutrients by the plants.
Aphids are important herbivores in natural and managed environments. We studied the response of aphids and their associated microbiota to the presence of the fungal endophyte Epichloë sp. LpTG‐3 strain AR37, and the AR37‐derived alkaloids in plants. We hypothesized that AR37 and/or AR37‐derived alkaloids would reduce the aphid performance, and that this reduction would be associated with endophyte‐mediated changes in the abundance, composition, and diversity of beneficial bacterial endosymbionts of aphids (e.g., Buchnera ). Plants of Lolium perenne associated with AR37 variants able (wild type and ∆ idtA ) and unable (∆ idtM ) to produce indole diterpene alkaloids were challenged with Rhopalosiphum padi aphids. We measured aphid population size, plant biomass, and the abundance, composition and diversity of the aphid's bacterial microbiota. The presence of AR37 increased the resistance of plants against R. padi aphids via the production of indole diterpene alkaloids, and this effect was independent of the plant biomass. The endophyte‐mediated reduction in aphid performance was not associated with changes in the abundance, composition and diversity of the insect's bacterial microbiota. However, we cannot rule out that the reduction in aphid performance could be associated with a putative endophyte effect on the bacterial provision of benefits to aphids. Our study highlighted the protective role of endophyte‐derived indole diterpene alkaloids against aphids. Further investigations will be needed to determine if there is a link between the endophyte‐mediated aphid resistance and the integrity of the insect's bacterial microbiota.
PURPOSE:The effects of in-home environmental exposures (IHEEs) on asthma are challenging to examine in populations because information on asthma triggers is usually absent. We leveraged data from electronic health records (EHRs) to investigate the associations of residential cockroach and rodent exposures with lung function among children with asthma. METHODS:We merged clinical pulmonary function test data from EHRs for children with asthma from a large safety net hospital in the Northeast United States with publicly available geospatial data matched to patient addresses. Predicted presence of key IHEE asthma triggers, cockroaches and rodents, were included as main exposures and housing parcel features and census tract characteristics were included as potential confounders in a sensitivity analysis. We fit latent Bayesian hierarchical models of percent predicted forced expiratory volume in one second (FEV1%). RESULTS:The study population of 1070 children had a mean age of 10.2 years and 75 % identified as Black, many living in historically segregated neighborhoods. In models adjusted for individual characteristics, we observed 2.26 (95 % credible interval, 95 %CrI: - 3.72, - 0.79) and 2.58 (95 %CrI: - 4.54, - 0.66) percentage points (pp) lower FEV1% from a one-unit increase in the log-odds of the probability of cockroach and rodent presence, respectively. The association with lung function increased in magnitude for cockroach exposure but attenuated for rodent exposure in sensitivity analyses. CONCLUSIONS:IHEEs were associated with worse lung function among children with asthma in a safety net population. The observed associations underscore how injustices in housing and neighborhood characteristics contribute to asthma morbidity.
Analyses of occurrences of residential burglary in urban areas have shown that crime rates are not spatially homogeneous: rates vary across the network of city streets, resulting in some areas being far more susceptible to crime than others. The explanation for why a certain segment of the city experiences high crime may be different than why a neighboring area experiences high crime. Motivated by the importance of understanding spatial patterns such as these, we consider a statistical model of burglary defined on the street network of Boston, Massachusetts. Leveraging ideas from functional data analysis, our proposed solution consists of a generalized linear model with vertex-indexed covariates, allowing for an interpretation of the covariate effects at the street level. We employ a regularization procedure cast as a prior distribution on the regression coefficients under a Bayesian setup so that the predicted responses vary smoothly according to the connectivity of the city. We introduce a novel variable selection procedure, examine computationally efficient methods for sampling from the posterior distribution of the model parameters, and demonstrate the flexibility of our proposed modeling structure. The resulting model and interpretations provide insight into the spatial network patterns and dynamics of residential burglary in Boston.
Exsheathment is crucial in the transition from free-living to parasitic phase for most strongyle nematode species. A greater understanding of this process could help in developing new parasitic control methods. This study aimed to identify commonalities in response to exsheathment triggers (heat acclimation, CO2 and pH) in a wide range of species (Haemonchus contortus, Trichostrongylus spp., Cooperia spp., Oesophagostomum spp., Chabertia ovina, and members of the subfamily Ostertagiinae) from sheep, cattle and farmed deer. The initial expectation of similarity in pH requirements amongst species residing within the same organ was not supported, with unexpected pH preferences for exsheathment of Trichostrongylus axei, Trichostrongylus vitrinus, Trichostrongylus colubriformis and Cooperia oncophora. We also found differences between species in their response to temperature acclimation, with higher exsheathment in response to heat shock observed for H. contortus, Ostertagia ostertagi, T. axei, T. vitrinus and Oesophagostomum sikae. Furthermore, some species showed poor exsheathment under all experimental conditions, such as Cooperia curticei and the large intestinal nematodes C. ovina and Oesophagostomum venulosum. Interestingly, there were some significant differences in response depending on the host from which the parasites were derived. The host species significantly impacted on the exsheathment response for H. contortus, Teladorsagia circumcincta, T. vitrinus and T. colubriformis. Overall, the data showed variability between nematode species in their response to these in vitro exsheathment triggers, highlighting the complexity of finding a common set of conditions for all species in order to develop a control method based on triggering the exsheathment process prematurely.
Model evaluation is of crucial importance in modern statistics application. The construction of ROC and calculation of AUC have been widely used for binary classification evaluation. Recent research generalizing the ROC/AUC analysis to multi-class classification has problems in at least one of the four areas: 1. failure to provide sensible plots 2. being sensitive to imbalanced data 3. unable to specify mis-classification cost and 4. unable to provide evaluation uncertainty quantification. Borrowing from a binomial matrix factorization model, we provide an evaluation metric summarizing the pair-wise multi-class True Positive Rate (TPR) and False Positive Rate (FPR) with one-dimensional vector representation. Visualization on the representation vector measures the relative speed of increment between TPR and FPR across all the classes pairs, which in turns provides a ROC plot for the multi-class counterpart. An integration over those factorized vector provides a binary AUC-equivalent summary on the classifier performance. Mis-clasification weights specification and bootstrapped confidence interval are also enabled to accommodate a variety of of evaluation criteria. To support our findings, we conducted extensive simulation studies and compared our method to the pair-wise averaged AUC statistics on benchmark datasets.
Accurate glycopeptide identification in mass spectrometry-based glycoproteomics is a challenging problem at scale. Recent innovation has been made in increasing the scope and accuracy of glycopeptide identifications, with more precise uncertainty estimates for each part of the structure. We present a dynamically adapting relative retention time model for detecting and correcting ambiguous glycan assignments that are difficult to detect from fragmentation alone, a layered approach to glycopeptide fragmentation modeling that improves N-glycopeptide identification in samples without compromising identification quality, and a site-specific method to increase the depth of the glycoproteome confidently identifiable even further. We demonstrate our techniques on a set of previously published datasets, showing the performance gains at each stage of optimization. These techniques are provided in the open-source glycomics and glycoproteomics platform GlycReSoft available at https://github.com/mobiusklein/glycresoft .
Aims: To determine whether evidence for infection with Theileria orientalis (Ikeda) could be identified in samples of commercial red deer (Cervus elaphus), horses, and working farm dogs in New Zealand. Methods: Blood samples were collected during October and November 2019 from a convenience sample of red deer (n = 57) at slaughter. Equine blood samples (n = 50) were convenience-sampled from those submitted to a veterinary pathology laboratory for routine testing in January 2020. Blood samples, collected for a previous study from a convenience sample of Huntaway dogs (n = 115) from rural regions throughout the North and South Islands of New Zealand between August 2018 and December 2020, were also tested. DNA was extracted and quantitative PCR was used to detect the T. orientalis Ikeda major piroplasm surface protein (MPSP) gene. A standard curve of five serial 10-fold dilutions of a plasmid carrying a fragment of the T. orientalis MPSP gene was used to quantify the number of T. orientalis organisms in the samples. MPSP amplicons obtained by end-point PCR on positive samples were isolated and subjected to DNA sequencing. The resulting sequences were compared to previously published T. orientalis sequences. Results: There were 6/57 (10%) samples positive for T. orientalis Ikeda from the deer and no samples positive for T. orientalis Ikeda from the working dogs or horses. The mean infection intensity for the six PCR-positive deer was 5.1 (min 2.2, max 12.4) T. orientalis Ikeda organisms/mu L. Conclusions and clinical relevance: Red deer can potentially sustain low infection intensities of T. orientalis Ikeda and could act as reservoirs of infected ticks. Further studies are needed to determine whether na & iuml;ve ticks feeding on infected red deer can themselves become infected. Abbreviations: Cq: Quantification cycle; LOQ: Limits of quantification; MPSP: Major piroplasm surface protein; qPCR: Quantitative polymerase chain reaction
Microbial interactions, which regulate the dynamics of eco- and agrosystems, can be harnessed to enhance antagonism against phytopathogenic fungi in agriculture. This study tests the hypothesis that plant growth-promoting rhizobacteria (PGPR) can also be potential biological control agents (BCAs). Antifungal activity assays against potentially phytopathogenic fungi were caried out using cultures and cell-free filtrates of nine PGPR strains previously isolated from agricultural soils. Cultures of Bacillus sp. BS36 inhibited the growth of Alternaria sp. AF12 and Fusarium sp. AF68 by 74 and 65%, respectively. Cell-free filtrates of the same strain also inhibited the growth of both fungi by 54 and 14%, respectively. Furthermore, the co-cultivation of Bacillus sp. BS36 with Pseudomonas sp. BS95 and the target fungi improved their antifungal activity. A subsequent metabolomic analysis using Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) identified fengycin- and surfactin-like lipopeptides (LPs) in the Bacillus sp. BS36 cell-free filtrates, which could explain their antifungal activity. The co-production of multiple families of LPs by Bacillus sp. BS36 is an interesting feature with potential practical applications. These results highlight the potential of the PGPR strain Bacillus sp. BS36 to work as a BCA and the need for more integrative approaches to develop biocontrol tools more accessible and adoptable by farmers.
Motivation Glycosylation elaborates the structures and functions of glycoproteins; glycoproteins are common post-translationally modified proteins and are heterogeneous and non-deterministically synthesized as an evolutionarily driven mechanism that elaborates the functions of glycosylated gene products. Glycoproteins, accounting for approximately half of all proteins, require specialized proteomics data analysis methods due to micro- and macro-heterogeneities as a given glycosite can be divided into several glycosylated forms, each of which must be quantified. Sampling of heterogeneous glycopeptides is limited by mass spectrometer speed and sensitivity, resulting in missing values. In conjunction with the low sample size inherent to glycoproteomics, a specialized toolset is needed to determine if observed changes in glycopeptide abundances are biologically significant or due to data quality limitations. Results We developed an R package, Relative Assessment of m/z Identifications by Similarity (RAMZIS), that uses similarity metrics to guide researchers to a more rigorous interpretation of glycoproteomics data. RAMZIS uses a permutation test to generate contextual similarity, which assesses the quality of mass spectral data and outputs a graphical demonstration of the likelihood of finding biologically significant differences in glycosylation abundance datasets. Investigators can assess dataset quality, holistically differentiate glycosites, and identify which glycopeptides are responsible for glycosylation pattern change. RAMZIS is validated by theoretical cases and a proof-of-concept application. RAMZIS enables comparison between datasets too stochastic, small, or sparse for interpolation while acknowledging these issues in its assessment. Using this tool, researchers will be able to rigorously define the role of glycosylation and the changes that occur during biological processes. Availability and implementation https://github.com/WillHackett22/RAMZIS.
Building energy use contributes to urban carbon dioxide (CO2) emissions while inadequate ventilation can yield indoor CO2 build up from human respiration. However, increasing ventilation rates can add to energy costs and climate burdens. Our objective was to quantify changes in emissions, energy, and financial cost when rooftop garden and ventilation upgrades are done simultaneously, with an opportunity to enhance plant growth from exhausted CO2. We measured indoor CO2 concentrations, calculated ventilation rates, and modeled five scenarios to assess these impacts. The indoor CO2 concentration maximum was 2210 ppm, median was 840 ppm, and 33% of the daytime was spent above 1000 ppm. The estimated ventilation rate was 4 L/s. . Our model calculations show that increasing ventilation to recommended levels (7 L/s) would increase total CO2 emissions, energy use, and cost (1-4%), but this could be counterbalanced by rooftop garden installation benefits, which yielded a net decrease of 23-46% in CO2 emissions, 12-13% in energy use, and 12-16% in cost. This novel integration of data collection and modeling provides support for the co-benefits of simultaneous improved installation ventilation systems and indoor CO2-enhanced rooftop gardens.
Rising ambient temperatures due to climate change will impact both indoor temperatures and heating and cooling utility costs. In traditionally colder climates, there are potential tradeoffs in how to meet the reduced heating and increased cooling demands, and issues related to lack of air conditioning (AC) access in older homes and among lower-income populations to prevent extreme heat exposure. We modeled a typical multi-family home in Boston (MA) in the building simulation program EnergyPlus to assess indoor temperature and energy consumption in current (2020) and projected future (2050) weather conditions. Selected households were those without AC (no AC), those who ran AC sometimes (some AC), and those with sufficient resources to run AC always (full AC). We considered stylized cooling subsidy policies that allowed households to move between groups, both independently and in conjunction with energy efficiency retrofits. Results showed that future weather conditions without policy changes yielded an increase in indoor summer temperatures of 2.1 °C (no AC), increased cooling demand (range: 34–50
We investigate a general matrix factorization for deviance-based data losses, extending the ubiquitous singular value decomposition beyond squared error loss. While similar approaches have been explored before, our method leverages classical statistical methodology from generalized linear models (GLMs) and provides an efficient algorithm that is flexible enough to allow for structural zeros via entry weights. Moreover, by adapting results from GLM theory, we provide support for these decompositions by (i) showing strong consistency under the GLM setup, (ii) checking the adequacy of a chosen exponential family via a generalized Hosmer-Lemeshow test, and (iii) determining the rank of the decomposition via a maximum eigenvalue gap method. To further support our findings, we conduct simulation studies to assess robustness to decomposition assumptions and extensive case studies using benchmark datasets from image face recognition, natural language processing, network analysis, and biomedical studies. Our theoretical and empirical results indicate that the proposed decomposition is more flexible, general, and robust, and can thus provide improved performance when compared to similar methods. To facilitate applications, an R package with efficient model fitting and family and rank determination is also provided.
Phosphorus (P) is an essential macronutrient for all life forms. Therefore, meeting the needs of a growing human population and their changing consumption patterns drastically intensified the use of mineral P fertilizers in agriculture. As a result, the current use of mineral P fertilizers causes severe negative economic, environmental and health impacts, which creates an urgent need for more sustainable agronomic practices capable of maintaining crop yields while improving P use efficiency. We consider that agronomic options that recycle/reuse the accumulated unavailable P (turn the unavailable P accumulated in the soil into P forms available for crop uptake) are an efficient strategy for food security, food production autonomy and sovereignty, and environmental sustainability. Here, we review P cycling in the soil and plant strategies to improve P acquisition, with special emphasis on the role of soil microbes as plant allies, namely their contribution to plant P acquisition directly through the production of organic acids and phosphatases, and indirectly through the production of phytohormones. Finally, we discuss why and how the use of soil microbes (mostly bacteria and fungi) with multiple modes of action may be the key to unlock soil P fractions unavailable for crop uptake, and highlight the benefits of combining: i) high-throughput sequencing; ii) new culturing methods to isolate and cultivate novel isolates; and iii) soil ecology experiments to develop multi-strain biofertilizers with diverse, complementary, and redundant modes of action in improving plant P acquisition and other benefits.
The binomial deviance and the SVM hinge loss functions are two of the most widely used loss functions in machine learning. While there are many similarities between them, they also have their own strengths when dealing with different types of data. In this work, we introduce a new exponential family based on a convex relaxation of the hinge loss function using softness and class-separation parameters. This new family, denoted Soft-SVM, allows us to prescribe a generalized linear model that effectively bridges between logistic regression and SVM classification. This new model is interpretable and avoids data separability issues, attaining good fitting and predictive performance by automatically adjusting for data label separability via the softness parameter. These results are confirmed empirically through simulations and case studies as we compare regularized logistic, SVM, and Soft-SVM regressions and conclude that the proposed model performs well in terms of both classification and prediction errors.
Heating and cooling requirement differences across climates not only have carbon emissions and energy efficiency implications but also impact indoor air quality (IAQ) and health. Energy and IAQ building simulation models help understand tradeoffs or co-benefits, but these have not been applied to evaluate climate zone or multi-family home differences. We modeled a four-story multi-family home in six U.S. climate zones and quantified energy, IAQ, and health outcomes with EnergyPlus, CONTAM, and a pediatric asthma systems science model. Pollutant sources included cooking and ambient. Outputs were daily PM2.5 and NO2 indoor concentrations, infiltration, energy for heating and cooling, and asthma exacerbations, which were compared across climate zones, apartment units, and resident behaviors. Daily ambient-sourced PM2.5 decreased and cooking-sourced PM2.5 increased with higher ambient temperatures. Infiltration air changes per hour were higher on the first versus the fourth floor and in colder climates. Window opening during cooking led to decreases in total pollutant concentrations (11%-18% for PM2.5 and 9%-15% for NO2 ), 3%-4% decreases in asthma exacerbations within climate zones, and minimal impacts on cooling, but led to increased heating demand (4%-8%). Our results demonstrate the influence of meteorology, multi-family building characteristics, and resident behavior on IAQ, energy, and health, focused on multi-zone methodology.