The Functional Autoregressive Model (FAR) generalizes the multivariate AR(1) model in Time Series Analysis to functional data. It serves as a historical foundational point in the study of functional time series and remains a fundamental and widely used model for dependent functional data. The process-the observed data-generated by the FAR model forms a Hilbert-valued Markov chain. This paper investigates the non-asymptotic prediction mean square error and derives a lower bound. This lower bound is established in the specific context of non-i.i.d. data and depends on the mixed smoothness of the functional time series and of the unknown correlation operator driving the FAR model. Instead of the standard functional PCA regularization, a ridge-type estimator is proposed, which avoids the preliminary estimation of the spectrum of the covariance sequence associated with the process. A non-asymptotic upper bound is derived for this estimate, which matches the lower bound up to multiplicative constants. Furthermore, a detailed study of the estimate's bias reveals connections between functional smoothness parameters and regularly/rapidly varying functions, which are common in extreme value theory. Simulation results corroborate the theoretical main theorems.
This study calls for a broadening of the perspective on academic success. While passing exams is an essential objective of higher education, it should not overshadow another important objective which is the development of students' skills, such as becoming curious, autonomous and reflective in the learning process. This study used Academic Performance in Exams (APE) and Deep Approach to Learning (DAL) as measures related to these two objectives. The aim was to identify and compare the factors that may influence APE and DAL. The study was conducted on first-year students (2011) at a French university. It was based on a random forest algorithm and took into account a wide range of factors belonging to different dimensions: demographics, social background, educational background, context of the educational programme, behavioural engagement, social environment, psychological and cognitive characteristics. The results show that the most important factors in predicting APE are the educational programme undertaken, student's educational background and parents' occupation. DAL was not found to be an important factor in APE. Regarding the prediction of DAL, the results point to the predominant weight of intrinsic motivation and the important weight of elaborated epistemic beliefs. In contrast, demographics and behavioural engagement were found to have negligible weight in predicting both APE and DAL. These findings raise questions about the type of success that is valued in the first year of university and call for reflection on assessment methods. They also allow the identification of levers that teachers can activate to support first year students.
Understanding the regulatory mechanisms that govern gene expression is crucial for deciphering cellular functions. Transcription factors (TFs) play a key role in regulating gene expression. In particular TF combinatorial interactions (TFCI) are now thought to largely shape genomic transcriptional responses, but predicting TFCI per se is still a difficult task. Single-cell RNA sequencing (scRNA-seq) has emerged as a powerful tool providing a whole new readout of gene regulatory effects. In this study, we propose a machine learning approach utilizing Classification and Regression Trees (CART) for predicting TFCI in >110k scRNA-seq data points yielded from Arabidopsis thaliana root. The proposed methodology provides a valuable tool for pointing to new TFCI mechanisms and could advance our understanding of Gene Regulatory Networks’ functioning. ### Competing Interest Statement The authors have declared no competing interest.
The problem of missing heritability requires the consideration of genetic interactions among different loci, called epistasis. Current GWAS statistical models require years to assess the entire combinatorial epistatic space for a single phenotype. We propose Next-Gen GWAS (NGG) that evaluates over 60 billion single nucleotide polymorphism combinatorial first-order interactions within hours. We apply NGG to Arabidopsis thaliana providing two-dimensional epistatic maps at gene resolution. We demonstrate on several phenotypes that a large proportion of the missing heritability can be retrieved, that it indeed lies in epistatic interactions, and that it can be used to improve phenotype prediction.
We propose in this work to derive a CLT in the functional linear regression model. The main difficulty is due to the fact that estimation of the functional parameter leads to a kind of ill-posed inverse problem. We consider estimators that belong to a large class of regularizing methods and we first show that, contrary to the multivariate case, it is not possible to state a CLT in the topology of the considered functional space. However, we show that we can get a CLT for the weak topology under mild hypotheses and in particular without assuming any strong assumptions on the decay of the eigenvalues of the covariance operator. Rates of convergence depend on the smoothness of the functional coefficient and on the point in which the prediction is made.
The first Genome Wide Association Studies (GWAS) shed light on the concept of missing heritability. It constitutes a mystery with transcending consequences from plant to human genetics. This mystery lies in the fact that a large proportion of phenotypes are not explained by unique or simple genomic modifications. One has to invoke genetic interactions among different loci, also known as epistasis, to partly account for it. However, current GWAS statistical models are moderately scalable, very sensitive to False Discovery Rate (FDR) corrections and, even combined with High Performance Computing (HPC), they can take years to evaluate for a full combinatorial epistatic space for a single phenotype. Here we propose a modeling approach, named Next-Gen GWAS (NGG) that evaluates, within hours, >60 billions of single nucleotide polymorphism (SNP) combinatorial first-order interactions, on a reasonable computer power. We first benchmark NGG on state of the art GWAS model results, and applied this to Arabidopsis thaliana providing 2D epistatic maps at gene resolution. We demonstrate on several phenotypes that a large proportion of the missing heritability can i) be retrieved with this modeling approach, ii) indeed lies in epistatic interactions and iii) can be used to improve phenotype prediction.
This paper proposes MpIC, an on-manifold derivation of the probabilistic Iterative Correspondence (pIC) algorithm, which is a stochastic version of the original Iterative Closest Point. It is developed in the context of autonomous underwater karst exploration based on acoustic sonars. First, a derivation of pIC based on the Lie group structure of SE(3) is developed. The closed-form expression of the covariance modeling the estimated rigid transformation is also provided. In a second part, its application to 3D scan matching between acoustic sonar measurements is proposed. It is a prolongation of previous work on elevation angle estimation from wide-beam acoustic sonar While the pIC approach proposed is intended to be a key component in a Simultaneous Localization and Mapping framework, this paper focuses on assessing its viability on a unitary basis. As ground truth data in karst aquifer are difficult to obtain, quantitative experiments are carried out on a simulated karst environment and show improvement compared to previous state-of-the-art approach. The algorithm is also evaluated on a real underwater cave dataset demonstrating its practical applicability.
The autoregressive Hilbertian model (ARH) was introduced in the early 90's by Denis Bosq. It was the subject of a vast literature and gave birth to numerous extensions. The model generalizes the classical multidimensional autoregressive model, widely used in Time Series Analysis. It was successfully applied in numerous fields such as finance, industry, biology. We propose here to compare the classical prediction methodology based on the estimation of the autocorrelation operator with a neural network learning approach. The latter is based on a popular version of Recurrent Neural Networks : the Long Short Term Memory networks. The comparison is carried out through simulations and real datasets.
The autoregressive Hilbertian model (ARH) was introduced in the early 90's by Denis Bosq. It was the subject of a vast literature and gave birth to numerous extensions. The model generalizes the classical multidimensional autoregressive model, widely used in Time Series Analysis. It was successfully applied in numerous fields such as finance, industry, biology. We propose here to compare the classical prediction methodology based on the estimation of the autocorrelation operator with a neural network learning approach. The latter is based on a popular version of Recurrent Neural Networks : the Long Short Term Memory networks. The comparison is carried out through simulations and real datasets.
We propose a new methodology to perform mineralogic inversion from wellbore logs based on a Bayesian linear regression model. Our method essentially relies on three steps. The first step makes use of Approximate Bayesian Computation (ABC) and selects from the Bayesian generator a set of candidates-volumes corresponding closely to the wellbore data responses. The second step gathers these candidates through a density-based clustering algorithm. A mineral scenario is assigned to each cluster through direct mineralogical inversion, and we provide a confidence estimate for each lithological hypothesis. The advantage of this approach is to explore all possible mineralogy hypotheses that match the wellbore data. This pipeline is tested on both synthetic and real datasets.
This article provides an overview of the basic theory and applications of linear processes for functional data, with particular emphasis on results published from 2000 to 2008. It first considers centered processes with values in a Hilbert space of functions before proposing some statistical models that mimic or adapt the scalar or finite-dimensional approaches for time series. It then discusses general linear processes, focusing on the invertibility and convergence of the estimated moments and a general method for proving asymptotic results for linear processes. It also describes autoregressive processes as well as two issues related to the general estimation problem, namely: identifiability and the inverse problem. Finally, it examines convergence results for the autocorrelation operator and the predictor, extensions for the autoregressive Hilbertian (ARH) model, and some numerical aspects of prediction when the data are curves observed at discrete points.
Particle swarm optimization algorithm is a stochastic meta-heuristic solving global optimization problems appreciated for its efficacity and simplicity. It consists in a swarm of particles interacting among themselves and searching the global optimum. The trajectory of the particles has been well-studied in a deterministic case and more recently in a stochastic context. Assuming the convergence of PSO, we proposed here two CLT for the particles corresponding to two kinds of convergence behavior. These results can lead to build confidence intervals around the local minimum found by the swarm or to the evaluation of the risk. A simulation study confirms these properties.
Pyrolitic lignin was modified through two methods. First, it was grafted with polylactide chains via a solvent-free process by ring-opening polymerization of l-lactide using calcium hydride as a catalyst. The efficiency of grafting was determined by infra-red, nuclear magnetic resonance and time-of-flight secondary ion mass spectrometry analyses. Then, lignin particles were oxygen plasma-treated and immersed in l-lactide solution. Infra-red and X-ray photoelectron spectroscopy revealed that chains bearing ester groups similar to that of lactide were covalently grafted onto the lignin. Composite cast films based on poly(l-lactide) matrix containing ungrafted lignin (lignin/PLLA), chemically-grafted lignin copolymer (PLA-g-lignin/PLLA) and plasma-treated lignin (plasma-treated lignin/PLLA) were investigated. Differential scanning calorimetry and dynamical mechanical analyses revealed that plasma-treated lignin preserved the crystalline structure of PLLA matrix and had a significant reinforcing effect compared with lignin and PLA-g-lignin. Results were compared with literature data.
Inferring transcriptional gene regulatory networks from transcriptomic datasets is a key challenge of systems biology, with potential impacts ranging from medicine to agronomy. There are several techniques used presently to experimentally assay transcription factors to target relationships, defining important information about real gene regulatory networks connections. These techniques include classical ChIP-seq, yeast one-hybrid, or more recently, DAP-seq or target technologies. These techniques are usually used to validate algorithm predictions. Here, we developed a reverse engineering approach based on mathematical and computer simulation to evaluate the impact that this prior knowledge on gene regulatory networks may have on training machine learning algorithms. First, we developed a gene regulatory networks-simulating engine called FRANK (Fast Randomizing Algorithm for Network Knowledge) that is able to simulate large gene regulatory networks (containing 104 genes) with characteristics of gene regulatory networks observed in vivo. FRANK also generates stable or oscillatory gene expression directly produced by the simulated gene regulatory networks. The development of FRANK leads to important general conclusions concerning the design of large and stable gene regulatory networks harboring scale free properties (built ex nihilo). In combination with supervised (accepting prior knowledge) support vector machine algorithm we (i) address biologically oriented questions concerning our capacity to accurately reconstruct gene regulatory networks and in particular we demonstrate that prior-knowledge structure is crucial for accurate learning, and (ii) draw conclusions to inform experimental design to performed learning able to solve gene regulatory networks in the future. By demonstrating that our predictions concerning the influence of the prior-knowledge structure on support vector machine learning capacity holds true on real data (Escherichia coli K14 network reconstruction using network and transcriptomic data), we show that the formalism used to build FRANK can to some extent be a reasonable model for gene regulatory networks in real cells.
The effects of polylactide-graft-cellulose nanocrystals on the thermal and mechanical properties of poly(l-lactide) matrices were investigated. Cellulose nanocrystals (CNCs) were grafted with polylactide chains via a solvent-free process by ring-opening polymerization of 1-lactide using magnesium hydride as a catalyst. The efficiency of grafting was determined by infra-red, X-ray photoelectron spectroscopy and nuclear magnetic resonance analyses. X-ray diffraction analyses showed that the crystalline nature of the CNCs was preserved. Nanocomposites based on poly(l-lactide) matrix containing ungrafted nanocrystals (PLLA/CNCs) and grafted nanocrystals (PLLA/PLLA-g-CNCs) were investigated. DSC revealed that the grafted nanocrystals exhibited a strong influence on the crystallinity of the nanocomposites, inducing a significant enhancement of the mechanical properties of PLLA/PLLA-g-CNCs compared with PLLA/CNCs material. The role played by the polylactide grafted layer on the interaction between the CNCs and PLLA matrix was revealed by mechanical analyses in the solid and molten states. (C) 2016 Elsevier Ltd. All rights reserved.
Functional linear regression has recently attracted considerable interest. Many works focus on asymptotic inference. In this paper we consider in a non asymptotic framework a simple estimation procedure based on functional Principal Regression. It revolves in the minimization of a least square contrast coupled with a classical projection on the space spanned by the m first empirical eigenvectors of the covariance operator of the functional sample. The novelty of our approach is to select automatically the crucial dimension m by minimization of a penalized least square contrast. Our method is based on model selection tools. Yet, since this kind of methods consists usually in projecting onto known non-random spaces, we need to adapt it to empirical eigenbasis made of data-dependent - hence random - vectors. The resulting estimator is fully adaptive and is shown to verify an oracle inequality for the risk associated to the prediction error and to attain optimal minimax rates of convergence over a certain class of ellipsoids. Our strategy of model selection is finally compared numerically with cross-validation.
The principal component analysis (PCA) is a famous technique from multivariate statistics. It is frequently carried out in dimension reduction either for functional data or in a high dimensional framework. To that aim PCA yields the eigenvectors \(\left( \widehat{\varphi }_{i}\right) _{i}\) of the covariance operator of a sample of interest. Dimension reduction is obtained by projecting on the eigenspaces spanned by the \(\widehat{\varphi }_{i}\)’s usually endowed with nice properties in terms of optimal information. We focus on the empirical eigenprojectors in the functional PCA of a \(n\)-sample and prove several non asymptotic results. More specifically we provide an upper bound for their mean square risk. This rate does not depend on the rate of decrease of the eigenvalues which seems to be a new result. We also derive a lower bound on the risk. The latter matches the upper bound up to a \(\log n\) term. The results are applied in a nonparametric functional estimation model.
Hypothesis: The interfacial compatibility between hydrophilic cellulose and hydrophobic poly(L-lactide) film surfaces is dependent on the interactions and interlocking of the macromolecular chains of the uppermost layers of both polymers. Grafting or coating the cellulose surface with molecular structures similar to the lactide monomer or oligomer is expected to improve the compatibility. Therefore, it should be possible to enhance the adhesive properties.Experiments: Cellulose films were oxygen plasma treated and immersed in a L-lactide solution. The grafting was performed under various conditions (power, pressure, time). The treated cellulose and poly(L-lactide) films were hot-pressed, and the resulting bi-layer laminates were subjected to a peel test. Comparative experiments were performed with the bi-layer laminates prepared from the cellulose films coated with poly(L-lactide-graft-vinyl alcohol) copolymers.Findings: X-ray photoelectron spectroscopy, infra-red analyses and wettability measurements revealed that chains bearing ester groups similar to that of lactide were covalently grafted onto the cellulose. The possible grafting mechanism that was initiated by the ionic species from the surface is discussed. As a result, the peel strength to separate the cellulose and the poly(L-lactide) films increased significantly. A comparison with data in the literature highlights the formation of entanglements inside the interfacial zone showing the efficiency of the plasma treatment. (C) 2015 Elsevier Inc. All rights reserved.
Functional data analysis aims to study and model observations which are by nature not vectors but random curves. For instance, the observation over n days of the price of share A provides a dependent sample of size n of the random function “Daily rate for share A”. The same is true for the recording of temperatures at a given location or the electricity consumption of a country. Each of these curves is of course initially discretized and can therefore be represented by a vector. But the limitations of the classic multivariate approach in this context quickly become apparent if the discretization frequency is too low or too high, or if there are corrupted data. It is in this context that functional data modeling has been a very active research area over the last twenty years. There has been remarkable development in this field, leading to modern advances in information technology and the advent of “Big Data”. Specific tools from signal processing or functional analysis must often be implemented and added to the statistician’s large collection of tools to circumvent the theoretical and numerical problems which then appear.
Experimental conditions for the synthesis of poly(epsilon-caprolatone)-graft-poly(vinyl alcohol) (PCL-g-PVA) and poly(t-lactide)-graft-poly(vinyl alcohol) (PLLA-g-PVA) via a PVA/MgH2 macroinitiator were optimized. Heat stability of PVA (99% hydrolyzed) in t-lactide CL-LA) melt was studied by DSC, TGA to lower the reaction temperature and time in order to avoid degradation. Thus, no degradation of copolymers occurred and C-13 NMR study showed that t-lactide was incorporated in graft chains as isotactic poly(t-lactide) without racemisation for 140-160 degrees C, up to 25 h, L-LA/PVA ratio in the range 3-12. Such synthesis using nontoxic catalyst and reactants in a solvent-free medium can be described as environment-friendly. Amphiphilic behavior of copolymers was studied in aqueous solution according to their chemical structures: critical micelle concentration CMC (0.03-0.5 g L-1), average diameter of micelles (80-150 nm), surface tension ycrvic (46-63 mN m(-1)), hydrophilic to lipophilic balance HLB (Griffin definition 4-7), and also from cast films: surface energy Vs (36-50 mJ m(-2)). This environment-friendly process is suitable for the use of,graft PVA copolymers in biomedical area for investigations under micelle applications. (C) 2013 Elsevier Ltd. All rights reserved.