In multivariate time series analysis, understanding the underlying causal relationships among variables is often of interest for various applications. Directed acyclic graphs (DAGs) provide a powerful framework for representing causal dependencies. This paper proposes a novel Bayesian approach for modeling multivariate time series where conditional independencies and causal structure are encoded by a DAG. The proposed model allows structural properties such as stationarity to be easily accommodated. Given the application, we further extend the model for matrix-variate time series. We take a Bayesian approach to inference, and a “immersion-posterior” based efficient computational algorithm is developed. The posterior convergence properties of the proposed method are established along with two identifiability results for the unrestricted structural equation models. The utility of the proposed method is demonstrated through simulation studies and real data analysis.
A conventional approach to the extraction of latent components in a time series is to first model extreme values (including level shifts and seasonal outliers) as fixed effects, followed by their removal. Then the extreme-value adjusted series can be filtered using linear (Gaussian) techniques. A drawback is that identification of the epochs of extreme values is needed, and the uncertainty about this identification-as well as the removal of extremes-goes unmeasured. Alternatively, each outlier effect can be modeled as a particular type of latent stochastic process driven by heavy-tailed innovations; extraction of latent components then follows non-linear techniques and does not require identification of extreme epochs. We model monthly retail data impacted by the Covid-19 epidemic by incorporating additive outliers and level shifts as heavy-tailed latent processes, and estimate the unknown parameters through a Bayesian approach that utilizes Gibbs sampling. As a result, we can extract retail trends that incorporate stochastic level shifts and a full measure of the extraction uncertainty. An added benefit of the proposed approach is an estimate of a counterfactual trend following an extreme event. The posterior estimate of the counterfactual trend can be used to quantify the impact of an extreme event.
We study the structural changes in multivariate time-series by estimating and comparing stationary graphs for macroeconomic time series before and after an economic crisis such as the Great Recession. Building on a latent time series framework called Orthogonally-rotated Univariate Time-series (OUT), we propose a shared-parameter framework-the spOUT autoregressive model (spOUTAR)-that jointly models two related multivariate time series and enables coherent Bayesian estimation of their corresponding stationary precision matrices. This framework provides a principled mechanism to detect and quantify which conditional relationships among the variables changed, or formed following the crisis. Specifically, we study the impact of the Great Recession (December 2007-June 2009) that substantially disrupted global and national economies, prompting long-lasting shifts in macroeconomic indicators and their interrelationships. While many studies document its economic consequences, far less is known about how the underlying conditional dependency structure among economic variables changed as economies moved from pre-crisis stability through the shock and back to normalcy. Using the proposed approach to analyze U.S. and OECD macroeconomic data, we demonstrate that spOUTAR effectively captures recession-induced changes in stationary graphical structure, offering a flexible and interpretable tool for studying structural shifts in economic systems.
Time series data arising in many applications nowadays are high-dimensional. A large number of parameters describe features of these time series. We propose a novel approach to modeling a high-dimensional time series through several independent univariate time series, which are then orthogonally rotated and sparsely linearly transformed. With this approach, any specified intrinsic relations among component time series given by a graphical structure can be maintained at all time snapshots. We call the resulting process an Orthogonally-rotated Univariate Time series (OUT). Key structural properties of time series such as stationarity and causality can be easily accommodated in the OUT model. For Bayesian inference, we put suitable prior distributions on the spectral densities of the independent latent times series, the orthogonal rotation matrix, and the common precision matrix of the component times series at every time point. A likelihood is constructed using the Whittle approximation for univariate latent time series. An efficient Markov Chain Monte Carlo (MCMC) algorithm is developed for posterior computation. We study the convergence of the pseudo-posterior distribution based on the Whittle likelihood for the model's parameters upon developing a new general posterior convergence theorem for pseudo-posteriors. We find that the posterior contraction rate for independent observations essentially prevails in the OUT model under very mild conditions on the temporal dependence described in terms of the smoothness of the corresponding spectral densities. Through a simulation study, we compare the accuracy of estimating the parameters and identifying the graphical structure with other approaches. We apply the proposed methodology to analyze a dataset on different industrial components of the US gross domestic product between 2010 and 2019 and predict future observations.
We present a new design of a fast-neutron camera based on SiPM array readout. This design has advantages over previous optical-readout designs, such as compactness and modularity. In this Part I contribution we evaluate by Geant4 Monte-Carlo simulations the hierarchy and neutron-energy-dependence of the contribution of various fast-neutron interactions and secondary processes to the creation of light and their influence on the intrinsic spatial resolution and radiographic contrast in a 200 x 200 x 50 mm3 organic scintillator screen. Specifically, the contribution of proton recoils, delta electrons, carbon recoils, high-energy electrons, positrons, alpha particles and Cherenkov radiation to the total light and their influence on the intrinsic spatial resolution was evaluated in the 2.5-14-MeV neutron energy range. In a 50-mm-thick scintillator camera the limiting intrinsic spatial resolution was about 450 mu m for 14-MeV neutrons and appreciably better for lower neutron energies.
Dynamic risk prediction that incorporates longitudinal measurements of biomarkers is useful in identifying high-risk patients for better clinical management. Our work is motivated by the prediction of cervical precancers. Currently, Pap cytology is used to identify HPV-positive (HPV+) women at high-risk of cervical precancer, but cytology lacks accuracy and reproducibility. Molecular markers, like HPV DNA methylation, that are closely linked to the carcinogenic process show promise of improved risk stratification. We are interested in developing a dynamic risk model that uses all longitudinal biomarker information to improve precancer risk estimation. We propose a joint model to link both the continuous methylation biomarker and a binary cytology biomarker to the time to precancer outcome using shared random effects. The model uses a discretization of the time scale to allow for closed-form likelihood expressions, thereby avoiding potential high dimensional integration of the random effects. The method handles an interval-censored time-to-event outcome, due to intermittent clinical visits, incorporates sampling weights to deal with stratified sampling data and can provide immediate and five-year risk estimates that may inform clinical decision-making. Applying the method to longitudinally measured HPV methylation data improves risk stratification for triage of HPV+ women.
Integrin α5β1 is crucial for cell attachment and migration in development and tissue regeneration, and α5β1 binding proteins could have considerable utility in regenerative medicine and next-generation therapeutics. We use computational protein design to create de novo α5β1-specific modulating miniprotein binders, called NeoNectins, that bind to and stabilize the open state of α5β1. When immobilized onto titanium surfaces and throughout 3D hydrogels, the NeoNectins outperform native fibronectin and RGD peptide in enhancing cell attachment and spreading, and NeoNectin-grafted titanium implants outperformed fibronectin and RGD-grafted implants in animal models in promoting tissue integration and bone growth. NeoNectins should be broadly applicable for tissue engineering and biomedicine.
Dual-phase liquid-xenon time projection chambers (LXe TPCs) deploying a few tonnes of liquid are presently leading the search for WIMP dark matter. Scaling these detectors to 10-fold larger fiducial masses, while improving their sensitivity to low-mass WIMPs presents difficult challenges in detector design. Several groups are considering a departure from current schemes, towards either single-phase liquid-only TPCs, or dual-phase detectors where the electroluminescence region consists of patterned electrodes. Here, we discuss the possible use of Thick Gaseous Electron Multipliers (THGEMs) coated with a VUV photocathode and immersed in LXe as a building block in such designs. We focus on the transfer efficiencies of ionization electrons and photoelectrons emitted from the photocathode through the electrode holes, and show experimentally that efficiencies approaching 100 % can be achieved with realistic voltage settings. The observed voltage dependence of the transfer efficiencies is consistent with electron transport simulations once diffusion and charging-up effects are included.
Under a high-dimensional vector autoregressive (VAR) model, we propose a way of efficiently estimating both the stationary graph structure between the nodal time series and their temporal dynamics. The framework is then used to make inferences on the change in interdependencies between several economic indicators due to the impact of the Great Recession, the financial crisis that lasted from 2007 through 2009. There are several key advantages of the proposed framework; (1) it develops a reparametrized VAR likelihood that can be used in general high-dimensional VAR problems, (2) it strictly maintains causality of the estimated process, making inference on stationary features more meaningful and (3) it is computationally efficient due to the reduced rank structure of the parameterization. We apply the methodology to the seasonally adjusted quarterly economic indicators available in the FRED-QD database of the Federal Reserve. The analysis essentially confirms much of the prevailing knowledge about the impact of the Great Recession on different economic indicators. At the same time, it provides deeper insight into the nature and extent of the impact on the interplay of the different indicators. We also contribute to the theory of Bayesian VAR by showing the consistency of the posterior under sparse priors for the parameters of the reduced rank formulation of the VAR process.
We propose a method for simultaneously estimating a contemporaneous graph structure and autocorrelation structure for a causal high-dimensional vector autoregressive process (VAR). The graph is estimated by estimating the stationary precision matrix using a Bayesian framework. We introduce a novel parameterization that is convenient for jointly estimating the precision matrix and the autocovariance matrices. The methodology based on the new parameterization has several desirable properties. A key feature of the proposed methodology is that it maintains causality of the process in its estimates and also provides a fast feasible way for computing the reduced rank likelihood for a high-dimensional Gaussian VAR. We use sparse priors along with the likelihood under the new parameterization to obtain the posterior of the graphical parameters as well as that of the temporal parameters. An efficient Markov Chain Monte Carlo (MCMC) algorithm is developed for posterior computation. We also establish theoretical consistency properties for the high-dimensional posterior. The proposed methodology shows excellent performance in simulations and real data applications.
Discussion forums are a key component of online learning platforms, allowing learners to ask for help, provide help to others, and connect with others in the learning community. Analyzing patterns of forum usage and their association with course outcomes can provide valuable insight into how learners actually use discussion forums, and suggest strategies for shaping forum dynamics to improve learner experiences and outcomes. However, the fine-grained coding of forum posts required for this kind of analysis is a manually intensive process that can be challenging for large datasets, e.g., those that result from popular MOOCs. To address this issue, we propose an AI-assisted labeling process that uses advanced natural language processing techniques to train machine learning models capable of labeling a large dataset while minimizing human annotation effort. We fine-tune pretrained transformer-based deep learning models on category, structure, and emotion classification tasks. The transformer-based models outperform a more traditional baseline that uses support vector machines and a bag-of-words input representation. The transformer-based models also perform better when we augment the input features for an individual post with additional context from the post's thread (e.g., the thread title). We validate model quality through a combination of internal performance metrics, human auditing, and common-sense checks. For our Python MOOC dataset, we find that annotating approximately 1% of the forum posts achieves performance levels that are reliable for downstream analysis. Using labels from the validated AI models, we investigate the association of learner and course attributes with thread resolution and various forms of forum participation. We find significant differences in how learners of different age groups, gender, and course outcome status ask for help, provide help, and make posts with emotional (positive or negative) sentiment.
This chapter provides an overview of the luminescence properties of Group 9 (Co, Rh, Ir) complexes. The synthesis and photophysical properties of Co(II/III), Rh(I/III), and Ir(III) complexes have been discussed with various judiciously designed ligands that fine-tune the luminescent features of the complexes with different natures of origin of luminescence, e.g., 3MLCT, 3LMCT, 3ILCT, and 3LC. The priority is given to the complexes showing the potentials for applications in sensing of metal ions, photoredox catalysis, and solid-state lighting devices (organic light-emitting diodes, OLEDs and light-emitting electrochemical cells, LECs). While Ir(III) complexes show magnificent photophysical properties due to their high spin-orbit coupling constant (3909 cm−1) and often finds applications in all these categories, intense research is yet to be performed to induce bright luminescence (ΦPL > 50%) in 3 and 4d T-metal Co(II/III) and Rh(I/III) complexes and find their practical applications.
We study the integral of the Frobenius norm as a measure of the discrepancy between two multivariate spectra. Such a measure can be used to fit time series models, and ensures proximity between model and process at all frequencies of the spectral density. We develop new asymptotic results for linear and quadratic functionals of the periodogram, and apply the integrated Frobenius norm to fit time series models and test whether model residuals are white noise. The case of structural time series models is addressed, wherein co-integration rank testing is formally developed. Both applications are studied through simulation studies and time series data. The numerical results show that the proposed estimator can fit moderate- to large-dimensional structural timeseries in real time.
The RGD (Arg-Gly-Asp)-binding integrins αvβ6 and αvβ8 are clinically validated cancer and fibrosis targets of considerable therapeutic importance. Compounds that can discriminate between the two closely related integrin proteins and other RGD integrins, stabilize specific conformational states, and have sufficient stability enabling tissue restricted administration could have considerable therapeutic utility. Existing small molecules and antibody inhibitors do not have all of these properties, and hence there is a need for new approaches. Here we describe a method for computationally designing hyperstable RGD-containing miniproteins that are highly selective for a single RGD integrin heterodimer and conformational state, and use this strategy to design inhibitors of αvβ6 and αvβ8 with high selectivity. The αvβ6 and αvβ8 inhibitors have picomolar affinities for their targets, and >1000-fold selectivity over other RGD integrins. CryoEM structures are within 0.6-0.7Å root-mean-square deviation (RMSD) to the computational design models; the designed αvβ6 inhibitor and native ligand stabilize the open conformation in contrast to the therapeutic anti-αvβ6 antibody BG00011 that stabilizes the bent-closed conformation and caused on-target toxicity in patients with lung fibrosis, and the αvβ8 inhibitor maintains the constitutively fixed extended-closed αvβ8 conformation. In a mouse model of bleomycin-induced lung fibrosis, the αvβ6 inhibitor potently reduced fibrotic burden and improved overall lung mechanics when delivered via oropharyngeal administration mimicking inhalation, demonstrating the therapeutic potential of de novo designed integrin binding proteins with high selectivity.
The catalytic versatility of pentacoordinated iron is highlighted by the broad range of natural and engineered activities of heme enzymes such as cytochrome P450s, which position a porphyrin cofactor coordinating a central iron atom below an open substrate binding pocket. This catalytic prowess has inspired efforts to design de novo helical bundle scaffolds that bind porphyrin cofactors. However, such designs lack the large open substrate binding pocket of P450s, and hence, the range of chemical transformations accessible is limited. Here, with the goal of combining the advantages of the P450 catalytic site geometry with the almost unlimited customizability of de novo protein design, we design a high-affinity heme-binding protein, dnHEM1, with an axial histidine ligand, a vacant coordination site for generating reactive intermediates, and a tunable distal pocket for substrate binding. A 1.6 Å X-ray crystal structure of dnHEM1 reveals excellent agreement to the design model with key features programmed as intended. The incorporation of distal pocket substitutions converted dnHEM1 into a proficient peroxidase with a stable neutral ferryl intermediate. In parallel, dnHEM1 was redesigned to generate enantiocomplementary carbene transferases for styrene cyclopropanation (up to 93% isolated yield, 5000 turnovers, 97:3 e.r.) by reconfiguring the distal pocket to accommodate calculated transition state models. Our approach now enables the custom design of enzymes containing cofactors adjacent to binding pockets with an almost unlimited variety of shapes and functionalities.
Statistical problems often involve linear equality and inequality constraints on model parameters. Direct estimation of parameters restricted to general polyhedral cones, particularly when one is interested in estimating low dimensional features, may be challenging. We use a dual form parameterization to characterize parameter vectors restricted to lower dimensional faces of polyhedral cones and use the characterization to define a notion of 'sparsity' on such cones. We show that the proposed notion agrees with the usual notion of sparsity in the unrestricted case and prove the validity of the proposed definition as a measure of sparsity. The identifiable parameterization of the lower dimensional faces allows a generalization of popular spike-and-slab priors to a closed convex polyhedral cone. The prior measure utilizes the geometry of the cone by defining a Markov random field over the adjacency graph of the extreme rays of the cone. We describe an efficient way of computing the posterior of the parameters in the restricted case. We illustrate the usefulness of the proposed methodology for imposing linear equality and inequality constraints by using wearables data from the National Health and Nutrition Examination Survey (NHANES) actigraph study where the daily average activity profiles of participants exhibit patterns that seem to obey such constraints.
We develop appropriate Bayesian procedures to draw inference about the parameters under a multivariate normal model based on synthetic data. We consider two standard forms of synthetic data, generated under plug in sampling method and posterior predictive sampling method. In addition to point estimates of the mean vector and dispersion matrix, Bayesian credible sets for the mean vector and the generalized variance are also provided under both the scenarios. The analysis in the case when some (partial) features are sensitive and need to be hidden is also briey indicated. Vol. 23(2), November, 2023, pp 1-18
BACKGROUND:Customized fetal growth charts assume birthweight at term to be normally distributed across the population with a constant coefficient of variation at earlier gestational ages. Thus, standard deviation used for computing percentiles (e.g., 10th, 90th) is assumed to be proportional to the customized mean, although this assumption has never been formally tested. METHODS:In a secondary analysis of NICHD Fetal Growth Studies-Singletons (12 U.S. sites, 2009-2013) using longitudinal sonographic biometric data (n = 2288 pregnancies), we investigated the assumptions of normality and constant coefficient of variation by examining behavior of the mean and standard deviation, computed following the Gardosi method. We then created a more flexible model that customizes both mean and standard deviation using heteroscedastic regression and calculated customized percentiles directly using quantile regression, with an application in a separate study of 102, 012 deliveries, 37-41 weeks. RESULTS:Analysis of term optimal birthweight challenged assumptions of proportionality and that values were normally distributed: at different mean birthweight values, standard deviation did not change linearly with mean birthweight and the percentile computed with the normality assumption deviated from empirical percentiles. Composite neonatal morbidity and mortality rates in relation to birthweight < 10th were higher for heteroscedastic and quantile models (10.3% and 10.0%, respectively) than the Gardosi model (7.2%), although prediction performance was similar among all three (c-statistic 0.52-0.53). CONCLUSIONS:Our findings question normality and constant coefficient of variation assumptions of the Gardosi customization method. A heteroscedastic model captures unstable variance in customization characteristics which may improve detection of abnormal growth percentiles. TRIAL REGISTRATION:ClinicalTrials.gov identifier: NCT00912132.
We present a first prototype for a multiplicity counter of fast neutrons and y rays, based on plastic scintillators coupled to silicon photomultiplier arrays, where the SiPMs are read out individually by the TOFPET2 ASIC providing time-stamps and charge information. We demonstrate the capabilities of our counter by measuring the Singles-, Doubles-and Triples-event rates of neutrons and y rays (without pulse shape discrimination), for a 252Cf source with a nominal activity of 4.9 mu Ci, and comparing them with a GEANT4 simulation of the process.