Current causal discovery approaches require restrictive model assumptions in the absence of interventional data to ensure structure identifiability. These assumptions often do not hold in real-world applications leading to a loss of guarantees and poor performance in practice. Recent work has shown that, in the bivariate case, Bayesian model selection can greatly improve performance by exchanging restrictive modelling for more flexible assumptions, at the cost of a small probability of making an error. Our work shows that this approach is useful in the important multivariate case as well. We propose a scalable algorithm leveraging a continuous relaxation of the discrete model selection problem. Specifically, we employ the Causal Gaussian Process Conditional Density Estimator (CGP-CDE) as a Bayesian non-parametric model, using its hyperparameters to construct an adjacency matrix. This matrix is then optimised using the marginal likelihood and an acyclicity regulariser, giving the maximum a posteriori causal graph. We demonstrate the competitiveness of our approach, showing it is advantageous to perform multivariate causal discovery without infeasible assumptions using Bayesian model selection.
This work develops a Bayesian non-parametric approach to signal separation where the signals may vary according to latent variables. Our key contribution is to augment Gaussian Process Latent Variable Models (GPLVMs) for the case where each data point comprises the weighted sum of a known number of pure component signals, observed across several input locations. Our framework allows arbitrary non-linear variations in the signals while being able to incorporate useful priors for the linear weights, such as summing-to-one. Our contributions are particularly relevant to spectroscopy, where changing conditions may cause the underlying pure component signals to vary from sample to sample. To demonstrate the applicability to both spectroscopy and other domains, we consider several applications: a near-infrared spectroscopy dataset with varying temperatures, a simulated dataset for identifying flow configuration through a pipe, and a dataset for determining the type of rock from its reflectance.
With the rise in engineered biomolecular devices, there is an increased need for tailor-made biological sequences. Often, many similar biological sequences need to be made for a specific application meaning numerous, sometimes prohibitively expensive, lab experiments are necessary for their optimization. This paper presents a transfer learning design of experiments workflow to make this development feasible. By combining a transfer learning surrogate model with Bayesian optimization, we show how the total number of experiments can be reduced by sharing information between optimization tasks. We demonstrate the reduction in the number of experiments using data from the development of DNA competitors for use in an amplification-based diagnostic assay. We use cross-validation to compare the predictive accuracy of different transfer learning models, and then compare the performance of the models for both single objective and penalized optimization tasks.
Gene expression has great potential to be used as a clinical diagnostic tool. However, despite the progress in identifying these gene expression signatures, clinical translation has been hampered by a lack of purpose-built. readily deployable testing platforms. We have developed Competitive Amplification Networks. CANs to enable analysis of an entire gene expression signature in a single PCR reaction. CANs consist of natural and synthetic amplicons that compete for shared primers during amplification, forming a reaction network that leverages the molecular machinery of PCR. These reaction components are tuned such that the final fluorescent signal from the assay is exactly calibrated to the conclusion of a statistical model. In essence, the reaction acts as a biological computer, simultaneously detecting the RNA targets, interpreting their level in the context of the gene expression signature, and aggregating their contributions to the final diagnosis. We illustrate the clinical validity of this technique, demonstrating perfect diagnostic agreement with the gold-standard approach of measuring each gene independently. Crucially, CAN assays are compatible with existing qPCR instruments and workflows. CANs hold the potential to enable rapid deployment and massive scalability of gene expression analysis to clinical laboratories around the world, in highly developed and low-resource J settings alike. ![Figure][1] ### Competing Interest Statement JPG and MMS are listed as inventors on a patent application describing the technology presented here, and hold founding shares in Signatur Biosciences, Inc, a company which seeks to commercialize this technology. [1]: pending:yes
Monitoring air pollution is of vital importance to the overall health of the population. Unfortunately, devices that can measure air quality can be expensive, and many cities in low and middle-income countries have to rely on a sparse allocation of them. In this paper, we investigate the use of Gaussian Processes for both nowcasting the current air-pollution in places where there are no sensors and forecasting the air-pollution in the future at the sensor locations. In particular, we focus on the city of Kampala in Uganda, using data from AirQo's network of sensors. We demonstrate the advantage of removing outliers, compare different kernel functions and additional inputs. We also compare two sparse approximations to allow for the large amounts of temporal data in the dataset.
There is a growing trend in molecular and synthetic biology of using mechanistic (non machine learning) models to design biomolecular networks. Once designed, these networks need to be validated by experimental results to ensure the theoretical network correctly models the true system. However, these experiments can be expensive and time consuming. We propose a design of experiments approach for validating these networks efficiently. Gaussian processes are used to construct a probabilistic model of the discrepancy between experimental results and the designed response, then a Bayesian optimization strategy used to select the next sample points. We compare different design criteria and develop a stopping criterion based on a metric that quantifies this discrepancy over the whole surface, and its uncertainty. We test our strategy on simulated data from computer models of biochemical processes.
BACKGROUND:Peripheral nerves carry afferent and efferent signals between the central nervous system and the periphery of the body. When nerves are strained above physiological levels, conduction blocks occur, resulting in debilitating loss of motor and sensory function. Understanding the effects of strain on nerve function requires knowledge of the multi-scale mechanical behaviour of the tissue, and how this is transferred to the cellular environment.NEW METHOD:The aim of this work was to establish a technique to measure the partitioning of strain between tissue and axons in axially loaded peripheral nerves. This was achieved by staining extracellular domains of sodium channels clustered at nodes of Ranvier, without altering tissue mechanical properties by fixation or permeabilisation.RESULTS:Stained nerves were imaged by multi-photon microscopy during in situ tensile straining, and digital image correlation was used to measure axonal strain with increasing tissue strain. Strain was partitioned between tissue and axon scales by an average factor of 0.55.COMPARISONS WITH EXISTING METHODS:This technique allows non-invasive probing of cell-level strain within the physiological tissue environment.CONCLUSIONS:This technique can help understand the mechanisms behind the onset of conduction blocks in injured peripheral nerves, as well as to evaluate changes in multi-scale mechanical properties in diseased nerves.