This paper studies the problem of forecasting general stochastic processes using a path-dependent extension of the Neural Jump ODE (NJ-ODE) framework \citep{herrera2021neural}. While NJ-ODE was the first framework to establish convergence guarantees for the prediction of irregularly observed time series, these results were limited to data stemming from It\^o-diffusions with complete observations, in particular Markov processes, where all coordinates are observed simultaneously. In this work, we generalise these results to generic, possibly non-Markovian or discontinuous, stochastic processes with incomplete observations, by utilising the reconstruction properties of the signature transform. These theoretical results are supported by empirical studies, where it is shown that the path-dependent NJ-ODE outperforms the original NJ-ODE framework in the case of non-Markovian data. Moreover, we show that PD-NJ-ODE can be applied successfully to classical stochastic filtering problems and to limit order book (LOB) data.
In this paper, we study the extension of Neural Jump ODEs to infinite-dimensional function spaces. In particular, the underlying process X now takes values in L^2(Ξ, ℝ^d_X) instead of ℝ^d_X and the Operator NJ-ODE approximates the optimal predictor of this process by producing a representative of the conditional expectation. The NJ-ODE model is a framework for online learning the optimal prediction of continuous-time stochastic processes, given discrete, possibly irregular and incomplete past observations. In a series of works, this model has been extended to deal with generic path-dependent processes, with observation noise and dependent observations, with long-term predictions, and with input-output systems. However, throughout all of these works, the underlying processes were restricted to be finite-dimensional. In particular, function-valued problems, like yield curve or volatility surface predictions, could only be handled through discretization, which inherently leads to a loss of information. In this work, we build on ideas from Neural Operator methods that allow us to extend the NJ-ODE framework to an infinite-dimensional output process. To prove convergence of the NJ-ODE to the optimal prediction process, we develop a new approximation strategy that also generalizes previous works in the finite-dimensional setting by considerably weakening the assumptions.
RNA splicing modulators, a new class of small molecules with the potential to modify the protein expression levels, have quickly been translated into clinical trials. These compounds hold promise for treating neurodegenerative disorders, including branaplam for lowering huntingtin levels in Huntington’s disease. However, the VIBRANT-HD trial was terminated due to the emergence of peripheral neuropathy. Here, we describe the complex mechanism whereby branaplam activates p53, induces nucleolar stress in human induced pluripotent stem cell (iPSC)-derived motor neurons (iPSC-MN), and thereby enhanced expression of the neurotoxic p53-target gene BBC3. On the cellular level, branaplam disrupts neurite integrity, reflected by elevated neurofilament light chain levels. These findings illustrate the complex pharmacology of RNA splicing modulators with a small therapeutic window between lowering huntingtin levels and the clinically relevant off-target effect of neuropathy. Comprehensive toxicological screening in human stem cell models can complement pre-clinical testing before advancing RNA-targeting drugs to clinical trials.
We propose MoE-F — a formalized mechanism for combining N pre-trained expert Large Language Models (LLMs) in online time-series prediction tasks by adaptively forecasting the best weighting of LLM predictions at every time step. Our mechanism leverages the conditional information in each expert's running performance to forecast the best combination of LLMs for predicting the time series in its next step. Diverging from static (learned) Mixture of Experts (MoE) methods, our approach employs time-adaptive stochastic filtering techniques to combine experts. By framing the expert selection problem as a finite state-space, continuous-time Hidden Markov model (HMM), we can leverage the Wohman-Shiryaev filter. Our approach first constructs N parallel filters corresponding to each of the N individual LLMs. Each filter proposes its best combination of LLMs, given the information that they have access to. Subsequently, the N filter outputs are optimally aggregated to maximize their robust predictive power, and this update is computed efficiently via a closed-form expression, thus generating our ensemble predictor.Our contributions are:- **(I)** the MoE-F algorithm — deployable as a plug-and-play filtering harness,- **(II)** theoretical optimality guarantees of the proposed filtering-based gating algorithm (via optimality guarantees for its parallel Bayesian filtering and its robust aggregation steps), and- **(III)** empirical evaluation and ablative results using state-of-the-art foundational and MoE LLMs on a real-world _Financial Market Movement_ task where MoE-F attains a remarkable 17% absolute and 48.5% relative F1 measure improvement over the next best performing individual LLM expert predicting short-horizon market movement based on streaming news. Further, we provide empirical evidence of substantial performance gains in applying MoE-F over specialized models in the _long-horizon time-series forecasting_ domain. Code available on github: https://github.com/raeidsaqur/moe-f
In this work, we explore how Neural Jump ODEs (NJODEs) can be used as generative models for Itô processes. Given (discrete observations of) samples of a fixed underlying Itô process, the NJODE framework can be used to approximate the drift and diffusion coefficients of the process. Under standard regularity assumptions on the Itô processes, we prove that, in the limit, we recover the true parameters with our approximation. Hence, using these learned coefficients to sample from the corresponding Itô process generates, in the limit, samples with the same law as the true underlying process. Compared to other generative machine learning models, our approach has the advantage that it does not need adversarial training and can be trained solely as a predictive model on the observed samples without the need to generate any samples during training to empirically approximate the distribution. Moreover, the NJODE framework naturally deals with irregularly sampled data with missing values as well as with path-dependent dynamics, allowing to apply this approach in real-world settings. In particular, in the case of path-dependent coefficients of the Itô processes, the NJODE learns their optimal approximation given the past observations and therefore allows generating new paths conditionally on discrete, irregular, and incomplete past observations in an optimal way.
Detecting anomalies in irregularly sampled multi-variate time-series is challenging, especially in data-scarce settings. Here we introduce an anomaly detection framework for irregularly sampled time-series that leverages neural jump ordinary differential equations (NJODEs). The method infers conditional mean and variance trajectories in a fully path dependent way and computes anomaly scores. On synthetic data containing jump, drift, diffusion, and noise anomalies, the framework accurately identifies diverse deviations. Applied to infant gut microbiome trajectories, it delineates the magnitude and persistence of antibiotic-induced disruptions: revealing prolonged anomalies after second antibiotic courses, extended duration treatments, and exposures during the second year of life. We further demonstrate the predictive capabilities of the inferred anomaly scores in accurately predicting antibiotic events and outperforming diversity-based baselines. Our approach accommodates unevenly spaced longitudinal observations, adjusts for static and dynamic covariates, and provides a foundation for inferring microbial anomalies induced by perturbations, offering a translational opportunity to optimize intervention regimens by minimizing microbial disruptions.
Neural Jump ODEs model the conditional expectation between observations by neural ODEs and jump at arrival of new observations. They have demonstrated effectiveness for fully data-driven online forecasting in settings with irregular and partial observations, operating under weak regularity assumptions. This work extends the framework to input-output systems, enabling direct applications in online filtering and classification. We establish theoretical convergence guarantees for this approach, providing a robust solution to L 2 L<^>{2} -optimal filtering. Empirical experiments highlight the model's superior performance over classical parametric methods, particularly in scenarios with complex underlying distributions. These results emphasize the approach's potential in time-sensitive domains such as finance and health monitoring, where real-time accuracy is crucial.
Brain organoids derived from human pluripotent stem cells (hPSCs) hold immense potential for modeling neurodevelopmental processes and disorders. However, their experimental variability and undefined organoid selection criteria for analysis hinder reproducibility. As part of the Bavarian ForInter consortium, we generated 72 brain organoids from distinct hPSC lines. We conducted a comprehensive analysis of their morphological and cellular characteristics at an early stage of their development. In our assessment, the Feret diameter emerged as a reliable, single parameter that characterizes brain organoid quality. Transcriptomic analysis of our organoid identified the abundance of unintended mesodermal differentiation as a major confounder of unguided brain organoid differentiation, correlating with Feret diameter. High-quality organoids consistently displayed a lower presence of mesenchymal cells. These findings provide a framework for enhancing brain organoid standardization and reproducibility, underscoring the need for morphological quality controls and considering the influence of mesenchymal cells on organoid-based modeling.
The Path-dependent Neural Jump ODE (PD-NJ-ODE) is a model for online prediction of generic (possibly non-Markovian) stochastic processes with irregular (in time) and potentially incomplete (with respect to coordinates) observations. It is a model for which convergence to the L^2-optimal predictor, which is given by the conditional expectation, is established theoretically. Thereby, the training of the model is solely based on a dataset of realizations of the underlying stochastic process, without the need of knowledge of the law of the process. In the case where the underlying process is deterministic, the conditional expectation coincides with the process itself. Therefore, this framework can equivalently be used to learn the dynamics of ODE or PDE systems solely from realizations of the dynamical system with different initial conditions. We showcase the potential of our method by applying it to the chaotic system of a double pendulum. When training the standard PD-NJ-ODE method, we see that the prediction starts to diverge from the true path after about half of the evaluation time. In this work we enhance the model with two novel ideas, which independently of each other improve the performance of our modelling setup. The resulting dynamics match the true dynamics of the chaotic system very closely. The same enhancements can be used to provably enable the PD-NJ-ODE to learn long-term predictions for general stochastic datasets, where the standard model fails. This is verified in several experiments.
Autophagy and lysosomal pathways are involved in the cell entry of SARS-CoV-2 virus. To infect the host cell, the spike protein of SARS-CoV-2 binds to the cell surface receptor angiotensin-converting enzyme 2 (ACE2). To allow the fusion of the viral envelope with the host cell membrane, the spike protein has to be cleaved. One possible mechanism is the endocytosis of the SARS-CoV-2-ACE2 complex and subsequent cleavage of the spike protein, mainly by the lysosomal protease cathepsin L. However, detailed molecular and dynamic insights into the role of cathepsin L in viral cell entry remain elusive. To address this, HeLa cells and iPSC-derived alveolarspheres were treated with recombinant SARS-CoV-2 spike protein, and the changes in mRNA and protein levels of cathepsins L, B, and D were monitored. Additionally, we studied the effect of cathepsin L deficiency on spike protein internalization and investigated the influence of the spike protein on cathepsin L promoters in vitro. Furthermore, we analyzed variants in the genes coding for cathepsin L, B, D, and ACE2 possibly associated with disease progression using data from Regeneron's COVID Results Browser and our own cohort of 173 patients with COVID-19, exhibiting a variant of ACE2 showing significant association with COVID-19 disease progression. Our in vitro studies revealed a significant increase in cathepsin L mRNA and protein levels following exposure to the SARS-CoV-2 spike protein in HeLa cells, accompanied by elevated mRNA levels of cathepsin B and D in alveolarspheres. Moreover, an increase in cathepsin L promoter activity was detected in vitro upon spike protein treatment. Notably, the knockout of cathepsin L resulted in reduced internalization of the spike protein. The study highlights the importance of cathepsin L and lysosomal proteases in the SARS-CoV-2 spike protein internalization and suggests the potential of lysosomal proteases as possible therapeutic targets against COVID-19 and other viral infections.
The robust PCA of covariance matrices plays an essential role when isolating key explanatory features. The currently available methods for performing such a low-rank plus sparse decomposition are matrix specific, meaning, those algorithms must re-run for every new matrix. Since these algorithms are computationally expensive, it is preferable to learn and store a function that nearly instantaneously performs this decomposition when evaluated. Therefore, we introduce Denise, a deep learning-based algorithm for robust PCA of covariance matrices, or more generally, of symmetric positive semidefinite matrices, which learns precisely such a function. Theoretical guarantees for Denise are provided. These include a novel universal approximation theorem adapted to our geometric deep learning problem and convergence to an optimal solution to the learning problem. Our experiments show that Denise matches state-of-the-art performance in terms of decomposition quality, while being approximately $2000\times$ faster than the state-of-the-art, principal component pursuit (PCP), and $200 \times$ faster than the current speed-optimized method, fast PCP.
The Path-Dependent Neural Jump Ordinary Differential Equation (PD-NJ-ODE) is a model for predicting continuous-time stochastic processes with irregular and incomplete observations. In particular, the method learns optimal forecasts given irregularly sampled time series of incomplete past observations. So far the process itself and the coordinate-wise observation times were assumed to be independent and observations were assumed to be noiseless. In this work we discuss two extensions to lift these restrictions and provide theoretical guarantees as well as empirical examples for them. In particular, we can lift the assumption of independence by extending the theory to much more realistic settings of conditional independence without any need to change the algorithm. Moreover, we introduce a new loss function, which allows us to deal with noisy observations and explain why the previously used loss function did not lead to a consistent estimator.
Power utilities, especially those that generate electricity by burning fossil fuels, produce significant amounts of carbon emissions. Mitigation of CO2-equivalent (CO(2)e) emissions can be achieved by replacing power plants with renewable power installations and by adopting carbon-sequestration technologies. Physical upgrades are expensive, but carbon taxes, or the purchase of certificates and allowances on a voluntary carbon market, can be costly, too. Carbon costs may increasingly become a threatening liability for power utilities, eating into profits and undermining the financial viability of emission-intensive electricity generation. Thus, we consider an asset-and-liability, structural firm model to investigate the creditworthiness of a generic power utility. The utility's assets dynamics are driven by the financial returns generated from the sold electricity for a set tariff which is modeled by a simple stochastic process. The liabilities not only depend on fuel, running, and depreciation costs, but also on the costs of CO(2)e emissions. As a case study, we consider Eskom, the South African power utility. We show the evolution of Eskom's default probability under various fuel mix plans and technologies (as per SA's Integrated Resources Plan (IRP) 2019), and under the Network for Greening the Financial System (NGFS) carbon price scenarios. The obtained results and insights present a trying path ahead, especially for carbon-intensive power utilities.
Robust utility optimization enables an investor to deal with market uncertainty in a structured way, with the goal of maximizing the worst-case outcome. In this work, we propose a generative adversarial network (GAN) approach to (approximately) solve robust utility optimization problems in general and realistic settings. In particular, we model both the investor and the market by neural networks (NN) and train them in a mini-max zero-sum game. This approach is applicable for any continuous utility function and in realistic market settings with trading costs, where only observable information of the market can be used. A large empirical study shows the versatile usability of our method. Whenever an optimal reference strategy is available, our method performs on par with it and in the (many) settings without known optimal strategy, our method outperforms all other reference strategies. Moreover, we can conclude from our study that the trained path-dependent strategies do not outperform Markovian ones. Lastly, we uncover that our generative approach for learning optimal, (non-) robust investments under trading costs generates universally applicable alternatives to well known asymptotic strategies of idealized settings.
Biallelic loss of SPG11 function constitutes the most frequent cause of complicated autosomal recessive hereditary spastic paraplegia (HSP) with thin corpus callosum, resulting in progressive multisystem neurodegeneration. While the impact of neuroinflammation is an emerging and potentially treatable aspect in neurodegenerative diseases and leukodystrophies, the role of immune cells in SPG11–HSP patients is unknown. Here, we performed a comprehensive immunological characterization of SPG11–HSP, including examination of three human postmortem brain donations, immunophenotyping of patients’ peripheral blood cells and patient-specific induced pluripotent stem cell-derived microglia-like cells (iMGL). We delineate a previously unknown role of innate immunity in SPG11–HSP. Neuropathological analysis of SPG11–HSP patient brain tissue revealed profound microgliosis in areas of neurodegeneration, downregulation of homeostatic microglial markers and cell-intrinsic accumulation of lipids and lipofuscin in IBA1 + cells. In a larger cohort of SPG11–HSP patients, the ratio of peripheral classical and intermediate monocytes was increased, along with increased serum levels of IL-6 that correlated with disease severity. Stimulation of patient-specific iMGLs with IFNγ led to increased phagocytic activity compared to control iMGL as well as increased upregulation and release of proinflammatory cytokines and chemokines, such as CXCL10. On a molecular basis, we identified increased STAT1 phosphorylation as mechanism connecting IFNγ-mediated immune hyperactivation and SPG11 loss of function. STAT1 expression was increased both in human postmortem brain tissue and in an Spg11 –/– mouse model. Application of an STAT1 inhibitor decreased CXCL10 production in SPG11 iMGL and rescued their toxic effect on SPG11 neurons. Our data establish neuroinflammation as a novel disease mechanism in SPG11–HSP patients and constitute the first description of myeloid cell/ microglia activation in human SPG11–HSP. IFNγ/ STAT1-mediated neurotoxic effects of hyperreactive microglia upon SPG11 loss of function indicate that immunomodulation strategies may slow down disease progression.
We propose an optimal iterative scheme for federated transfer learning, where a central planner has access to datasets ${\cal D}_1,\dots,{\cal D}_N$ for the same learning model $f_{\theta}$. Our objective is to minimize the cumulative deviation of the generated parameters $\{\theta_i(t)\}_{t=0}^T$ across all $T$ iterations from the specialized parameters $\theta^\star_{1},\ldots,\theta^\star_N$ obtained for each dataset, while respecting the loss function for the model $f_{\theta(T)}$ produced by the algorithm upon halting. We only allow for continual communication between each of the specialized models (nodes/agents) and the central planner (server), at each iteration (round). For the case where the model $f_{\theta}$ is a finite-rank kernel regression, we derive explicit updates for the regret-optimal algorithm. By leveraging symmetries within the regret-optimal algorithm, we further develop a nearly regret-optimal heuristic that runs with $\mathcal{O}(Np^2)$ fewer elementary operations, where $p$ is the dimension of the parameter space. Additionally, we investigate the adversarial robustness of the regret-optimal algorithm showing that an adversary which perturbs $q$ training pairs by at-most $\varepsilon>0$, across all training sets, cannot reduce the regret-optimal algorithm's regret by more than $\mathcal{O}(\varepsilon q \bar{N}^{1/2})$, where $\bar{N}$ is the aggregate number of training pairs. To validate our theoretical findings, we conduct numerical experiments in the context of American option pricing, utilizing a randomly generated finite-rank kernel.