We construct the least superharmonic majorant of a continuous function g on the d-dimensional unit ball (d ≥ 2) via a canonical sequential scheme. While classical theory identifies this majorant with the value function of the optimal stopping problem for Brownian motion absorbed at the domain boundary, no comparable constructive approximation scheme has been available. We introduce branched harmonic majorants, obtained by arranging classical harmonic functions on smoothly bounded domains in a finite, depth-indexed branching structure, and prove two main results. First, the optimal stopping region is identified as the contact set between the gain function g and the pointwise infimum of this family; the value function is recovered as the expected gain at the first exit time from the non-contact set. This yields a multidimensional generalisation of the Dynkin–Yushkevich concave-envelope theorem in which affine functions are replaced by branched harmonic majorants. Second, truncation in the branching depth produces a decreasing sequence of envelopes that converges pointwise to this infimum, yielding an explicit approximation scheme not present in classical formulations. Analytically, the branching structure relaxes the global majorisation constraint to a local constraint imposed on a decreasing sequence of non-contact sets, yielding a representation of the Perron envelope in terms of harmonic functions on smoothly bounded domains. Probabilistically, the construction corresponds to the sequential composition of stopping times and overcomes the localisation obstruction arising from the thinness of Brownian paths in dimensions d ≥ 2.
We use the geometry of functions associated with martingales under nonlinear expectations to solve risk-sensitive Markovian optimal stopping problems. Generalising the linear case due to Dynkin and Yushkievich (1969), the value function is the majorant or pointwise infimum of those functions which dominate the gain function. An emphasis is placed on the geometry of the majorant and pathwise arguments, rather than exploiting convexity, positive homogeneity or related analytical properties. An algorithm is provided to construct the value function at the computational cost of a two-dimensional search.
The RangL project hosted by The Alan Turing Institute aims to encourage the wider uptake of reinforcement learning by supporting competitions relating to real-world dynamic decision problems. This article describes the reusable code repository developed by the RangL team and deployed for the 2022 Pathways to Net Zero Challenge, supported by the UK Net Zero Technology Centre. The winning solutions to this particular Challenge seek to optimize the UK's energy transition policy to net zero carbon emissions by 2050. The RangL repository includes an OpenAI Gym reinforcement learning environment and code that supports both submission to, and evaluation in, a remote instance of the open source EvalAI platform as well as all winning learning agent strategies. The repository is an illustrative example of RangL's capability to provide a reusable structure for future challenges.
We formulate a probabilistic Markov property in discrete time under a dynamic risk framework with minimal assumptions. This is useful for recursive solutions to risk-sensitive versions of dynamic optimisation problems such as optimal prediction, where at each stage the recursion depends on the whole future. The property holds for standard measures of risk used in practice, and is formulated in several equivalent versions including a representation via acceptance sets, a strong version, and a dual representation.
We solve non-Markovian optimal switching problems in discrete time on an infinite horizon, when the decision-maker is risk-aware and the filtration is general, and establish existence and uniqueness of solutions for the associated reflected backward stochastic difference equations. An example application to hydropower planning is provided.
We present and numerically analyse the Basin Hopping with Skipping (BH-S) algorithm for stochastic optimisation. This algorithm replaces the perturbation step of basin hopping (BH) with a so-called skipping mechanism from rare-event sampling. Empirical results on benchmark optimisation surfaces demonstrate that BH-S can improve performance relative to BH by encouraging non-local exploration, that is, by hopping between distant basins.
This paper investigates large fluctuations of Locational Marginal Prices (LMPs) in wholesale energy markets caused by volatile renewable generation profiles. Specifically, we study events of the form $\mathbb{P} \Big ( \mathbf{LMP} \notin \prod_{i=1}^n [\alpha_i^-, \alpha_i^+] \Big),$ where $\mathbf{LMP}$ is the vector of LMPs at the $n$ power grid nodes, and $\boldsymbol{\alpha}^-,\boldsymbol{\alpha}^+\in\mathbb{R}^n$ are vectors of price thresholds specifying undesirable price occurrences. By exploiting the structure of the supply-demand matching mechanism in power grids, we look at LMPs as deterministic piecewise affine, possibly discontinuous functions of the stochastic input process, modeling uncontrollable renewable generation. We utilize techniques from large deviations theory to identify the most likely ways for extreme price spikes to happen, and to rank the nodes of the power grid in terms of their likelihood of experiencing a price spike. Our results are derived in the case of Gaussian fluctuations and are validated numerically on the IEEE 14-bus test case.
We perform a rare-event study on a simulated power system in which grid-scale batteries provide both regulation and emergency frequency control ancillary services. Using a model of random power disturbances at each bus, we employ the skipping sampler, a Markov Chain Monte Carlo algorithm for rare-event sampling, to build conditional distributions of the power disturbances leading to two kinds of instability: frequency excursions outside the normal operating band, and load shedding. Potential saturation in the benefits, and competition between the two services, are explored as the battery maximum power output increases. This article is part of the theme issue 'The mathematics of energy systems'.
The urgent need to decarbonize energy systems gives rise to many challenging areas of interdisciplinary research, bringing together mathematicians, physicists, engineers and economists. Renewable generation, especially wind and solar, is inherently highly variable and difficult to predict. The need to keep power and energy systems balanced on a second-by-second basis gives rise to problems of control and optimization, together with those of the management of liberalized energy markets. On the longer time scales of planning and investment, there are problems of physical and economic design. The papers in the present issue are written by some of the participants in a programme on the mathematics of energy systems which took place at the Isaac Newton Institute for Mathematical Sciences in Cambridge from January to May 2019-see http://www.newton.ac.uk/event/mes. This article is part of the theme issue 'The mathematics of energy systems'.
The RangL project encourages the wider uptake of reinforcement learning by supporting a cloud-native platform that hosts industrially focused sequential decision problems. In this opinion piece, members of the RangL team make the case for the potential of competition platforms to facilitate collaboration between academia and industry, and to drive progress in real-world applications of reinforcement learning, over the coming years.
This paper addresses the problem of 'quickest possible' online transient stability assessment, by minimizing the decision time of combined event detection that might lead to a system split and unstable generator group prediction, from real-time wide area power system measurements. More importantly it does so by respecting predefined probabilistic error constraints for the prediction. The statistical theory of optimal detection is applied, firstly to choose the detection threshold and secondly to select a flexible assessment time, after using probabilistic neural networks to provide a temporal representation of the data. On simulated wide area measurements from the interconnected New England test system and New York power system this approach is between two and three times faster on average than strategies based on fixed assessment times, despite having comparable error rates.
We aim to improve upon the exploration of the general-purpose random walk Metropolis algorithm when the target has non-convex support $A \subset \mathbb{R}^d$, by reusing proposals in $A^c$ which would otherwise be rejected. The algorithm is Metropolis-class and under standard conditions the chain satisfies a strong law of large numbers and central limit theorem. Theoretical and numerical evidence of improved performance relative to random walk Metropolis are provided. Issues of implementation are discussed and numerical examples, including applications to global optimisation and rare event sampling, are presented.
In the nonzero-sum setting, we establish a connection between Nash equilibria in games of optimal stopping (Dynkin games) and generalised Nash equilibrium problems (GNEP). In the Dynkin game this reveals novel equilibria of threshold type and of more complex types, and leads to novel uniqueness and stability results.
We aim to analyse a Markovian discrete-time optimal stopping problem for a risk-averse decision maker under model ambiguity. In contrast to the analytic approach based on transition risk mappings, a probabilistic setting is introduced based on novel concepts of regular conditional risk mapping and Markov update rule. To accommodate model ambiguity we introduce appropriate notions of history-consistent updating and of transition consistency for risk mappings on nested probability spaces.
Following disturbances to a power system triggering emergency responses such as protection or load/generation shedding, several factors affect the way in which these responses may cascade through the network. Beyond deterministic factors such as network topology, in this paper we aim to quantify the effect of correlations in power disturbances. These arise, for example, from common weather patterns causing correlated forecast errors in renewable generation. Our results suggest that for highly connected networks, the cascade size distribution is bimodal and positively correlated disturbances have the benefit of reducing cascade size. For a fixed network the latter relationship is observed to be stronger when emergency responses are rare, which is consistent with the mathematical theory of large deviations.
In this paper we study continuous-time two-player zero-sum optimal switching games on a finite horizon. Using the theory of doubly reflected BSDEs with interconnected barriers, we show that this game has a value and an equilibrium in the players' switching controls.
In this paper we provide a complete theoretical analysis of a two-dimensional degenerate nonconvex singular stochastic control problem. The optimisation is motivated by a storage-consumption model in an electricity market, and features a stochastic real-valued spot price modelled by Brownian motion. We find analytical expressions for the value function, the optimal control, and the boundaries of the action and inaction regions. The optimal policy is characterised in terms of two monotone and discontinuous repelling free boundaries, although part of one boundary is constant and the smooth fit condition holds there.
As decarbonisation progresses and conventional thermal generation gradually gives way to other technologies including intermittent renewables, there is an increasing requirement for system balancing from new and also fast-acting sources such as battery storage. In the deregulated context, this raises questions of market design and operational optimisation. In this paper, we assess the real option value of an arrangement under which an autonomous energy-limited storage unit sells incremental balancing reserve. The arrangement is akin to a perpetual American swing put option with random refraction times, where a single incremental balancing reserve action is sold at each exercise. The power used is bought in an energy imbalance market (EIM), whose price we take as a general regular one-dimensional diffusion. The storage operator's strategy and its real option value are derived in this framework by solving the twin timing problems of when to buy power and when to sell reserve. Our results are illustrated with an operational and economic analysis using data from the German Amprion EIM.
This paper studies the optimal control of a commercial building's thermostatic load during off-peak hours as an ancillary service to the power grid. It provides an algorithmic framework that commercial buildings can implement to cost-effectively increase their electricity demand at night while they are unoccupied, instead of using standard inflexible setpoint control. Consequently, there is minimal or no impact on user comfort, while the building manager gains an additional income stream from providing the ancillary service. By introducing a novel benefit-cost ratio of ancillary service payment to night-time price of electricity, we are able to study the building's capability to provide a service that is both useful to the power grid and profitable to the building manager. Numerical results show that there can be an economic incentive to participate even if the payment rate for the ancillary service is less than the price of electricity.
The frequency stability of power systems is increasingly challenged by various types of disturbance. In particular, the increasing penetration of renewable energy sources is increasing the variability of power generation while reducing system inertia against disturbances. In this paper we explore how this could give rise to rate of change of frequency (RoCoF) violations. Correlated and non-Gaussian power disturbances, such as may arise from renewable generation, have been shown to be significant in power system security analysis. We therefore introduce ghost sampling which, given any unconditional distribution of disturbances, efficiently produces samples conditional on a violation occurring. Our goal is to address questions such as "which generator is most likely to be disconnected due to a RoCoF violation?" or "what is the probability of having simultaneous RoCoF violations, given that a violation occurs?"