In this article, we consider a class of stochastic control problems, which have been widely used in optimal foraging theory and financial modeling. The optimal state process has two distinct dynamics, characterized by two pairs of drift and diffusion coefficients, depending on whether it takes values bigger or smaller than a threshold value. Adopting a perturbation-type approach, we find an expression for the potential measure of the optimal state process. We then obtain an expression for the transition density of the optimal state process by inverting the associated Laplace transform. Properties, including the stationary distribution of the optimal state process, are discussed. Finally, an expression of the value function and the optimal control is given for such stochastic control problems. As an application, we transform the continuous-time two-armed bandit problem in finance into our control framework to derive the optimal investment strategy, which maximizes the probability of reaching a specified threshold.
In the study of probability theory,the limit theory holds a pivotal position.Under the axiomatic system of Kolmogorov,from the initial Bernoulli's law of large numbers to the general form of the central limit theorem,the assumption of independent and identically distributed(IID)for random variables is crucial in the development of probability theory.At this time,the central limit theorem and the corresponding normal distribution occupy a"central"position in probability theory.However,when random variables do not satisfy the conditions of independent and identical distribution,mathematicians have found that the asymptotic distribution is likely no longer a normal distribution.Therefore,exploring the limited distribution of the non-IID case has been an open question for more than two hundred years,and has attracted the attention of many scholars.In addition,from the perspective of the economic market,the Kolmogorov probability theory framework can well quantify the inherent laws of market operation,but it is very difficult to accurately depict the external impact of human behavior on the economic market by applying the Kolmogorov axiomatic system.For example,the three famous paradoxes in the economic indicate that the probability theory under the Kolmogorov axiomatic system has obvious limitations in practical application,and many uncertain phenomena cannot be accurately modeled using linear probability and linear expectations.This has promoted the development of probability from a linear to a nonlinear framework.With the rise of nonlinear probability,establishing generalized conclusions on the asymptotic distribution of the sum of random variables under the framework of nonlinear expectations and non-IID cases has become a possibility.In this paper,we mainly introduce the development history of limit theory under the classical framework,as well as the main results of laws of large numbers,central limit theorems,and laws of the iterated logarithm under the framework of nonlinear expectations.
The econometric theory and methods of financial risk are critical technical issues that consistently concern both the economic and mathematical communi- ties. Traditional risk measurement theories and methods are based on classical probability and statistics. However, frequent financial crises have demonstrated significant flaws in conventional risk measurement theories and methods, which cannot accurately depict the risks and uncertainties in incomplete financial mar- kets. Instead, the nonlinear expectation seems more suitable when depicting the uncertain phenomenon in economics and finance. This paper will introduce the latest developments in this field, including nonlinear expectation, the non-linear central limit theorems, nonlinear normal distributions, and limit theorems in the quantum realm.
This paper proposes a framework for the global optimization of a possibly multimodal continuous function in a bounded rectangular domain. We first show that global optimization is equivalent to an optimal (sampling) strategy formation in a two-armed decision model with known distributions, based on the strategic law of large numbers we establish. There are many optimal strategies in general. We show that a concrete strategy using the sign of the partial gradient of the unique solution to a parabolic partial differential equation (PDE) is asymptotically optimal. Motivated by these results, we propose a class of Strategic Monte Carlo Optimization (SMCO) algorithms, which uses a simple strategy that makes coordinate-wise two-armed decisions based on the signs of the partial gradient (or practically the first difference) of the objective function, without the need of solving PDEs. Under some sufficient conditions, we establish that our SMCO algorithm converges to a local optimizer from a single starting point, and to a global optimizer under a growing set of starting points. Extensive numerical studies demonstrate the suitability of our SMCO algorithms for global optimization well beyond the theoretical guarantees established herein. For a wide range of deterministic and random test functions with challenging landscapes (multimodal, nondifferentiable, discontinuous), our SMCO algorithms perform robustly well, even in high-dimensional ([Formula: see text]) settings. In fact, our algorithms outperform many state-of-the-art global optimizers, as well as local algorithms (with the same set of starting points as ours).
Traditional drug development poses significant financial and temporal costs, whereas drug repurposing emerges as a cost-effective and efficient alternative. As large-scale biological networks proliferate, computational drug repurposing has become feasible, yet accurately capturing intricate heterogeneous network structures remains a persistent challenge. To address this challenge, we introduced a novel approach, called DRQuantum: Drug Repurposing via Quantum walks. Unlike random walks, quantum walks dispense with independence and harness quantum entanglement to simultaneously explore multiple paths, enabling faster traversal of networks. Moreover, DRQuantum accounts for both the local and global network structures. In this study, we constructed a heterogeneous multi-layer network by integrating drug-drug, disease-disease and protein-protein interaction networks. We then employed quantum walks to learn low-dimensional feature representations of nodes in these heterogeneous networks, ultimately inferring candidate drugs for repurposing beyond their original indications. Consequently, we observed that DRQuantum outperforms traditional drug repurposing methods in terms of AUROC, AUPRC and accuracy. Additionally, case studies for several specific diseases further validate the practical utility of our proposed method.
This study addresses the problem of stable acoustic relay assignment in an underwater acoustic network. Unlike the objectives of most existing literature, two distinct objectives, namely classical stable arrangement and ambiguous stable arrangement, are considered. To achieve these stable arrangements, a laser chaos-based multi-processing learning (LC-ML) method is introduced to efficiently obtain high throughput and rapidly attain stability. In order to sufficiently explore the relay's decision-making, this method uses random numbers generated by laser chaos to learn the assignment of relays to multiple source nodes. This study finds that the laser chaos-based random number and multi-processing in the exchange process have a positive effect on higher throughput and strong adaptability with environmental changing over time. Meanwhile, ambiguous cognitions result in the stable configuration with less volatility compared to accurate ones. This provides a practical and useful method and can be the basis for relay selection in complex underwater environments.
The 1936 Mills Futurity slot machine had the feature that, if a player loses 10 times in a row, the 10 lost coins are returned. Ethier and Lee (2010) studied a generalized version of this machine, with 10 replaced by deterministic parameter J. They established the Parrondo effect for a hypothetical two-armed machine with the Futurity award. Specifically, arm A and arm B, played individually, are asymptotically fair, but when alternated randomly (the so-called random mixture strategy), the casino makes money in the long run. They also considered the nonrandom periodic pattern strategy for patterns with r As and s Bs (e.g., ABABB if r=2 and s=3). They established the Parrondo effect if r+s divides J, and conjectured it in four other situations, including the case J=2 with r >= 1 and s >= 1. We prove the conjecture in the latter case.
In this paper we propose the mathematical model for a vane with stochastic rotation. By backward stochastic differential equation and the Feynman–Kac formula, we obtain the representation of solution to the proposed convection–diffusion equation with symmetric initial condition through the transition probability density of stochastic differential equation. Moreover the explicit probability density functions of solution with t=1 and x=0 is obtained, which shows that the distribution of k<0 is bimodal distribution with double maximum. By comparing to the experiment, it is found that the angular velocity with stochastic bidirectional is the binormal distribution with k<0, thus our model can describe the rotation of a vane with stochastic rotation very well.
This paper proposes a unified framework for the global optimization of a continuous function in a bounded rectangular domain. Specifically, we show that: (1) under the optimal strategy for a two-armed decision model, the sample mean converges to a global optimizer under the Strategic Law of Large Numbers, and (2) a sign-based strategy built upon the solution of a parabolic PDE is asymptotically optimal. Motivated by this result, we propose a class of Strategic Monte Carlo Optimization (SMCO) algorithms, which uses a simple strategy that makes coordinate-wise two-armed decisions based on the signs of the partial gradient of the original function being optimized over (without the need of solving PDEs). While this simple strategy is not generally optimal, we show that it is sufficient for our SMCO algorithm to converge to local optimizer(s) from a single starting point, and to global optimizers under a growing set of starting points. Numerical studies demonstrate the suitability of our SMCO algorithms for global optimization, and illustrate the promise of our theoretical framework and practical approach. For a wide range of test functions with challenging optimization landscapes (including ReLU neural networks with square and hinge loss), our SMCO algorithms converge to the global maximum accurately and robustly, using only a small set of starting points (at most 100 for dimensions up to 1000) and a small maximum number of iterations (200). In fact, our algorithms outperform many state-of-the-art global optimizers, as well as local algorithms augmented with the same set of starting points as ours.
We consider a class of stochastic control problems which has been widely used in optimal foraging theory. The state processes have two distinct dynamics, characterized by two pairs of drift and diffusion coefficients, depending on whether it takes values bigger or smaller than a threshold value. Adopting a perturbation type approach, we find an expression for potential measure of the optimal state process. We then obtain an expression for the transition density of the optimal state process by inverting the associated Laplace transform. Properties including the stationary distribution of the optimal state process are discussed. Finally, the expression of the value function is given for this class stochastic control problems.
In the classical two-armed bandit (TAB) problems, the goal of maximizing the total return specifies the optimal decision rule and fosters the development of abundant practicable methods, such as &greedy and upper confidence bound for TAB procedures. However, the existence of an arched reward function for the total return breaks the previously updated algorithms, like traditional myopic strategy, and prompts the generation of novel computation and theory for maximizing its expectation. Here we develop a special class of Bayesian two-armed bandit problems by imposing a prior probability on which arm has a greater expectation of the returns and propose a myopic strategy to specify the decision rule. We prove that the proposed myopic strategy is optimal under the structure of the arched reward function by exploring the dynamic programming principle. This article also demonstrates that arched functions are extensions of the monotone functions including the classical linear functions. Meanwhile, we establish a corresponding law of large numbers for our myopic strategy. Simulation studies pose supportive evidence that the newly proposed strategy performs well.
This paper studies a sequential decision problem where payoff distributions are known and where the riskiness of payoffs matters. Equivalently, it studies sequential choice from a repeated set of independent lotteries. The decision-maker is assumed to pursue strategies that are approximately optimal for large horizons. By exploiting the tractability afforded by asymptotics, conditions are derived characterizing when specialization in one action or lottery throughout is asymptotically optimal and when optimality requires intertemporal diversification. The key is the constancy or variability of risk attitude, that is, the decision-maker’s risk/reward tradeoff.
Prospective application of quantum technologies for reinforcement learning (RL) is an exciting surge in quantum fields. Quantum random number generators (QRNGs) produce high‐frequency random bits, with advantages of true randomness and device independence over pseudo‐random numbers. This work explores a new approach, called the quantum random numbers for multi‐armed bandit (QRN‐MAB) algorithm, for multi‐user access in wireless communication system upon random bits. The primary objective of the algorithm is to attain a stable assignment state without further exchange. QRN‐MAB utilizes random bits to learn channel features and conducts concurrent exchange to attain stability. This work finds that the intrinsic randomness property used in QRN‐MAB enables arrangement to maintain high accuracy and outperforms other classical algorithms. Additionally, the algorithm exhibits strong adaptability when the environment changes over time and quantum random numbers are advantageous over other pseudo‐random methods in achieving the target. This work provides an effective way for quantum technologies applications in RL and unfolds a promising avenue to stabilize assignment among multiple users.
AbstractThe myopic strategy is one of the most important strategies when studying bandit problems. In 2018, Nouiehed and Ross put forward a conjecture about Feldman’s bandit problem (J. Appl. Prob. (2018) 55, 318–324). They proposed that for Bernoulli two-armed bandit problems, the myopic strategy stochastically maximizes the number of wins. In this paper we consider the two-armed bandit problem with more general distributions and utility functions. We confirm this conjecture by proving a stronger result: if the agent playing the bandit has a general utility function, the myopic strategy is still optimal if and only if this utility function satisfies reasonable conditions.
BACKGROUND:Breast cancer (BC) risk-stratification tools for Asian women that are highly accurate and can provide improved interpretation ability are lacking. We aimed to develop risk-stratification models to predict long- and short-term BC risk among Chinese women and to simultaneously rank potential non-experimental risk factors. METHODS:The Breast Cancer Cohort Study in Chinese Women, a large ongoing prospective dynamic cohort study, includes 122,058 women aged 25-70 years old from the eastern part of China. We developed multiple machine-learning risk prediction models using parametric models (penalized logistic regression, bootstrap, and ensemble learning), which were the short-term ensemble penalized logistic regression (EPLR) risk prediction model and the ensemble penalized long-term (EPLT) risk prediction model to estimate BC risk. The models were assessed based on calibration and discrimination, and following this assessment, they were externally validated in new study participants from 2017 to 2020. RESULTS:The AUC values of the short-term EPLR risk prediction model were 0.800 for the internal validation and 0.751 for the external validation set. For the long-term EPLT risk prediction model, the area under the receiver operating characteristic curve was 0.692 and 0.760 in internal and external validations, respectively. The net reclassification improvement index of the EPLT relative to the Gail and the Han Chinese Breast Cancer Prediction Model (HCBCP) models for external validation was 0.193 and 0.233, respectively, indicating that the EPLT model has higher classification accuracy. CONCLUSIONS:We developed the EPLR and EPLT models to screen populations with a high risk of developing BC. These can serve as useful tools to aid in risk-stratified screening and BC prevention.
This paper applies the method of backward stochastic differential equations (BSDEs) to study the bang–bang optimal stochastic control problem, where the optimal control is of the feedback form and the final cost functional is given by a symmetric function. In addition to obtaining the existence of the optimal control, we also give the explicit representation of the optimal control and the optimal value function of the stochastic control problem by the explicit solution of nonlinear BSDEs with symmetric terminal condition.
As laser chaos has been proven to be a robust tool to solve the multi-armed bandit (MAB) problem, this study investigates the problem of multiuser dynamic channel assignment using laser chaos in cognitive radio networks with K -orthogonal channels and M secondary users. A novel dynamic channel assignment algorithm with laser chaos series for multiple users, named parallel processing learning with laser chaos (PPL-LC) algorithm, is proposed to efficiently address two main objectives: stable channel assignment and fuzzy stable channel assignment. The latter objective accounts for the realistic scenario where users have fuzzy preferences and do not necessarily pursue the best preference. The PPL-LC algorithm uses the randomness properties of laser chaos to learn the assignment of channels to multiple users without any limitations on the number of channels, which has not been considered in existing laser chaos algorithms. Moreover, the PPL-LC is equipped with parallel processing channel selections, resulting in higher throughput and stronger adaptability with environmental changes over time than comparison algorithms, such as distributed stable strategy learning and coordinated stable marriage MAB algorithms. Finally, numerical examples are presented to demonstrate the performance of the PPL-LC algorithm.
This paper studies a multi-armed bandit problem where the decision-maker is loss averse, in particular she is risk averse in the domain of gains and risk loving in the domain of losses. The focus is on large horizons. Consequences of loss aversion for asymptotic (large horizon) properties are derived in a number of analytical results. The analysis is based on a new central limit theorem for a set of measures under which conditional variances can vary in a largely unstructured history-dependent way subject only to the restriction that they lie in a fixed interval.
In this paper we consider the continuous time two-armed bandit (TAB) problem where the slot machine has two different arms in the sense that the two arms have different expected rewards and variances. We explore the optimal distribution of rewards for two-armed bandit problems, and obtain the explicit distribution function as well as the searching rules of optimal strategy. As a by-product, we find two new counter-intuitive phenomena in nonlinear probability framework (optimal strategic framework). The first is that the combination of losing arm and winning arm can make the winning arm achieve a greater coverage probability to win expected reward, which is also referred to "good + bad = better". The discovery implies that the traditional advice of always pursuing the arm with larger expected reward (i.e., stay on a winner rule) is not optimal in the probability framework. The second is that the combination sequence out of two independent and normal distribution-based arms is not normally distributed if the two arms are different, which is straightforward understood as "mutually independent normal + nor mal = unnormal". Furthermore, we provide the optimal sequential strategy to construct the "combination" arm and numerically examine the underlying mechanism. (c) 2022 Elsevier B.V. All rights reserved.