
Quantile-parameterized distributions (QPDs) are widely used in decision analysis because expert judgments are naturally expressed in terms of quantiles. Existing QPD families such as the Metalog offer considerable flexibility but lack analytic monotonicity guarantees, requiring ex post numerical repair. This paper introduces the QFlex distribution, a new QPD system constructed entirely from monotone transformations of valid quantile functions. QFlex interleaves powers of exponential, reflected-exponential, and centered-uniform quantile bases, yielding a flexible expansion that remains monotone under simple coefficient conditions. We show that QFlex has a generically full-rank design matrix with universal full rank through order K = 6, can interpolate any finite set of quantile assessments, and converges to the target quantile function as the number of terms increases. When all coefficients are nonnegative, QFlex is strictly increasing and unimodal; additional modes arise only in a structured and controllable manner. A comprehensive comparison across approximately 3,500 Pearson-system distributions demonstrates that QFlex generally matches or exceeds the accuracy of the Metalog distribution at moderate orders while offering explicit analytic monotonicity conditions. A synthetic example illustrates QFlex’s practical advantage: valid fits can be enforced directly through simple coefficient constraints, avoiding the complex post-fit monotonicity repair required by the Metalog distribution.
The kidney exchange problem (KEP) determines a set of planned kidney transplants, an exchange plan, for a pool of nondirected donors and biologically incompatible patient-donor pairs (PDPs) that maximizes weighted transplant quantity. Exchange pool members with fewer compatible donors experience bias, being excluded from exchange plans more often. To reduce bias, KEP optimization considers fairness at the group level, which obscures individual differences, or at the individual level, which is resource intensive. To bridge the gap between these approaches, we develop a model, individually fair outcomes for the KEP (IFO-KEP), which uses a novel prioritization schema, to individually prioritize PDPs, based on their probability of being included in fairness-agnostic optimization, within a hierarchical optimization approach balancing planned transplant quantity and fairness. Analysis compares IFO-KEP to existing individual and group fairness KEP approaches in myopic contexts, using metrics related to exchange performance and measures of parity, and dynamic contexts, evaluating exchange pool dynamics over time. In myopic contexts, IFO-KEP generates KEP solutions that achieve a high degree of parity without decreasing expected utility. In dynamic testing, IFO-KEP achieves fairness without causing negative impacts on long-term exchange performance. IFO-KEP is a novel method to achieve individually fair KEP exchange plans and reduce the bias against PDPs with low access to a compatible donor, without high-resource requirements. Global growth in KEP exchange pools elevates the importance of our contributions.
Vaccination sites face the operational decision of determining the timing for notifying standby recipients ("jumpers") to claim doses that would otherwise expire. Such timing decisions arise broadly in health-security operations involving perishable medical resources and responses to emerging health threats. This study formulates and analyzes a simulation-based model to evaluate notification policies within a framework defined by dose supply and site throughput levels. The notification policy is represented as a threshold based on the fraction of a dose's remaining shelf life at which a jumper is alerted, and policy performance is measured using an objective function defined as the sum of average dose wait time and average requester wait time. We develop a sample-average approximation procedure to obtain performance bounds and optimality gaps, and subsequently relax key baseline assumptions through robustness analyses that examine time-varying requester arrivals modeled as a nonhomogeneous Poisson process, alternative jumper travel-time distributions, pooled jumper configurations, and alternative value weights capturing equity-efficiency preferences. Across the baseline and extended analyses, notification timing and system performance exhibit a nonmonotonic relationship. When supply is scarce, system performance varies little across notification thresholds, with later notification offering practical protection of priority access. Under balanced supply and demand, notifications near midshelf life typically perform well. When supply is abundant, higher-throughput sites benefit from earlier notification to reduce the risk of expiration. In many scenarios, a range of policies yields statistically indistinguishable performance, indicating robust near-best policies. The study provides a health decision analysis framework to help vaccination sites select notification thresholds tailored to their supply conditions and operational capacity, offering practical guidance for balancing equitable access with the efficient use of expiring healthcare resources.
Predictions guide important prevention responses, from treating patients in hospitals to pretreating roads before snowstorms. Recent advances in machine learning and artificial intelligence have accelerated improvements in prediction accuracy. However, it is unclear how these improvements reshape preventive strategies and resource allocation. We develop a framework for forecast-based prevention, extending canonical loss-prevention models to explicitly incorporate prediction-based information updates. Our theoretical analysis provides three key insights with practical implications. First, improved predictions shift prevention toward more intense but less frequent responses. Second, as predictions resolve more uncertainty, risk preferences matter less in determining optimal loss prevention, resulting in greater convergence of preventive strategies. Third, under identifiable conditions, average prevention spending may decline as prediction skill rises, especially for actions with elastic marginal benefits. These results highlight the importance of aligning preventive strategies and resource allocation with evolving prediction capabilities.
Roger Cooke demonstrates that, under fat-tailed uncertainty, aggregation of expert point forecasts can fail dramatically, challenging the validity of the wisdom-of-crowds approach. In certain circumstances—for example, in extreme natural hazards—disregarding fat-tailed risk distributions can create an acute danger to public safety. This commentary reflects on Cooke’s findings from the perspective of applied volcanology and decision support in high-stakes situations. Empirical evidence from structured expert judgment panels indicates that forecast uncertainties often exhibit extreme tail behavior, undermining the assumptions that justify averaging. In such settings, aggregation of full probability distributions—particularly when performance-weighted—offers a more robust framework. Drawing on experiences from volcanic crisis management and contrasting these with shortcomings observed in the UK COVID-19 response, this paper argues that effective decision analysis requires formal quantification of uncertainty, multidisciplinary integration, and mathematically coherent aggregation methods.
We develop a model of cybersecurity in a supply chain economy populated by multiple firms organized into two tiers of suppliers and retailers—and by cybercriminals. Suppliers and retailers form a network in which each supplier may be linked to several retailers. The length of each supplier-retailer edge represents the number of access points a cybercriminal can exploit to inflict damage; exploiting more access points brings the attacker “closer” to the firm and increases expected loss. The cybercriminal may be of two types, low cost or high cost, drawn by nature and unobserved by firms, which know only the distribution over types. Attackers allocate effort both to penetrating a firm’s defenses and to evading detection. We focus on ex ante prevention and detection choices and do not model false alarm responses as a separate decision stage because firms in our framework do not condition any actions on realized alarm signals. Accordingly, false alarms are absorbed into the effective cost of detection rather than modeled as a distinct decision. Within this structure, we derive equilibrium properties under four organizational settings: when firms act in isolation, when they share information vertically across tiers, when they share information horizontally within tiers, and when cybersecurity is coordinated by a central planner.
Identifying a supplier's capability in delivering high-quality consignments is critical for a manufacturer who might incur a noncontractable hidden cost of poor quality if the selected supplier is of low-capability type. A supplier's investment in quality improvement through defect reduction can be a useful signal. However, the manufacturer is misled by the signal if the effect of preventive investment in reducing the cost of quality (COQ) is ignored. We analyze a signaling game between a manufacturer and a supplier, factoring in the decrease in COQ through a reduction in appraisal and failure costs resulting from investment in quality improvement. If the manufacturer is a payoff maximizer, then a high-capability supplier lacks incentive to signal type. However, surplus sharing, induced by the manufacturer's inequity aversion, incentivizes the highcapability supplier to signal type through investment in defect prevention. Separating perfect Bayesian equilibrium (PBE) exists if the capability gap between types of suppliers, or the noncontractable hidden cost to the manufacturer, is not too small. The likelihood of separating PBE increases with an increase in the capability gap and the supplier's equitable share of the surplus. Our analysis helps manufacturers avoid the hidden cost of poor quality by identifying a high-capability-type supplier correctly, factoring in the effects of investment in the supplier's appraisal and failure costs. Finally, our numerical simulation provides operational guidelines to the manufacturer for designing an effective signaling game, while ensuring an equitable distribution of supply chain surplus.
Proper loss functions are used in decision analysis to elicit probability estimates that reflect an individual's beliefs. In this paper, we consider a decision framework in which the probability estimates come from a statistical or machine learning model. We regard the probability estimates themselves as random variables with respect to the underlying distribution of the data set from which the estimates were derived. We derive a decision-tailored proper loss function, and we show that the generalized variance (based on a generalized bias-variance decomposition that we also derive in the paper) of the probability estimates under this decision-tailored loss quantifies the uncertainty in the probability estimates in a way that is relevant to the decision problem. In particular, points with low decision-tailored generalized variance correspond to points whose optimal decisions are robust with respect to the distribution of the probability estimate at that point, whereas high decision-tailored generalized variance corresponds to higher uncertainty under the distribution of the probability estimate.
Although utility functions are a basic component of decision analysis, there are a variety of functional forms that can be used. For small decisions, the choice might not change the decision. But different utility functions can provide vastly different recommendations for large decisions. There are qualitative recommendations on which utility functional form to use. But there is no exact answer to how large uncertainties can be before modeling risk preferences or modeling risk preference as a function of wealth is required. By maximizing the certain equivalent error between different utility functions, this paper analyzes and provides guidance into how large uncertainties can become relative to a decision maker's wealth before risk aversion should be modeled and when risk aversion needs to be modeled as a function of wealth.
Point estimates plus margins of error communicate statistical information to nonstatisticians. Although descriptive of bell-shaped probability distributions, they can be misleading for uniform or bi-modal distributions. Cooke shows that forecast errors are often described by fat-tailed distributions. Although the fat-tailed Cauchy distribution is bell shaped, the Bayesian combination of multiple Cauchy forecasts can be bell shaped, multimodal, or have a flat maximum. When forecast variation is intermediate, the combined forecast will only be bell shaped when the number of forecasts is odd. In this case, point estimates with margins of error can be informative.
We study the effectiveness of judgmental forecast aggregation in environments in which structural breaks cause sudden shifts in the mean demand. In such scenarios, judgmental forecasts are often biased, depending on whether demand shifts upward or downward. We demonstrate that asymmetric trimming of judgmental forecasts can increase the likelihood of bracketing in the presence of structural breaks, whereas symmetric trimming performs better in stable environments. Given that the timing and directionality of structural breaks are typically unpredictable, we propose two forecast combination methods that select the appropriate trimming rule by dynamically adapting to structural changes observed. The break perception method draws on subjective judgments regarding the occurrence of a structural break to determine which trimming rule should be chosen. The past performance method, on the other hand, relies on historical forecast errors of various trimming rules as a criterion for selection. We test their performance using both artificially generated data and real-world data on the daily U.S. COVID-19 death toll. Our results highlight that these combination methods consistently outperform other aggregation approaches. We discuss the implications of our findings for managerial practice.
This is not primarily a methods paper, but rather a paper about the limitations of methods. Using a large data set of expert forecasts with realizations, it emerges that expert point forecasts vary widely, and the distribution of point forecast errors appears to be very fat tailed. This paper uses simulated data to show that when samples from fattailed distributions are averaged, the results may not converge: extreme "outliers" arise with a frequency sufficient to disrupt any trends toward (apparent) convergence. It also shows that real-world expert judgment data from a number of different domains, collected over a period of decades, exhibit such fat-tailed behavior as regards point forecasts. Methods based on the "wisdom of crowds" that aggregate point forecasts of multiple experts are unlikely to yield good performance in such cases, regardless of the precise method by which judgments are weighted and aggregated. Moreover, increasing the number of experts used in such studies does not improve performance but rather degrades it, as does increasing the "diversity" of expert panels. However, when experts provide probabilistic forecasts, both the number of experts and diversity are helpful for the panel's overall statistical accuracy, which in turn correlates positively with lower forecast error. The moral of the story is that point estimates elicited from multiple experts should not be aggregated. Decision analysts should instead rely on aggregating probability distributions elicited from experts rather than point estimates, as is already considered best practice. Weighting based on the experts' measured performance yields even better results.
AI-augmented decision-making (AIADM) aims to leverage the computational power of machine learning (ML) models to assist humans in their decision-making processes. In many such systems, especially for complex tasks like medical image classification, ML models are often trained on large datasets annotated by humans. Neglecting to account for human decision-making biases when constructing these labeled datasets can lead to biased datasets, and subsequently models trained on such datasets can inherit the biases. We propose a novel approach to developing AIADM systems that aims to overcome these challenges by harnessing human uncertainty. Our approach has three elements: we collect subjective judgments from human annotators, we calibrate those subjective judgments, and we use the recalibrated subjective judgments to create probabilistic (i.e., soft) labels, which the AI decision aid is then trained on. We evaluate our methods through two studies using data from DiagnosUs, a crowdsourcing platform for medical image annotation. Across multiple training datasets, we assess how our proposed methods impact three key properties of AI decision aids: accuracy, calibration, and alignment with human uncertainty. We refer to these properties as the AIADM tri-criteria. Our results show that ML models trained on recalibrated soft labels are more accurate and better aligned with expert judgments. We also observe a tradeoff between ML calibration and alignment with human uncertainty. These findings highlight the value of capturing and correcting human uncertainty in ML training data and the need to consider the tri-criteria when developing AI systems.
Classic notions of stochastic dominance have integer degrees. Recent studies have imposed distinct preference conditions for refinement, resulting in a range of fractionaldegree stochastic dominance rules. However, preference conditions are generally not mutually exclusive, making it challenging to establish a strict criterion for rule selection. To address this, we establish fractional-degree stochastic dominance rules based exclusively on invariance laws under a general condition that is applicable to all intermediate utility sets. This approach enables practitioners to rely solely on the mutual compatibility and exclusivity of invariance properties to compare and select the appropriate rules. We illustrate the usefulness of our approach through an application to the problem of mutual fund selection.
In this paper, we present an overview of solutions to real option valuation (ROV) problems, using a simple oil field development project as an example. We categorize the solution tools into two major groups: learning and planning. Least squares Monte Carlo (LSM) represents the planning method, whereas reinforcement learning (RL) represents the learning approach. We discuss the application of each method in detail to evaluate the effectiveness of RL, a state-of-the-art technique in machine learning, compared with conventional solution methods such as LSM in the context of ROV problems. RL, as a state-of-the-art method for sequential decision making (SDM), has demonstrated strong success in solving complex problems in finance, including trading, option pricing, and portfolio management, where decisions are continuous and adaptive. However, our findings suggest that whereas RL has the potential to handle ROV, it is often too sophisticated and unnecessary for the typical structure of ROV and managerial flexibility analyses, where simpler methods such as LSM are usually sufficient. Particularly given the common characteristics of problems in the ROV context, as exemplified by the case study here, the features that highlight RL's strengths and were key to its development are not present in the ROV context. Therefore, careful consideration is needed when choosing the appropriate solution method.
This paper proposes a general framework for assessing the quality of inputs to a decision analysis provided by large language models (LLMs). The paper provides two benchmark case studies that focus on alternatives, preferences, and uncertainties related to a decision and are used to illustrate the proposed framework using ChatGPT version 4.0. The analysis uses the proposed framework and the data obtained to compare the efficacy of decision inputs provided by crowdsourcing from a group and those obtained from LLMs, with the relevance of inputs determined independently by a panel. The results show that (i) panel judgements about the relevance of inputs exhibited high correlation to one another; (ii) human groups performed better on generating alternatives, with higher rates of relevant alternatives; (iii) LLMs performed better on generating uncertainties, with higher rates of relevant alternatives; and (iv) human groups and LLMs performed similarly on generating preferences. These findings repeated across both subsets of data. Direct questions to participants about which input source they preferred resulted in a slight edge for artificial intelligence inputs. Although the benchmark case studies used ChatGPT version 4.0, the general framework applies to any LLM.
In this paper, we propose a continuum of asymptotic stochastic dominance criteria as a novel rule for comparing and ranking prospects in long term decision-making contexts. The new criteria encompass the established asymptotic first-degree and seconddegree stochastic dominance criteria, and between the two integer degrees, they effectively characterize the preferences of decision makers who are predominantly risk averse but do not categorically dislike all forms of risk, that is, those whose utility functions exhibit local convexities. To this end, we first introduce the concept of asymptotic fractional degree stochastic dominance, and then derive its equivalent conditions under the assumption of lognormal distributions. Furthermore, to enhance the tractability of asymptotic fractional degree stochastic dominance, under an additional condition on decision makers' utility functions that marginal utilities are finite, we introduce a variant of asymptotic fractional degree stochastic dominance, referred to as operational asymptotic fractional degree stochastic dominance, and derive its corresponding equivalent distributional characterizations. The (operational) asymptotic fractional degree stochastic dominance offers a more comprehensive criterion for ranking prospects in long term decision-making contexts. This study elucidates how various constraints on marginal utilities of decision makers shape the equivalent distributional conditions of asymptotic stochastic dominance criteria. Empirical examples also illustrate that the newly proposed (operational) asymptotic fractional degree can be effectively utilized in practice.
Large projects are complex undertakings that often involve lengthy execution, sizable budget, and substantial uncertainty. A large project is typically managed as a single, holistic program. This study compares the execution of a large project as a one-shot project (LT) with its execution via a series of smaller, possibly shorter-term, projects (SQ). We develop analytical models to evaluate the LT and SQ approaches pointing to the fundamental role of flexibility in project management, under varying conditions of project complexity, environmental dynamics, resource constraints, and uncertainty. Our findings indicate that LT is favored when the environment is stable, and SQ is preferred when the project budget is very high, project integration costs are low, and the rate of reduction over time in the suitability of the project outcome to its intended requirement is high. In addition, we show how technological uncertainty and uncertainty about the importance of the project's outcome at the time of project completion affect the choice between LT and SQ. The analysis of several real-world case studies highlights the importance of flexible execution in large projects within dynamic environments.