A distinctive aspect of human intelligence lies in its capacity to optimize decision-making in interactive settings based on the potential actions of others, a phenomenon known as strategic sophistication. This ability is pivotal in strategic interactions, where predicting others' actions, anticipating their beliefs about our possible actions, and making decisions based on such beliefs are essential. This competence holds particular relevance in negotiation, where achieving a compromise between demand and offer necessitates a thorough understanding of the counterpart's true objectives and the compromises they would be willing to accept. Despite extensive behavioral research revealing human behavioral flaws in strategic interactions, there's limited knowledge about the differences in strategic thinking abilities between expert negotiators and those not primarily involved in negotiation within their professional domain. Additionally, there's little understanding of the potential for enhancing strategic sophistication in expert negotiators and any variations in their ability to learn strategic behavior compared to individuals inexperienced in negotiation. In this innovative study, both non-expert individuals and expert negotiators are given the opportunity to learn through a specially designed app aimed at improving decision-making skills in complex strategic environments. Using a between-subject design, we compare the learning achieved by expert negotiators and non-expert individuals after extensive training with this app. Analysis of choice and belief elicitation data reveals that practicing with this app significantly increases participants' strategic abilities. Crucially, this improvement is more notable among expert negotiators, indicating a greater adaptability to feedback received during the training phase. Our findings offer new insights into the factors influencing learning in strategic interactions and shed light on differences in learning between expert negotiators and individuals with no experience in negotiation. Finally, our results emphasize the app's effectiveness in enhancing strategic skills across individuals, irrespective of their initial abilities.
Given the ubiquity of exploration in everyday life, researchers from many disciplines have developed methods to measure exploratory behaviour. There are therefore many ways to quantify and measure exploration. However, it remains unclear whether the different measures (i) have convergent validity relative to one another, (ii) capture a domain general tendency, and (iii) capture a tendency that is stable across time. In a sample of 678 participants, we found very little evidence of convergent validity for the behavioural measures (Hypothesis 1); most of the behavioural measures lacked sufficient convergent validity with one another or with the self-reports. In psychometric modelling analyses, we could not identify a good fitting model with an assumed general tendency to explore (Hypothesis 2); the best fitting model suggested that the different behavioural measures capture behaviours that are specific to the tasks. In a subsample of 254 participants who completed the study a second time, we found that the measures had stability across an 1 month timespan (Hypothesis 3). Therefore, although there were stable individual differences in how people approached each task across time, there was no generalizability across tasks, and drawing broad conclusions about exploratory behaviour from studies using these tasks may be problematic. The Stage 1 protocol for this Registered Report was accepted in principle on 2nd December 2022 https://doi.org/10.6084/m9.figshare.21717407.v1. The protocol, as accepted by the journal, can be found at https://doi.org/10.17605/OSF.IO/64QJU. Exploratory behaviours involve a trade-off between exploration and exploitation. Here the authors investigate exploration behaviour across different domains and whether tendency to explore is stable over time.
Increased strategic flexibility and exploration have been long promoted as essential for survival and profitability in highly dynamic environments. Nonetheless, managers seem to frequently defy this idea when making business decisions by frequently committing to various strategic decisions despite environmental dynamism. We investigate if there are types of dynamic environments in which constraining decision-making through commitment - understood as sticking to a choice for a given time horizon, without the possibility of changing it within that time span - is beneficial for long-term performance. Results from a computational study show that dynamic environments characterized by low magnitude of change favor constrained decisions. The findings contribute to the literature on individual and organizational decision making in dynamic environments and set the stage for empirical research on this topic.
Research in intertemporal decisions shows that people value future gains less than equivalent but immediate gains by a factor known as the discount rate (i.e., people want a premium for waiting to receive a reward). A robust phenomenon in intertemporal decisions is the finding that the discount rate is larger for small gains than for large gains, termed the magnitude effect. However, the psychological underpinnings of this effect are not yet fully understood. One explanation proposes that intertemporal choices are driven by comparisons of features of the present and future choice options (e.g., information on rewards). According to this explanation, the hypothesis is that the magnitude effect is stronger when the absolute difference between present and future rewards is emphasized, compared to when their relative difference is emphasized. However, this hypothesis has only been tested using one task (the two-choice paradigm) and only for gains (i.e., not losses). It's therefore unclear whether the findings that support the hypothesis can be generalized to different methodological paradigms (e.g., preference matching) and to the domain of losses. To address this question, we conducted experiments using the preference-matching method whereby the premium amounts that people could ask for were framed in terms of either currencies (emphasizing absolute differences) or percentages (emphasizing relative differences). We thus tested the robustness of the evidence in support of the hypothesis that percent framing, relative to currency framing, attenuates the magnitude effect in the domain of gains (Studies 1, 2, and 3) and in the domain of losses (Study 1, 3, and 4). The data were heavily skewed and the assumption of equal variances was violated. Therefore, in place of parametric statistical tests, we calculated and interpreted parametric and nonparametric standardized and unstandardized effect size estimates and their confidence intervals. Overall, the results support the hypothesis.
We propose an experimental eye-tracking study to test how strategic sophistication is shaped by experience in 3x3 two-person normal-form games. Although strategic sophistication has been shown to be linked to a variety of endogenous and exogenous factors, little is known about how it is affected by previous interactive decisions. We show that complete feedback in previous games can significantly enhance strategic sophistication, and that games that in principle provide equivalent learning opportunities lead instead to substantially different learning outcomes. Specifically, only repeated play with feedback of games that emphasize strategic interdependence significantly enhances strategic learning, producing an increase in the frequency of equilibrium play and a shift of attention to the incentives of the counterpart. Moreover, we find that the type of learning underlying newly gained strategic skills can vary substantially across players. Whereas some players eventually learn to visually analyze the payoff matrix consistently with equilibrium reasoning, others appear to use experience with previous interactions to devise simple heuristics of play. Our results have implications for theoretical and computational modeling of learning. (C) 2021 Elsevier Inc. All rights reserved.
Is there a general tendency to explore that connects search behaviour across different domains? Although the experimental evidence collected so far suggests an affirmative answer, this fundamental question about human behaviour remains open. A feasible way to test the domain-generality hypothesis is that of testing the so-called priming hypothesis: priming explorative behaviour in one domain should subsequently influence explorative behaviour in another domain. However, only a limited number of studies have experimentally tested this priming hypothesis, and the evidence is mixed. We tested the priming hypothesis in a registered report. We manipulated explorative behaviour in a spatial search task by randomly allocating people to search environments with resources that were either clustered together or dispersedly distributed. We hypothesized that, in a subsequent anagram task, participants who searched in clustered spatial environments would search for words in a more clustered way than participants who searched in the dispersed spatial environments. The pre-registered hypothesis was not supported. An equivalence test showed that the difference between conditions was smaller than the smallest effect size of interest (d = 0.36). Out of several exploratory analyses, we found only one inferential result in favour of priming. We discuss implications of these findings for the theory and propose future tests of the hypothesis.
A robust phenomenon in intertemporal decisions—the magnitude effect—shows that people value future gains less than equivalent but immediate gains by a factor known as the discount rate (i.e., people want a premium for waiting to receive a reward). However, the psychological underpinnings of this effect are not yet fully understood. One explanation proposes that intertemporal choices are driven by comparisons of features of the present and future choice options (e.g., information on rewards). According to this explanation, the hypothesis is that the magnitude effect is stronger when the absolute difference between present and future rewards is emphasized, compared to when their relative difference is emphasized. However, this hypothesis has only been tested using one task (the two-choice paradigm) and only for gains (i.e., not losses). It’s therefore unclear whether the findings that support the hypothesis can be generalized to different methodological paradigms (e.g., preference matching) and to the domain of losses. To address this question, we conducted experiments using the preference-matching method whereby the premium amounts that people could ask for were framed in terms of either currencies (emphasizing absolute differences) or percentages (emphasizing relative differences). We thus tested the robustness of the evidence in support of the hypothesis that percent framing, relative to currency framing, attenuates the magnitude effect in the domain of gains (Studies 1, 2, and 3) and in the domain of losses (Study 1, 3, and 4). Study 5 ruled out floor effects as an alternative explanation for the results in the losses domain. Overall, the results support the hypothesis.
Human organizations are commonly characterized by a hierarchical chain of command that facilitates division of labor and integration of effort. Higher-level employees set the strategic frame that constrains lower-level employees who carry out the detailed operations serving to implement the strategy. Typically, strategy and operational decisions are carried out by different individuals that act over different timescales and rely on different kinds of information. We hypothesize that when such decision processes are hierarchically distributed among different individuals, they produce highly heterogeneous and strongly path-dependent joint learning dynamics. To investigate this, we design laboratory experiments of human dyads facing repeated joint tasks, in which one individual is assigned the role of carrying out strategy decisions and the other operational ones. The experimental behavior generates a puzzling bimodal performance distribution–some pairs learn, some fail to learn after a few periods. We also develop a computational model that mirrors the experimental settings and predicts the heterogeneity of performance by human dyads. Comparison of experimental and simulation data suggests that self-reinforcing dynamics arising from initial choices are sufficient to explain the performance heterogeneity observed experimentally.
Previous research suggests that human reaction to risky opportunities reflects two contradicting biases: “loss aversion”, and “limited level of reasoning” that leads to overconfidence. Rejection of attractive gambles is explained by loss aversion, while counterproductive risk seeking is attributed to limited level of reasoning. The current research highlights a shortcoming of this popular (but often implicit) “contradicting biases” assertion. Studies of “negative-sum betting games” reveal high rate of counterproductive betting even when limited level of reasoning and loss aversion imply no betting. The results reflect two reasons for the high betting rate: initial tendency to participate and slow learning. Under certain conditions, the observed betting rate was higher than the rate predicted under random choice even after 250 trials with immediate feedback. These results can be captured with a model that assumes a tendency to select strategies that have led to good outcomes in a small set of similar past experiences, and allows for an initial framing effect.
Review of previous research highlights 2 pairs of inconsistent reactions to rare events: (a) studies of probability judgment reveal conservatism (that implies overestimation of rare events) and overconfidence (that implies underestimation of rare events); (b) studies of choice behavior reveal overwe
Three studies are presented that compare decisions from experience in Denmark, Israel, and Taiwan. They focus on two change-related cultural differences suggested by previous research on dialectical vs. analytic approach to thinking. The first implies that East Asians are more likely to change their behavior over time (i.e., are less consistent), the second that they expect more changes in the environment. The results show that the "less consistency in the East" hypothesis has a high predictive value. This hypothesis accurately predicts a behavioral pattern that was documented in all three studies, as well as a non-trivial effect of limited feedback in Study 3: When feedback was limited to the obtained payoff, the participants from Taiwan exhibited less risk aversion than the Israeli. Analysis of the "expecting more changes in the East" hypothesis reveals mixed results. This hypothesis was supported in Study 2, which examined relatively complex multi-alternative multi-outcome tasks, but not in Studies 1 and 3, which examined simple two-alternative two-outcome choice tasks. A possible explanation for the different predictive value of the two examined hypotheses is discussed. (C) 2015 Elsevier B.V. All rights reserved.
OPINION article Front. Psychol., 17 February 2015Sec. Decision Neuroscience Volume 6 - 2015 | https://doi.org/10.3389/fpsyg.2015.00159
This paper revisits a recent study by Posen and Levinthal (2012) on the exploration/exploitation tradeoff for a multi-armed bandit problem, where the reward probabilities undergo random shocks. We show that their analysis suffers two shortcomings: it assumes that learning is based on stale evidence, and it overlooks the steady state. We let the learning rule endogenously discard stale evidence, and we perform the long run analyses. The comparative study demonstrates that some of their conclusions must be qualified.
Previous research documents two pairs of inconsistent reactions to rare events: 1) Studies of probability judgment reveal conservatism which implies overestimation of rare events, and overconfidence which implies underestimation of rare events. 2) Studies of choice behavior reveal overweighting of rare events in one-shot tasks, and the opposite bias in decisions from experience. The current analysis and experimental results demonstrate that the coexistence and relative importance of the four biases can be captured with simple models that share the assumption that judgments and decisions are made based on the information conveyed by small and noisy samples of past experiences.
This paper examines the effect of default options on choice behavior in experience-based decisions. To this end, we designed the "radio-button" experimental paradigm, in which participants are asked to set default options that remain effective until they decide to change them, and the outcomes from active and inactive options are determined and presented to participants every two seconds. Comparison of behavior in six basic decision problems run under the "clicking" and "radio-button" experimental paradigms reveals an unexpected result. We find that, although the basic properties of decisions from experience are robust to the option to rely on self-set defaults, the introduction of defaults reduces the tendency to prefer the status quo. We dub this behavioral pattern as the "do something" effect. (C) 2012 Elsevier B.V. All rights reserved.
The contribution by Yechiam and Telpaz (Y&T) published in Frontiers in Cognitive Science places it in a corpus of literature which bridges at least three different disciplines, i.e., psychology, economics, and neuroscience. The goal of this line of research is to explore the neurological and physiological underpinnings of one of the central topics in judgment and decision-making (JDM) research – choice behavior in decisions from experience. Y&T successfully contributes to this goal by demonstrating a novel effect that losses increase experimental participants’ arousal as measured by pupil dilatation, which in turn positively correlates with a risk aversion behavior. They hypothesize that participants’ attention is increased in decision problems involving losses, which trigger an innate prudent behavior in situations entailing danger and/or hazard. Interestingly, Y&T find that the nature of attention is not selective, i.e., when losses are present, participants are shown to devote more attention to the task as a whole rather than to the single negative outcomes, in contrast to Prospect Theory's loss aversion. YT Tom et al., 2007). These studies suggest that behavioral loss aversion in decisions from description reflects an asymmetric response to gain and losses in the neural system encoding for reward values (the ventromedial prefrontal cortex, orbitofrontal cortex, and ventral striatum). What makes Y&T's contribution particularly noteworthy is their mediating attentional hypothesis, which links physiological mechanisms to the psychological processes involved in experience-based decisions. One of the possible future developments from YT Holt and Laury, 2005). Specifically, it has been observed that participants’ degree of risk aversion increases significantly as actual positive payoffs are scaled up, and that this effect is negligible when payoffs are hypothetical. These findings provide an opportunity to widen the scope of the attentional hypothesis. Specifically, payoffs corresponding to large cash amounts might have the analogous effects of losses of increasing arousal and of triggering a higher level of risk aversion; whereas hypothetical payoffs might result in a substantial inhibition of attention. Therefore, the motivation implied by real stakes can be interpreted as one of the possible boundary conditions (see below) for Y&T's attentional hypothesis, giving rise to a question of the relative weight of attention and motivation in shaping risk attitudes. Y&T's report can also be contextualized within the wide literature on individual differences in reasoning, judgment and decision making (e.g., Stanovich and West, 2000) and their implications to the rationality debate. The prototypical finding in that literature is the correlation between cognitive ability and normative responding, with a strong emphasis on normative evaluation of rationality. This so-called “normativist” approach has recently been subject to criticism (Elqayam and Evans, 2011) as unhelpful in developing a psychological theory of human rationality. It is therefore noteworthy that Y&T take their individual differences work in a completely different direction, with what seems to be a purely ‘descriptivist’ approach, with no normativist connotations. As one reviewer of this manuscript put it, any behavior in this setting could be justified as ‘rational’. The behavioral patterns described vary qualitatively rather than quantitatively. This is typical of descriptivist approaches to cognitive variability higher mental processing (Evans and Elqayam, 2011). Given the dearth of such focus in higher mental processing, this is a welcome development. Lastly, a potentially significant issue here is the implications to risk aversion as originally portrayed in prospect theory (Kahneman and Tversky, 1979). One could argue that Y&T contribute to defining boundary conditions for Prospect Theory, by proposing an alternative explanation for specific settings in which Prospect Theory is not supported by empirical evidence1. Indeed, as a unified theory of risk aversion is not yet at hand, knowing the range of application of each of the existing theories is crucial. One reason that YT the algorithmic level, which has to do with processes (e.g., the calculator's software); and the implementational level, which explores the physical underpinnings of the system – its hardware/wetware characterization (e.g., the calculator's chip). Viewed in these terms, we see prospect theory as portraying behavior mainly on the computational (i.e., functional) level of analysis; or, as some authors put it – an “axiomatic” system (see Wakker, 2010). In contrast, YT Tom et al., 2007), and studies that combine several levels of analysis, as Y&T have done, are even rarer. This makes Y&T's contribution of particular interest to scholars of human thinking and decision making.