This study considers the practically important case of nonparametrically estimating heterogeneous average treatment effects that vary with a limited number of discrete and continuous covariates in a selection-on-observables framework where the number of possible confounders is very large. We propose a two-step estimator for which the first step is estimated by machine learning. We show that this estimator has desirable statistical properties such as consistency, asymptotic normality, and rate double robustness. In particular, we derive the coupled convergence conditions between the nonparametric and the machine-learning steps. We also show that estimating population average treatment effects by averaging the estimated heterogeneous effects is semiparametrically efficient. The resulting estimators are compared to other suggestions in the literature in Monte Carlo experiments that are inspired by real data. They are found to perform relatively better in most settings. The new estimators are applied to the empirical example of the effects of mothers' smoking during pregnancy on the birthweight of their babies.
It is valuable for any decision maker to know the impact of decisions (treatments) on average and for subgroups. The causal machine learning literature has recently provided tools for estimating group average treatment effects (GATE) to better describe treatment heterogeneity. This article addresses the challenge of interpreting such differences in treatment effects between groups while accounting for variations in other covariates. We propose a new parameter, the balanced group average treatment effect (BGATE), which measures a GATE with a specific distribution of a priori-determined covariates. By taking the difference between two BGATEs, we can analyze heterogeneity more meaningfully than by comparing two GATEs, as we can separate the difference due to the different distributions of other variables and the difference due to the variable of interest. The main estimation strategy for this parameter is based on double/debiased machine learning for discrete treatments in an unconfoundedness setting, and the estimator is shown to be N-consistent and asymptotically normal under standard conditions. We propose two additional estimation strategies: automatic debiased machine learning and a specific reweighting procedure. Last, we demonstrate the usefulness of these parameters in a small-scale simulation study and in an empirical example.
Active labor market programs are important instruments used by European employment agencies to help the unemployed find work. Investigating large administrative data on German long-term unemployed persons, we analyze the effectiveness of three job search assistance and training programs using Causal Machine Learning. Participants benefit from quickly realizing and long-lasting positive effects across all programs, with placement services being the most effective. For women, we find differential effects in various characteristics. Especially, women benefit from better local labor market conditions. We propose more effective data-driven rules for allocating the unemployed to the respective labor market programs that could be employed by decision-makers.
Online dating emerged as a key instrument for human mating. This research investigates the effect of sports activity on human mating by exploiting a unique data set from an online dating platform. We leverage advances in causal machine learning to estimate the causal effect of sports frequency on contact chances. We find that for male users, sport on a weekly basis increases the probability of receiving a first message from another user by 50%, relative to not doing sport. For female users, we do not find evidence for such an effect. In addition, for male users, the effect increases with higher income.
Fairness and interpretability play an important role in the adoption of decision-making algorithms across many application domains. These requirements are intended to avoid undesirable group differences and to alleviate concerns related to transparency. This paper proposes a framework that integrates fairness and interpretability into algorithmic decision making by combining data transformation with policy trees, a class of interpretable policy functions. The approach is based on pre-processing the data to remove dependencies between sensitive attributes and decision-relevant features, followed by a tree-based optimization to obtain the policy. Since data pre-processing compromises interpretability, an additional transformation maps the parameters of the resulting tree back to the original feature space. This procedure enhances fairness by yielding policy allocations that are pairwise independent of sensitive attributes, without sacrificing interpretability. Using administrative data from Switzerland to analyze the allocation of unemployed individuals to active labor market programs (ALMP), the framework is shown to perform well in a realistic policy setting. Effects of integrating fairness and interpretability constraints are measured through the change in expected employment outcomes. The results indicate that, for this particular application, fairness can be substantially improved at relatively low cost.
Active labor market policies are widely used by the Swiss government, enrolling over half of all unemployed individuals. This paper evaluates the effectiveness of Swiss programs in improving employment and earnings outcomes using causal machine learning and rich administrative data on unemployed individuals in 2014 and 2015, including detailed labor market histories and other covariates. The findings for Swiss citizens and immigrants with permanent residency indicate a small positive average effect of a Temporary Wage Subsidy program on employment and earnings in the third year after program start. In contrast, Basic Courses, such as job application training, exhibit negative effects on both outcomes over the same period. No significant impacts are found for Employment Programs conducted outside the regular labor market or for Training Courses such as language or computer classes. The programs are most effective for individuals with a non-EU migration background, while Temporary Wage Subsidies also benefit those with lower educational attainment. Finally, shallow policy trees provide practical guidance for improving the targeting of program assignments.
This article shows how coworker performance affects individual performance evaluation in a teamwork setting at the workplace. We use high-quality data on football matches to measure an important component of individual performance, shooting performance, isolated from collaborative effects. Employing causal machine learning methods, we address the assortative matching of workers and estimate both average and heterogeneous effects. There is substantial evidence for spillover effects in performance evaluations. Coworker shooting performance, meaningfully impacts both, manager decisions and third-party expert evaluations of individual performance. Our results underscore the significant role coworkers play in shaping career advancements and highlight a complementary channel, to productivity gains and learning effects, how coworkers impact career advancement. We characterize the groups of workers that are most and least affected by spillover effects and show that spillover effects are reference point dependent. While positive deviations from a reference point create positive spillover effects, negative deviations are not harmful for coworkers.
Decision making plays a pivotal role in shaping outcomes across various disciplines, such as medicine, economics, and business. This paper provides practitioners with guidance on implementing a decision tree designed to optimise treatment assignment policies through an interpretable and non-parametric algorithm. Building upon the method proposed by Zhou, Athey, and Wager (2023), our policy tree introduces three key innovations: a different approach to policy score calculation, the incorporation of constraints, and enhanced handling of categorical and continuous variables. These innovations enable the evaluation of a broader class of policy rules, all of which can be easily obtained using a single module. We showcase the effectiveness of our policy tree in managing multiple, discrete treatments using datasets from diverse fields. Additionally, the policy tree is implemented in the open-source Python package mcf (modified causal forest), facilitating its application in both randomised and observational research settings.
The mass production of bipolar plates for electrolyzers and fuel cells is a central step towards the realization of efficient and cost-effective energy systems of the future. However, current production processes are reaching their limits and can hardly realize the quantities that will soon be demanded, nor can they scale up to the required volumes. Particularly for the handling of half-plates and bipolar plates, major challenges are to be expected, especially with regard to production rates. Existing handling systems have restricted scalability and precision. Therefore, new stacking technologies are necessary, which have to be adaptable to the mechanical properties of the components and maintain tight tolerances during stacking to ensure hydrogen sealing for safety and efficiency. An important property in the handling of the plates is their limpness, which is distinguished by instability of the components as well as plastic deformation at low forces and moments. Therefore, the limp behavior of the components must be analyzed. To investigate the limpness of foil components, a flowfield is first formed using a 1.4404 stainless steel foil with a sheet thickness of 0.075 mm. Subsequently, the workpieces are analyzed in terms of their limp properties by means of a 3-point bending test.
In this paper we develop a new machine learning estimator for ordered choice models based on the random forest. The proposed Ordered Forest flexibly estimates the conditional choice probabilities while taking the ordering information explicitly into account. In addition to common machine learning estimators, it enables the estimation of marginal effects as well as conducting inference and thus provides the same output as classical econometric estimators. An extensive simulation study reveals a good predictive performance, particularly in settings with non-linearities and near-multicollinearity. An empirical application contrasts the estimation of marginal effects and their standard errors with an ordered logit model. A software implementation of the Ordered Forest is provided both in R and Python in the package orf available on CRAN and PyPI, respectively.
Increasing material costs, decreasing availability, and ever-higher demands on environmental compatibility and complexity require new strategies in the development and production of functional components. Consequently, a combined approach from the areas of design, material science, and manufacturing is mandatory, in order to meet the requirements. Reducing the number of parts, using lightweight materials and applying hybrid components with a multimaterial mix are possible solutions. Nevertheless, conventional joining operations like welding or riveting are reaching their limits in terms of material utilization, load-bearing capacity as well as versatility of the process. Thus, innovative and versatile joining by forming operations and process combinations are focus of current research. In this context, the innovative process of orbital forming had been investigated as a joining by forming operation to manufacture load-adapted hybrid functional components. By tilting of one tool component during the process, a radial material flow is generated, allowing the crimping of the two joining partners. Nevertheless, the load-bearing capacity in axial direction could be identified as limiting factor for a possible application. Therefore, the aim of this investigation is the development of a fundamental process understanding on the influence of a novel geometrical adaption of the joint on the resulting load bearing capacity. The influence of varying geometrical proportions of the joint on the quality is evaluated, considering the form filling, the geometrical properties of the components as well as the maximum transmittable axial load. As joining partners, the dual-phase steel DP600 and the aluminum alloy EN AW-5754 with a thickness of 2.0 mm are used.
We estimate the transmission of the pandemic shock in 2020 to the residential and commercial real estate market by causal machine learning, using granular data for Germany. We exploit differences in the incidence of Covid infections and short-time work at the municipal level for the identification of epidemiological and economic effects of the pandemic. We find that (i) a larger incidence of Covid infections temporarily reduced rents for retail real estate; (ii) a larger incidence of short-time work temporarily reduced rents of office real estate; (iii) the pandemic increased prices, particularly in the top price segment of commercial real estate.
Uncovering causal effects at various levels of granularity provides substantial value to decision makers. Comprehensive machine learning approaches to causal effect estimation allow to use a single causal machine learning approach for estimation and inference of causal mean effects for all levels of granularity. Focusing on selection-on-observables, this paper compares three such approaches, the modified causal forest (mcf), the generalized random forest (grf), and double machine learning (dml). It also provides proven theoretical guarantees for the mcf and compares the theoretical properties of the approaches. The findings indicate that dml-based methods excel for average treatment effects at the population level (ATE) and group level (GATE) with few groups, when selection into treatment is not too strong. However, for finer causal heterogeneity, explicitly outcome-centred forest-based approaches are superior. The mcf has three additional benefits: (i) It is the most robust estimator in cases when dml-based approaches underperform because of substantial selectivity; (ii) it is the best estimator for GATEs when the number of groups gets larger; and (iii), it is the only estimator that is internally consistent, in the sense that low-dimensional causal ATEs and GATEs are obtained as aggregates of finer-grained causal parameters.
We introduce and prove the validity of nonparametric bootstrap procedures for the approximation of the sampling distribution of pair or one-to-many propensity score matching estimators. Unlike the conventional bootstrap, the proposed bootstrap approach does not construct bootstrap samples by randomly resampling from the observations with uniform weights. Instead, it constructs the bootstrap approximation by randomly resampling from the martingale representation for matching estimators. Finally, we also conduct a simulation study in which the nonparametric bootstrap performs well even when the sample size is relatively small.
The application of forming operations instead of conventional cutting processes is a practical approach to increase material efficiency and functional integration, thus addressing lightweight design. In this context, the innovative process class of Sheet-Bulk Metal Forming is presented, which combines the advantage of bulk and sheet metal forming. One of the assigned processes for the manufacturing of functional components with different form elements is orbital forming. Within previous investigations, the major challenge could be identified as a control of the material flow. Since process parameters, like an increased forming force, are not sufficient to avoid process failures in form of wrinkles, other measures have to be taken into account. Especially when applying precipitation hardenable aluminum alloys, one possibility is the application of a local short-term heat treatment. By locally reversing the hardening effect of the precipitation clusters, the interaction between softened and still hard areas can be used to attain the desired material flow. Although this procedure is established for conventional sheet metal forming processes and was furthermore investigated for the orbital forming of components from the aluminum alloy EN AW-6016, the influence on the forming of high-strength aluminum from the 7xxx-series is still content of current research. Especially due to a reduced formability at room temperature, this alloy is predominantly formed at elevated temperatures. By introducing the established method of a local short-term heat treatment on 7xxx alloys, a contribution towards a possible forming at room temperature is made. Therefore, this work focuses on the development of a tailored heat treatment strategy to control the material flow during orbital forming. Specimens out of the high-strength aluminum alloy EN AW-7075 with a thickness of 2.0 mm in condition T6 are heat-treated and consequently cold formed. The results are quantified by a geometry-based analysis. The resulting strain distribution is used to verify the material flow proportions.
Based on administrative data of unemployed in Belgium, we estimate the labour market effects of three training programmes at various aggregation levels using Modified Causal Forests, a causal machine learning estimator. While all programmes have positive effects after the lock-in period, we find substantial heterogeneity across programmes and unemployed. Simulations show that "black-box" rules that reassign unemployed to programmes that maximise estimated individual gains can considerably improve effectiveness: up to 20% more (less) time spent in (un)employment within a 30 months window. A shallow policy tree delivers a simple rule that realizes about 70% of this gain.
We utilize data from 5,010 soccer games in the top two Swiss divisions between the 2005/06 and 2018/19 seasons. In these games, a referee can share the same linguistic area with one of the teams. Using referee-per-season fixed effects, we find that referees issue significantly more penalties, in the form of yellow cards, to teams that are not from the referee's linguistic area. We also find some evidence, in the highest level league only, that referees issue more red cards to teams that are not from their linguistic area and that away teams achieve fewer points when home teams share the same linguistic area with the referee. Our analyses suggest that referees' bias is likely to be subconscious and reflexive rather than being a deliberate act of discrimination.
Lightweight construction in modern car design leads to an increased usage of various aluminium semi-finished products. Besides sheet material, aluminium extrusion profiles are frequently used due to their high stiffness and variety of possible cross-sections. However, similar to sheet material, aluminium profiles exhibit limited formability in comparison to mild steel materials. One possibility to increase the forming limits of precipitation hardened aluminium alloys is the so-called Tailored Heat Treatment technology. By a local short-term heat treatment, the material is softened and the material flow can be controlled to reduce stresses in critical forming zones. The purposeful definition of the heat treatment zones is mandatory to improve the forming results. Therefore, numerical methods are necessary. In this investigation, a numerical process chain is presented. It combines the thermo-mechanical simulation of a local laser heat treatment with a subsequent bending process of the heat-treated profile using the alloy EN AW-6082. The temperature distribution, mechanical properties, and finally, the bending result of the numerical model are validated by experimental tests.
We estimate the transmission of the pandemic shock in 2020 to prices in the residential and commercial real estate market by causal machine learning, using new granular data at the municipal level for Germany. We exploit differences in the incidence of Covid infections or short-time work at the municipal level for identification. In contrast to evidence for other countries, we find that the pandemic had only temporary negative effects on rents for some real estate types and increased asset prices of real estate particularly in the top price segment of commercial real estate.