With reference to a binary outcome and a binary mediator, we derive identification bounds for natural effects under a reduced set of assumptions. Specifically, no assumptions about confounding are made that involve the outcome; we only assume no unobserved exposure-mediator confounding as well as a pair of conditions termed Partially Constant Cross-World Dependence (PC-CWD) and Logit Constancy (LC). These assumptions pose fewer constraints on the counterfactual probabilities than the set of assumptions they replace. The proposed strategy permits to achieve interval identification of the total effect, which is no longer point identified under the considered set of assumptions. Our derivations are based on postulating a logistic regression model for the mediator as well as for the outcome. However, in both cases the functional form governing the dependence on the explanatory variables is allowed to be arbitrary, thereby resulting in a semi-parametric approach. To account for sampling variability, we provide delta-method approximations of standard errors to build uncertainty intervals from identification bounds. The method is compared to some alternative ones in a simulation study and then applied to a dataset gathered from a Spanish prospective cohort study, with the aim to evaluate whether the effect of smoking on lung cancer risk is mediated by the onset of pulmonary emphysema.
Abstract: Two procedures are proposed to assess sensitivity to a binary unobserved confounder for a binary outcome when the causal effect is expressed as a (log) odds ratio, as commonly arises in standard logistic modelling, particularly in case-control studies. The methods are based on graphical tools that visualize the extent to which an unobserved confounder could attenuate, nullify, or even reverse the estimated causal effect. The second procedure relies on a single sensitivity parameter, yielding a naturally bounded assessment that can be translated into an objective measure. Connections with Cornfield's conditions on relative risks are presented, thereby enlarging the circumstances where the proposed procedures can be applied and establishing a link that allows extensions to polytomous confounders.
In statistical analysis, Cochran's formula plays a crucial role in disentangling the relationships between marginal and conditional regression coefficients. However, its results and implications are valid only within the linear case. Despite this, due to its simplicity and interpretability, practitioners often continue to use Cochran's formula also outside linear models. With reference to binary outcome models, we derived the approximated expression of the marginal regression coefficient when marginalization is performed over a continuous covariate and show that it mimics Cochran's formula under certain simplifying assumptions. We initially postulate a logistic link function and then show how it can be generalized. We then explore the implications of this formula in the context of sensitivity analysis and causal mediation analysis, thereby enlarging the number of circumstances where explicit parametric formulations can be used to evaluate causal direct and indirect effects, otherwise computed via numerical integration. Simulations show that our proposed estimators perform equally well as others based on numerical methods and that the additional interpretability of the explicit formulas does not compromise their precision.
As a result of greater demand from funders and governing bodies to assess their 'impact for money', the need for aggregate aid effectiveness for development agencies has become increasingly prominent. Measurement of aid effectiveness or corporate impact, however, requires attribution, i.e. the capacity to causally attribute observed impacts from an investment project to the project alone, as well as counterfactual-based impact evaluations. As the prevalence of impact evaluations grows among development agencies, attributable aggregate development effectiveness to measure institution-wide results is a necessary component of any institution's impact evaluation agenda to ensure accountability, identify performance gaps, and areas for improvement. 'Intelligent' aggregation additionally requires three key elements, namely a critical mass of impact evaluations representing the investment portfolio of the agency in question, a methodology for aggregation, and a universe of projects from where the projects evaluated were randomly drawn. In this paper, a novel methodology based on meta-analytic techniques, selection models and projection methods is proposed along with a number of systematic analyses that adjust for the possible presence of selection bias, a crucial factor to take into account while estimating aggregate development effectiveness.
In regression models with missing outcomes, selection bias can arise when the missingness mechanism depends on the outcome itself. This proposal focuses on an extension of the Heckman model to a setting where the outcome is binary and both the selection process and the outcome are modeled through logistic regression. A correction term analogous to the inverse Mills' ratio is derived based on relative risks. Under given assumptions, such a strategy provides an effective tool for bias correction in the presence of informative missingness.
This study investigates how environmental, social, and governance (ESG) scores influence credit ratings in the banking sector, using mediation analysis to explore their role as intermediaries between financial indicators and creditworthiness. Findings reveal that Environmental and Social scores positively impact credit ratings, while Governance does not play any role. Environmental and Social scores appear negatively influenced by financial metrics such as net interest margin (NIM) and return on assets (ROA), suggesting that banks are yet to consider sustainability as a profitable investment strategy. Indeed, when decomposing the effect of NIM on Credit Rating, via mediation analysis, the negative indirect effect of a lower Social score seems to be negligible when compared to the positive direct impact on Credit Rating. However, this is not the case for the Environmental score, which seems to capture all relevant information to form the Credit Rating. Overall, the commitment of rating agencies to account for sustainability factors seems to be supported, possibly triggering a shift in banks' long-term investment strategies.
A methodological framework to address self-selection is presented that combines previous results on stratified case-control sampling designs with a newly developed theory on secondary outcome analyses. The derivations are then applied to data coming from a single-centre longitudinal prospective cohort study on Post and Long Covid syndrome, conducted at Luigi Sacco University Hospital in Milan from May 2020 to October 2022.
This short note is a commentary on a 2024 article by Mathur and Shpitser in the Journal, with the aim to enlarge the class of graphs for which the conditional average treatment effect is nonparametrically identified, by allowing the outcome to be on the pathway between the treatment and the selection indicator. A first straightforward generalization is possible when (1) the outcome $Y$ is binary, and (2) the population prevalence of $Y$ is known a priori or can be made the object of a sensitivity analysis. Furthermore, identification of the effect is possible also for $Y$ having any nature, provided that a selection bias breaking node $V$ exists and the population prevalence of $V$ is known.
In recent years, attention toward Environmental, Social and Governance (ESG) issues has become increasingly important in the investment decision-making process, prompting interest of investors, companies, regulators and researchers on the possible relationships between financial performances and sustainable variables. With the aim to increase our understanding of these relationships, we use a graphical modeling approach on the MSCI and Bloomberg sustainable dataset for years from 2017 to 2021. Our analysis shows that companies with a higher level of compliance with ESG standards have lower assets’ volatility than others and are not penalized in terms of returns. Furthermore, the increasing level of mandatory disclosure within the European area, induced by the current regulation, has reduced the strength of the positive relationship between Disclosure Score and ESG Score. Moreover, the negative relationship between ESG Score and volatility remains consistent across temporal and geographic areas.
In recent years, the leveraged loan market has experienced considerable growth, with the covenant-lite loan being the predominant agreement. The goal of this research is to assess whether the covenant-lite type reduces or increases the probability of default. Mediation analysis allows us to decompose the effect of balance sheet indicators on a default event into direct and indirect effects, the latter mediated by the covenant-lite. Results show that the covenant-lite is granted to borrower with a greater profitability. In turn, all other conditions being equal, this agreement plays a role in making a default event less likely, giving rise to a significant indirect effect.
In recent years, the context of the banking system,characterised by expansive monetary policies, has boosted the investments in leveraged loans. The COVID-19 pandemic brought the first real slowdown of the global economy since the financial crisis of 2007-08, and the growth of the leveraged loan market has been subject to significant attention from the competent authorities. Banks have remained solid despite the adverse outlook, however, the banking landscape continues to be impacted by the uncertainty relating to the evolution of the pandemic. The original sample for this paper, made up of leveraged loans, combines instrument-specific information with information on financial borrowing and the composition of the syndicate of banks/lenders. The aim of the paper is to identify a systemic risk indicator that takes into account the concentration of credit risk within each bank. For this purpose, using an Mquantile regression, it is possible to obtain an indicator (Mquantile coefficient) for each bank that varies between 0 and 1, where higher values indicate the greater presence of risky leveraged loans in that specific bank. Combined with an indicator of loan sharing between banks, this also allows a graphical representation of the network of banks in this specific market.
BACKGROUND:Long-term sequelae of SARS-CoV-2 infection, namely long COVID syndrome, affect about 10% of severe COVID-19 survivors. This condition includes several physical symptoms and objective measures of organ dysfunction resulting from a complex interaction between individual predisposing factors and the acute manifestation of disease. We aimed at describing the complexity of the relationship between long COVID symptoms and their predictors in a population of survivors of hospitalization for severe COVID-19-related pneumonia using a Graphical Chain Model (GCM). METHODS:96 patients with severe COVID-19 hospitalized in a non-intensive ward at the "Santa Maria" University Hospital, Terni, Italy, were followed up at 3-6 months. Data regarding present and previous clinical status, drug treatment, findings recorded during the in-hospital phase, presence of symptoms and signs of organ damage at follow-up were collected. Static and dynamic cardiac and respiratory parameters were evaluated by resting pulmonary function test, echocardiography, high-resolution chest tomography (HRCT) and cardiopulmonary exercise testing (CPET). RESULTS:Twelve clinically most relevant factors were identified and partitioned into four ordered blocks in the GCM: block 1 - gender, smoking, age and body mass index (BMI); block 2 - admission to the intensive care unit (ICU) and length of follow-up in days; block 3 - peak oxygen consumption (VO2), forced expiratory volume at first second (FEV1), D-dimer levels, depression score and presence of fatigue; block 4 - HRCT pathological findings. Higher BMI and smoking had a significant impact on the probability of a patient's admission to ICU. VO2 showed dependency on length of follow-up. FEV1 was related to the self-assessed indicator of fatigue, and, in turn, fatigue was significantly associated with the depression score. Notably, neither fatigue nor depression depended on variables in block 2, including length of follow-up. CONCLUSIONS:The biological plausibility of the relationships between variables demonstrated by the GCM validates the efficacy of this approach as a valuable statistical tool for elucidating structural features, such as conditional dependencies and associations. This promising method holds potential for exploring the long-term health repercussions of COVID-19 by identifying predictive factors and establishing suitable therapeutic strategies.
With reference to a stratified case-control procedure based on a binary variable of primary interest, we derive the expression of the distortion induced by the sampling design on the parameters of the logistic model of a secondary variable. This is particularly relevant when performing mediation analysis (possibly in a causal framework) with stratified case-control data in settings where both the outcome and the mediator are binary. Our identification result opens the way to M-estimation and Maximum Likelihood estimation. We then conduct a simulation study showing the gain in efficiency of the estimators of both the outcome and mediator model parameters w.r. to existing methods, based on weighting. As an illustrative example, we reanalyze a German case-control dataset in order to investigate whether the effect of reduced immunocompetency on listeriosis onset is mediated by the intake of gastric acid suppressors.
BACKGROUND:The severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is responsible for the COVID-19 pandemic and so it is crucial the right evaluation of viral infection. According to the Centers for Disease Control and Prevention (CDC), the Real-Time Reverse Transcription PCR (RT-PCR) in respiratory samples is the gold standard for confirming the disease. However, it has practical limitations as time-consuming procedures and a high rate of false-negative results. We aim to assess the accuracy of COVID-19 classifiers based on Arificial Intelligence (AI) and statistical classification methods adapted on blood tests and other information routinely collected at the Emergency Departments (EDs).METHODS:Patients admitted to the ED of Careggi Hospital from April 7th-30th 2020 with pre-specified features of suspected COVID-19 were enrolled. Physicians prospectively dichotomized them as COVID-19 likely/unlikely case, based on clinical features and bedside imaging support. Considering the limits of each method to identify a case of COVID-19, further evaluation was performed after an independent clinical review of 30-day follow-up data. Using this as a gold standard, several classifiers were implemented: Logistic Regression (LR), Quadratic Discriminant Analysis (QDA), Random Forest (RF), Support Vector Machine (SVM), Neural Networks (NN), K-nearest neighbor (K-NN), Naive Bayes (NB).RESULTS:Most of the classifiers show a ROC >0.80 on both internal and external validation samples but the best results are obtained applying RF, LR and NN. The performance from the external validation sustains the proof of concept to use such mathematical models fast, robust and efficient for a first identification of COVID-19 positive patients. These tools may constitute both a bedside support while waiting for RT-PCR results, and a tool to point to a deeper investigation, by identifying which patients are more likely to develop into positive cases within 7 days.CONCLUSIONS:Considering the obtained results and with a rapidly changing virus, we believe that data processing automated procedures may provide a valid support to the physicians facing the decision to classify a patient as a COVID-19 case or not.
By exploiting the theory of skew-symmetric distributions, we generalise existing results in sensitivity analysis by providing the analytic expression of the bias induced by marginalization over an unobserved continuous confounder in a logistic regression model. The expression is approximated and mimics Cochran's formula under some simplifying assumptions. Other link functions and error distributions are also considered. A simulation study is performed to assess its properties. The derivations can also be applied in causal mediation analysis, thereby enlarging the number of circumstances where simple parametric formulations can be used to evaluate causal direct and indirect effects. Standard errors of the causal effect estimators are provided via the first-order Delta method. Simulations show that our proposed estimators perform equally well as others based on numerical methods and that the additional interpretability of the explicit formulas does not compromise their precision. The new estimator has been applied to measure the effect of humidity on upper airways diseases mediated by the presence of common aeroallergens in the air.
The attention to sustainable finance has dramatically increased in the recent past. In Europe, the perceived relevance of financial sustainability is mainly due to the European Commission’s commitment to integrate Environmental, Social and Governance (ESG) parameters into all aspects of the financial system. Our objective is to investigate the existence and extent of the impact of compliance with ESG on Credit Rating of European companies. With reference to the financial sector, we use mediation analysis to disentangle the effect of standard balance sheet indicators (measuring the stability and leverage of a company) on Credit Ratings, into a direct one and an indirect one, this second mediated by the ESG rating. Given the different implications of the three aspects of the overall score, Environment, Social and Governance, we apply the methodology to each score separately considered. Results show that only Governance plays a role, with a significant indirect effect that is, however, negligible in magnitude. Data are provided by Thomson Reuters.
The decomposition of the overall effect of a treatment into direct and indirect effects is here investigated with reference to a recursive system of binary random variables. We show how, for the single mediator context, the marginal effect measured on the log odds scale can be written as the sum of the indirect and direct effects plus a residual term that vanishes under some specific conditions. We then extend our definitions to situations involving multiple mediators and address research questions concerning the decomposition of the total effect when some mediators on the pathway from the treatment to the outcome are marginalized over. Connections to the counterfactual definitions of the effects are also made. Data coming from an encouragement design on students’ attitude to visit museums in Florence, Italy, are reanalyzed. The estimates of the defined quantities are reported together with their standard errors to compute p values and form confidence intervals.
With reference to a single mediator context, this brief report presents a model-based strategy to estimate counterfactual direct and indirect effects when the response variable is ordinal and the mediator is binary. Postulating a logistic regression model for the mediator and a cumulative logit model for the outcome, the exact parametric formulation of the causal effects is presented, thereby extending previous work that only contained approximated results. The identification conditions are equivalent to the ones already established in the literature. The effects can be estimated by making use of standard statistical software and standard errors can be computed via a bootstrap algorithm. To make the methodology accessible, routines to implement the proposal in R are presented in the Appendix. A natural effect model coherent with the postulated data generating mechanism is also derived.
With reference to causal mediation analysis, a parametric expression for natural direct and indirect effects is derived for the setting of a binary outcome with a binary mediator, both modelled via a logistic regression. The proposed effect decomposition operates on the odds ratio scale and does not require the outcome to be rare. It generalizes the existing ones, allowing for interactions between both the exposure and the mediator and the confounding covariates. The derived parametric formulae are flexible, in that they readily adapt to the two different natural effect decompositions defined in the mediation literature. In parallel with results derived under the rare outcome assumption, they also outline the relationship between the causal effects and the correspondent pathway-specific logistic regression parameters, isolating the controlled direct effect in the natural direct effect expressions. Formulae for standard errors, obtained via the delta method, are also given. An empirical application to data coming from a microfinance experiment performed in Bosnia and Herzegovina is illustrated.
We derive the exact formula linking the parameters of marginal and conditional logistic regression models with binary mediators when no conditional independence assumptions can be made. The formula has the appealing property of being the sum of terms that vanish whenever parameters of the conditional models vanish, thereby recovering well-known results as particular cases. It also permits the disentangling of direct and indirect effects as well as quantifying the distortion induced by the omission of relevant covariates, opening the way to sensitivity analysis. As the parameters of the conditional models are multiplied by terms that are always bounded, the derivations may also be used to construct reasonable bounds on the parameters of interest when relevant intermediate variables are unobserved. We assume that, conditionally on a set of covariates, the data-generating process can be represented by a directed acyclic graph. We also show how the results presented here lead to the extension of path analysis to a system of binary random variables.