2015•American Journal of EpidemiologyRequires access

Re: “Multivariable Mendelian Randomization: The Use of Pleiotropic Genetic Variants to Estimate Causal Effects”

Stephen Burgess, Frank Dudbridge, Simon G. Thompson

Open publisher page 1,009 citations

Abstract

In the manuscript “Multivariable Mendelian Randomization: the Use of Pleiotropic Genetic Variants to Estimate Causal Effects” (1), we commented on an analysis method recently used in the literature (2), which we referred to as a “regression-based method.” We presented simulation results showing that the method had some weaknesses—in particular, not estimating the same parameter as that from a 2-stage least-squares analysis of individual-level data, and giving widely varying results when there were causal effects between the risk factors. A specific criticism of the method is that the uncertainty in the summarized associations (beta coefficients) used in the method is ignored. We recently noticed a simple modification which can be made to the regression-based method that results in estimates which closely approximate those from a 2-stage least-squares analysis and remain stable when there are causal effects between the risk factors. In this letter, we describe this modification of the regression-based method (we refer to the modified method as a “weighted regression-based method”) and repeat the simulations from the main paper using this modified method. We assume that data are available on the associations of J uncorrelated genetic variants with each of K risk factors, such that the association estimate of variant j with risk factor k is Xkj, with standard error σXkj, and the association estimate of variant j with the outcome is Yj, with standard error σYj. This notation matches that of our original article (1). The weighted regression-based method is performed by regression of the association estimates Yj on each of the Xkj in a multivariable weighted regression model, using the σYj−2 as weights (inverse-variance weights) (3). The regression model should have no intercept. Estimates from the weighted regression-based method for the simulation study in the original article (1) are given in Supplementary Data. These tables replicate parts of Tables 1 and 2 from the original publication (1). Mean estimates from the weighted regression-based method equal those from the 2-stage least-squares method to within 0.001 in all scenarios considered. Mean standard errors, standard deviations of estimates, and the empirical power at a 5% significance level from the weighted regression-based method are also comparable to those from the other methods (Supplementary Data). This contrasts with estimates from the unweighted regression-based method, which differed substantially from those from a 2-stage least-squares method and had reduced power to detect a causal effect (see Table 1 in the original publication (1)). When there are causal effects between the risk factors (Supplementary Data), the weighted regression-based method gives estimates that do not depend greatly on these causal effects (similarly to the other analysis methods). The main criticism of the unweighted regression-based method is that the uncertainties in the genetic association estimates are not accounted for in the analysis. Through inclusion of the inverse-variance weights, evidence from more precisely estimated genetic associations receives more weight in the analysis. It would also be possible to weight the analyses using the inverse variance of the genetic associations with any of the risk factors; provided that the association estimates for that risk factor corresponding to each variant are based on the same sample size, the weighted estimates should be the same, although the standard errors would be incorrect. In conclusion, we stand by our criticism of the unweighted regression-based method and the results of our original study (1), but we have now demonstrated a simple weighting of the regression-based method which produces estimates based on summarized data that are similar to those from the established 2-stage least-squares method based on individual-level data. Conflict of interest: none declared.

About this research paper

What this paper is about

In the manuscript “Multivariable Mendelian Randomization: the Use of Pleiotropic Genetic Variants to Estimate Causal Effects” (1), we commented on an analysis method recently used in the literature (2), which we referred to as a “regression-based method.” We presented simulation results showing that the method had some weaknesses—in particular, not estimating the same parameter as that from a 2-stage least-squares analysis of individual-level data, and giving widely varying results when there were causal effects between the risk factors. A specific criticism of the method is that the uncertainty in the summarized associations (beta coefficients) used in the method is ignored. We recently noticed a simple modification which can be made to the regression-based method that results in estimates which closely approximate those from a 2-stage least-squares analysis and remain stable when there are causal effects between the risk factors. In this letter, we describe this modification of the regression-based method (we refer to the modified method as a “weighted regression-based method”) and repeat the simulations from the main paper using this modified method. We assume that data are available on the associations of J uncorrelated genetic variants with each of K risk factors, such that the association estimate of variant j with risk factor k is Xkj, with standard error σXkj, and the association estimate of variant j with the outcome is Yj, with standard error σYj. This notation matches that of our original article (1). The weighted regression-based method is performed by regression of the association estimates Yj on each of the Xkj in a multivariable weighted regression model, using the σYj−2 as weights (inverse-variance weights) (3). The regression model should have no intercept. Estimates from the weighted regression-based method for the simulation study in the original article (1) are given in Supplementary Data. These tables replicate parts of Tables 1 and 2 from the original publication (1). Mean estimates from the weighted regression-based method equal those from the 2-stage least-squares method to within 0.001 in all scenarios considered. Mean standard errors, standard deviations of estimates, and the empirical power at a 5% significance level from the weighted regression-based method are also comparable to those from the other methods (Supplementary Data). This contrasts with estimates from the unweighted regression-based method, which differed substantially from those from a 2-stage least-squares method and had reduced power to detect a causal effect (see Table 1 in the original publication (1)). When there are causal effects between the risk factors (Supplementary Data), the weighted regression-based method gives estimates that do not depend greatly on these causal effects (similarly to the other analysis methods). The main criticism of the unweighted regression-based method is that the uncertainties in the genetic association estimates are not accounted for in the analysis. Through inclusion of the inverse-variance weights, evidence from more precisely estimated genetic associations receives more weight in the analysis. It would also be possible to weight the analyses using the inverse variance of the genetic associations with any of the risk factors; provided that the association estimates for that risk factor corresponding to each variant are based on the same sample size, the weighted estimates should be the same, although the standard errors would be incorrect. In conclusion, we stand by our criticism of the unweighted regression-based method and the results of our original study (1), but we have now demonstrated a simple weighting of the regression-based method which produces estimates based on summarized data that are similar to those from the established 2-stage least-squares method based on individual-level data. Conflict of interest: none declared.

Why it matters

OpenAlex reports 1009 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

In the manuscript “Multivariable Mendelian Randomization: the Use of Pleiotropic Genetic Variants to Estimate Causal Effects” (1), we commented on an analysis method recently used in the literature (2), which we referred to as a “regression-based method.” We presented simulation results showing that the method had some weaknesses—in particular, not estimating the same parameter as that from a 2-stage least-squares analysis of individual-level data, and giving widely varying results when there were causal effects between the risk factors. A specific criticism of the method is that the uncertainty in the summarized associations (beta coefficients) used in the method is ignored. We recently noticed a simple modification which can be made to the regression-based method that results in estimates which closely approximate those from a 2-stage least-squares analysis and remain stable when there are causal effects between the risk factors. In this letter, we describe this modification of the regression-based method (we refer to the modified method as a “weighted regression-based method”) and repeat the simulations from the main paper using this modified method. We assume that data are available on the associations of J uncorrelated genetic variants with each of K risk factors, such that the association estimate of variant j with risk factor k is Xkj, with standard error σXkj, and the association estimate of variant j with the outcome is Yj, with standard error σYj. This notation matches that of our original article (1). The weighted regression-based method is performed by regression of the association estimates Yj on each of the Xkj in a multivariable weighted regression model, using the σYj−2 as weights (inverse-variance weights) (3). The regression model should have no intercept. Estimates from the weighted regression-based method for the simulation study in the original article (1) are given in Supplementary Data. These tables replicate parts of Tables 1 and 2 from the original publication (1). Mean estimates from the weighted regression-based method equal those from the 2-stage least-squares method to within 0.001 in all scenarios considered. Mean standard errors, standard deviations of estimates, and the empirical power at a 5% significance level from the weighted regression-based method are also comparable to those from the other methods (Supplementary Data). This contrasts with estimates from the unweighted regression-based method, which differed substantially from those from a 2-stage least-squares method and had reduced power to detect a causal effect (see Table 1 in the original publication (1)). When there are causal effects between the risk factors (Supplementary Data), the weighted regression-based method gives estimates that do not depend greatly on these causal effects (similarly to the other analysis methods). The main criticism of the unweighted regression-based method is that the uncertainties in the genetic association estimates are not accounted for in the analysis. Through inclusion of the inverse-variance weights, evidence from more precisely estimated genetic associations receives more weight in the analysis. It would also be possible to weight the analyses using the inverse variance of the genetic associations with any of the risk factors; provided that the association estimates for that risk factor corresponding to each variant are based on the same sample size, the weighted estimates should be the same, although the standard errors would be incorrect. In conclusion, we stand by our criticism of the unweighted regression-based method and the results of our original study (1), but we have now demonstrated a simple weighting of the regression-based method which produces estimates based on summarized data that are similar to those from the established 2-stage least-squares method based on individual-level data. Conflict of interest: none declared.

Key concepts: Mendelian randomization, Risk factor, Randomization, Medicine, Multivariable calculus, Randomized controlled trial, Bioinformatics, Internal medicine

Related papers

Back to paper searchBrowse research topicsOriginal source
Re: “Multivariable Mendelian Randomization: The Use of Pleiotropic Genetic Variants to Estimate Causal Effects” — Research Paper | ScholarLens