arXiv Science⌕ Search

arXiv subjects

George Davey Smith

Publications and source records attributed to George Davey Smith.

7 recordsLinked to original sources

The INSIDE assumption under all positive coding: interpretation and partial empirical assessment

Mendelian randomisation (MR) implemented through instrumental variable (IV) analysis is a popular strategy for strengthening causal inference in observational studies. A key assumption for several MR estimators, including MR-Egger regression, is the INstrument Strength Independent of Direct Effect (INSIDE) assumption. However, there is no established empirical test for assessing the plausibility of this assumption. Moreover, INSIDE depends on how genetic variants are coded (i.e., on the choice of the effect allele), which is often arbitrary and therefore hampers assessing the plausibility of this assumption on substantive grounds. In this paper, we show that the all-positive coding scheme (i.e., for all variants, choosing the allele positively associated with the exposure as the effect allele), which is typically used in MR-Egger, is equivalent to a coding-invariant model that can be given a natural interpretation because the direct effect parameters under this coding scheme are in the same direction as the bias of individual-variant ratio estimators. Moreover, using both theoretical arguments and simulations, we show that, under commonly assumed data-generating models in the MR methodological literature, heteroscedasticity of instrument-outcome coefficients according to instrument-exposure coefficients is a feature of at least some types of INSIDE violation, indicating that heteroscedasticity tests could contribute to assessing the plausibility of the INSIDE assumption. We further highlight specific cases where the test would not work. We illustrate its application by re-analysing a real dataset assessing the causal effect of large particle high density lipoprotein cholesterol on age-related macular degeneration.

stat.ME↗

Towards more plausible point-identifying assumptions in two-sample Mendelian randomization

Two-sample Mendelian randomization (MR) is a widely applied methodology in epidemiology. In two-sample MR, summary data (typically, regression coefficients and standard errors) quantifying the association between multiple genetic variants and the exposure and the outcome are used in an instrumental variable framework aimed at estimating the causal effect of the exposure on the outcome. Most two-sample MR methods were developed under data-generating models where the association of for each candidate genetic instrument with the exposure, as well as the causal effect of the exposure on the outcome, are constant in the additive scale. These assumptions are useful because they imply that, had all genetic variants been valid IVs, they would all estimate the same causal parameter - namely, the constant causal effect. We refer to this condition as summary-level homogeneity. However, these are rather strong homogeneity conditions which may raise concerns about the plausibility of these methods in practice. In this paper, we show that summary-level homogeneity is implied by the following conditions: the causal effect is additive linear, but not necessarily constant across, all strata of the population; and uncorrelatedness between heterogeneity in the causal effect and in the association between each genetic variant and the exposure. Under these conditions, typical two-sample MR methods can be interpreted as estimators of the average causal effect. These results clarify that point-identifying assumptions required for two-sample MR methods are weaker than previously anticipated, which contributes to their plausibility and interpretation in at least some practical applications.

stat.ME↗

Empirically assessing the plausibility of unconfoundedness in observational studies

The possibility of unmeasured confounding is one of the main limitations for causal inference from observational studies. There are different methods for (partially) empirically assessing the plausibility of unconfoundedness. However, most currently available methods require (at least partial) assumptions about the confounding structure, which may be difficult to know in practice. In this paper we describe a simple strategy for empirically assessing the plausibility of conditional unconfoundedness (i.e., whether the candidate adjustment set of covariates suffices for confounding adjustment) which does not require any explicit assumptions about the confounding structure, relying instead on assumptions related to temporal ordering between covariates, exposure and outcome (which can be guaranteed by design) and selection into the study. The proposed method essentially relies on testing the association between a subset of the covariates included in the adjustment set (those associated with the exposure, given all other covariates) and the outcome conditional on the remaining covariates and the exposure. We describe the assumptions underlying the method, provide proofs, use simulations to corroborate the theory and illustrate the method with an applied example assessing the causal effect of delivery mode and intelligence quotient measured in adulthood using data from the 1982 Pelotas (Brazil) birth cohort. We also discuss the implications of measurement error and some important limitations of the suggested approach.

stat.ME↗

Woolf et als GWAS by subtraction is not useful for cross-generational Mendelian randomization studies

Mendelian randomization (MR) is an epidemiological method that can be used to strengthen causal inference regarding the relationship between a modifiable environmental exposure and a medically relevant trait and to estimate the magnitude of this relationship1. Recently, there has been considerable interest in using MR to examine potential causal relationships between parental phenotypes and outcomes amongst their offspring. In a recent issue of BMC Research Notes, Woolf et al (2023) present a new method, GWAS by subtraction, to derive genome-wide summary statistics for paternal smoking and other paternal phenotypes with the goal that these estimates can then be used in downstream (including two sample) MR studies. Whilst a potentially useful goal, Woolf et al. (2023) focus on the wrong parameter of interest for useful genome-wide association studies (GWAS) and downstream cross-generational MR studies, and the estimator that they derive is neither efficient nor appropriate for such use.

q-bio.QM↗

Almost exact Mendelian randomization

Mendelian randomization (MR) is a natural experimental design based on the random transmission of genes from parents to offspring. However, this inferential basis is typically only implicit or used as an informal justification. As parent-offspring data becomes more widely available, we advocate a different approach to MR that is exactly based on this natural randomization, thereby formalizing the analogy between MR and randomized controlled trials. We begin by developing a causal graphical model for MR which represents several biological processes and phenomena, including population structure, gamete formation, fertilization, genetic linkage, and pleiotropy. This causal graph is then used to detect biases in population-based MR studies and identify sufficient confounder adjustment sets to correct these biases. We then propose a randomization test in the within-family MR design using the exogenous randomness in meiosis and fertilization, which is extensively studied in genetics. Besides its transparency and conceptual appeals, our approach also offers some practical advantages, including robustness to misspecified phenotype models, robustness to weak instruments, and elimination of bias arising from population structure, assortative mating, dynastic effects, and horizontal pleiotropy. We conclude with an analysis of a pair of negative and positive controls in the Avon Longitudinal Study of Parents and Children. The accompanying R package can be found at https://github.com/matt-tudball/almostexactmr.

stat.ME↗

Homogeneity in the instrument-treatment association is not sufficient for the Wald estimand to equal the average causal effect for a binary instrument and a continuous exposure

Background: Interpreting instrumental variable results often requires further assumptions in addition to the core assumptions of relevance, independence, and the exclusion restriction. Methods: We assess whether instrument-exposure additive homogeneity renders the Wald estimand equal to the average derivative effect (ADE) in the case of a binary instrument and a continuous exposure. Results: Instrument-exposure additive homogeneity is insufficient for ADE identification when the instrument is binary, the exposure is continuous and the effect of the exposure on the outcome is non-linear on the additive scale. For a binary exposure, the exposure-outcome effect is necessarily additive linear, so the homogeneity condition is sufficient. Conclusions: For binary instruments, instrument-exposure additive homogeneity identifies the ADE if the exposure is also binary. Otherwise, additional assumptions (such as additive linearity of the exposure-outcome effect) are required.

stat.ME↗

How to estimate heritability, a guide for epidemiologists

Traditionally, heritability has been estimated using family-based methods such as twin studies. Advancements in molecular genomics have facilitated the development of alternative methods that utilise large samples of unrelated or related individuals. Yet, specific challenges persist in the estimation of heritability such as epistasis, assortative mating and indirect genetic effects. Here, we provide an overview of common methods applied in genetic epidemiology to estimate heritability i.e., the proportion of phenotypic variation explained by genetic variation. We provide a guide to key genetic concepts required to understand heritability estimation methods from family-based designs (twin and family studies), genomic designs based on unrelated individuals (LD score regression, GREML), and family-based genomic designs (Sibling regression, GREML-KIN, Trio-GCTA, MGCTA, RDR). For each method, we describe how heritability is estimated, the assumptions underlying its estimation, and discuss the implications when these assumptions are not met. We further discuss the benefits and limitations of estimating heritability within samples of unrelated individuals compared to samples of related individuals. Overall, this article is intended to help the reader determine the circumstances when each method would be appropriate and why.

stat.AP↗