Mixed-model association for biobank-scale datasets

Po-Ru Loh; Gleb Kichaev; Steven Gazal; Armin P Schoech; Alkes L Price

doi:10.1038/s41588-018-0144-6

Mixed-model association for biobank-scale datasets

Nat Genet. 2018 Jul;50(7):906-908. doi: 10.1038/s41588-018-0144-6.

Authors

Po-Ru Loh^{1

2}, Gleb Kichaev³, Steven Gazal^{4

5}, Armin P Schoech^{4

5

6}, Alkes L Price^{7

8

9}

Affiliations

¹ Division of Genetics, Department of Medicine, Brigham and Women's Hospital and Harvard Medical School, Boston, MA, USA. poruloh@broadinstitute.org.
² Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA. poruloh@broadinstitute.org.
³ Bioinformatics Interdepartmental Program, University of California, Los Angeles, Los Angeles, CA, USA.
⁴ Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
⁵ Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, USA.
⁶ Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA.
⁷ Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA. aprice@hsph.harvard.edu.
⁸ Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, USA. aprice@hsph.harvard.edu.
⁹ Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA. aprice@hsph.harvard.edu.

Abstract

Biobank-based genome-wide association studies are enabling exciting insights in complex trait genetics, but much uncertainty remains over best practices for optimizing statistical power and computational efficiency in GWAS while controlling confounders. Here, we introduce a much faster version of our BOLT-LMM Bayesian mixed model association method—capable of running analyses of the full UK Biobank cohort in a few days on a single compute node—and show that it produces highly powered, robust test statistics when run on all 459K European samples (retaining related individuals). When used to conduct a GWAS for height in UK Biobank, BOLT-LMM achieved power equivalent to linear regression on 650K samples—a 93% increase in effective sample size versus the common practice of analyzing unrelated British samples using linear regression (UK Biobank documentation; Bycroft et al. bioRxiv). Across a broader set of 23 highly heritable traits, the total number of independent GWAS loci detected increased from 5,839 to 10,759, an 84% increase. We recommend the use of BOLT-LMM (retaining related individuals) for biobank-scale analyses, and we have publicly released BOLT-LMM summary association statistics for the 23 traits analyzed as a resource for all researchers.

Publication types

Letter
Research Support, N.I.H., Extramural
Research Support, Non-U.S. Gov't

MeSH terms

Biological Specimen Banks*
Datasets as Topic
Genome, Human*
Genome-Wide Association Study
Humans
Linear Models
United Kingdom

Abstract

Publication types

MeSH terms

Grants and funding