Abstract
Calculating optimal polygenic risk scores (PRS) across diverse ancestries, particularly in admixed populations, is necessary to enable equitable genetic research and clinical translation. However, the relatively low representation of admixed populations in both discovery and fine-tuning individual-level datasets limits PRS development for admixed populations. Under the assumption that the most informative PRS weight for a homogeneous sample, which can be approximated by a data point in the ancestry continuum space, varies linearly in that space, we introduce a Genetic Distance-assisted PRS Combination Pipeline for Diverse Genetic Ancestries (DiscoDivas) to interpolate a harmonized PRS for diverse, especially admixed, ancestries, leveraging multiple PRS weights fine-tuned within single-ancestry samples and the genetic ancestry continuum information. DiscoDivas treats ancestry as a continuous variable and does not require shifting between different models when calculating PRS for different ancestries. We generated PRS with DiscoDivas and the current conventional method, i.e. fine-tuning multiple GWAS PRS using the matched or similar ancestry sample, for simulated datasets and large-scale biobank datasets (UK Biobank [UKBB] N=415,402, Mass General Brigham Biobank N=53,306, All of Us N=245,394) and compared our method with the conventional method with quantitative traits and complex disease traits. DiscoDivas generated a harmonized PRS of the accuracy comparable to or higher than the conventional approach, with the greatest advantage exhibited in admixed samples: DiscoDivas PRS for admixed samples was more statistically accurate than the PRS fine-tuned in matched or similar ancestry sample in 12 out of 16 simulated scenarios and was statistically equivalent in the remaining four scenarios; when tested with quantitative trait data in UKBB, DiscoDivas increased the PRS accuracy of admixed sample by 5% on average; yet no statistical difference was observed when tested for binary traits in UKBB where ancestry-matched data was available. For the single ancestry samples, the accuracy of DiscoDivas PRS and PRS fine-tuned in match samples was similar. In summary, our method DiscoDivas yields a harmonized PRS of robust accuracy for individuals across the genetic ancestry spectrum, including where ancestry-matched training data may be incomplete.
Competing Interest Statement
Disclosures: P.N. reports research grants from Allelica, Amgen, Apple, Boston Scientific, Genentech / Roche, and Novartis, personal fees from Allelica, Apple, AstraZeneca, Blackstone Life Sciences, Creative Education Concepts, CRISPR Therapeutics, Eli Lilly & Co, Esperion Therapeutics, Foresite Capital, Foresite Labs, Genentech / Roche, GV, HeartFlow, Magnet Biomedicine, Merck, Novartis, TenSixteen Bio, and Tourmaline Bio, equity in Bolt, Candela, Mercury, MyOme, Parameter Health, Preciseli, and TenSixteen Bio, and spousal employment at Vertex Pharmaceuticals, all unrelated to the present work.
Funding Statement
P.N. is supported by NHGRI U01HG011719 and NHLBI R01HL127564. N.C is supported by NHGRI U01HG011719 and R01HG010480.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
The access to biobank data (UK Biobank, Mass General Brigham Biobank, and All of US Research Program) were gained upon application. The simulated genotype data based on 1000 Genomes were downloaded from https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/COXHAP; The resource of summary statistics GWAS datasets used in this project to generate PRS was listed in the supplementary file.
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Footnotes
Funding: P.N. is supported by NHGRI U01HG011719 and NHLBI R01HL127564. N.C is supported by NHGRI U01HG011719 and R01HG010480.
Disclosures: P.N. reports research grants from Allelica, Amgen, Apple, Boston Scientific, Genentech / Roche, and Novartis, personal fees from Allelica, Apple, AstraZeneca, Blackstone Life Sciences, Creative Education Concepts, CRISPR Therapeutics, Eli Lilly & Co, Esperion Therapeutics, Foresite Capital, Foresite Labs, Genentech / Roche, GV, HeartFlow, Magnet Biomedicine, Merck, Novartis, TenSixteen Bio, and Tourmaline Bio, equity in Bolt, Candela, Mercury, MyOme, Parameter Health, Preciseli, and TenSixteen Bio, and spousal employment at Vertex Pharmaceuticals, all unrelated to the present work.
Data Availability
The access to biobank data (UK Biobank, Mass General Brigham Biobank, and All of US Research Program) were gained upon application. The simulated genotype data based on 1000 Genomes was downloaded from https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/COXHAP; The resource of summary statistics GWAS datasets used in this project to generate PRS was listed in the supplementary file. The scripts of running DiscoDivas and other supporting files can be found at https://github.com/YunfengRuan/DiscoDivas