Secure discovery of genetic relatives across large-scale and distributed genomic data sets

Matthew M. Hong; David Froelicher; Ricky Magner; Victoria Popic; Bonnie Berger; Hyunghoon Cho

doi:10.1101/gr.279057.124

Secure discovery of genetic relatives across large-scale and distributed genomic data sets

¹Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, Massachusetts 02139, USA;
²Broad Institute of the Massachusetts Institute of Technology and Harvard, Cambridge, Massachusetts 02142, USA;
³Department of Mathematics, Massachusetts Institute of Technology, Cambridge, Massachusetts 02139, USA;
⁴Department of Biomedical Informatics and Data Science, Yale University, New Haven, Connecticut 06510, USA

↵5 These authors contributed equally to this work.

Corresponding authors: vpopic{at}broadinstitute.org, bab{at}mit.edu, hoon.cho{at}yale.edu

Abstract

Finding relatives within a study cohort is a necessary step in many genomic studies. However, when the cohort is distributed across multiple entities subject to data-sharing restrictions, performing this step often becomes infeasible. Developing a privacy-preserving solution for this task is challenging owing to the burden of estimating kinship between all the pairs of individuals across data sets. We introduce SF-Relate, a practical and secure federated algorithm for identifying genetic relatives across data silos. SF-Relate vastly reduces the number of individual pairs to compare while maintaining accurate detection through a novel locality-sensitive hashing (LSH) approach. We assign individuals who are likely to be related together into buckets and then test relationships only between individuals in matching buckets across parties. To this end, we construct an effective hash function that captures identity-by-descent (IBD) segments in genetic sequences, which, along with a new bucketing strategy, enable accurate and practical private relative detection. To guarantee privacy, we introduce an efficient algorithm based on multiparty homomorphic encryption (MHE) to allow data holders to cooperatively compute the relatedness coefficients between individuals and to further classify their degrees of relatedness, all without sharing any private data. We demonstrate the accuracy and practical runtimes of SF-Relate on the UK Biobank and All of Us data sets. On a data set of 200,000 individuals split between two parties, SF-Relate detects 97% of third-degree or closer relatives within 15 h of runtime. Our work enables secure identification of relatives across large-scale genomic data sets.

Footnotes

[Supplemental material is available for this article.]
Article published online before print. Article, supplemental material, and publication date are at https://www.genome.org/cgi/doi/10.1101/gr.279057.124.
Freely available online through the Genome Research Open Access option.

Received February 16, 2024.
Accepted July 31, 2024.

This article, published in Genome Research, is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.