Parente2: A fast and accurate method for detecting identity by descent

  1. Serafim Batzoglou1
  1. Stanford University
  1. * Corresponding author; email: serafim{at}cs.stanford.edu

Abstract

Identity-by-descent (IBD) inference is the problem of establishing a genetic connection between two individuals through a genomic segment that is inherited by both individuals from a recent common ancestor. IBD inference is an important preceding step in a variety of population genomic studies, ranging from demographic studies to linking genomic variation with phenotype and disease. The problem of accurate IBD detection has become increasingly challenging with the availability of large collections of human genotypes and genomes: given a cohort's size, a quadratic number of pairwise genome comparisons must be performed. Therefore, computation time and the false discovery rate can also scale quadratically. To enable accurate and efficient large-scale IBD detection, we present Parente2, a novel method for detecting IBD segments. Parente2 is based on an embedded log-likelihood ratio and uses a model that accounts for linkage disequilibrium by explicitly modeling haplotype frequencies. Parente2 operates directly on genotype data without the need to phase data prior to IBD inference. We evaluate Parente2's performance through extensive simulations using real data, and show that it provides substantially higher accuracy compared to previous state-of-the-art, while maintaining high computational efficiency.

  • Received February 5, 2014.
  • Accepted September 30, 2014.

This manuscript is Open Access.

This article, published in Genome Research, is available under a Creative Commons License (Attribution-NonCommercial 4.0 International license), as described at http://creativecommons.org/licenses/by-nc/4.0/.

OPEN ACCESS ARTICLE
ACCEPTED MANUSCRIPT

Preprint Server