Method

Resolving missing human polymorphic inversions and other complex variants from ultralong read data

    • 1Institut de Biotecnologia i de Biomedicina, Universitat Autònoma de Barcelona, Bellaterra 08193, Barcelona, Spain;
    • 2Research Programme on Biomedical Informatics (GRIB), Hospital del Mar Research Institute, Barcelona 08003, Spain;
    • 3Departament de Genètica i de Microbiologia, Universitat Autònoma de Barcelona, Bellaterra 08193, Barcelona, Spain;
    • 4Àrea d'Arquitectura i de Tecnologia de Computadors, Universitat Autònoma de Barcelona, Bellaterra 08193, Barcelona, Spain;
    • 5ICREA, Barcelona 08010, Spain
    • Present addresses: 6Department of Biochemistry and Molecular Biology, University of Valencia, Burjassot 46100, València, Spain; 7Departamento de Ciencias Básicas, Facultad de Medicina, Universidad de la Frontera, Temuco 4780000, Chile; 8Laboratorio de Bioinformática y Microbiología Aplicada, Centro de Excelencia en Medicina Traslacional, Universidad de la Frontera, Temuco 4780000, Chile
Published September 1, 2026. https://doi.org/10.1101/gr.280867.125
Download PDF Cite Article Permissions Share
cover of Genome Research Vol 36 Issue 9
Current Issue:

Abstract

Inversions are a unique type of balanced structural variants (SVs) with important consequences in multiple organisms. However, despite considerable effort, these and other complex SVs remain poorly characterized because of the presence of large repeats. New techniques are finally allowing us to identify the full spectrum of human inversions, but the number of individuals analyzed is still quite limited. Here, we take advantage of Oxford Nanopore Technologies (ONT) long reads to characterize an exhaustive catalog of 612 candidate inversions between 197 bp and 4.4 Mb of length, and flanked by <190 kb long inverted repeats (IRs). To that end, we have developed a bioinformatic package to identify inversion alleles reliably from long-read data. Next, using a combination of different DNA extraction, library preparation, and ONT sequencing protocols, we show that ultralong reads (50–100 kb) and adaptive sampling are an efficient method to detect most human inversions. Lastly, by analyzing ONT data from 54 diverse individuals, 87%–99% of the inversions can be genotyped in each sample, depending mainly on read and IR length and genome coverage. Both orientations have been observed for 155 of the analyzed regions (frequency 0.01–0.49), which triples the number of polymorphic IR-mediated inversions studied in detail so far. Moreover, we have found more than 300 additional independent SVs in the studied regions and resolved several complex rearrangements. Therefore, our work provides an accurate benchmark of those inversions that typically escape most analyses, and it demonstrates the potential of nanopore sequencing to characterize missing human genomic variation.

Loading
Loading
Loading
Back to top