Resource

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow

    • 1Department of Genetics, University of North Carolina, Chapel Hill, North Carolina 27599-7264, USA;
    • 2Institute of Human Genetics, University Hospital Schleswig-Holstein, University of Luebeck, 23562 Lübeck, Germany;
    • 3Berlin Institute of Health at Charité–Universitätsmedizin Berlin, 10117 Berlin, Germany;
    • 4Department of Biostatistics and Bioinformatics, Duke University Medical School, Durham, North Carolina 27710, USA;
    • 5Center for Advanced Genomic Technologies, Duke University, Durham, North Carolina 27708, USA;
    • 6Department of Bioengineering and Therapeutic Sciences, University of California San Francisco, San Francisco, California 94143, USA;
    • 7Institute for Human Genetics, University of California San Francisco, San Francisco, California 94143-0794, USA;
    • 8Department of Genetics, Yale School of Medicine, New Haven, Connecticut 06510, USA;
    • 9Department of Biostatistics, University of North Carolina, Chapel Hill, North Carolina 27599-7420, USA
Published August 11, 2026. https://doi.org/10.1101/gr.281462.125
Download PDF Cite Article Permissions Share
cover of Genome Research Vol 36 Issue 8
Current Issue:

Abstract

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Loading
Loading
Loading
Back to top