Method

Viral haplotype reconstruction from long reads with virCHap

    • Shandong University
Published September 23, 2026. https://doi.org/10.1101/gr.281975.126
Download PDF Please log-in to or register for your personal account in order to access PDF Cite Article Permissions Share
cover of Genome Research Vol 36 Issue 9
Current Issue:

Abstract

Resolving genomes at the haplotype level for viral populations is crucial for understanding the prevalence of viral diseases and for the development of effective therapeutic treatments. However, viral haplotype reconstruction still presents challenges, such as an unknown number of strains, high inter-strain similarity, repetitive regions, and difficulties in abundance estimation. Here, we develop virCHap, a new reference-based haplotype phasing algorithm for viruses, which applies graph partitioning followed by iteratively quantifiable cluster merging on long-read sequencing data. Benchmarking on simulated and real datasets demonstrates that virCHap outperforms current tools in terms of recall, accurate abundance estimates and read clustering accuracy. On the simulated large-genome VZV experiment, virCHap has a 96% recall, 13.6% higher than the second-best method, and has the most accurate abundance estimates. On a real 5-strain PVY dataset, virCHap has a precision exceeding 94.4%, a recall of over 97%, and a read clustering accuracy of 93.4%, outperforming the second-best method by 33%. On a real 6-strain SARS-CoV-2 dataset, virCHap achieves >96.7% accuracy, and the most accurate abundance estimates within the spike gene.

Loading
Loading
Back to top