Abstract
Partial order alignment (POA) has emerged as a fundamental component in long-read error correction, assembly and pangenomics. However, conventional POA algorithms are limited by high time and memory requirements, making them inefficient for large-scale datasets. Here, we present minipoa, a fast and memory-efficient POA tool that incorporates seed-chain-align heuristics, adaptive or static banding strategies, and single-instruction multiple-data optimizations. Minipoa achieves up to a 5-fold speedup over abPOA, reduces memory usage by up to 16-fold, and improves correction accuracy, while maintaining strong performance on both Pacific Biosciences and Oxford Nanopore Technologies simulated datasets, and can be readily integrated into existing long-read error correction and assembly workflows. In multiple sequence alignment datasets, minipoa demonstrates highly competitive computational efficiency and alignment accuracy, achieving Total Column scores up to 2.5-fold higher than MAFFT in low-similarity scenarios. Moreover, minipoa enables multiple sequence alignment of megabase-long genomes and million-sequence datasets, demonstrated by 342 Mycobacterium tuberculosis sequences and one million SARS-CoV-2 sequences respectively. Collectively, minipoa is well positioned to become a cornerstone in the era of large-scale pangenomics.