Method

Read-level genotyping of short tandem repeats using long reads and single-nucleotide variation with STRkit

    • 1Canadian Centre for Computational Genomics, Montréal, Québec H3A 0G1, Canada;
    • 2Department of Human Genetics, McGill University, Montréal, Québec H3A 0G1, Canada;
    • 3Genomic Medicine Center, Children's Mercy Hospital and Research Institute, Kansas City, Missouri 64108, USA;
    • 4Victor Phillip Dahdaleh Institute of Genomic Medicine at McGill University, Montréal, Québec H3A 0G1, Canada
Published January 15, 2026. Vol 36 Issue 3, pp. 578-588. https://doi.org/10.1101/gr.280766.125
Download PDF Cite Article Permissions Share
cover of Genome Research Vol 36 Issue 9
Current Issue:

Abstract

Variation in short tandem repeats (STRs) is implicated in Mendelian disease and complex traits but can be difficult to resolve with short-read genome sequencing. We present STRkit, a software package for genotyping STRs using long-read sequencing (LRS) that uses proximate single-nucleotide variants to improve genotyping accuracy without a priori haplotype information. We show that STRkit has unique strengths versus other methods: It can use data from both major LRS technologies (Pacific Biosciences HiFi [PacBio] and Oxford Nanopore Technologies [ONT]) to output both allele- and read-level copy number and sequence; it performs best in benchmarking with F1 scores of 0.9631 and 0.9544 with PacBio and ONT data, respectively; it achieves higher rates of Mendelian consistency than other genotyping tools; and it is open source software. STRkit's features open up new possibilities for association testing, assessing patterns of STR inheritance and better understanding the functional effects of these notable repeat elements.


Short tandem repeats (STRs), or microsatellites, are repetitive genomic elements found in both eukaryotes and prokaryotes in which a small motif, from 1 or 2 to 6–13 bp long (International Human Genome Sequencing Consortium 2001; Ellegren 2004; Chiu et al. 2021), is repeated several times in a row. These regions cover around 5% of the human genome (Horton et al. 2023), are often multiallelic, and are significant contributors to total genomic variation (Ellegren 2004; Hannan 2018; Liao et al. 2023; English et al. 2025). STRs typically have high rates of mutation, usually via stepwise repeat expansion or contraction (Ellegren 2004; Fan and Chu 2007).

Variation in STRs is associated with more than 60 inherited and sporadic diseases in humans (Gall-Duncan et al. 2022). In Huntington's disease, fragile X syndrome (FXS), and others, copy number expansion beyond a threshold causes the disease, and further expansion is associated with phenotypic severity (Igarashi et al. 1992; Duyao et al. 1993; Brinkman et al. 1997; Matsuura et al. 2000; Blauw et al. 2012; Bragg et al. 2017). The relationship between copy number and disease severity is not always linear; for example, individuals with between 85 and 89 CCG copies (in the same repeat that causes FXS with 200 or more copies) have a higher risk of developing fragile X–associated primary ovarian insufficiency (FXPOI) than those with more or fewer repeats in the FXPOI-pathogenic repeat range of 55 to 199 copies (Allen et al. 2021). STR motif composition can also affect phenotype: insertion of a stretch of a noncanonical motif into a noncoding STR in the DAB1 gene is responsible for a form of spinocerebellar ataxia (Seixas et al. 2017). Beyond expansion detection, precisely determining copy number and motif composition in disease-causing STR expansions may be required to fully quantify their relationship with phenotype (Rajan-Babu et al. 2024).

Other phenotypic relationships with STRs are not limited to a single locus: Autism is associated with a genome-wide increase in rare STR expansions (Trost et al. 2020) and small de novo STR mutations (Mitra et al. 2021). Expansion and somatic instability in STRs feature in several cancers (Yamamoto and Imai 2019; Erwin et al. 2023). STRs are implicated in gene expression and regulation (Gymrek et al. 2016; Hannan 2018; Cui et al. 2025), including by acting as binding sites for transcription factors (Horton et al. 2023). Direct STR genotyping at a genome scale is thus valuable; STRs cannot always be imputed using SNVs (Gymrek et al. 2016; Saini et al. 2018), as SNV–STR linkage disequilibrium (LD) is lower than SNV–SNV pair LD (Willems et al. 2014; Press et al. 2018). Examination of STR variation in the human genome may help to address the “missing heritability” problem, explaining a portion of complex trait heritability not yet identified (Hannan 2018).

However, STRs are difficult to characterize with short-read sequencing (SRS) technologies owing to often long repeat tracts and lack of flanking sequence within reads (Narzisi and Schatz 2015; Mousavi et al. 2019), affecting mappability to a reference genome. Some SRS STR genotyping tools can estimate larger-than-read-length allele sizes (Dolzhenko et al. 2017; Dolzhenko et al. 2019; Mousavi et al. 2019) but cannot capture variation in motif structure nor precisely genotype these long expansions (Mousavi et al. 2019). In contrast, modern long-read sequencing (LRS) technologies, like Oxford Nanopore Technologies (ONT) nanopore sequencing with newer chemistry or Pacific Biosciences (PacBio) high-fidelity circular consensus sequencing (HiFi), can give accurate reads of sizes in the tens of kilobases (Wenger et al. 2019; Gustafson et al. 2024; Tanudisastro et al. 2024). These reads span pathogenic expansions that exceed SRS read length (Hannan 2018) and include more flanking sequence. These flanking sequences may include single-nucleotide variants (SNVs) that could be used to cluster STR-spanning reads into alleles, as SNVs are the most reliably called form of variation with LRS (Wenger et al. 2019; Gustafson et al. 2024).

Several LRS STR genotyping tools have been recently developed, including Straglr (Chiu et al. 2021), TRGT (Dolzhenko et al. 2024), LongTR (Ziaei Jam et al. 2024), and STRdust (De Coster et al. 2024). These methods take standard BAM/CRAM alignments as input and output variants as VCF files. LongTR and TRGT have been tested on whole-genome catalogs of STRs, meaning they should be suitable for association testing and novel expansion discovery, among other whole-genome-scale applications. However, they do not take full advantage of LRS data, or they have additional restrictions which limit their use: TRGT can use proximate SNVs to aid genotyping but is license-restricted to only work with PacBio data; Straglr does not output consensus sequences for alleles; and neither LongTR nor STRdust outputs read-level data, limiting their use for motif composition analysis or detecting within-allele somatic instability.

Here, we aim to develop and evaluate a new STR genotyping software, STRkit, which uses STR-proximate SNV information contained in long reads to accurately genotype and locally phase STRs, while giving researchers as much data as possible for downstream proband, trio, and population-level analyses, including STR consensus sequences and read-level STR copy number and sequence data. By using standard benchmarking tools and publicly available data from multiple LRS platforms, as well as a set of 30 trios sequenced using PacBio HiFi technology, we will evaluate our new software's effectiveness in capturing STR variation, including pathogenic expansions, and compare its genotyping and computational performance against other long-read STR genotyping tools.

Results

The STRkit software package

We created an STR genotyping method, STRkit, as an easily installable Python package. Our algorithm takes aligned LRS BAM or CRAM files as input, with an optional realignment step for mismapped reads (Methods) (Fig. 1A). It then determines STR allele copy number and optionally genotypes proximate SNVs, whose locations are extracted from a user-provided VCF catalog file from, for example, dbSNP (Sherry et al. 2001). For diploid alleles, depending on what SNV data are present, STRkit either uses a Gaussian mixture model (GMM) allele-calling approach with copy numbers, segregates alleles based on SNV haplotype groups, or uses a mix of both information sources. STRkit can also genotype haploid alleles for, for example, loci on the sex chromosomes of human XY or XO individuals. If two STRs have common SNV haplotype groups, they will be locally phased. If a phased alignment file is provided, STRkit can use read haplotype data instead. Genotypes are then reported in either VCF, TSV, or JSON formats. JSON output can be explored in an included web-based visualization tool (Fig. 1B). Our software package also includes a Mendelian inheritance (MI) analysis tool for benchmarking callers or detecting candidate de novo mutations. We compared the features of STRkit against those of other LRS STR genotyping packages to highlight its unique features, including the MI tool, SNV incorporation and output during STR genotyping, and read-level copy number output (Table 1).

Figure 1.

Flowchart of the STRkit genotyping and visualization workflow. (A) The genotyping workflow (also see Methods). Our approach includes an optional SNV incorporation step, which can find heterozygous SNVs proximate to STRs and use them to cluster reads to call STR alleles. (B) Users can explore the generated report in the STRkit visualization tool. The read count histogram shows the read-level distribution of copy numbers for the locus. Here, a pathogenic expansion in the HTT gene is shown. The k-mer distribution plots show motif-sized k-mer sequence diversity among read STR sequences for each allele peak.

578f01
Table 1.

Feature comparison of long-read sequencing STR genotyping tools

FeatureSTRkitLongTRTRGTStraglrSTRdustNotes
Copy number output
Allele size confidence intervals
Allele consensus sequence outputStraglr v1.5.2+ VCFs use symbolic alternate alleles.
Beyond read-length genotypingPartialStraglr can use reads that cover only part of a repeat to support alleles, but it does not have a size model that incorporates in-repeat reads like some short-read STR genotypers.
De novo proximate SNV phasing✓ +outputSTRkit and TRGT can use heterozygous SNVs to cluster STR alleles; STRkit can output them to VCF.
Existing phased SNV incorporation✓✓
Haplotagged alignment file supportAll but Straglr can use phased read data from, for example, WhatsHap (Martin et al. 2016), to call STRs.
Methylation handling✓✓
ONT read supportExplicitly forbiddenTRGT theoretically works on ONT data but is forbidden by the software license.
Read-level dataPartialSTRkit: JSON output with read-level peak ID+sequence data; TRGT: overlapping reads in BAM; Straglr: TSV output with read-level copy numbers; STRdust: read-level sequence output in VCF.
Mendelian inheritance calculation tool✓✓STRkit includes a tool to output loci that do not respect Mendelian inheritance in a set of trio VCFs.
Free and open-source software licenseYes (GPLv3)Yes (GPLv2)Yes (GPLv3)Yes (MIT)TRGT's license restricts it to be only used with PacBio sequencing data, and the software cannot be forked and subsequently redistributed.
Multithreading/processing

[i] Comparison of features of STRkit with those of recent long-read sequencing STR genotyping tools: LongTR, TRGT, Straglr, and STRdust. Features with double check marks are unique to that particular tool.

Precise STR genotyping with long reads using proximate SNVs

To evaluate the performance of STRkit, we first obtained PacBio HiFi reads at ∼32× coverage, ONT R10 simplex reads that we subsetted to ∼32× coverage, and ONT R10 duplex reads at ∼12× coverage for the Genome-in-a-Bottle (GIAB) HG002 individual from the GIAB Ashkenazi trio. We then aligned these reads to the GRCh38 (hg38) reference genome (see Methods). Next, we genotyped about 900,000 STR regions using STRkit running in two modes: one mode incorporating SNV data to help call and locally phase alleles, and one mode without SNVs or other haplotype tag data.

Our benchmark STR regions are derived from an approximately 1.7 million tandem repeat benchmark variant set developed for HG002 (English et al. 2025) and released as an official benchmark by the GIAB consortium. The full benchmark covers ∼8.1% of the hg38 human reference genome, but not all of these benchmark regions are STRs; some have homopolymers or larger TRs, which we excluded to provide more precise estimates of performance (see Methods). We also removed regions with more than one tandem repeat because the tool we used to evaluate performance on the benchmark, Truvari (English et al. 2022), cannot handle proximate unphased variants. Our final set of 914,676 benchmark STRs covers ∼0.87% of hg38.

To evaluate STRkit (with and without SNV incorporation) against other tools, we also genotyped our benchmark set of STR loci using LongTR (Ziaei Jam et al. 2024), Straglr (Chiu et al. 2021), STRdust (De Coster et al. 2024), and TRGT (Dolzhenko et al. 2024). All tools were configured to not use any prior phasing data from read haplotype tags or phased sample SNVs. We then assessed their respective performance using Truvari and Laytr (English et al. 2025), except for Straglr, whose output is not compatible with Truvari. These results are summarized in Table 2. With HiFi data, STRkit with SNV incorporation performed best overall in most metrics (accuracy = 0.9767, F1 = 0.9633), although TRGT achieved marginally better recall (0.9697 vs. 0.9552 for STRkit). STRkit without SNVs performed second best in terms of precision, accuracy, and F1 score with HiFi. With ONT simplex data, we see similar rankings in caller performance, with slightly worse absolute performance across the board versus HiFi. Here, STRkit with SNVs (accuracy = 0.9709, F1 = 0.9544) performs best, followed by STRkit without SNVs (accuracy = 0.9641, F1 = 0.9440), LongTR (accuracy = 0.9613, F1 = 0.9354), and STRdust (accuracy = 0.8237, F1 = 0.7581). With low-coverage ONT duplex data, we again lose absolute performance in metrics versus both higher-coverage sequencing technology configurations, with the same tool ranking as with ONT simplex: STRkit with SNV incorporation performs best (accuracy = 0.9435, F1 = 0.9051), followed by STRkit without SNVs (accuracy = 0.9373, F1 = 0.8952) and then the other tools.

Table 2.

Performance metrics for long-read STR genotypers on the HG002 benchmark subset

CallerPrecisionRecallAccuracyF1Avg. TruScore% called
HG002, PacBio HiFi, ∼32× coverage
 STRkit0.97170.95450.97660.963198.8599.88%
 STRkit (-SNV)0.97010.94870.97420.959398.6299.87%
 LongTR0.95700.95410.97380.955598.8099.63%
 StraglrN/A97.18%
 STRdust0.79110.87900.88620.832795.7299.84%
 TRGT0.94740.96750.97220.957498.8099.87%
HG002, ONT R10 Simplex, ∼32× coverage
 STRkit0.95940.94940.97090.954498.4499.90%
 STRkit (-SNV)0.94880.93940.96410.944098.1499.88%
 LongTR0.93120.93970.96130.935498.2499.70%
 StraglrN/A96.86%
 STRdust0.67010.87270.82370.758194.8099.63%
HG002, ONT R10 Duplex, ∼12× coverage
 STRkit0.90980.90040.94350.905196.0399.78%
 STRkit (-SNV)0.89800.89240.93730.895295.6699.76%
 LongTR0.86820.88930.92960.878695.7199.59%
 StraglrN/A96.86%
 STRdust0.70590.84600.84430.769794.1199.63%

[i] Performance metrics for long-read STR genotypers on an STR-only subset of HG002 variants regions from the Genome-in-a-Bottle tandem repeats v1.0 benchmark, as measured by Truvari: precision, recall, accuracy, F1 score, and Truvari's TruScore allele similarity metric. All metrics except TruScore are only for variants with an allele size difference of ≥5 bp versus hg38. Straglr VCFs are not compatible with Truvari, and TRGT cannot use ONT data owing to a licensing restriction. Underlines indicate the best value for a metric within the sequencing technology.

Truvari also reports a compound variant quality score, called the “TruScore,” combining overlap, size similarity, and sequence similarity between a given call and the benchmark into an overall score (English et al. 2022). With HiFi data, STRkit with SNVs achieved the best average TruScore of 98.85, closely followed by LongTR and TRGT with 98.80, STRkit without SNVs with 98.62 and beating STRdust's average score (95.72) by a larger margin. With ONT simplex and duplex data, we observe the same relative ordering in TruScore.

Using the Laytr tool, we stratified performance metrics (F1 score, precision, recall) by the locus-maximum absolute change (Δ) in allele size from the reference genome (Supplemental Fig. S1). Most STR Δ allele sizes are small (95.5% of allele size changes are within 50 bp of the hg38 reference genome), and that is where STRkit with HiFi data shows a slight but measurable improvement over other tools in terms of F1 score and precision. With ONT simplex and duplex data for small and intermediate allele size changes versus the reference (within 200 bp of hg38), STRkit performed markedly better versus other tools. With larger Δ allele size, all tools performed worse with both sequencing technologies, with TRGT appearing to perform the best with these loci when it could be used (PacBio only). It was with these larger Δ allele sizes that STRkit's SNV incorporation helps its genotyping performance the most, although only in the case of the low-coverage ONT duplex data did STRkit with SNVs perform the best overall with medium-large Δ allele sizes (300–1000 bp).

Our evaluation of STRkit without the SNV incorporation step confirmed that using SNV data improves STR genotyping. All metrics output by Truvari for STRkit improved when the SNV step was included, for both HiFi and both types of ONT data (Table 2). For most loci in the call set (∼68%–78% depending on individual and technology), STRkit used information from at least one heterozygous SNV for clustering reads into alleles (Supplemental Table S1).

Separately, we confirmed the accuracy of STRkit's proximate SNV calling and local phasing step by comparing SNV output to a phased small variant benchmark set published by the GIAB consortium (see Methods) (Wagner et al. 2022). Across the trio, we found an incorrect call rate of 0.008% (HiFi 32×), 0.025% (ONT simplex 32×), or 0.126% (ONT duplex 12×), considering only SNVs shared between the STRkit output and the benchmark. To look for phasing errors, we took phase sets with more than one SNV and compared phased SNV calls with the benchmark variant set. Here, we found a phase-set error rate (with one or more “flipped” SNVs) of 0.31% (HiFi), 0.34% (ONT simplex), and 0.51% (ONT duplex). These error rates are detailed in Supplemental Table S2.

High-quality genotypes track STR allele inheritance in trios

STRkit also includes a tool for calculating the rate of MI of alleles from parents to offspring, in terms of copy number and allele sequence/length. Loci are considered to follow MI if allele sequences from the child can be found in the parents. This process also identifies loci that do not meet expectations of MI in the trio: sites of potential de novo mutation. Versions of this metric have been used to evaluate other STR genotyping software, for example, GangSTR (Mousavi et al. 2019). On the Ashkenazi trio with HiFi data, STRkit achieves a sequence-wise (seq.) MI rate of 97.56% (97.98% in terms of seq. length, 99.04% in terms of seq. length ±1bp). Without SNVs, STRkit achieves a lower seq. MI rate of 93.64% (95.58% seq. len., 98.42% ±1 bp). LongTR and TRGT again perform almost as well as STRkit with HiFi data; LongTR achieves a seq. MI rate of 94.52% (94.85% seq. len., 97.54% ±1 bp), and TRGT achieves a seq. MI rate of 94.77% (97.19% seq. len., 98.98% ±1 bp). On the same trio with ONT simplex data, rates of MI are slightly lower across the board, with STRkit (with SNVs) once again performing the best in seq. MI rate (94.44%), seq. length MI (94.99%), and seq. len. MI ±1 bp (97.96%), closely followed by LongTR and STRkit without SNVs. With both sequencing technologies, incorporating SNVs to the read clustering process improves allele sequence quality with STRkit. With both HiFi and ONT simplex data, Straglr and STRdust achieve significantly lower rates of MI versus other tools. These results are summarized in Supplemental Table S3.

We then stratified sequence-wise MI by reference genome locus length (Fig. 2). The bulk of loci are small, which is again where STRkit has the clearest advantage over other methods when incorporating SNVs with both HiFi and ONT simplex. Above ∼200 bp in length, the rates drop off, as STRs become harder to sequence accurately and are more likely to accumulate mutations in their passage from parent to child.

Figure 2.

Rates of sequence Mendelian inheritance (MI), and number of trio-called loci, by reference genome locus length. Expectations of MI are met if both child allele sequences can be seen in the two parents (one in the mother and one in the father). The vast majority of STR regions in the catalog are below ∼200 bp long, at which STRkit (with SNV incorporation) achieves the highest rate.

578f02

Finally, to test the callers’ performance beyond the benchmark trio, we examined MI in a set of 30 trios sequenced using PacBio HiFi data from the Genomic Answers for Kids (GA4K) cohort (Cohen et al. 2022). Most (26/30) trios tested were sequenced at low depth (<10×; parental coverage as low as 3×), two at medium depth (10×–15×), and two at high depth (∼37× in the proband, ∼22× in the parents). Here, we see that STRkit (with or without SNVs) and TRGT outperform LongTR, STRdust, and Straglr (Supplemental Fig. S2), with TRGT appearing to perform slightly better than STRkit with copy number–wise MI in the two high-coverage trios, and no significant differences between STRkit and TRGT with low-coverage trios or with sequence/sequence length–wise MI.

Computational performance of SNV-informed STR genotyping

Most genotyping tools tested (LongTR, STRkit, and TRGT) can analyze the filtered HG002 benchmark catalog in under a day measured in core-minutes using ∼32× coverage PacBio HiFi/ONT simplex data or ∼12× ONT R10 duplex data (Supplemental Table S4). STRkit is slower than TRGT and LongTR but is on the same order of magnitude (in the hundreds of core-minutes irrespective of technology); in contrast, Straglr is in the thousands of core-minutes, and STRdust is the slowest overall, requiring thousands of core-minutes for HiFi/low-coverage ONT and tens of thousands of core-minutes for the 32× ONT simplex samples. Where possible, we used eight cores of an AMD EPYC 7532 processor, although LongTR does not support using multiple CPU cores. With HiFi data, TRGT was the fastest overall. For both types of ONT data, LongTR was fastest overall, although STRkit was the fastest program supporting multicore parallelization. STRkit consumes more memory than TRGT or LongTR with HiFi data on average but consumes less than STRdust and Straglr, and the SNV incorporation step increases its memory usage. LongTR is the most memory-efficient STR caller for the GIAB trio samples sequenced with both ONT sequencing data sets.

Read-level genotyping of targeted LRS data captures known instability in a pathogenic expansion

We ran STRkit and other tandem repeat callers on targeted CCS sequencing data of pathogenic repeat expansions in the HTT and FMR1 genes (Methods) (Supplemental Table S5). We ran STRkit using its targeted sequencing mode, without SNV incorporation as the targeted sequencing data lack sufficient flanking sequence to make use of proximate SNVs. The callers successfully detected the repeat expansions in almost all cases, except for one HTT expansion missed by Straglr and one FMR1 expansion missed by STRdust. The genotypes overall were highly concordant.

Many of the tools tested also provide some degree of read-level output, allowing further expansion analysis. STRkit and Straglr both output read-level copy numbers, which we can use to examine an instance of somatic instability in this data set. De Luca et al. (2021) found that one of the expansion samples tested, NA20253, has three copy number peaks, namely, mosaic alleles, which they measured at 107, 134, and 175 CAG repeats, with the first being the most prominent. Figure 3 shows the read-level data for this expanded allele from STRkit's output, along with annotations showing the peaks found by De Luca et al. (2021); we observe three read peaks that mirror those found using their triplet repeat–primed PCR assay, but slightly shifted toward a higher copy number. Evidence of these three peaks is also visible in the read-level copy number data output by Straglr.

Figure 3.

HTT HD-causing STR allele copy number distribution in targeted sequencing of the NA20253 sample. Vertical lines show the three expansion copy number peaks found by De Luca et al. (2021), which fall near read peaks in read-level output: two mosaic alleles, at 134 and 175 repeats, with the main allele peak (with the majority of reads) at 107 repeats.

578f03

STRkit also can output read-level motif-sized k-mer data, which can be used to explore noncanonical motif content in expansions. We used this output to visualize canonical and noncanonical motifs in FMR1 expansion–containing samples, showing small differences in noncanonical motif content between samples (Supplemental Fig. S3).

Discussion

Here, we present STRkit, a new tool for calling STR genotypes from LRS data. Our tool has a unique combination of features, filling a niche in the STR genotyping analytical space: an STR caller that can operate at a genome-wide scale with both major LRS technologies, precisely genotype both sequence and copy number, optionally incorporate proximate SNVs into the calling process and use them to phase STRs locally, and report read-level data (Table 1). Both the STR copy number and sequence are phenotypically relevant properties (Olson et al. 2023); however, LongTR and STRdust only output sequence, and Straglr (as of v1.5.2) outputs only copy number. Of the tested tools, all are independent of LRS technology except TRGT, which can only be used with PacBio data owing to a licensing restriction. Technology-independent tools are critical, because they enable repeatable research and community-inclusive development, allow for confirmation of findings across technologies, and allow the best genotyping approach to be evaluated and applied regardless of data source, for example, the cross-technology benchmarking we performed to evaluate STRkit. Our benchmarking, except for a set of trios used in our MI analysis, also used open data, tools, and scripts (see Methods).

STRkit, when incorporating proximate SNVs into the genotyping process, demonstrates state-of-the-art performance in several benchmark metrics with both PacBio HiFi data and modern nanopore sequencing data (ONT R10 simplex and duplex). Although performance gains are modest with HiFi data versus the next best-performing tool, LongTR (Table 2), maximizing precision and accuracy has direct benefits for genome-wide analysis, in which the statistical power in correlation tests would be reduced by lower-quality genotypes. At similar depths of sequencing coverage, HiFi data outperform ONT data, a pattern that has been observed in other long-read variant call benchmarking (Harvey et al. 2023). With ONT data, our performance gains were larger, achieving a ∼2%–3% higher F1 score (for high-coverage simplex and low-coverage duplex, respectively) than LongTR and narrowing the performance gap between ONT and HiFi data for calling STRs. Our SNV incorporation process allows clustering STR-overlapping reads into alleles without relying solely on copy number. This aids our genotyping approach to a small but measurable degree (Table 2), and for most loci, we can incorporate SNV information into the STR calling process (Supplemental Table S1), although it may not be the primary explanation for STRkit's improved benchmark metrics versus other tools, especially with ONT data (Table 2; Supplemental Fig. S1). SNV data captured by STRkit may prove useful for more than genotyping; in downstream analysis, they could aid, for example, trio phasing or STR mutation modeling, although our current approach only calls heterozygous STR-proximate SNVs. When we assessed our SNV calls against a small variant benchmark for this trio, we found very low error rates in terms of incorrect calls or phasing errors (Supplemental Table S2), supporting their useability in downstream analysis. STRkit can also function without SNV data, which can be useful for targeted sequencing data sets in which SNV-containing flanking regions are not present, as well as for reference genomes (human or otherwise) in which an SNV catalog is not available. It also may be helpful for discovering low-coverage alleles (e.g., certain expansions), to avoid discarding reads without sufficient flanking information.

When metrics are faceted by allele size difference (Δ allele size) versus the reference genome, tools generally performed well until Δ allele size reached ∼2.5 kb or longer (Supplemental Fig. S1), except for a noticeable dip in genotyping quality in the 3–500 bp range. Manual inspection indicates that some of these may be regions with noncanonical motifs or non-tandem-repeat sequences embedded within a tandem repeat locus, which may impact STRkit's modeling of candidate repeat tract sequences or result in reads being filtered out to a greater degree versus other tools; future work could implement alternate allele clustering systems that do not assume an entirely tandem repeat sequence structure.

Of the five callers, STRkit, LongTR, and TRGT best accommodated the large catalog of approximately 900,000 loci, with relatively quick runtimes (in the 100 sec of core-minutes) and low memory requirements (Supplemental Table S4). Our general approach to genotyping does add a computational overhead, as it is slower to genotype the benchmark catalog versus LongTR and TRGT. We believe that this is because of our additional read-level data collection and copy number determination approach; future performance improvements may be achievable from optimizing these procedures. STRdust and Straglr both had significantly longer runtimes (1000 sec of core-minutes), and Straglr used an order of magnitude more memory than the other tools. STRkit is the only one of the three most performant tools (in terms of both genotype quality and computation speed) that can report full read-level copy number data, meaning it may be useful for genome-wide somatic instability detection. LongTR's lack of parallel processing support means it avoids overhead, at the cost of longer per-sample runtime, whereas other methods were evaluated using eight CPU cores, requiring interprocess communication. STRkit's shared SNV context, used for local phasing of STR genotypes, results in slower-than-linear performance gains with more CPU cores, trading off additional computation for increased genotype quality. Reducing the overhead of this interprocess communication could yield future performance improvements.

In our trio analysis for MI, we quantified the proportion of STR loci in which both alleles in the HG002 child were found in the parents. STRkit had the best performance across the four versions of the metric we tested in this trio (Supplemental Table S3). We also analyzed rates of MI in a set of 30 trios sequenced with HiFi, most of which were sequenced at low depths of coverage (Supplemental Fig. S2); low coverage almost certainly resulted in allelic dropout, but in general, STRkit and TRGT outperform the other tools tested here. MI is limited as a benchmarking metric, however, because systematic errors such as off-by-one errors that affect both parent and child may not be adequately considered. Despite this, our high MI rates, when combined with high absolute scores on the benchmark (which are less affected by these forms of systematic error, because they are compared to a truth set), indicate that STRkit should be the most accurate currently available tool for genome-scale STRs genotyping and potentially for detecting the most common category of de novo STR mutation events: ±1 repeat unit changes. However, without a more complete mutation detection procedure with statistical confidence measures to control for false discovery, we cannot claim that all loci that do not respect MI are true de novo mutations. Future benchmarking could also include more robust trio-aware approaches using, for example, trio assemblies or trio-aware phasing data, or could involve developing simulated benchmark data, for example, by incorporating an STR error model into a simulation tool like PBSIM3 (Ono et al. 2022).

So far, we have discussed genome-wide benchmarking results from about 900,000 tandem repeat loci. Very few STR loci have known pathogenic expanded forms. Gall-Duncan et al. (2022) reported 63 as of 2021. STRkit provides a unique feature-set useful for close examination of specific expansion loci. To demonstrate this, we used its targeted sequencing mode (which changes read resampling behavior for use with targeted sequencing) to genotype STRs with known expansions in the HTT or FMR1 genes (Supplemental Table S5), for which, like other tools, it correctly detects the expanded allele. Unlike other tools, STRkit can both generate a whole-allele consensus sequence and report read-level copy number and sequence data, including motif-sized k-mer distributions (Fig. 1B; Supplemental Fig. S3), at the same time. STRkit does not directly detect somatic instability, unlike a tool such as prancSTR (Sehgal et al. 2024); however, read-level copy number data (reported by STRkit and Straglr) replicate a multipeak HTT mosaic expansion previously captured using a repeat-primed PCR technique by De Luca et al. (2021) (Fig. 3). The peaks visible in the sequencing data are slightly offset from the reported peaks in this expansion, possibly owing to the passage of the sampled cell line. This mosaic expansion revealed itself only with high-depth targeted sequencing and analysis of read-level data, demonstrating that read-level data can be crucial for characterizing expanded alleles.

Beyond somatic instability detection, there are other STR genotyping features STRkit does not yet implement. Methylation reporting using HiFi's base methylation data, for example, remains exclusive to TRGT. Our calling approach also has limitations: Reads are primarily clustered based on motif copy number if SNV or haplotype information is not available; this is useful for obtaining copy number confidence intervals but may miss more subtle allelic sequence variation. In the absence of SNV data, our tool may also miss variation when faced with large allele differences versus the hg38 reference genome to a greater degree than other methods (Supplemental Fig. S1). Every tool tested here, with the exception of Straglr, cannot genotype alleles beyond the length of a long read (Table 1); in the vast majority of cases, this will not pose issues, but this can cause reduced sequencing coverage or complete allele dropout with some very large expansions that have been observed in, for example, individuals affected by SCA10 (Matsuura et al. 2000). This can be mitigated through the use of ultra-LRS technology, for example, ultralong nanopore reads (Jain et al. 2022); long-read genotyping tools could also employ longer-than-read-length allele-calling methods similar to those used by short-read STR genotyping tools like GangSTR (Mousavi et al. 2019) or ExpansionHunter (Dolzhenko et al. 2019). Finally, STRkit, like most of the other tools tested, is a catalog-based genotyper; it requires a list of loci with reference genome coordinates and a motif or IUPAC-coded motif pattern. This limits what can be genotyped to only what is already present in the reference, and it may not function well with nonreference motifs within known loci or may miss certain disease-causing expansions completely (Tanudisastro et al. 2024). Annotating an assembly or using a de novo approach such as ExpansionHunter Denovo (Dolzhenko et al. 2020) or Straglr without a provided catalog can find repeats that STRkit and others cannot.

STRs cover ∼5% of the human genome, are implicated in genetic disorders and transcriptional regulation, and are major contributors to overall genomic variation. STRs are important to genotype directly, as they cannot always be imputed with single-nucleotide variations (SNVs), but these elements are difficult to resolve with traditional sequencing methods and complex to analyze. They are multiallelic and defined by both their length and motif composition, rather than a simple binary genotype. STRkit uses accurate long-read data and a strategy that incorporates both SNVs and copy number genotyping into the STR genotyping process to achieve better performance in several metrics and production of local STR-SNV phasing data. The tool also includes read-level copy number and motif composition data for STR loci, whereas many other methods do not, enabling analysis of intra-allele variation in STR copy number and motif composition. The approach fills a niche in the STR genotyping software ecosystem and opens up new possibilities for association testing, assessing patterns of STR inheritance, and better understanding the functional effects of these repeat elements.

Methods

STR genotyping approach

The general genotyping approach used in STRkit is as follows:

  1. The user provides a BAM or CRAM-formatted LRS alignment file, generated using an aligner such as minimap2 (Li 2018), as well as a BED-formatted catalog of STRs to genotype. Optionally, they also provide a VCF with a list of SNVs to genotype if proximate to an STR.

  2. For each STR in the catalog, we optionally redetermine the copy number of the STR in the reference genome to allow for imperfect reference motif copies and adjust the catalog coordinates to correct for small boundary errors.

  3. The set of all reads spanning each STR region (including flanking sequences on either side for realignment) is extracted using the PySAM library (v0.22.0; https://github.com/pysam-developers/pysam).

  4. Low-quality or badly aligned reads are filtered out. If realignment is enabled, badly aligned reads may be realigned to the reference genome.

  5. For each read, a series of candidate STR tracts of varying copy number, starting with a candidate closest in length to the aligned read's STR segment size, is generated using the cataloged motif. Each of these candidate STR tracts is aligned to the read using the parasail library's semiglobal alignment algorithm (Daily 2016), incorporating 5′ and 3′ flanking sequences to “anchor” the STR candidate and allow for complete STR tract deletion. The best-scoring alignment candidate is kept for each read in question. After we have finished this process, we have the approximate copy numbers for each read. If the read-level k-mer functionality is enabled, motif-sized k-mers are computed for each read.

  6. If an SNV VCF catalog is provided: For each read, SNVs that differ from the reference are collected; these are then filtered to those that occur in many reads, are of high base-level quality, and can split the reads into two groups, that is, likely a heterozygous SNV and not a sequencing error.

  7. For each read, a resampling weight is generated, corresponding to the estimated inverse probability of observing a read encompassing an STR tract of the same size. This process adds weight to reads with longer STR tracts. Given random read localization on a reference and a read size distribution, longer STR alleles are less likely to be entirely “captured” by a read; in this process, allele size bias is corrected for, and long expansion tracts can be better captured. Reads (with associated copy numbers and SNV calls) are then resampled a number of times (defaulting to 100), according to the resampling weight.

  8. From here, multiple possibilities are available to call allele peaks:

    1. If the locus is haploid (e.g., on the Y sex chromosome of an XY individual), a Gaussian distribution is fitted to the copy numbers.

    2. If there is haplotype information in the alignment file and ‐‐use-hp is specified, the haplotype information for the reads is used if available. Otherwise, one of the next three methods is used.

    3. If there is a lot of informational SNVs for the locus, SNVs are used to split the reads into allele peaks, because SNV assessment is generally very accurate in NGS data. A distance matrix is calculated, and an agglomerative clustering function from the scikit-learn library (Pedregosa et al. 2011) is applied to group the reads into alleles.

    4. If there is some SNV information available, SNV data are used, but copy number is incorporated into a compound distance function, in which SNV differences are weighed more heavily (5:1 or 10:1) versus copy number differences. Agglomerative clustering is used here as well.

    5. If no or limited SNV information is available, a GMM from scikit-learn is applied to each resampling to derive estimates of allele copy numbers via a bootstrap process. Each allele starts with a two-peak GMM; if one peak is assigned very little explanatory weight after fitting, we switch to a single-peak GMM (i.e., a homozygous model).

      If the organism is diploid, reads are assigned to peaks based on SNV genotypes (if available) or bootstrapped estimates of peak mean, standard deviation, and weight: the complete set of parameters characterizing a peak in a GMM. If the peaks overlap to a high degree and no informative SNV genotypes are available, reads are instead randomly assigned to peaks.

  9. Final copy number genotype estimates, consensus sequences, and copy number confidence intervals are computed. If the k-mer functionality is enabled, motif-sized k-mers are computed for each allele.

  10. A final report is generated with STR genotypes, parameters used, peak assignments (and peak assignment method used) for each read, copy numbers determined for each read, standard deviations for peaks, and optionally

    1. Read-level SNV calls and

    2. Read- and/or peak-level motif-size k-mer quantification.

Benchmarking genome-scale tandem repeat genotyping

We obtained unaligned HiFi LRS data for the HG002 individual from PacBio (https://downloads.pacbcloud.com/public/revio/2022Q4/). We obtained ONT R10 stereo duplex LRS data for HG002 from the Human Pangenomics Project (https://human-pangenomics.s3.amazonaws.com/index.html?prefix=submissions/0CB931D5-AE0C-4187-8BD8-B3A9C9BFDADE--UCSC_HG002_R1041_Duplex_Dorado/Dorado_v0.1.1/stereo_duplex/). We obtained ONT R10 simplex LRS data for HG002 from EPI2ME (https://42basepairs.com/browse/s3/ont-open-data/giab_2025.01/). We aligned all these read-sets to the UCSC hg38 analysis set reference genome (https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/analysisSet/hg38.analysisSet.fa.gz) using minimap2 v2.28, with the map-hifi and lr:hq presets for the HiFi and ONT R10 (simplex and duplex) data, respectively.

English et al. (2025) have published a genome-wide tandem repeat benchmark via the GIAB consortium for the HG002 Ashkenazim individual. We obtained version 1.0 of this benchmark from the NCBI (https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/AshkenazimTrio/HG002_NA24385_son/TandemRepeats_v1.0/GRCh38/) and corresponding benchmark region annotations from Zenodo (https://doi.org/10.5281/zenodo.6930201). This benchmark includes 17,84,804 regions, some with multiple tandem repeat loci defined in the same region. The authors extended the Truvari software (English et al. 2022) to support benchmarking region-based STR calls. For benchmarking, we used Truvari version 5.3.0. We filtered out some complex or non-STR tandem repeat definitions from their benchmark region catalog, using the following steps:

  1. Motifs were normalized (e.g., CACACA → CA).

  2. Homopolymers and tandem repeats with motif length >10 bp were removed.

  3. Benchmark regions with more than one tandem repeat were removed, because unphased proximate haplotypes cannot be correctly evaluated with Truvari.

Additionally, homopolymers were filtered from the HG002 benchmark variant VCF to mitigate erroneous false negatives when homopolymers were present near other STR regions. To combine and visualize the benchmarking results after running Truvari, we used Laytr (commit 9cc1910; https://github.com/ACEnglish/laytr). Applying this filtering, we ended up with a catalog of 9,14,676 loci.

For benchmarking STRkit, we chose four recently developed general-purpose LRS STR genotyping tools, all of which take standard BAM/CRAM-formatted alignment files and can output VCFs, for use with Truvari. The set of tools, with their versions used for benchmarking, are as follows:

The scripts that we used to benchmark these tools, alongside the parameters used for benchmarking, can be found in the tool batch scripts located at GitHub (https://github.com/davidlougheed/strkit_paper/tree/main/2_giab_calls). This folder also contains scripts used to generate Supplemental Figure S1 and compute the summary data found in Table 2 and the Supplemental Tables.

To calculate the approximate STRkit SNV call and phasing error rates, we compared (1) all SNV calls with a phased version of the small variant benchmark set for the HG002 individual (Wagner et al. 2022; obtained from https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/AshkenazimTrio/HG002_NA24385_son/NISTv4.2.1/GRCh38/SupplementaryFiles/HG002_GRCh38_1_22_v4.2.1_benchmark_hifiasm_v11_phasetransfer.vcf.gz) and (2) multi-SNV phase sets with phased SNVs in this benchmark set. SNV errors were counted as either a “false heterozygote” (called by STRkit as heterozygous but homozygous in the benchmark set) or “other errors” (nonmatching calls). A phase set was counted as having an error if at least one SNV in the phase set had a flipped genotype versus the benchmark.

Assessing computational performance

All STR genotyping tools were run on the Digital Research Alliance of Canada's Narval cluster, using eight cores of an AMD EPYC 7532 processor or one core when parallel processing was not supported (i.e., LongTR). Compute time was normalized into CPU minutes by multiplying runtime by the number of cores used.

Calculating MI in call sets

We obtained unaligned HiFi reads for the Ashkenazi trio from PacBio (https://downloads.pacbcloud.com/public/revio/2022Q4/) and unaligned ONT R10 simplex reads for this trio from EPI2ME (https://42basepairs.com/browse/s3/ont-open-data/giab_2025.01/). We aligned both read sets to the UCSC hg38 analysis set reference genome using minimap2 v2.28 with the map-hifi and lr:hq presets for HiFi and ONT data, respectively.

We used trio data for 30 trios (with a mixture of low and high sequencing coverage) from the GA4K data to perform additional MI analysis. GA4K data are available under controlled access via the NCBI database of Genotypes and Phenotypes (dbGaP; https://dbgap.ncbi.nlm.nih.gov/home/) and hosted on the AnVIL platform under the accession number phs002206.v5.p1.

In STRkit, we include a module named strkit mi to calculate rates of trio MI in autosomal loci from the output of all STR genotyping tools we benchmarked. STRkit can calculate MI in terms of binary “yes/no” copy number inheritance, inheritance ±1 repeat unit, parent–offspring 95% confidence interval overlap, exact sequence inheritance, sequence length, and sequence length ±1 bpr. For MI calculation, the tool performs the following steps:

  1. For each locus in the STR call set for the child in the trio, check if a corresponding call can be found in both parent call sets. If not, skip the locus.

  2. If the locus is in a user-specified exclusion set, for example, for removing known de novo mutation, skip it.

  3. Add this locus to the set S of seen loci for this trio.

  4. If (slightly different from step 1) a call failed in any of the three individuals, skip the locus (after it has been added to set S).

  5. If the calls for the locus are consistent with MI (for each version of the MI metric X ∈ {strict, ±1 repeat, 95% CI, sequence, seq. len., seq. len. ±1 bp}), add one to a counter cX of consistent loci for metric X. Some versions of the metric are not supported by some of the genotyping methods; LongTR and STRdust do not provide copy numbers or 95% copy number confidence intervals, and Straglr does not provide 95% copy number confidence intervals or consensus allele sequences.

  6. After all loci have been processed, return the total rate of MI as c/|S|.

Genotyping expansions from targeted sequencing data

We accessed a data set of PacBio targeted sequencing of expansions at https://downloads.pacbcloud.com/public/dataset/RepeatExpansionDisorders_NoAmp/.

The same STR genotyping tool versions used in our benchmarking were compared with this targeted disease expansion sequencing data set. We specified sex chromosome configurations for each sample to genotype the FMR1 locus. Batch scripts with specific parameters, as well as the script used to generate Figure 3, are available at GitHub (https://github.com/davidlougheed/strkit_paper/tree/main/4_pathogenic_exp). Coriell genotypes were manually collected from https://www.coriell.org/ Oct 4, 2022.

Code availability

The source code for the STRkit package is available at GitHub (https://github.com/davidlougheed/strkit/) and at Zenodo (https://doi.org/10.5281/zenodo.12689906) and is obtainable as a Python package in the Python package index (PyPI) under the name strkit. The analysis scripts from this manuscript are available at GitHub (https://github.com/davidlougheed/strkit_paper/) and as Supplemental Code. STRkit version 0.23.0, used for the analyses in this manuscript, is available as Supplemental Code.

Competing interest statement

The authors declare no competing interests.

Acknowledgments

We thank Rob Eveleigh and Mathieu Bourgey for providing valuable instruction on long-read data and benchmarking approaches, José Héctor Gálvez López for his important feedback on our genotyping approach, Elia Afanasiev for his help increasing computational performance and for proofreading the manuscript, Boriana Apostolov and Stephen C. Lougheed for proofreading the manuscript, Nicolas Buitrago for his work on the Rust parasail bindings, and Adam C. English for his work and support with the Truvari and Laytr tools. We acknowledge Calcul Quebec, SecureData4Health, and the Digital Research Alliance of Canada for computing resources. This work was supported by the Natural Sciences and Engineering Research Council of Canada (RGPIN-2024-05329) and the Canadian Institutes of Health Research (PJT-191707). G.B. is supported by a Canada Research Chair Tier 1 award and a Fonds de Recherche du Québec - Santé (FRQ-S), Distinguished Research Scholar award.

Author contributions: D.R.L. developed the software package and performed the benchmarking analyses. G.B. and D.R.L. conceived of the project. T.P. gave guidance and insight during the development of the project. G.B. supervised the project. All authors read and approved the final manuscript.

Notes

[3] Supplementary material [Supplemental material is available for this article.]

[4] Article published online before print. Article, supplemental material, and publication date are at https://www.genome.org/cgi/doi/10.1101/gr.280766.125.

References

  1. Allen EG, Charen K, Hipp HS, Shubeck L, Amin A, He W, Nolin SL, Glicksman A, Tortora N, McKinnon B, 2021. Refining the risk for fragile X–associated primary ovarian insufficiency (FXPOI) by FMR1 CGG repeat size. Genet Med 23: 1648–1655. 10.1038/s41436-021-01177-y
  2. Benson G. 1999. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res 27: 573–580. 10.1093/nar/27.2.573
  3. Blauw HM, van Rheenen W, Koppers M, Van Damme P, Waibel S, Lemmens R, van Vught PWJ, Meyer T, Schulte C, Gasser T, 2012. NIPA1 polyalanine repeat expansions are associated with amyotrophic lateral sclerosis. Hum Mol Genet 21: 2497–2502. 10.1093/hmg/dds064
  4. Bragg DC, Mangkalaphiban K, Vaine CA, Kulkarni NJ, Shin D, Yadav R, Dhakal J, Ton M-L, Cheng A, Russo CT, 2017. Disease onset in X-linked dystonia-parkinsonism correlates with expansion of a hexameric repeat within an SVA retrotransposon in TAF1. Proc Natl Acad Sci 114: E11020–E11028. 10.1073/pnas.1712526114
  5. Brinkman RR, Mezei MM, Theilmann J, Almqvist E, Hayden MR. 1997. The likelihood of being affected with Huntington disease by a particular age, for a specific CAG size. Am J Hum Genet 60: 1202–1210.
  6. Chiu R, Rajan-Babu I-S, Friedman JM, Birol I. 2021. Straglr: discovering and genotyping tandem repeat expansions using whole genome long-read sequences. Genome Biol 22: 224. 10.1186/s13059-021-02447-3
  7. Cohen ASA, Farrow EG, Abdelmoity AT, Alaimo JT, Amudhavalli SM, Anderson JT, Bansal L, Bartik L, Baybayan P, Belden B, 2022. Genomic answers for children: dynamic analyses of >1000 pediatric rare disease genomes. Genet Med 24: 1336–1348. 10.1016/j.gim.2022.02.007
  8. Cui Y, Arnold FJ, Li JS, Wu J, Wang D, Philippe J, Colwin MR, Michels S, Chen C, Sallam T, 2025. Multi-omic quantitative trait loci link tandem repeat size variation to gene regulation in human brain. Nat Genet 57: 369–378. 10.1038/s41588-024-02057-2
  9. Daily J. 2016. Parasail: SIMD C library for global, semi-global, and local pairwise sequence alignments. BMC Bioinformatics 17: 81. 10.1186/s12859-016-0930-z
  10. De Coster W, Hoijer I, Bruggeman I, D'Hert S, Melin M, Ameur A, Rademakers R. 2024. Visualization and analysis of medically relevant tandem repeats in nanopore sequencing of control cohorts with pathSTR. Genome Res 34: 2074–2080. 10.1101/gr.279265.124
  11. De Luca A, Morella A, Consoli F, Fanelli S, Thibert JR, Statt S, Latham GJ, Squitieri F. 2021. A novel triplet-primed PCR assay to detect the full range of trinucleotide CAG repeats in the huntingtin gene (HTT). Int J Mol Sci 22: 1689. 10.3390/ijms22041689
  12. Dolzhenko E, van Vugt JJFA, Shaw RJ, Bekritsky MA, van Blitterswijk M, Narzisi G, Ajay SS, Rajan V, Lajoie BR, Johnson NH, 2017. Detection of long repeat expansions from PCR-free whole-genome sequence data. Genome Res 27: 1895–1903. 10.1101/gr.225672.117
  13. Dolzhenko E, Deshpande V, Schlesinger F, Krusche P, Petrovski R, Chen S, Emig-Agius D, Gross A, Narzisi G, Bowman B, 2019. ExpansionHunter: a sequence-graph-based tool to analyze variation in short tandem repeat regions. Bioinformatics 35: 4754–4756. 10.1093/bioinformatics/btz431
  14. Dolzhenko E, Bennett MF, Richmond PA, Trost B, Chen S, van Vugt JJFA, Nguyen C, Narzisi G, Gainullin VG, Gross AM, 2020. ExpansionHunter Denovo: a computational method for locating known and novel repeat expansions in short-read sequencing data. Genome Biol 21: 102. 10.1186/s13059-020-02017-z
  15. Dolzhenko E, English A, Dashnow H, De Sena Brandine G, Mokveld T, Rowell WJ, Karniski C, Kronenberg Z, Danzi MC, Cheung WA, 2024. Characterization and visualization of tandem repeats at genome scale. Nat Biotechnol 42: 1606–1614. 10.1038/s41587-023-02057-3
  16. Duyao M, Ambrose C, Myers R, Novelletto A, Persichetti F, Frontali M, Folstein S, Ross C, Franz M, Abbott M, 1993. Trinucleotide repeat length instability and age of onset in Huntington's disease. Nat Genet 4: 387–392. 10.1038/ng0893-387
  17. Ellegren H. 2004. Microsatellites: simple sequences with complex evolution. Nat Rev Genet 5: 435–445. 10.1038/nrg1348
  18. English AC, Menon VK, Gibbs RA, Metcalf GA, Sedlazeck FJ. 2022. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biol 23: 271. 10.1186/s13059-022-02840-6
  19. English AC, Dolzhenko E, Ziaei Jam H, McKenzie SK, Olson ND, De Coster W, Park J, Gu B, Wagner J, Eberle MA, 2025. Analysis and benchmarking of small and large genomic variants across tandem repeats. Nat Biotechnol 43: 431–442. 10.1038/s41587-024-02225-z
  20. Erwin GS, Gürsoy G, Al-Abri R, Suriyaprakash A, Dolzhenko E, Zhu K, Hoerner CR, White SM, Ramirez L, Vadlakonda A, 2023. Recurrent repeat expansions in human cancer genomes. Nature 613: 96–102. 10.1038/s41586-022-05515-1
  21. Fan H, Chu J-Y. 2007. A brief review of short tandem repeat mutation. Genomics Proteomics Bioinformatics 5: 7–14. 10.1016/S1672-0229(07)60009-6
  22. Gall-Duncan T, Sato N, Yuen RKC, Pearson CE. 2022. Advancing genomic technologies and clinical awareness accelerates discovery of disease-associated tandem repeat sequences. Genome Res 32: 1–27. 10.1101/gr.269530.120
  23. Gustafson JA, Gibson SB, Damaraju N, Zalusky MPG, Hoekzema K, Twesigomwe D, Yang L, Snead AA, Richmond PA, De Coster W, 2024. High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive catalog of human genetic variation. Genome Res 34: 2061–2073. 10.1101/gr.279273.124
  24. Gymrek M, Willems T, Guilmatre A, Zeng H, Markus B, Georgiev S, Daly MJ, Price AL, Pritchard JK, Sharp AJ, 2016. Abundant contribution of short tandem repeats to gene expression variation in humans. Nat Genet 48: 22–29. 10.1038/ng.3461
  25. Hannan AJ. 2018. Tandem repeats mediating genetic plasticity in health and disease. Nat Rev Genet 19: 286–298. 10.1038/nrg.2017.115
  26. Harvey WT, Ebert P, Ebler J, Audano PA, Munson KM, Hoekzema K, Porubsky D, Beck CR, Marschall T, Garimella K, 2023. Whole-genome long-read sequencing downsampling and its effect on variant-calling precision and recall. Genome Res 33: 2029–2040. 10.1101/gr.278070.123
  27. Horton CA, Alexandari AM, Hayes MGB, Marklund E, Schaepe JM, Aditham AK, Shah N, Suzuki PH, Shrikumar A, Afek A, 2023. Short tandem repeats bind transcription factors to tune eukaryotic gene expression. Science 381: eadd1250. 10.1126/science.add1250
  28. Igarashi S, Tanno Y, Onodera O, Yamazaki M, Sato S, Ishikawa A, Miyatani N, Nagashima M, Ishikawa Y, Sahashi K, 1992. Strong correlation between the number of CAG repeats in androgen receptor genes and the clinical onset of features of spinal and bulbar muscular atrophy. Neurology 42: 2300–2302. 10.1212/WNL.42.12.2300
  29. International Human Genome Sequencing Consortium. 2001. Initial sequencing and analysis of the human genome. Nature 409: 860–921. 10.1038/35057062
  30. Jain C, Rhie A, Hansen NF, Koren S, Phillippy AM. 2022. Long-read mapping to repetitive reference sequences using Winnowmap2. Nat Methods 19: 705–710. 10.1038/s41592-022-01457-8
  31. Kalman L, Johnson MA, Beck J, Berry-Kravis E, Buller A, Casey B, Feldman GL, Handsfield J, Jakupciak JP, Maragh S, 2007. Development of genomic reference materials for Huntington disease genetic testing. Genet Med 9: 719–723. 10.1097/GIM.0b013e318156e8c1
  32. Li H. 2018. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34: 3094–3100. 10.1093/bioinformatics/bty191
  33. Liao W-W, Asri M, Ebler J, Doerr D, Haukness M, Hickey G, Lu S, Lucas JK, Monlong J, Abel HJ, 2023. A draft human pangenome reference. Nature 617: 312–324. 10.1038/s41586-023-05896-x
  34. Martin M, Patterson M, Garg S, Fischer SO, Pisanti N, Klau GW, Schöenhuth A, Marschall T. 2016. WhatsHap: fast and accurate read-based phasing. bioRxiv 10.1101/085050
  35. Matsuura T, Yamagata T, Burgess DL, Rasmussen A, Grewal RP, Watase K, Khajavi M, McCall AE, Davis CF, Zu L, 2000. Large expansion of the ATTCT pentanucleotide repeat in spinocerebellar ataxia type 10. Nat Genet 26: 191–194. 10.1038/79911
  36. Mitra I, Huang B, Mousavi N, Ma N, Lamkin M, Yanicky R, Shleizer-Burko S, Lohmueller KE, Gymrek M. 2021. Patterns of de novo tandem repeat mutations and their role in autism. Nature 589: 246–250. 10.1038/s41586-020-03078-7
  37. Mousavi N, Shleizer-Burko S, Yanicky R, Gymrek M. 2019. Profiling the genome-wide landscape of tandem repeat expansions. Nucleic Acids Res 47: e90. 10.1093/nar/gkz501
  38. Narzisi G, Schatz MC. 2015. The challenge of small-scale repeats for indel discovery. Front Bioeng Biotechnol 3: 8. 10.3389/fbioe.2015.00008
  39. Olson ND, Wagner J, Dwarshuis N, Miga KH, Sedlazeck FJ, Salit M, Zook JM. 2023. Variant calling and benchmarking in an era of complete human genome sequences. Nat Rev Genet 24: 464–483. 10.1038/s41576-023-00590-0
  40. Ono Y, Hamada M, Asai K. 2022. PBSIM3: a simulator for all types of PacBio and ONT long reads. NAR Genom Bioinform 4: lqac092. 10.1093/nargab/lqac092
  41. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, 2011. Scikit-learn: machine learning in Python. J Mach Learn Res 12: 2825–2830.
  42. Press MO, McCoy RC, Hall AN, Akey JM, Queitsch C. 2018. Massive variation of short tandem repeats with functional consequences across strains of Arabidopsis thaliana. Genome Res 28: 1169–1178. 10.1101/gr.231753.117
  43. Rajan-Babu I-S, Dolzhenko E, Eberle MA, Friedman JM. 2024. Sequence composition changes in short tandem repeats: heterogeneity, detection, mechanisms and clinical implications. Nat Rev Genet 25: 476–499. 10.1038/s41576-024-00696-z
  44. Saini S, Mitra I, Mousavi N, Fotsing SF, Gymrek M. 2018. A reference haplotype panel for genome-wide imputation of short tandem repeats. Nat Commun 9: 4397. 10.1038/s41467-018-06694-0
  45. Sehgal A, Jam HZ, Shen A, Gymrek M. 2024. Genome-wide detection of somatic mosaicism at short tandem repeats. Bioinformatics 40: btae485. 10.1093/bioinformatics/btae485
  46. Seixas AI, Loureiro JR, Costa C, Ordóñez-Ugalde A, Marcelino H, Oliveira CL, Loureiro JL, Dhingra A, Brandão E, Cruz VT, 2017. A pentanucleotide ATTTC repeat insertion in the non-coding region of DAB1, mapping to SCA37, causes spinocerebellar ataxia. Am J Hum Genet 101: 87–103. 10.1016/j.ajhg.2017.06.007
  47. Sherry ST, Ward M-H, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. 2001. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res 29: 308–311. 10.1093/nar/29.1.308
  48. Tanudisastro HA, Deveson IW, Dashnow H, MacArthur DG. 2024. Sequencing and characterizing short tandem repeats in the human genome. Nat Rev Genet 25: 460–475. 10.1038/s41576-024-00692-3
  49. Trost B, Engchuan W, Nguyen CM, Thiruvahindrapuram B, Dolzhenko E, Backstrom I, Mirceta M, Mojarad BA, Yin Y, Dov A, 2020. Genome-wide detection of tandem DNA repeats that are expanded in autism. Nature 586: 80–86. 10.1038/s41586-020-2579-z
  50. Wagner J, Olson ND, Harris L, Khan Z, Farek J, Mahmoud M, Stankovic A, Kovacevic V, Yoo B, Miller N, 2022. Benchmarking challenging small variants with linked and long reads. Cell Genom 2: 100128. 10.1016/j.xgen.2022.100128
  51. Wenger AM, Peluso P, Rowell WJ, Chang P-C, Hall RJ, Concepcion GT, Ebler J, Fungtammasan A, Kolesnikov A, Olson ND, 2019. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat Biotechnol 37: 1155–1162. 10.1038/s41587-019-0217-9
  52. Willems T, Gymrek M, Highnam G, Mittelman D, Erlich Y. 2014. The landscape of human STR variation. Genome Res 24: 1894–1904. 10.1101/gr.177774.114
  53. Yamamoto H, Imai K. 2019. An updated review of microsatellite instability in the era of next-generation sequencing and precision medicine. Semin Oncol 46: 261–270. 10.1053/j.seminoncol.2019.08.003
  54. Ziaei Jam H, Zook JM, Javadzadeh S, Park J, Sehgal A, Gymrek M. 2024. LongTR: genome-wide profiling of genetic variation at tandem repeats from long reads. Genome Biol 25: 176. 10.1186/s13059-024-03319-2
Loading
Loading
Loading
Loading
Back to top