Abstract
Accurate cell-type annotation is essential for revealing the dynamic, cell-type-specific accessibility of regulatory elements from single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data. However, unlike the more mature single-cell RNA-sequencing (scRNA-seq) cell-type annotation workflows, scATAC-seq cell-type annotation remains challenging owing to extreme sparsity, high dimensionality, the scarcity of labeled scATAC references, and pronounced batch effects across data sets. To enable annotation without relying on extensive scATAC labels, we introduce CARA, a cross-modal Bayesian framework that transfers cell-type knowledge from scRNA-seq to scATAC-seq. CARA simultaneously enables cell-type annotation, batch correction, and detection of cell types absent from the reference by combining pretraining on scRNA-seq data and semisupervised learning on target scATAC-seq data, along with distribution alignment, dynamic class weighting, and novel cell-type detection. Across diverse benchmark data sets, CARA consistently outperforms baseline methods, remaining robust to batch effects. CARA's cross-modal alignment and robust annotation strategy preserve fine-grained lineage structure, enabling reconstruction of the hematopoietic differentiation trajectory. Through multidimensional uncertainty and latent-space clustering, CARA identifies novel, rare, or aberrant populations outside the reference cell-type space, providing candidates for further biological validation and perturbation. Using an omic-specific generative framework and distribution alignment, CARA delivers accurate knowledge transfer and detects novel cell types in single-cell DNA methylation data, demonstrating generalization capability to gene-centric modalities. Ultimately, CARA offers a powerful and flexible solution for cross-modal cell-type annotation in complex single-cell settings, facilitating the discovery of novel cell types and mechanistic insight into cell-type-specific regulation across diverse analyses.