Abstract
Accurate cell type annotation is essential for revealing the dynamic, cell type-specific accessibility of regulatory elements from single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data. However, unlike the more mature single-cell RNA-sequencing (scRNA-seq) cell type annotation workflows, scATAC-seq cell type annotation remains challenging due to extreme sparsity, high dimensionality, the scarcity of labelled scATAC references, and pronounced batch effects across datasets. To enable annotation without relying on extensive scATAC labels, we introduce CARA, a cross-omics Bayesian framework that transfers cell type knowledge from scRNA-seq to scATAC-seq. CARA simultaneously enables cell type annotation, batch correction, and detection of cell types absent from the reference by combining pretraining on scRNA-seq data and semisupervised learning on target scATAC-seq data, along with distribution alignment, dynamic class weighting, and novel cell type detection. Across diverse benchmark datasets, CARA consistently outperforms baseline methods, remaining robust to batch effects. CARA's cross-modal alignment and robust annotation strategy preserve fine-grained lineage structure, enabling reconstruction of the hematopoietic differentiation trajectory. Through multidimensional uncertainty and latent-space clustering, CARA identifies novel, rare, or aberrant populations outside the reference cell type space, providing candidates for further biological validation and perturbation. Using an omics-specific generative framework and distribution alignment, CARA delivers accurate knowledge transfer and detects novel cell types in single-cell DNA methylation data, demonstrating seamless extensibility to new modalities. Ultimately, CARA offers a powerful and flexible solution for cross-modal cell type annotation in complex single-cell settings, facilitating the discovery of novel cell types and mechanistic insight into cell type-specific regulation across diverse analyses.