Overview of scPSS. (A) At first, the query (Q1, Q2) and reference data sets (R1, R2, R3) are concatenated together, and then principal component analysis is done. The concatenated data sets are then integrated using the Harmony method. This removes batch-specific effects on these data sets by adjusting the principal component (PC) values. (B) The Euclidean distance of a cell to the k-th nearest reference cell in PC space is used as its pathological shift score. The distances of the reference cells to their neighboring reference cells are considered as the null distribution. We can determine the P-value of each query cell distance belonging to the reference null distribution to get a significance measure for pathological progression. Cells with P-values below a specified threshold are considered significantly pathological.
