Figure 1.

NS-Forest version 2.0 workflow. The method begins with a cell-by-gene expression matrix with cluster assignments for each cell (A). This clustered expression matrix is used to generate binary classification models for each cell cluster using the random forest machine learning method. Features are extracted from the model and ranked by Gini Index (B). Top features are filtered by expression level to remove negative markers (C) before being reranked by Binary Expression Score (D,E). Decision branch expression level cutoffs are derived from decision tree analysis for the most binary features (F) and F-beta score used as an objective function to evaluate the discriminatory power of all permutations of selected markers (G).

1767f01