Abstract
Minimizing experimental noise is integral to robust data generation in single-cell omics. The current standard for avoiding batch effects during sample processing is barcode- or hashtag-assisted combining of different experimental treatments into one pool, allowing all samples to be subject to the technical protocols uniformly. The final datapoints for each treatment group are then computationally separated based on the original hashtag labels. Clearly, whereas hashtagging all groups and pooling them in a single well is expected to minimize batch effects, the procedure can also lead to a loss of cells that cannot be confidently decoded during the computational demultiplexing step. Here, we examine four alternate experimental designs, namely compound, reference, chain, and confounded, that could be used instead of a single-pool approach and quantify the batch effects as well as cell loss in each case. We find a linear relationship - the percentage of cells lost is double the number of hashtags used in the experiment. We use these analyses to identify experimental designs that can successfully mitigate batch effects while minimizing multiplexing, hence the cell loss, in each well. While a reference design offers the best overall performance, this study can help individual investigators choose particular approaches that are best suited for their biological questions.