Skip to content

Latest commit

 

History

History
33 lines (23 loc) · 2.48 KB

File metadata and controls

33 lines (23 loc) · 2.48 KB

Workflow overview

This workflow is a best-practice workflow for preprocessing counts from single cell RNA sequencing data. The workflow is built using snakemake and follows the Preprocessing and visualization section of the Single Cell Best Practices.

It consists of the following steps:

  1. Convert input data to zarr format.
  2. Filter low-quality barcodes with scanpy.
  3. Correct for ambient RNA contamination with SoupX.
  4. Detect doublets with scDblFinder.
  5. Normalize counts with scanpy.

Running the workflow

To configure the workflow run, go through the provided config/config.yaml entry by entry and adjust them where necessary. After an initial run, we recommend going through the quality control plots mentioned there and double-checking that the provided threshold values for filtering make sense.

Input data

To specify the input, provide a config/sample_sheet.tsv file with the following layout:

sample_id raw_counts_path format
cellranger_1 ../path/to/raw_feature_bc_matrix/ 10x_mtx
kallisto_bustools_1 ../path/to/adata.h5ad h5ad
alevin_fry_1 ../path/to/quants_mat.mtx mtx

Here, the columns are: