Skip to content

Cell Ranger 9.0.1 User Guide

Overview

Cell Ranger is a set of analysis pipelines from 10x Genomics that processes single-cell RNA sequencing data from the Chromium platform. It performs sample demultiplexing, barcode processing, single-cell 3' or 5' gene counting, V(D)J transcript sequence assembly, and multi-omics analysis.

Version: 9.0.1 Category: Bioinformatics / Single-Cell Genomics Official Documentation: https://www.10xgenomics.com/support/software/cell-ranger

Loading the Module

module load cellranger/9.0.1

Check loaded environment:

module list
which cellranger
cellranger --version

Main Commands

cellranger count

Count gene expression from single-cell RNA-seq data.

Basic usage:

cellranger count --id=sample_id \
                 --transcriptome=/path/to/refdata \
                 --fastqs=/path/to/fastqs \
                 --sample=sample_name

Common options: - --id: Unique run ID (output directory name) - --transcriptome: Path to Cell Ranger reference transcriptome - --fastqs: Path to directory containing FASTQ files - --sample: Sample name(s) from FASTQ filenames - --expect-cells: Expected number of recovered cells - --chemistry: Assay configuration (auto-detected by default) - --localcores: Number of cores to use (default: all available) - --localmem: GB of memory to use (default: 90% of system)

cellranger aggr

Aggregate data from multiple Cell Ranger runs.

Usage:

cellranger aggr --id=aggregated \
                --csv=aggregation.csv

aggregation.csv format:

library_id,molecule_h5
sample1,/path/to/sample1/outs/molecule_info.h5
sample2,/path/to/sample2/outs/molecule_info.h5

cellranger reanalyze

Re-run secondary analysis (dimensionality reduction, clustering, etc.).

Usage:

cellranger reanalyze --id=reanalysis \
                     --matrix=/path/to/filtered_feature_bc_matrix.h5 \
                     --params=params.csv

cellranger vdj

Assemble V(D)J transcripts from single-cell data.

Usage:

cellranger vdj --id=sample_vdj \
               --reference=/path/to/vdj_reference \
               --fastqs=/path/to/fastqs \
               --sample=sample_name

cellranger multi

Analyze Gene Expression, Feature Barcode, and/or V(D)J data together.

Usage:

cellranger multi --id=multi_sample \
                 --csv=multi_config.csv

cellranger mkref

Build a Cell Ranger-compatible reference from FASTA and GTF files.

Usage:

cellranger mkref --genome=genome_name \
                 --fasta=/path/to/genome.fa \
                 --genes=/path/to/genes.gtf

cellranger mkgtf

Filter a GTF file for Cell Ranger compatibility.

Usage:

cellranger mkgtf input.gtf output.gtf \
                 --attribute=gene_biotype:protein_coding

Running on the Cluster

Interactive Job (Testing/Small Datasets)

srun --nodes=1 --cpus-per-task=16 --mem=64G --time=4:00:00 --pty bash

module load cellranger/9.0.1

cellranger count --id=test_run \
                 --transcriptome=/sw/cellranger/refdata-gex-GRCh38-2024-A \
                 --fastqs=/path/to/fastqs \
                 --sample=test_sample \
                 --localcores=16 \
                 --localmem=60

Batch Job (Production Runs)

Create a Slurm submission script cellranger_count.sh:

#!/bin/bash
#SBATCH --job-name=cellranger
#SBATCH --output=cellranger_%j.out
#SBATCH --error=cellranger_%j.err
#SBATCH --nodes=1
#SBATCH --cpus-per-task=32
#SBATCH --mem=128G
#SBATCH --time=24:00:00

module purge
module load cellranger/9.0.1

SAMPLE_ID="sample_001"
TRANSCRIPTOME="/sw/cellranger/refdata-gex-GRCh38-2024-A"
FASTQ_DIR="/path/to/fastqs"

cellranger count --id=${SAMPLE_ID} \
                 --transcriptome=${TRANSCRIPTOME} \
                 --fastqs=${FASTQ_DIR} \
                 --sample=${SAMPLE_ID} \
                 --localcores=${SLURM_CPUS_PER_TASK} \
                 --localmem=120

echo "Cell Ranger count completed for ${SAMPLE_ID}"

Submit the job:

sbatch cellranger_count.sh

Multi-Sample Processing

For processing multiple samples in parallel:

#!/bin/bash
#SBATCH --job-name=cellranger_array
#SBATCH --output=cellranger_%A_%a.out
#SBATCH --error=cellranger_%A_%a.err
#SBATCH --array=1-10
#SBATCH --nodes=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=64G
#SBATCH --time=12:00:00

module load cellranger/9.0.1

# Sample list file (one sample ID per line)
SAMPLE=$(sed -n "${SLURM_ARRAY_TASK_ID}p" sample_list.txt)

cellranger count --id=${SAMPLE} \
                 --transcriptome=/sw/cellranger/refdata-gex-GRCh38-2024-A \
                 --fastqs=/path/to/fastqs \
                 --sample=${SAMPLE} \
                 --localcores=${SLURM_CPUS_PER_TASK} \
                 --localmem=60

Reference Genomes

Cell Ranger requires pre-built reference transcriptomes. Common references should be stored in:

/sw/cellranger/references/

Available References

Check with your system administrator for available references, or download from: https://www.10xgenomics.com/support/software/cell-ranger/downloads

Common references: - refdata-gex-GRCh38-2024-A - Human (GRCh38/hg38) - refdata-gex-GRCm39-2024-A - Mouse (GRCm39/mm39) - refdata-gex-GRCh38-and-mm10-2024-A - Human + Mouse barnyard

Building Custom References

cellranger mkref --genome=custom_genome \
                 --fasta=genome.fa \
                 --genes=genes.gtf \
                 --nthreads=16

Output Structure

After running cellranger count, outputs are in the --id directory:

sample_id/
├── outs/
│   ├── web_summary.html              # QC metrics summary
│   ├── metrics_summary.csv           # Key metrics in CSV
│   ├── filtered_feature_bc_matrix/   # Filtered count matrix (cells only)
│   │   ├── barcodes.tsv.gz
│   │   ├── features.tsv.gz
│   │   └── matrix.mtx.gz
│   ├── filtered_feature_bc_matrix.h5 # HDF5 format count matrix
│   ├── raw_feature_bc_matrix/        # Unfiltered matrix (all barcodes)
│   ├── analysis/                     # Secondary analysis results
│   │   ├── clustering/
│   │   ├── diffexp/
│   │   ├── pca/
│   │   ├── tsne/
│   │   └── umap/
│   ├── molecule_info.h5              # Per-molecule information
│   ├── possorted_genome_bam.bam      # Aligned reads
│   ├── possorted_genome_bam.bam.bai
│   └── cloupe.cloupe                 # Loupe Browser file
└── SC_RNA_COUNTER_CS/                # Pipeline internal files

Key QC Metrics

Important metrics to check in web_summary.html:

  1. Number of Cells: Should match expected cell count
  2. Mean Reads per Cell: Typically 20,000-50,000 for 3' gene expression
  3. Median Genes per Cell: Higher is generally better (varies by cell type)
  4. Sequencing Saturation: >80% is good for most applications
  5. Valid Barcodes: Should be >75%
  6. Q30 Bases in Barcode/UMI/Read: Should be >80%
  7. Reads Mapped to Genome: Should be >70%
  8. Reads Mapped to Transcriptome: Should be >60%

Downstream Analysis

Load count matrices into analysis tools:

Seurat (R)

library(Seurat)
data <- Read10X(data.dir = "sample_id/outs/filtered_feature_bc_matrix/")
seurat_obj <- CreateSeuratObject(counts = data)

Scanpy (Python)

import scanpy as sc
adata = sc.read_10x_h5("sample_id/outs/filtered_feature_bc_matrix.h5")

Loupe Browser

Open cloupe.cloupe file with 10x Genomics Loupe Browser for interactive visualization.

Resource Requirements

Typical Requirements by Dataset Size

Sample Type Cells Reads Cores Memory Time
Small 1-5K 50M 8 32 GB 2-4h
Medium 5-10K 200M 16 64 GB 4-8h
Large 10-20K 500M 32 128 GB 8-16h
Very Large >20K >1B 32+ 256 GB 24h+

Note: Cell Ranger scales well with more cores. Using 16-32 cores significantly reduces runtime.

Troubleshooting

"No input FASTQs were found"

  • Check FASTQ path is correct
  • Ensure --sample matches FASTQ filenames exactly
  • FASTQ files must follow Illumina naming: SampleName_S1_L001_R1_001.fastq.gz

Low valid barcodes (<75%)

  • Check chemistry version matches data (--chemistry flag)
  • Verify correct library type (3' vs 5' Gene Expression)

Low cells detected

  • Adjust --expect-cells parameter
  • Check sequencing quality in web_summary.html
  • May indicate low cell concentration or failed library prep

Out of memory errors

  • Increase --localmem (must be less than physical RAM)
  • Request more memory in Slurm script (#SBATCH --mem=)
  • Reduce --localcores to free memory

Pipeline failures

Check detailed logs:

cat sample_id/_log

Best Practices

  1. Always check web_summary.html after each run for QC metrics
  2. Use appropriate reference - ensure genome build matches your experiment
  3. Estimate cell numbers - provide --expect-cells for better cell calling
  4. Monitor disk space - Cell Ranger generates large intermediate files
  5. Save molecule_info.h5 - required for aggregation and reanalysis
  6. Use aggregation for cross-sample normalization rather than simple concatenation
  7. Test with small datasets first to verify pipeline parameters

Support and Documentation

  • Official Documentation: https://www.10xgenomics.com/support/software/cell-ranger
  • Release Notes: https://www.10xgenomics.com/support/software/cell-ranger/downloads
  • Community Forum: https://www.10xgenomics.com/support/community
  • Local Support: Contact your cluster system administrators

Version History

  • 9.0.1 (Current): Latest version with improved algorithms and compatibility
  • 7.1.0 (Previous): Legacy version on old cluster
  • See release notes for detailed changes: https://www.10xgenomics.com/support/software/cell-ranger/latest/release-notes

Installation Location: /sw/cellranger/cellranger-9.0.1 Module File: /opt/modulefiles/cellranger/9.0.1.lua Last Updated: 2025-10-12