Single-cell RNA-seq analysis report

Example deliverable · dataset: public 10x Genomics PBMC 3k · every value computed by the accompanying script

1. Summary

2,700cells in raw matrix
2,638cells after QC
817median genes / cell
2,197median UMIs / cell
6clusters
0.0%unassigned cells

2. Quality control

MetricValue
Cells before QC2,700
Cells after QC2,638 (2.3% removed)
Genes retained13,714 of 32,738
Median genes per cell817
Median UMIs per cell2,197
Median mitochondrial fraction2.03%
Filters appliedgenes/cell 200–2,500; mitochondrial < 5%; gene in ≥ 3 cells
Figure 1. Per-cell distributions before filtering.
Figure 1. Per-cell distributions before filtering.

3. Clustering

Figure 2. UMAP embedding coloured by Leiden cluster.
Figure 2. UMAP embedding coloured by Leiden cluster.

4. Cell-type composition

Cell typeCellsShare
CD4+ T1,19345.2%
CD14+ Monocyte63524.1%
NK42216.0%
B34012.9%
Dendritic351.3%
Platelet130.5%
Figure 3. UMAP coloured by assigned cell type.
Figure 3. UMAP coloured by assigned cell type.

5. Marker genes

Figure 4. Top marker genes per cluster.
Figure 4. Top marker genes per cluster.

6. Independent verification of the annotation

A reference-based classifier was run without access to the manual labels. Agreement is reported per cluster; disagreements would be listed here and reflected in more conservative labels.

ClusterCellsMarker-basedAutomated (blind)Within-clusterTop markers
01,193CD4+ TTcm/Naive helper T cells100.0%LDHB, RPS12, RPS25, RPS3, RPS27
1422NKCD16+ NK cells100.0%NKG7, CST7, CTSW, B2M, GZMA
2340BB cells100.0%CD74, CD79A, HLA-DRA, CD79B, HLA-DPB1
3635CD14+ MonocyteClassical monocytes100.0%FTL, FTH1, TYROBP, LYZ, CST3
413PlateletMegakaryocytes/platelets100.0%PF4, GNG11, SDPR, PPBP, NRGN
535DendriticDC100.0%HLA-DPA1, HLA-DPB1, HLA-DRA, HLA-DRB1, HLA-DQA1

7. Methods (ready for a manuscript)

Raw counts were processed in Python with Scanpy. Cells with fewer than 200 detected genes, more than 2,500 detected genes, or more than 5% of counts from mitochondrial genes were excluded, leaving 2,638 of 2,700 cells (2.3% removed); genes detected in fewer than three cells were discarded, leaving 13,714 genes. Counts were normalised to 10,000 per cell and log1p-transformed. 1,838 highly variable genes were selected (mean expression 0.0125–3, minimum dispersion 0.5); total counts and mitochondrial fraction were regressed out and the data scaled to unit variance with values clipped at 10. Principal component analysis was computed on the scaled matrix and a neighbourhood graph built from the first 30 components with 10 neighbours per cell, followed by UMAP embedding and Leiden clustering at resolution 0.5, yielding 6 clusters. Cluster marker genes were identified with a Wilcoxon rank-sum test. Cell identities were assigned from canonical marker panels and verified independently with CellTypist (Immune_All_Low model) run without access to the manual labels; the two assignments agreed for 6 of 6 clusters at lineage level. Analysis code, package versions and parameters accompany this report.

8. Reproducibility

This report was generated from the accompanying analysis script. Re-running it on the same input reproduces every number and figure above. Package versions are recorded alongside the script; the processed object is delivered with the report so downstream analyses can start from it rather than from raw counts.

9. Limitations