Example deliverable · dataset: public 10x Genomics PBMC 3k · every value computed by the accompanying script
1. Summary
2,700cells in raw matrix
2,638cells after QC
817median genes / cell
2,197median UMIs / cell
6clusters
0.0%unassigned cells
2. Quality control
Metric
Value
Cells before QC
2,700
Cells after QC
2,638 (2.3% removed)
Genes retained
13,714 of 32,738
Median genes per cell
817
Median UMIs per cell
2,197
Median mitochondrial fraction
2.03%
Filters applied
genes/cell 200–2,500; mitochondrial < 5%; gene in ≥ 3 cells
Figure 1. Per-cell distributions before filtering.
3. Clustering
Figure 2. UMAP embedding coloured by Leiden cluster.
4. Cell-type composition
Cell type
Cells
Share
CD4+ T
1,193
45.2%
CD14+ Monocyte
635
24.1%
NK
422
16.0%
B
340
12.9%
Dendritic
35
1.3%
Platelet
13
0.5%
Figure 3. UMAP coloured by assigned cell type.
5. Marker genes
Figure 4. Top marker genes per cluster.
6. Independent verification of the annotation
A reference-based classifier was run without access to the manual labels. Agreement is reported per
cluster; disagreements would be listed here and reflected in more conservative labels.
Cluster
Cells
Marker-based
Automated (blind)
Within-cluster
Top markers
0
1,193
CD4+ T
Tcm/Naive helper T cells
100.0%
LDHB, RPS12, RPS25, RPS3, RPS27
1
422
NK
CD16+ NK cells
100.0%
NKG7, CST7, CTSW, B2M, GZMA
2
340
B
B cells
100.0%
CD74, CD79A, HLA-DRA, CD79B, HLA-DPB1
3
635
CD14+ Monocyte
Classical monocytes
100.0%
FTL, FTH1, TYROBP, LYZ, CST3
4
13
Platelet
Megakaryocytes/platelets
100.0%
PF4, GNG11, SDPR, PPBP, NRGN
5
35
Dendritic
DC
100.0%
HLA-DPA1, HLA-DPB1, HLA-DRA, HLA-DRB1, HLA-DQA1
7. Methods (ready for a manuscript)
Raw counts were processed in Python with Scanpy. Cells with fewer than 200 detected genes, more than 2,500 detected genes, or more than 5% of counts from mitochondrial genes were excluded, leaving 2,638 of 2,700 cells (2.3% removed); genes detected in fewer than three cells were discarded, leaving 13,714 genes. Counts were normalised to 10,000 per cell and log1p-transformed. 1,838 highly variable genes were selected (mean expression 0.0125–3, minimum dispersion 0.5); total counts and mitochondrial fraction were regressed out and the data scaled to unit variance with values clipped at 10. Principal component analysis was computed on the scaled matrix and a neighbourhood graph built from the first 30 components with 10 neighbours per cell, followed by UMAP embedding and Leiden clustering at resolution 0.5, yielding 6 clusters. Cluster marker genes were identified with a Wilcoxon rank-sum test. Cell identities were assigned from canonical marker panels and verified independently with CellTypist (Immune_All_Low model) run without access to the manual labels; the two assignments agreed for 6 of 6 clusters at lineage level. Analysis code, package versions and parameters accompany this report.
8. Reproducibility
This report was generated from the accompanying analysis script. Re-running it on the same input
reproduces every number and figure above. Package versions are recorded alongside the script; the processed object
is delivered with the report so downstream analyses can start from it rather than from raw counts.
9. Limitations
A single sample cannot support a comparison between conditions; results here are descriptive.
Doublet detection and ambient-RNA correction are not applied in this example run and would be included in a
project deliverable.
Cluster count depends on the resolution parameter; the value used is stated in the methods above.