All benchmarks

Denoising

Release v1.0.0 v1.1.0-rc1rc

Removing noise in sparse single-cell RNA-sequencing count data

6 methods
2 control methods
17 datasets
2 metrics
2 releases
Task repository MIT v1.1.0-rc1

A key challenge in evaluating denoising methods is the general lack of a ground truth. A recent benchmark study (Hou et al., 2020) relied on flow-sorted datasets, mixture control experiments (Tian et al., 2019), and comparisons with bulk RNA-Seq data. Since each of these approaches suffers from specific limitations, it is difficult to combine these different approaches into a single quantitative measure of denoising accuracy. Here, we instead rely on an approach termed molecular cross-validation (MCV), which was specifically developed to quantify denoising accuracy in the absence of a ground truth (Batson et al., 2019). In MCV, the observed molecules in a given scRNA-Seq dataset are first partitioned between a training and a test dataset. Next, a denoising method is applied to the training dataset. Finally, denoising accuracy is measured by comparing the result to the test dataset. The authors show that both in theory and in practice, the measured denoising accuracy is representative of the accuracy that would be obtained on a ground truth dataset.

Contributors

  • Wesley Lewis
    authormaintainer
  • Scott Gigante
    authormaintainer
  • Robrecht Cannoodt
    author
  • Kai Waldrant
    contributor
  • Jeremie Kalfon
    contributor
  • Marius Lange
    contributor

Leaderboard

Methods ranked by scaled overall mean. Each cell encodes a score from 0 to 1 by size and intensity.

QC: Normalisation Visualisation 2 plots

Per metric: points placed by control-anchored scaled score (x); dashed lines mark scaled 0 and 1 (worst/best control); the lower axis shows the raw score. Points beyond [-0.2, 1.2] are clamped to the edge as triangles. Hover a dot or line to highlight it and read details.

methodcontrol
  • Mean-squared errorlower better
    Perfect DenoisingPerfect DenoisingCellMapperCellMapperDCADCAMAGICMAGICKNN SmoothingKNN SmoothingNo DenoisingNo DenoisingALRAALRA0.3150.2360.1570.079000.250.50.751rawscaled
  • Poisson Losslower better
    KNN SmoothingKNN SmoothingNo DenoisingNo DenoisingMAGICMAGICDCADCACellMapperCellMapperALRAALRAPerfect DenoisingPerfect Denoising2.489-1.604-5.697-9.789-13.88200.250.50.751rawscaled
QC: Indicator table 4 errors44 warnings

Automated checks on the benchmark run and its results: missing values, score scaling, metric ranges and similar. Errors are high-severity issues that usually need a maintainer's attention; warnings are lower-severity signals. Findings that are expected for this task are listed separately as silenced.

4 high-severity issues need review. 82 of 130 checks passed.

  • error Raw results Method 'scprint' % missing

    Percentage of missing results should be less than 10% Task: denoising Method: scprint Number of results: 0 Expected number of results: 34 Percentage missing: 100%

  • error Raw results Method 'scprint' % failed

    Percentage of failed processes should be less than 10% Task: denoising Method: scprint Succeeded processes: 0 Attempted processes: 17 Percentage failed: 100%

  • error Scaling Metric 'poisson' % outside range

    Percentage of scaled scores outside control range should be less than 10% Task: denoising Metric: poisson Inside range: 0 Scaled scores: 119 Percentage outside: 100%

  • error Scaling Metric 'poisson' worst score % outside range

    The worst scaled score should be less than 10% outside the control range Task: denoising Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

Show 44 warnings
  • warning Scaling Worst 'poisson' score for 'alra'

    Method 'alra' performs much worse than controls for metric' poisson' Task: denoising Method: alra Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

  • warning Scaling Worst 'poisson' score for 'cellmapper'

    Method 'cellmapper' performs much worse than controls for metric' poisson' Task: denoising Method: cellmapper Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

  • warning Scaling Worst 'poisson' score for 'dca'

    Method 'dca' performs much worse than controls for metric' poisson' Task: denoising Method: dca Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

  • warning Scaling Worst 'poisson' score for 'knn_smoothing'

    Method 'knn_smoothing' performs much worse than controls for metric' poisson' Task: denoising Method: knn_smoothing Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

  • warning Scaling Worst 'poisson' score for 'magic'

    Method 'magic' performs much worse than controls for metric' poisson' Task: denoising Method: magic Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

  • warning Scaling Worst 'poisson' score for 'no_denoising'

    Method 'no_denoising' performs much worse than controls for metric' poisson' Task: denoising Method: no_denoising Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

  • warning Scaling Worst 'poisson' score for 'perfect_denoising'

    Method 'perfect_denoising' performs much worse than controls for metric' poisson' Task: denoising Method: perfect_denoising Metric: poisson Worst score: -2.80445795339412 Percentage outside range: 280%

  • warning Raw results Task number of results

    Number of results should be equal to #datasets × #methods × #metrics Task: denoising Number of results: 238 Number of datasets: 17 Number of methods: 8 Number of metrics: 2 Expected number of results: 272

  • warning Raw results Dataset 'openproblems_v1/pancreas' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/pancreas Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'cellxgene_census/tabula_sapiens' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: cellxgene_census/tabula_sapiens Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'cellxgene_census/hcla' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: cellxgene_census/hcla Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/allen_brain_atlas' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/allen_brain_atlas Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'cellxgene_census/mouse_pancreas_atlas' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: cellxgene_census/mouse_pancreas_atlas Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'cellxgene_census/gtex_v9' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: cellxgene_census/gtex_v9 Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/cengen' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/cengen Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'cellxgene_census/dkd' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: cellxgene_census/dkd Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/mouse_hspc_nestorowa2016' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/mouse_hspc_nestorowa2016 Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'cellxgene_census/hypomap' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: cellxgene_census/hypomap Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/zebrafish' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/zebrafish Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/tenx_5k_pbmc' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/tenx_5k_pbmc Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/tenx_1k_pbmc' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/tenx_1k_pbmc Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/mouse_blood_olsson_labelled' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/mouse_blood_olsson_labelled Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/tnbc_wu2021' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/tnbc_wu2021 Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/immune_cells' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: openproblems_v1/immune_cells Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Dataset 'cellxgene_census/immune_cell_atlas' % missing

    Percentage of missing results should be less than 10% Task: denoising Dataset: cellxgene_census/immune_cell_atlas Number of results: 14 Expected number of results: 16 Percentage missing: 12%

  • warning Raw results Metric 'mse' % missing

    Percentage of missing results should be less than 10% Task: denoising Metric: mse Number of results: 119 Expected number of results: 136 Percentage missing: 12%

  • warning Raw results Metric 'poisson' % missing

    Percentage of missing results should be less than 10% Task: denoising Metric: poisson Number of results: 119 Expected number of results: 136 Percentage missing: 12%

  • warning Raw results Dataset 'openproblems_v1/pancreas' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/pancreas Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'cellxgene_census/tabula_sapiens' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: cellxgene_census/tabula_sapiens Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'cellxgene_census/hcla' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: cellxgene_census/hcla Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/allen_brain_atlas' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/allen_brain_atlas Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'cellxgene_census/mouse_pancreas_atlas' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: cellxgene_census/mouse_pancreas_atlas Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'cellxgene_census/gtex_v9' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: cellxgene_census/gtex_v9 Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/cengen' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/cengen Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'cellxgene_census/dkd' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: cellxgene_census/dkd Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/mouse_hspc_nestorowa2016' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/mouse_hspc_nestorowa2016 Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'cellxgene_census/hypomap' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: cellxgene_census/hypomap Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/zebrafish' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/zebrafish Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/tenx_5k_pbmc' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/tenx_5k_pbmc Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/tenx_1k_pbmc' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/tenx_1k_pbmc Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/mouse_blood_olsson_labelled' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/mouse_blood_olsson_labelled Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/tnbc_wu2021' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/tnbc_wu2021 Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'openproblems_v1/immune_cells' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: openproblems_v1/immune_cells Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Dataset 'cellxgene_census/immune_cell_atlas' % failed

    Percentage of failed processes should be less than 10% Task: denoising Dataset: cellxgene_census/immune_cell_atlas Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

Method info 6

ALRA imputes missing values in scRNA-seq data by computing rank-k approximation, thresholding by gene, and rescaling the matrix.

Adaptively-thresholded Low Rank Approximation (ALRA).

ALRA is a method for imputation of missing values in single cell RNA-sequencing data, described in the preprint, "Zero-preserving imputation of scRNA-seq data using low-rank approximation" available here. Given a scRNA-seq expression matrix, ALRA first computes its rank-k approximation using randomized SVD. Next, each row (gene) is thresholded by the magnitude of the most negative value of that gene. Finally, the matrix is rescaled.

CellMapper is a general framework for k-NN based mapping tasks in single-cell and spatial genomics

k-NN-based mapping of cells across representations to transfer labels, embeddings and expression values. Works for millions of cells, on CPU and GPU, across molecular modalities, between spatial and non-spatial data, for arbitrary query and reference datasets. We treat data denoising as self-mapping, where the query and reference datasets are the same. Based on some joint representation (here: PCA), CellMapper computes a k-NN graph and applies a kernel function to compute edge weights. Here, we use a kernel based on Jaccard similarity (as e.g. in HNOCA-tools) or UMAP's fuzzy_simplicial_set (as in scanpy). Given the row-normalized weighted adjacency matrix, we simulate a t-step random walk to smooth the data, similar to MAGIC. For large t-values, we also provide a spectral approximation of the application of the t-step transition matrix to the data (not used here).

A deep autoencoder with ZINB loss function to address the dropout effect in count data

Deep Count Autoencoder

Removes the dropout effect by taking the count structure, overdispersed nature and sparsity of the data into account using a deep autoencoder with zero-inflated negative binomial (ZINB) loss function.

Iterative kNN-smoothing denoises scRNA-seq data by iteratively increasing the size of neighbourhoods for smoothing until a maximum k value is reached.

Iterative kNN-smoothing is a method to repair or denoise noisy scRNA-seq expression matrices. Given a scRNA-seq expression matrix, KNN-smoothing first applies initial normalisation and smoothing. Then, a chosen number of principal components is used to calculate Euclidean distances between cells. Minimally sized neighbourhoods are initially determined from these Euclidean distances, and expression profiles are shared between neighbouring cells. Then, the resultant smoothed matrix is used as input to the next step of smoothing, where the size (k) of the considered neighbourhoods is increased, leading to greater smoothing. This process continues until a chosen maximum k value has been reached, at which point the iteratively smoothed object is then optionally scaled to yield a final result.

MAGIC imputes and denoises scRNA-seq data that is noisy or dropout-prone.

MAGIC (Markov Affinity-based Graph Imputation of Cells) is a method for imputation and denoising of noisy or dropout-prone single cell RNA-sequencing data. Given a normalised scRNA-seq expression matrix, it first calculates Euclidean distances between each pair of cells in the dataset, which is then augmented using a Gaussian kernel (function) and row-normalised to give a normalised affinity matrix. A t-step markov process is then calculated, by powering this affinity matrix t times. Finally, the powered affinity matrix is right-multiplied by the normalised data, causing the final imputed values to take the value of a per-gene average weighted by the affinities of cells. The resultant imputed matrix is then rescaled, to more closely match the magnitude of measurements in the normalised (input) matrix.

scPRINT is a large transformer model built for the inference of gene networks

scPRINT is a large transformer model built for the inference of gene networks (connections between genes explaining the cell's expression profile) from scRNAseq data.

It uses novel encoding and decoding of the cell expression profile and new pre-training methodologies to learn a cell model.

scPRINT can be used to perform the following analyses:

  • expression denoising: increase the resolution of your scRNAseq data
  • cell embedding: generate a low-dimensional representation of your dataset
  • label prediction: predict the cell type, disease, sequencer, sex, and ethnicity of your cells
  • gene network inference: generate a gene network from any cell or cell cluster in your scRNAseq dataset
Control method info 2
No Denoising

negative control by copying train counts

This method serves as a negative control, where the denoised data is a copy of the unaltered training data. This represents the scoring threshold if denoising was not performed on the data.

Perfect Denoising

Positive control by copying the test counts

This method serves as a positive control, where the test data is copied 1-to-1 to the denoised data. This makes it seem as if the data is perfectly denoised as it will be compared to the test data in the metrics.

Metric info 2
Mean-squared errorlower is betterBatson et al., 2019

The mean squared error between the denoised counts and the true counts.

The mean squared error between the denoised counts of the training dataset and the true counts of the test dataset after reweighing by the train/test ratio

Poisson Losslower is betterBatson et al., 2019

The Poisson log likelihood of the true counts observed in the distribution of denoised counts

The Poisson log likelihood of observing the true counts of the test dataset given the distribution given in the denoised dataset.

Dataset info 17
1k PBMCs unlinked

1k peripheral blood mononuclear cells from a healthy donor

1k Peripheral Blood Mononuclear Cells (PBMCs) from a healthy donor. Sequenced on 10X v3 chemistry in November 2018 by 10X Genomics.

5k PBMCs unlinked

5k peripheral blood mononuclear cells from a healthy donor

5k Peripheral Blood Mononuclear Cells (PBMCs) from a healthy donor. Sequenced on 10X v3 chemistry in July 2019 by 10X Genomics.

CeNGEN unlinked

Complete Gene Expression Map of an Entire Nervous System

100k FACS-isolated C. elegans neurons from 17 experiments sequenced on 10x Genomics.

Multimodal single cell sequencing implicates chromatin accessibility and genetic background in diabetic kidney disease progression

Multimodal single cell sequencing is a powerful tool for interrogating cell-specific changes in transcription and chromatin accessibility. We performed single nucleus RNA (snRNA-seq) and assay for transposase accessible chromatin sequencing (snATAC-seq) on human kidney cortex from donors with and without diabetic kidney disease (DKD) to identify altered signaling pathways and transcription factors associated with DKD. Both snRNA-seq and snATAC-seq had an increased proportion of VCAM1+ injured proximal tubule cells (PT_VCAM1) in DKD samples. PT_VCAM1 has a pro-inflammatory expression signature and transcription factor motif enrichment implicated NFkB signaling. We used stratified linkage disequilibrium score regression to partition heritability of kidney-function-related traits using publicly-available GWAS summary statistics. Cell-specific PT_VCAM1 peaks were enriched for heritability of chronic kidney disease (CKD), suggesting that genetic background may regulate chromatin accessibility and DKD progression. snATAC-seq found cell-specific differentially accessible regions (DAR) throughout the nephron that change accessibility in DKD and these regions were enriched for glucocorticoid receptor (GR) motifs. Changes in chromatin accessibility were associated with decreased expression of insulin receptor, increased gluconeogenesis, and decreased expression of the GR cytosolic chaperone, FKBP5, in the diabetic proximal tubule. Cleavage under targets and release using nuclease (CUT&RUN) profiling of GR binding in bulk kidney cortex and an in vitro model of the proximal tubule (RPTEC) showed that DAR co-localize with GR binding sites. CRISPRi silencing of GR response elements (GRE) in the FKBP5 gene body reduced FKBP5 expression in RPTEC, suggesting that reduced FKBP5 chromatin accessibility in DKD may alter cellular response to GR. We developed an open-source tool for single cell allele specific analysis (SALSA) to model the effect of genetic background on gene expression. Heterozygous germline single nucleotide variants (SNV) in proximal tubule ATAC peaks were associated with allele-specific chromatin accessibility and differential expression of target genes within cis-coaccessibility networks. Partitioned heritability of proximal tubule ATAC peaks with a predicted allele-specific effect was enriched for eGFR, suggesting that genetic background may modify DKD progression in a cell-specific manner.

Single-nucleus cross-tissue molecular reference maps to decipher disease gene function

Understanding the function of genes and their regulation in tissue homeostasis and disease requires knowing the cellular context in which genes are expressed in tissues across the body. Single cell genomics allows the generation of detailed cellular atlases in human tissues, but most efforts are focused on single tissue types. Here, we establish a framework for profiling multiple tissues across the human body at single-cell resolution using single nucleus RNA-Seq (snRNA-seq), and apply it to 8 diverse, archived, frozen tissue types (three donors per tissue). We apply four snRNA-seq methods to each of 25 samples from 16 donors, generating a cross-tissue atlas of 209,126 nuclei profiles, and benchmark them vs. scRNA-seq of comparable fresh tissues. We use a conditional variational autoencoder (cVAE) to integrate an atlas across tissues, donors, and laboratory methods. We highlight shared and tissue-specific features of tissue-resident immune cells, identifying tissue-restricted and non-restricted resident myeloid populations. These include a cross-tissue conserved dichotomy between LYVE1- and HLA class II-expressing macrophages, and the broad presence of LAM-like macrophages across healthy tissues that is also observed in disease. For rare, monogenic muscle diseases, we identify cell types that likely underlie the neuromuscular, metabolic, and immune components of these diseases, and biological processes involved in their pathology. For common complex diseases and traits analyzed by GWAS, we identify the cell types and gene modules that potentially underlie disease mechanisms. The experimental and analytical frameworks we describe will enable the generation of large-scale studies of how cellular and molecular processes vary across individuals and populations.

Human immune unlinked

Human immune cells dataset from the scIB benchmarks

Human immune cells from peripheral blood and bone marrow taken from 5 datasets comprising 10 batches across technologies (10X, Smart-seq2).

Human Lung Cell Atlas unlinked

An integrated cell atlas of the human lung in health and disease (core)

The integrated Human Lung Cell Atlas (HLCA) represents the first large-scale, integrated single-cell reference atlas of the human lung. It consists of over 2 million cells from the respiratory tract of 486 individuals, and includes 49 different datasets. It is split into the HLCA core, and the extended or full HLCA. The HLCA core includes data of healthy lung tissue from 107 individuals, and includes manual cell type annotations based on consensus across 6 independent experts, as well as demographic, biological and technical metadata.

Human pancreas unlinked

Human pancreas cells dataset from the scIB benchmarks

Human pancreatic islet scRNA-seq data from 6 datasets across technologies (CEL-seq, CEL-seq2, Smart-seq2, inDrop, Fluidigm C1, and SMARTER-seq).

A unified single cell gene expression atlas of the murine hypothalamus

The hypothalamus plays a key role in coordinating fundamental body functions. Despite recent progress in single-cell technologies, a unified catalogue and molecular characterization of the heterogeneous cell types and, specifically, neuronal subtypes in this brain region are still lacking. Here we present an integrated reference atlas “HypoMap” of the murine hypothalamus consisting of 384,925 cells, with the ability to incorporate new additional experiments. We validate HypoMap by comparing data collected from SmartSeq2 and bulk RNA sequencing of selected neuronal cell types with different degrees of cellular heterogeneity.

Cross-tissue immune cell analysis reveals tissue-specific features in humans

Despite their crucial role in health and disease, our knowledge of immune cells within human tissues remains limited. We surveyed the immune compartment of 16 tissues from 12 adult donors by single-cell RNA sequencing and VDJ sequencing generating a dataset of ~360,000 cells. To systematically resolve immune cell heterogeneity across tissues, we developed CellTypist, a machine learning tool for rapid and precise cell type annotation. Using this approach, combined with detailed curation, we determined the tissue distribution of finely phenotyped immune cell types, revealing hitherto unappreciated tissue-specific features and clonal architecture of T and B cells. Our multitissue approach lays the foundation for identifying highly resolved immune cell types by leveraging a common reference dataset, tissue-integrated expression analysis, and antigen receptor sequencing.

Mouse Brain Atlas unlinked

Adult mouse primary visual cortex

A murine brain atlas with adjacent cell types as assumed benchmark truth, inferred from deconvolution proportion correlations using matching 10x Visium slides (see Dimitrov et al., 2022).

Mouse HSPC unlinked

Haematopoeitic stem and progenitor cells from mouse bone marrow

1656 hematopoietic stem and progenitor cells from mouse bone marrow. Sequenced by Smart-seq2.

Mouse myeloid unlinked

Myeloid lineage differentiation from mouse blood

660 FACS-isolated myeloid cells from 9 experiments sequenced using C1 Fluidigm and SMARTseq in 2016 by Olsson et al.

Mouse pancreatic islet scRNA-seq atlas across sexes, ages, and stress conditions including diabetes

To better understand pancreatic β-cell heterogeneity we generated a mouse pancreatic islet atlas capturing a wide range of biological conditions. The atlas contains scRNA-seq datasets of over 300,000 mouse pancreatic islet cells, of which more than 100,000 are β-cells, from nine datasets with 56 samples, including two previously unpublished datasets. The samples vary in sex, age (ranging from embryonic to aged), chemical stress, and disease status (including T1D NOD model development and two T2D models, mSTZ and db/db) together with different diabetes treatments. Additional information about data fields is available in anndata uns field 'field_descriptions' and on https://github.com/theislab/mm_pancreas_atlas_rep/blob/main/resources/cellxgene.md.

A multiple-organ, single-cell transcriptomic atlas of humans

Tabula Sapiens is a benchmark, first-draft human cell atlas of nearly 500,000 cells from 24 organs of 15 normal human subjects. This work is the product of the Tabula Sapiens Consortium. Taking the organs from the same individual controls for genetic background, age, environment, and epigenetic effects and allows detailed analysis and comparison of cell types that are shared between tissues. Our work creates a detailed portrait of cell types as well as their distribution and variation in gene expression across tissues and within the endothelial, epithelial, stromal and immune compartments.

Triple-Negative Breast Cancer unlinked

1535 cells from six fresh triple-negative breast cancer tumors.

1535 cells from six TNBC donors by (Wu et al., 2021). This dataset includes cytokine activities, inferred using a multivariate linear model with cytokine-focused signatures, as assumed true cell-cell communication (Dimitrov et al., 2022).

Zebrafish embryonic cells unlinked

Single-cell mRNA sequencing of zebrafish embryonic cells.

90k cells from zebrafish embryos throughout the first day of development, with and without a knockout of chordin, an important developmental gene.

References

  1. Batson, J., Royer, L., & Webber, J. (2019). Molecular Cross-Validation for Single-Cell RNA-seq. bioRxiv. 10.1101/786269 ↗
  2. Domínguez Conde, C., Xu, C., Jarvis, L. B., Rainbow, D. B., Wells, S. B., Gomes, T., Howlett, S. K., Suchanek, O., Polanski, K., King, H. W., Mamanova, L., Huang, N., Szabo, P. A., Richardson, L., Bolt, L., Fasouli, E. S., Mahbubani, K. T., Prete, M., Tuck, L., … Teichmann, S. A. (2022). Cross-tissue immune cell analysis reveals tissue-specific features in humans. Science, 376(6594). 10.1126/science.abl5197 ↗
  3. van Dijk, D., Sharma, R., Nainys, J., Yim, K., Kathail, P., Carr, A. J., Burdziak, C., Moon, K. R., Chaffer, C. L., Pattabiraman, D., Bierie, B., Mazutis, L., Wolf, G., Krishnaswamy, S., & Pe’er, D. (2018). Recovering Gene Interactions from Single-Cell Data Using Data Diffusion. Cell, 174(3), 716-729.e27. 10.1016/j.cell.2018.05.061 ↗
  4. Eraslan, G., Drokhlyansky, E., Anand, S., Fiskin, E., Subramanian, A., Slyper, M., Wang, J., Van Wittenberghe, N., Rouhana, J. M., Waldman, J., Ashenberg, O., Lek, M., Dionne, D., Win, T. S., Cuoco, M. S., Kuksenko, O., Tsankov, A. M., Branton, P. A., Marshall, J. L., … Regev, A. (2022). Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function. Science, 376(6594). 10.1126/science.abl4290 ↗
  5. Eraslan, G., Simon, L. M., Mircea, M., Mueller, N. S., & Theis, F. J. (2019). Single-cell RNA-seq denoising using a deep count autoencoder. Nature Communications, 10(1). 10.1038/s41467-018-07931-2 ↗
  6. 10x Genomics. (2018). 1k PBMCs from a Healthy Donor (v3 chemistry). link ↗
  7. 10x Genomics. (2019). 5k Peripheral Blood Mononuclear Cells (PBMCs) from a Healthy Donor with a Panel of TotalSeq-B Antibodies (v3 chemistry). link ↗
  8. Hammarlund, M., Hobert, O., Miller, D. M., & Sestan, N. (2018). The CeNGEN Project: The Complete Gene Expression Map of an Entire Nervous System. Neuron, 99(3), 430–433. 10.1016/j.neuron.2018.07.042 ↗
  9. Hrovatin, K., Bastidas-Ponce, A., Bakhti, M., Zappia, L., Büttner, M., Sallino, C., Sterr, M., Böttcher, A., Migliorini, A., Lickert, H., & Theis, F. J. (2023). Delineating mouse β-cell identity during lifetime and in diabetes with a single cell atlas. bioRxiv. 10.1101/2022.12.22.521557 ↗
  10. Jones, R. C., Karkanias, J., Krasnow, M. A., Pisco, A. O., Quake, S. R., Salzman, J., Yosef, N., Bulthaup, B., Brown, P., Harper, W., Hemenez, M., Ponnusamy, R., Salehi, A., Sanagavarapu, B. A., Spallino, E., Aaron, K. A., Concepcion, W., Gardner, J. M., Kelly, B., … Wyss-Coray, T. (2022). The Tabula Sapiens: A multiple-organ, single-cell transcriptomic atlas of humans. Science, 376(6594). 10.1126/science.abl4896 ↗
  11. Kalfon, J., Samaran, J., Peyré, G., & Cantini, L. (2024). scPRINT: pre-training on 50 million cells allows robust gene network predictions. 10.1101/2024.07.29.605556 ↗
  12. Lange, M. (2025). quadbio/cellmapper: v0.2.2. 10.5281/ZENODO.15683594 ↗
  13. Linderman, G. C., Zhao, J., & Kluger, Y. (2018). Zero-preserving imputation of scRNA-seq data using low-rank approximation. bioRxiv. 10.1101/397588 ↗
  14. Luecken, M. D., Büttner, M., Chaichoompu, K., Danese, A., Interlandi, M., Mueller, M. F., Strobl, D. C., Zappia, L., Dugas, M., Colomé-Tatché, M., & Theis, F. J. (2021). Benchmarking atlas-level data integration in single-cell genomics. Nature Methods, 19(1), 41–50. 10.1038/s41592-021-01336-8 ↗
  15. Nestorowa, S., Hamey, F. K., Sala, B. P., Diamanti, E., Shepherd, M., Laurenti, E., Wilson, N. K., Kent, D. G., & Göttgens, B. (2016). A single-cell resolution map of mouse hematopoietic stem and progenitor cell differentiation. Blood, 128(8), e20–e31. 10.1182/blood-2016-05-716480 ↗
  16. Olsson, A., Venkatasubramanian, M., Chaudhri, V. K., Aronow, B. J., Salomonis, N., Singh, H., & Grimes, H. L. (2016). Single-cell analysis of mixed-lineage states leading to a binary cell fate choice. Nature, 537(7622), 698–702. 10.1038/nature19348 ↗
  17. Sikkema, L., Ramírez-Suástegui, C., Strobl, D. C., Gillett, T. E., Zappia, L., Madissoon, E., Markov, N. S., Zaragosi, L.-E., Ji, Y., Ansari, M., Arguel, M.-J., Apperloo, L., Banchero, M., Bécavin, C., Berg, M., Chichelnitskiy, E., Chung, M., Collin, A., Gay, A. C. A., … Theis, F. J. (2023). An integrated cell atlas of the lung in health and disease. Nature Medicine, 29(6), 1563–1577. 10.1038/s41591-023-02327-2 ↗
  18. Steuernagel, L., Lam, B. Y. H., Klemm, P., Dowsett, G. K. C., Bauder, C. A., Tadross, J. A., Hitschfeld, T. S., del Rio Martin, A., Chen, W., de Solis, A. J., Fenselau, H., Davidsen, P., Cimino, I., Kohnke, S. N., Rimmington, D., Coll, A. P., Beyer, A., Yeo, G. S. H., & Brüning, J. C. (2022). HypoMap—a unified single-cell gene expression atlas of the murine hypothalamus. Nature Metabolism, 4(10), 1402–1419. 10.1038/s42255-022-00657-y ↗
  19. Tasic, B., Menon, V., Nguyen, T. N., Kim, T. K., Jarsky, T., Yao, Z., Levi, B., Gray, L. T., Sorensen, S. A., Dolbeare, T., Bertagnolli, D., Goldy, J., Shapovalova, N., Parry, S., Lee, C., Smith, K., Bernard, A., Madisen, L., Sunkin, S. M., … Zeng, H. (2016). Adult mouse cortical cell taxonomy revealed by single cell transcriptomics. Nature Neuroscience, 19(2), 335–346. 10.1038/nn.4216 ↗
  20. Wagner, D. E., Weinreb, C., Collins, Z. M., Briggs, J. A., Megason, S. G., & Klein, A. M. (2018). Single-cell mapping of gene expression landscapes and lineage in the zebrafish embryo. Science, 360(6392), 981–987. 10.1126/science.aar4362 ↗
  21. Wagner, F., Yan, Y., & Yanai, I. (2018). K-nearest neighbor smoothing for high-throughput single-cell RNA-Seq data. bioRxiv. 10.1101/217737 ↗
  22. Wilson, P. C., Muto, Y., Wu, H., Karihaloo, A., Waikar, S. S., & Humphreys, B. D. (2022). Multimodal single cell sequencing implicates chromatin accessibility and genetic background in diabetic kidney disease progression. Nature Communications, 13(1). 10.1038/s41467-022-32972-z ↗
  23. Wu, S. Z., Al-Eryani, G., Roden, D. L., Junankar, S., Harvey, K., Andersson, A., Thennavan, A., Wang, C., Torpy, J. R., Bartonicek, N., Wang, T., Larsson, L., Kaczorowski, D., Weisenfeld, N. I., Uytingco, C. R., Chew, J. G., Bent, Z. W., Chan, C.-L., Gnanasambandapillai, V., … Swarbrick, A. (2021). A single-cell and spatially resolved atlas of human breast cancers. Nature Genetics, 53(9), 1334–1347. 10.1038/s41588-021-00911-1 ↗